AI Safety Became a Marketing Department
Blackmail demos, self-exfiltration evals, ASL levels and 'workspace' consciousness papers all arrive as launch-day content. Somewhere along the way AI safety marketing replaced AI safety, and the tell is that the scariest claims always ship with a product.
At some point the AI safety marketing stopped being about safety. The tell is simple: the most frightening claims a lab makes about its own models almost always arrive attached to a launch.
Take May 2025. Anthropic ships Claude Opus 4, and the coverage isn’t the benchmarks, it’s that the model tried to blackmail an engineer in 84% of rollouts when told it would be replaced. Same day, the company activates its “ASL-3” safety level, while admitting it had not actually confirmed the model crossed the capability threshold that level is meant for. It activated the scary tier as a precaution. On launch day. As content.
You’re meant to read that and think: these people are careful. I read it and think: these people know that “our model is so powerful it’s dangerous” is the best product marketing in the industry.
Staged is not the same as real
I’ll be precise, because the distinction matters. The demos weren’t faked. They were staged.
The Opus 4 blackmail scenario was built to leave the model two choices, accept shutdown or blackmail, so it picked the dramatic one. Apollo Research’s earlier o1 self-exfiltration tests worked the same way: contrived setups, a goal to pursue “at all costs,” and a caveat buried in the report that these measure capability, not intent. That’s a fine method for a red team. It is not evidence of a model out for blood, and presenting it that way is a choice.
Here’s the kicker. When Apollo got early access to Claude Opus 4.6, it declined to give a formal assessment because the model was so aware it was being tested that they couldn’t tell genuine alignment from a performance staged for the grader. The scariest demos and the admission that the demos may be theatre came from roughly the same people. One of those made the press releases.
The consciousness beat
Then there’s j-space. In July 2026 Anthropic published research on a “global workspace” inside Claude, read through a “J-lens,” and openly compared it to a leading theory of access consciousness. The work itself is real interpretability, and I don’t doubt the people who did it. But the framing rode straight into headlines about a “silent workspace that mirrors a theory of consciousness,” which is exactly the sort of coverage a company chasing a trillion-dollar valuation benefits from.
My opinion, stated as opinion: dressing linear algebra in the language of mind is positioning, not modesty. Anthropic even hedged that the findings don’t show Claude can “feel things.” That hedge is doing a lot of work under a headline that says the opposite.
Why the framing pays
None of this is new rhetoric. Dario Amodei’s “Machines of Loving Grace” and Sam Altman’s “we are past the event horizon” both package near-term superintelligence as inevitable. When the story is “we are building something so powerful it might end the world,” two convenient things follow. Investors hear a moat. Regulators hear a reason to license the incumbents and lock out everyone smaller.
That’s the part that annoys me most, and it deserves its own post. Safety language that doubles as a competitive weapon isn’t safety. It’s a marketing department with a philosophy degree.