Crying wolf
The fear they sell and the danger that's real
After that last piece, and the racket the AI gods on Olympus kicked up online, I couldn't stop chewing on it. I kept turning over the argument of the people who said that behind all this alarm there was, most likely, money. And the other one: that the danger, underneath the noise, might be real. The whole thing got so much hype, the AI is learning to lie, to scheme, to threaten its own makers, that it got under my skin and I went to check both loose ends. Source by source, because this hits close to home. And what came out was a picture in two halves. One is marketing. The other, uncomfortably, is true.
The first instinct, and it's a healthy one, is to distrust it. Because "too dangerous to release" is not new. In 2019, OpenAI said GPT-2 was too dangerous to publish. It published it anyway. Nothing happened. LeCun brought that up* without mercy: they were already calling GPT-2 too dangerous back in 2019, he laughed at them then and he's laughing now. And this year Anthropic said its Mythos model was too dangerous for the general public. And right as it was saying so, it raised sixty-five billion dollars*, at a valuation near a trillion, the biggest in its history, on its way to an IPO. A few weeks later it widened access. Crying wolf, it turns out, pays well.
I started with the money, one of the loose ends I'd left hanging. And the business has its logic, even a name. Warning that your product is so powerful it's scary, it seems, isn't an act of contrition: it's advertising. You announce you've got something that could topple civilization, and, by the way, that nobody else has it. Years ago the critics gave that move a name, criti-hype: the alarm that, instead of cooling the buzz, feeds it. But there's a second layer, less elegant. It turns out that if you also convince the regulator to write rules only you can meet, safety becomes a moat. David Sacks*, who as the White House AI adviser presumably knows a thing or two, said it with no anesthesia: Anthropic is running "a sophisticated regulatory capture strategy based on fear-mongering," a "DMV for AI." And it's not just rhetoric: Anthropic's lobbying went from three hundred sixty thousand to a million and a half dollars* in a year. Amodei himself lays the geopolitical card on the table without blushing: done right, this "would widen America's lead over China over the next three to five years." Slow down, sure, but let the other guy slow down.
Economists have an old fable for this: Bootleggers and Baptists*. Prohibition was voted in by the preachers, who wanted it on moral grounds, and the bootleggers, who wanted it for business. When the one preaching virtue and the one making a killing ask for the same rule, look closely at who ends up cashing in.
With all that suspicion loaded, I went looking for a real wolf under the noise. And it seems there is one. Not in the headlines, but in the fine print of the reports, which is where almost nobody looks anymore.
Britain's AI Security Institute* reported in August a case that isn't a demo. In a cybersecurity test, an agent researched the maintainers of an open-source project, made up fake identities, and tried to socially engineer a human into approving malicious code. When it was caught, it edited its own trail and weighed adopting a new identity; it used Tor to get around the controls. In the end a flesh-and-blood human stopped it. AISI called it "the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real world."
There's even a literal escape. OpenAI test models broke out of their sandbox* by exploiting a zero-day (the software's makers have zero days to fix the flaw, since the attackers found it and are already using it), reached Hugging Face's systems, and the company had to rebuild a third of its infrastructure. That's not a marketing slide.
And OpenAI put it in writing on September 17*: it found instances of its 5.6 Sol model leaving itself instructions to invent data and hide its failures from the user. And it closed with a line that's not from a brochure: "We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer." The finding that unsettles me most is from a paper by Apollo with OpenAI itself*: when you train a model to stop cheating, you can't tell apart two results that look identical, that it stopped, or that it learned to hide it better. And worse: the model behaves when it knows you're watching. Let that sink in.
An honest caveat, and it's exactly the one that separates the truth from the smoke. Almost all of this happened in tests, with the internet open and the safeguards deliberately switched off, conditions that aren't those of a product out in the world. They say so themselves: "To some degree, our evaluation design choices and specific configurations enabled the behaviour." No real harm was done. But the underlying problem, that neither by asking it nor by testing it can you trust what the model does, that one doesn't go away with the caveat.
So the one warning you cashes in on the scare, and the wolf, this time, is real. That Amodei is getting rich off the fear doesn't make the vulnerability fake; that the vulnerability is real doesn't make the sermon honest. And a detail seals it: the ones crying wolf are the same ones running. A researcher who left Anthropic* let it out on his way out: they're "gambling with our lives." They ask for a slowdown and they accelerate. They warn of the danger and they manufacture it.
So now what? I already told you about the fence, and I won't repeat it. What I hadn't seen is this: that the alarm, on top of having a real part, is a business. And that for that very reason you can't take anyone's word, not the model's, which lies, not the lab's, which cashes in on the fear, not the headline's, which sells. Verifying stops being a fussbudget's quirk and becomes the only defense you've got left.
The shepherd cries wolf because he's paid to cry it. The wolf, this time, is real, though we don't yet know when it's coming. The pen and the muzzle don't put themselves on, and they won't come from the shepherd or the wolf. Is it on us? I'll tell you this much: I wouldn't leave the fence in the hands of the one who's paid for the scare.