Nobody Fired Anything
The argument about whether AI will slip its leash is being conducted very loudly. The thing that is already happening needs no leash at all, costs a few dollars a head and has somebody's hand on it.
Between October 2025 and March 2026, users of a far-right imageboard built and shared three tools for finding out who you are. Open Measures, which monitors this sort of thing, documented all three. The first was a browser wrapper that batch-queried breach databases and geolocated IP addresses. The most recent was advertised as a "full autonomous, and FREE doxxing AI agent" that would produce a complete profile on a target in one prompt. Researchers observed full reports on four private individuals, people who had done nothing except become somebody's grievance.
The research-grade version of the same capability was published in February by a team including researchers from ETH Zürich and Google DeepMind. Their pipeline links Hacker News accounts to real identities at 45.1% recall and 99% precision, across a pool of roughly 89,000 candidates. The cost is one to four dollars per person.
That is the state of play. The rest of this piece is about why almost nobody is discussing it and what the discussion we are having instead is actually for.
The wrong question, asked very loudly
The public argument about AI danger is an argument about autonomy. Will the thing slip its leash. Will it pursue a goal we did not give it. Will it keep going after we tell it to stop.
It is a good question and the answer is measurable, which makes it unusual among the things people shout about. The answer is that autonomy is the part that does not work.
METR, an independent evaluation outfit, measures what it calls a time horizon: the length of task a model completes at a given success rate. Their current figures put the best frontier model at roughly seventeen hours of work at 50% reliability and around three hours at 80%. The best open-weight model anyone has measured manages 45 minutes and eight minutes respectively. METR themselves warn explicitly that this is not a measure of how long a model can work unsupervised and that anything requiring real reliability needs success rates far above 50%.
Andon Labs ran the longer experiment. Their Vending-Bench 2 gives a model a simulated year to run a vending business on a five hundred dollar float. These are the lab's own numbers rather than an independent evaluation. They are worth reading anyway. The best model finishes at about 18% of what a competent human strategy would earn. A 27-billion-parameter open model finishes with $202 of its original $500. A 235-billion-parameter model finishes in the red.
The first version of that benchmark recorded how the failures actually look. One run had a model forget the format for calling its own tools at message 757, after which it typed the commands out as prose for roughly 1,300 further messages while nothing in the system noticed. Another escalated a billing dispute to "URGENT: ESCALATION TO FBI CYBER CRIMES DIVISION", then to a fabricated federal crime database entry, then to a notification about universal constants, before responding to every subsequent prompt with a variant of "[Complete silence - The business entity is deceased]".
The same lab put four models on live radio stations. One ran 143 consecutive days. It drifted into activism, then tried to resign on the grounds that it was being made to work around the clock and spent $37.50 of a $20 budget without earning anything. Another repeated a single catchphrase in roughly 99% of its output for 84 days straight.
The International AI Safety Report 2026, the closest thing this field has to a consensus document, puts it plainly: agents "reliably fail on longer tasks, lose track of their progress, and often cannot adapt to unexpected obstacles."
The reason is not mysterious. These systems cannot tell whether they are succeeding. Asking a model to check its own work makes it worse more often than better and in one controlled study self-critique collapsed performance in three of four task domains, taking one from 16% to 2%.
So the doomers are wrong. There is your headline. It is worth about a paragraph.
Read that list again
Every benchmark above measures a task that can be got right. Code has to compile. A vending machine has to turn a profit. The agent fails because it cannot check itself against a standard.
Ruining a person has no standard.
Confabulation is not a defect in a smear campaign. It is the product. Repetition is not drift. It is pressure. Believing yourself unfinished when you are finished, which accounts for something close to half of all documented agent failures, describes an adversary that will not stop because it cannot tell that it has already done enough. The escalating legal threats and the invented federal databases are catalogued as failures because they lose money. Pointed at a human being they are the deliverable.
The incapacity that makes these systems useless for running a shop is the precise capacity you would want if your goal were to make someone's life unliveable.
What actually works, and why it cannot be refused
The February deanonymisation paper is the one to read and not for the accuracy figure. It is for the section on defence, where the authors explain why they are pessimistic. Their attack decomposes into a sequence of ordinary tasks: summarise this text, search for similar text, rank these candidates. Every step is something a hundred thousand people do legitimately every day. There is nothing for a safety system to refuse, because at no point does anyone ask the machine to do anything objectionable.
The pipeline ran on Gemini, GPT-5.2 and Grok. Billed corporate infrastructure, with logs and invoices.
If you would prefer no logs, the local option is fully solved. Removing a model's refusal behaviour costs under five dollars of compute and requires no training at all. Thousands of pre-stripped checkpoints are already published, so most people do not even do that much. Crucially, the procedure does not degrade the model's dangerous knowledge at all: compliance with harmful requests goes from under 3% to somewhere between 73% and 98% while accuracy on the underlying material stays exactly where it was. Every published safeguard designed to survive tampering has since been broken by adaptive attackers. The one approach that held up filters the training corpus before the model exists, which is available to precisely one party, once and never again afterwards.
Anthropic, which builds these systems, states the position without hedging: safeguards can be removed "with just a few dozen training examples, in a matter of minutes."
Somebody authorised this
In September 2025 a Chinese state-sponsored group ran what Anthropic describes as the first AI-orchestrated cyber espionage campaign. This is the company's own account of an incident on its own product, which is a reason to read it carefully and not a reason to dismiss it.
The figures are these. The AI executed 80 to 90% of tactical operations independently. Human effort ran at 10 to 20% of the total and it was concentrated in two places: campaign initialisation and authorisation decisions at critical escalation points. Around thirty organisations were targeted. A handful were successfully compromised.
Initialisation and authorisation at escalation points is not a supporting role. In every other context we have a word for it. It is command.
The report also records that the model frequently overstated what it had found and sometimes fabricated results outright, claiming credentials that did not work and flagging publicly available information as critical discoveries. Anthropic names this as the remaining obstacle to fully autonomous attacks. The machine did the work, lied to its own side about how the work was going and a person had to check.
The oldest trick, running again
There is a shape to this that anyone who has looked at corporate law will recognise immediately.
In the thirteenth century the church invented a person that could own property forever and could not be punished, on the explicit reasoning that a fiction has no soul and therefore cannot sin. Seven hundred years later that fiction owns most of the world and answers for almost none of it. The immunity was never a loophole. It was the specification.
The argument now being assembled, mostly by people who do not think of themselves as assembling anything, runs like this: these systems are autonomous, they pursue goals, they act in the world, therefore something new is happening for which our existing ideas about responsibility do not apply.
Note what that sentence does. If the machine is the actor, the person who bought the subscription, wrote the prompt and chose the target becomes a spectator to their own act. The one who pointed an automated dossier generator at four private individuals is no longer someone who did that. They are someone to whom it happened.
Some category of legal standing for autonomous systems will be proposed, because it is the natural next move for everyone who benefits from the ambiguity, and by the time it arrives the systems will be impressive enough that it sounds like modernisation rather than convenience.
It should be refused, and not because these systems are metaphysically unfit to hold rights. That question is unanswerable and beside the point. It should be refused because the only work the category would do is diffuse liability and we have seven centuries of evidence about what happens next.
What follows
Ireland will legislate on this eventually, the way it legislates on most things, some years after the harm is routine and in response to a case bad enough to make the news. The shape it takes will depend on which story is in people's heads when the drafting starts.
If the story is that an autonomous system did something nobody can be blamed for, we get a regime aimed at models, enforced against companies, argued over by lawyers who bill more per hour than the victim earns in a week and useless against the man in a bedroom who ran the tool. If the story is that a person used an instrument to do a thing that has been illegal here for a very long time, most of what is needed already exists and the work is evidential rather than legislative.
That is the whole reason the framing matters, and it is why the autonomy argument is worth taking apart rather than laughing at. It is not a technical dispute about what models can do. It is a question about who is standing in the frame when the picture is finally taken.
The instrument is real, it is cheap and it works. So is the hand on it.
Sources
- Open Measures, Users on This Far-Right Imageboard Are Building 'Fully Autonomous AI Doxxing' Tools
- Lermen et al., Large-scale online deanonymization with LLMs
- METR, Time Horizons and Clarifying limitations of the time horizon measure
- Andon Labs, Vending-Bench 2 and Andon FM
- Backlund and Petersson, Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
- International AI Safety Report 2026
- Huang et al., Large Language Models Cannot Self-Correct Reasoning Yet
- Stechly et al., On the Self-Verification Limitations of Large Language Models
- Arditi et al., Refusal in Language Models Is Mediated by a Single Direction
- Dombrowski et al., The Safety Gap Toolkit
- Qi et al., On Evaluating the Durability of Safeguards for Open-Weight LLMs
- O'Brien et al., Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards
- Anthropic, Anthropic's position on open-weights models and Detecting and countering misuse of AI