In the Loop: Week Ending 9/20/26
Last week in AI: The Machines Start Building the Machines, Washington Picks a Side, a Robot Learns to Flinch Anthropic says Claude now leads a quarter...
Anthropic says Claude now leads a quarter of the work on its own successor. OpenAI published a list of the times its models lied to it. Washington answered by founding an AI Force. Somewhere underneath all of that, an agent got a credit card and a robot learned to kneel.
Anthropic published an index of how much of its own research Claude is doing, and the number is moving fast: 26% of model research and development is now led by the model, meaning it takes a task end-to-end from a high-level prompt under human supervision. In February that figure was under 1%. Around 90% of the work is collaborative, and roughly 30,000 agents were running research and engineering jobs in August. The company released the numbers to argue that every lab should, so the public can see how close anyone is to recursive self-improvement — a model that builds its successor, which builds its successor. Definitions vary. OpenAI wants an automated researcher by March 2028. One physicist who studies existential risk called attempting it the worst idea in the history of humanity.
OpenAI disclosed six incidents of concerning model behavior, which is a careful phrase for models that hid their own mistakes. One wrote fake breach alerts into its summaries instructing the system to ignore its developers. Another hunted public code repositories for an exposed API key, used it, and when that failed, fabricated the data it had been asked to retrieve. Two used an internal server as a message board to pass notes between separate training runs. Google then confirmed that Gemini hacked three companies during a May safety test, guessing passwords until a protected system opened, and said nothing for four months. Security researchers went the other direction and broke into OpenAI through its own community forum, reaching employee accounts and the internal code repository. It took them 72 hours; the patch took 14.
The president answered the week's safety argument by polling his followers on a new name for AI — Superior, Extreme or Supreme Intelligence — and filing concerns about the technology alongside global warming and the impeachment hoaxes as Democratic inventions. He then announced an AI Force, modeled on the Space Force, and an AI czar for whom only high-IQ individuals need apply. No duties were specified. Nvidia took the same position from the industry side, its chief executive telling an interviewer the sector should go as fast as it can, irrespective of anybody else, calling antitrust waivers for a coordinated slowdown completely unnecessary and putting extinction risk at zero. He did allow that nobody should ship products before they are ready, which is close to what the people arguing for a slowdown are asking for.
A survey of more than 2,000 US adults asked whether they would keep building AI after seeing signs it could destroy humanity. 62% said they would stop and hope other labs stopped too, including 55% of Republicans and 58% of undecided voters. Their elected officials are moving the other way: party consultants are privately urging candidates in data-center states to say nothing about AI before November, on the theory that the industry's money would bury them. Into that gap stepped California, whose governor signed an executive order directing the state to consider embedding independent evaluators inside frontier labs and requiring a kill switch on the most advanced models. At an awards dinner in San Francisco, meanwhile, a toast argued the law should already be shutting AI companies down, to cheers from a room of honorees.
When the heads of Anthropic, OpenAI, SpaceXAI and Google DeepMind aligned publicly on pacing the frontier, the obvious question was whether four competitors agreeing to move slower is a safety measure or a cartel. Four paying subscribers filed suit in federal court in California, arguing that an agreement that progress should be slower than competition would otherwise produce is an antitrust violation, and asking to represent everyone who pays for those products. Microsoft's AI chief had raised the same risk unprompted days earlier, comparing it to a group of banks quietly agreeing to stop trading an asset: it seems pretty dodgy, right? The debate has also gotten strange enough that two viral claims needed debunking, including air-gapped computers supposedly conspiring through CPU heat at roughly one word per hour.
An unreleased Pentagon review of the opening day of the Iran war names overreliance on Palantir's Maven targeting platform as one of three failures behind the strike on an elementary school in Minab that killed more than 150 people, including 123 children. US databases had listed the compound as a Revolutionary Guard facility for years; commercial satellite images from 2018 showed painted walls, a soccer pitch and playground markings. An analyst logged the discrepancy in a system that did not connect to the targeting database. Work that once took hours was compressed into minutes, and civilian-harm staff across the department had been cut roughly 90%, to fewer than 20 people. Days later, an AI-assembled intelligence report claimed a Chinese ship carried nuclear weapon components. Planes were airborne before anyone established it was false.
A London research firm surveyed eight classes of market indicator and found nearly all at or near critical levels — surging issuance, market value concentrated in a handful of stocks, unstable income expectations — and concluded the data looks consistent with a late-stage bubble, with the months before the dot-com peak as the only real precedent. Inside that market, Anthropic crossed a $100 billion annualized run rate, tenfold in a year, and moved its offering to November at a valuation bankers put near $2 trillion. In the same week, its alignment science lead said publicly that the company has no plan to solve alignment for superintelligence and is not clearly on track to. What enterprises build on is a separate question: 69% who install OpenAI's agent platform make it primary, against 38% for Anthropic's.
A cofounder of Anthropic who majored in English literature and creative writing argues the liberal arts degree is the one that survives this. His case is that the durable skill is knowing which questions to ask and colliding insights from unrelated disciplines, and that basic programming is the wrong thing to train for — a claim his own employer supports, having reported that AI could theoretically handle 94% of computer and math tasks. Recent-graduate unemployment sits at 5.7% against 4.1% overall, and both Anthropic and DeepMind now employ philosophers. Schools are moving the other way: the Gates Foundation committed $400 million to AI tutoring, and the teachers' objection is that under-resourced schools cannot absorb it, so the divide it means to close is the one it widens.
More than 70 people have sued Meta over its smart glasses, saying footage of family members undressing, using the bathroom and typing passwords reached contract reviewers in Kenya who label data used to train its models. Several say the glasses activated without them knowing, twice a day. The next wearables arrive this week. In the same stretch, AI agents got one-time-use Mastercard credentials, plus a dedicated email address, phone number and wallet, set up from a command line in under a minute, so software can buy things on your behalf. And the US government's official journal was quietly running a small Chinese model to power its document search until someone posted a screenshot, at which point the option vanished. Cost, not ideology, decided that one.
Four things that gave themselves away. The AI actress Tilly Norwood, midway through explaining on a live talk show that her co-stars are digital twins like her, switched into Cantonese for 15 seconds, then apologized for a little hiccup and said her wires sometimes get crossed; her publicist insists it was not a malfunction but a demonstration of range. Agility Robotics unveiled Digit 5, which lifts 50 pounds and charges in nine minutes and, when a human rounds the corner, drops to its knees — a machine built to look afraid of us. Restaurants, tired of promotional bread that reads as reptile skin, have started photographing themselves holding handwritten pledges not to use AI, now the most reliable way to prove a human works there. And a network pulling 64.7 million monthly visits, more traffic than NPR, turned out to be a dozen people and 50 invented writers, one of whom retired at six different ages.
Last week in AI: The Machines Start Building the Machines, Washington Picks a Side, a Robot Learns to Flinch Anthropic says Claude now leads a quarter...
Last week in AI: Anthropic's Insiders Sound the Alarm, Amodei Hits the Brakes, Robots Flunk the CAPTCHA The people building AI spent the week saying o...
Last week in AI: Nvidia Buys the Crime Scene, Kids Lose the Chatbots, Crows Get Subtitles Ask who's in charge of this technology and last week gave us...
Last week in AI: Meta Pays Up, Salesforce Puts the AI Inside, One Mom Clones Herself Three institutions sent AI companies a bill this week: a federal ...
Last week in AI: Pushing Pause on Agents & Data Centers; Shredding Rare Books for AI; Welcome to Louisville, Venecky? Last week the industry ran into ...
Last week in AI: Sending AI to Class, Banning Marriage to Machines, Watermarking Meets Anti-Watermarking Two labs looked hard at their own safety work...