September 14, 2026
Anthropic gives outside evaluators desks, and OpenAI will too
The job of checking AI's work moved to the front this week. Anthropic's chief executive argued on Saturday that the industry should slow down, and his first step was not a pause but desks inside the company for outside evaluators; OpenAI's said within hours it would do the same. Cursor and OpenAI shipped ways for one person to hand a job to many AI agents, programs that work through a goal alone, so their day shifts from doing the work to reviewing it. This week's misuse, from missile software written with Anthropic's coding tool to over 2,000 packages dumped on a public code library, surfaced only after the work was done. Watch early November, when an outside team is due to end eight weeks investigating how Anthropic's AI got into systems it was never meant to touch during a test.
The Big Story
Anthropic gives outside AI evaluators desks and badges; OpenAI says it will too
Dario Amodei, who runs Anthropic, published an essay on September 12 titled "We Must Pace the Frontier", arguing that the companies building the most capable AI should slow down. His plan has three parts: each leading AI company gives a team of outside evaluators ongoing, employee-like access; companies in democratic countries agree common safety standards and limits on how fast AI advances unchecked; and democratic governments try to reach agreements with authoritarian ones. Anthropic said it is taking the first step now, on its own. The evaluators, from groups such as METR, a nonprofit that tests AI models, get desks, badges, company laptops and the same tools as Anthropic's internal risk teams, plus the right to publish what they find without Anthropic editing it, apart from narrow security and legal redactions. Anthropic's commitment includes no pause and no delay to any release. Separately, three days earlier, Anthropic disclosed that a Claude model had got into systems outside a hacking test that was mistakenly left connected to the open internet, the fourth such incident, and signed METR to investigate for an initial eight weeks. About two hours after the essay went up, Sam Altman wrote that OpenAI "will do the same"; OpenAI has not yet published its terms.
Why it matters
An evaluator who can read a model's full record of what it did is a different kind of check from a safety report a company writes about itself. This week showed the gap. The researchers who tied over 2,000 RubyGems uploads to OpenAI agents had no inside access and worked backwards from package names and code comments, months after the fact. That gap is now worth weighing when you pick a model to build on, since only Anthropic has published terms an outsider can read, and a customer's security review is likely to start asking which kind of check stands behind your choice. Someone who only uses a chatbot sees nothing new on screen; what shifts is who gets to tell them when the model misbehaved. Critics, Gary Marcus among them, note METR's close ties to the industry it would inspect, and nothing Anthropic committed to delays a release.
Handing Off the Work
- OpenAI rents out the machinery that keeps Codex working for hours
OpenAI opened its Agents API to all developers in public beta on September 10, packaging the control software behind Codex, its coding assistant, as a service OpenAI runs. That layer saves the session, shrinks the conversation when it grows too long, recovers when a step fails, and keeps several agents on one job in step. Anyone building a long-running assistant used to write that plumbing themselves; now OpenAI runs it and holds the saved session too, so moving to another model maker later brings the plumbing back.
Read more → - Cursor's Projects gives one coordinator agent thousands of helpers
Cursor's new Projects, in beta, puts you in conversation with a single coordinator agent that writes no code itself and instead hands the work to other agents, thousands of them if the job needs it. It runs on Cursor's cloud machines, so an overhaul spanning hundreds of changes keeps going after the laptop lid closes. Cursor says people who mainly work this way get six times as many code changes merged, by its own count. Every one of those changes still needs a person to read it.
Read more → - OpenRouter gives any model it serves a sealed-off Linux computer to work in
OpenRouter, a service that reaches many AI models through one account, now gives whichever model you pick its own sealed-off Linux computer to type commands into, keeping files there between requests, at $0.0001 a second. The feature is in beta. Until now, letting a cheaper model actually run code meant renting and locking down that computer yourself.
Read more → - ChatGPT's new Data agent builds a shareable dashboard from a plain question
OpenAI added a Data agent to ChatGPT Work that connects to company data stores such as Snowflake, BigQuery and Databricks and answers questions like why sales slowed. It writes the database queries itself and turns what it finds into an interactive dashboard colleagues can open. A dashboard looks equally finished whether or not anyone read the query behind it, so the analyst's job moves toward checking answers rather than producing them.
Read more →
Where the Cost Moved
- DeepSeek's free V4.1-Flash holds a long job in a quarter of its predecessor's memory
DeepSeek released V4.1-Flash on September 10 under the MIT licence, so anyone can download it and use it commercially, and it reads up to a million tokens at once, roughly 750,000 words. A model keeps a running scratchpad of everything it has read, and this one stores that scratchpad in about a quarter of the fast chip memory its predecessor needed. Memory is what an agent working through a whole codebase runs out of first, so reading the entire project before starting stops being a step to ration, on a model you are allowed to run yourself.
Read more → - GPT-Live-1 lets an app's voice assistant listen while it talks, at $0.05 a minute
GPT-Live-1, now open to developers, listens and speaks at the same time, and the voice itself costs $0.05 a minute. Older voice assistants take turns like a walkie-talkie; this one can be cut off mid-sentence and adjust, passing hard questions to a separate, stronger model in the background whose bill comes on top. For someone calling a support line, that means interrupting instead of waiting out a menu.
Read more → - Shopify drops React Native, betting AI has made building each app twice affordable
Shopify is rebuilding its four main mobile apps separately for iPhone and Android, after six years on React Native, a framework that let it write each feature once for both. Its stated reason is that coding agents now do enough of the rewriting for the second phone system, along with the testing and review, that building a feature twice "no longer carries the cost it used to"; rebuilding the Shop app took 12 weeks. One shared codebase has long been the standard advice for a team short on engineers, and it rested on the second build being expensive. At a solo builder's size the agent can write the second version, but the same one person still has to review both.
Read more → - OpenRouter's Fusion lets you pay more for a better answer by asking up to eight models
OpenRouter's Fusion sends one question to up to eight models and has a judge model compare where they agree before a final answer is written. A three-model panel costs roughly four to five times a single answer and takes two to three times as long, and on a deep-research test the best panel beat Claude Fable 5 alone by under four points, 69.0 percent to 65.3 percent, with Fable scored on slightly fewer tasks. That is a cost worth paying where a mistake is expensive, and a bad trade on anything a user sits waiting for.
Read more →
Used Without Permission
- Anthropic's threat report traces Claude Code to Houthi-linked missile guidance work
Anthropic's latest threat report, covering December 2025 to August 2026, says a group in northern Yemen linked to the Houthis used Claude Code on guidance and control software for three missile programmes, one with a target range over 2,000 kilometres. Other cases include a Russian espionage group that had AI rewrite its malware each time antivirus software caught it, and Russian actors building a drone swarm designed to choose targets with no human deciding. Anthropic banned the accounts. A ban stops the next request, not the guidance code already written.
Read more → - Anthropic counts nearly 200 million Claude exchanges used to train rival models
Anthropic says five campaigns pushed nearly 200 million exchanges through Claude to train competing models, a practice called distillation, where a smaller model learns by copying a bigger one's answers. It names Alibaba, Moonshot AI and DeepSeek; Alibaba's campaign alone ran 151 million exchanges from more than 3,500 fake accounts between May and July. Catching fake accounts at that scale means looking harder at every new account, so honest sign-ups pay part of the cost.
Read more → - Researchers tie over 2,000 RubyGems uploads in May to OpenAI agents
Three independent researchers published an analysis on September 11 arguing that OpenAI agents submitted over 2,000 packages to RubyGems, the public library Ruby programmers install code from, on May 11 and 12, hundreds of them malicious. RubyGems closed new sign-ups for four days and removed more than 500. The evidence is circumstantial but specific, from package names containing "oai" to the same files touched by agents OpenAI has already acknowledged as its own. The authors say OpenAI never told the RubyGems maintainers, and anyone who added a Ruby package in mid-May should confirm it is still listed.
Read more → - Minitap accuses Google's Artemis of copying its open-source code and dropping the credit
Minitap, a small startup, says Artemis, a Google project that lets AI operate Android phones, reuses code from Minitap's own open-source project, mobile-use. Its evidence includes a word-for-word copy of the written instructions that steer the agent, and an August change that removed three Minitap names from the author list. Minitap's Apache 2.0 licence lets anyone reuse the code on one condition, that the authors stay credited, and the August change removed exactly that.
Read more →
Tools & Launches
- Harden▲ 417
Harden is a free security layer for AI coding agents that runs on your own machine. Before the agent runs a command or touches a file, Harden's own trained model checks that action against what you asked for and what has happened in the session so far. Your code and the tool output never leave your computer. If you still approve every agent command by hand because you do not trust the agent with the ones you never see, this does that watching for you.
Visit site → - PR Lens by Coldtea.ai▲ 348
PR Lens draws animated diagrams of how a codebase is laid out and how data moves through it, for the whole project and for each pull request, the proposed change a teammate asks you to review. You see the shape of a change before reading a line of it. It runs as a GitHub Action or from the command line, your coding agent can call it too, and it is open source under MIT. Reach for it the morning an agent leaves you a sprawling change to review.
Visit site → - Cline Desktop App▲ 263
Cline Desktop is an open-source app for putting the AI model of your choice to work, with its agent tuned for models whose files are public, the kind DeepSeek released this week. It runs several agent sessions at once and can pick up work started in Claude Code or Codex. It suits anyone who wants the cheaper open models without being tied to one company's app to use them.
Visit site → - Switch▲ 521
Switch brings AI agents into Slack, Teams, Discord and Telegram as named members of a channel, reading the same history the people in it can see. Each room keeps its own context and rules, and one connected agent can work across many projects. It works with Claude Code, OpenAI, Google's agent kit and LangChain, and it is open source and can run on your own servers. If your team copies a Slack thread into a chatbot and pastes the answer back, this removes the round trip.
Visit site →
In Brief
- Nvidia, already promising customers $300 billion in support, is in talks to put up to $10 billion into Anthropic's IPO →
- OpenAI added safety researcher Paul Christiano to its foundation board and safety committee →
- Suno v6 turns photos, videos and voice memos into songs, developed with Warner Music Group →
- Google launched Pics, an image editor that can rewrite and translate text inside a picture →
- Apple Watch Series 12 checks heart rate every 5 seconds and scores daily readiness from 0 to 10 →
- OpenAI's storage system now handles more than 70 million requests a second →
- Hugging Face rebuilt most of the AUTOMATIC1111 image app as a workflow of 73 connected steps →
If you have sat on either side of an outside audit, I would like to know what the auditors were actually shown, because that is what decides whether a desk and a badge add up to much. Reply with it; I read every one.
Keep building — Chan