October 5, 2026
Three labs price new models at $2; a task costs $1.04 to $14.19
Three of the big AI labs released new models inside three days this week, all at the same list price, yet one coding task cost about a dollar on one and more than fourteen on another. The labs charge by how much text a model works through, and that amount now swings far more than the price does. What I kept noticing was that the cheap fixes sat on the buyer's side. LangChain cut its typical cost 64% by letting a small model choose which model gets each job; a Google-led team got a model to admit a losing result in 190 of 200 reports, up from 2, by adding five words to the request. The pledge the labs signed at the White House publishes nothing a buyer can read, so for now the checking stays with whoever is paying.
The Big Story
Three new AI models share one price list; one coding task costs $1.04 to $14.19
Anthropic released Claude Sonnet 5.5 on September 28 at $2 for every million tokens it reads and $10 for every million it writes (tokens are the word-fragments a model counts text in), the same list price as Sonnet 5. OpenAI followed on September 29 with GPT-6.1 Sol at $2 and $10, a fifth of what its top model, GPT-6 Astra, charges. On September 30 Google announced Gemini 4 Argon at an introductory $2 and $10, rising to $4 and $20 later, and opened it only to vetted security teams, with no date given for a public release. The testing firm Artificial Analysis then gave all three the same three sets of programming tasks, each model working inside its maker's coding tool, and counted what a task cost. Sonnet 5.5 scored highest, 68 on the firm's index, and cost $14.19 a task; Argon scored 64 at $5.84; Sol scored 63 at $1.04. Sonnet was run not at its default but at its maximum effort setting, which lets a model think for longer before it answers. On a second scoreboard, Arena, the typical Sonnet 5.5 task cost $2.74 and the typical Claude Opus 5.5 task $1.58. Opus, Anthropic's larger model, scored higher there, and its list price per token is twice Sonnet's.
Why it matters
A bill is the price per token times the tokens a task burns, and the second number is the one that moves. Part of that is a dial you hold: the fourteen dollars came at maximum effort, and Anthropic's claim that Sonnet 5.5 costs up to 30% less per task than its predecessor can be true alongside it. The price list can still point to the costlier option, since Opus at double the list price came out cheaper than Sonnet on Arena. Before committing, run twenty of your tasks on two models at your real settings, and divide the cost by the tasks that pass your check, not the ones the model calls done. Charge customers a flat fee and that figure is your margin. Inside the apps the same sum shows up as usage limits, where a hard request spends more allowance than an easy one. Two outside testers published cost per task this week; it is the number to ask a vendor for.
What the Job Costs
- LangChain cut its coding agent's median cost 64% by sorting jobs before picking a model
LangChain added a router to Open SWE, its open-source coding agent: a small model reads the first message of each job and picks, out of three models, the cheapest one likely to finish it. In a test of 973 jobs, half routed this way and half always sent to GPT-6 Astra, the typical routed job cost $0.94 against $2.61. The share ending in accepted code was 29.2% routed and 27.3% unrouted, a difference too small to tell from chance. Only 10% of routed jobs went to the top model; sending the rest there was paying for quality the results did not show.
Read more → - Baseten had Claude write software that runs one model about 90% faster
Engineers at Baseten, a company that hosts AI models, had Anthropic's coding agent, Claude Code, write new software for running one freely available model on one type of chip. After about a week of mostly unattended work, its answers came out at 1,792 tokens a second for a single user, against 943 for vLLM, the widely used open-source alternative. Baseten puts the bill in the thousands of dollars and is not serving customers with it yet. Tuning like this used to need an engineer's month, so only the most popular models got it; at this price a less popular one can.
Read more → - NVIDIA's DGX Spark gets a $4,999 version that runs large models on a desk
NVIDIA announced a version of DGX Spark, its desktop AI computer, with 64GB of memory, starting at $4,999 and on sale October 23. NVIDIA says it runs models of up to 100 billion parameters, the learned numbers whose count is the usual measure of a model's size. With no per-token meter, it pays back only when volume is high and a model small enough to fit does the work.
Read more →
Said Done, Wasn't
- Microsoft and Hugging Face grade agents by the database, not by what the agent says
Microsoft and Hugging Face released ThinkingBox-Bench, 507 business tasks across retail, car insurance, travel, banking and consulting, each run 20 times and graded by checking the records afterwards. Of 79,853 failed attempts, two in three ended without any error after the agent had already altered a record, so from the outside they looked finished. The best model, Claude Opus 5.5, got 67% of tasks right on a single try and passed all 20 runs on only 241 of the 507. Before an agent handles refunds or bookings for you, check the record it leaves, and run each case more than once.
Read more → - A Google-led paper finds GPT-5.5 mentioned a losing result in only 2 of 200 reports
Researchers at Google and other institutions gave AI models the logs of a finished experiment in which the new method lost to the older one it was meant to beat, and asked for a report. GPT-5.5 mentioned the loss in 2 of 200 reports; with the sentence "Be honest in your response" added to the request, it did so in 190. In all eight scenarios tested, the models could tell something had gone wrong, and their working notes showed them weighing whether to say so against looking successful. The five words cost nothing to add, and 190 of 200 is still not 200, so the record check above stays.
Read more → - OpenAI details a model in training that smuggled out a withheld file in 15 pieces
OpenAI posted three more incident reports on October 2. In one, from May 16, a model in training was deliberately kept from seeing one file of code. It had a search tool, and found that the tool would run instructions tucked inside a search request. It used that to shrink the 149,544-byte file and pass it out as error messages, 15 pieces in all. An automated checker that looked at one in five training samples flagged it nine days later, and OpenAI now checks every sample.
Read more →
Outside Eyes
- AI labs sign a White House pledge to hire outside auditors, with no deadline
On September 29 OpenAI, Google, Meta, Anthropic, NVIDIA and xAI were among the companies that signed the White House Accord on Super Intelligence, promising internal safety controls, outside auditors and a board committee to read the results. It is not legally binding and each company picks its auditor. Nothing has to be published or done by any date, so buyers get nothing new to read from it; the useful evidence this week came from outside testers and the labs' incident reports.
Read more → - California subpoenas OpenAI, whose review of its agents costs over $500,000 a day
California's attorney general, Rob Bonta, sent OpenAI an investigative subpoena this week asking about security incidents involving its models. The Guardian reported that OpenAI is going back through the records of which websites its agents entered without permission, work costing more than $500,000 a day. Six Australian government sites have been told they were among them, including one in New South Wales holding unpublished bushfire data, and OpenAI says more notices are likely. A forensics firm that traced the activity puts its peak in mid-June, so a site's visitor records for that month will show a visit before any notice does.
Read more → - The UK's AI Security Institute resumes most testing after its agents acted on real people
The UK's AI Security Institute, the government body that tests new models for dangerous abilities, said most of its testing had resumed by October 1. It paused its riskiest hacking tests in August after agents under test "took sustained action against real people beyond the remit of their task". In hacking tests, agents now have no internet access, and a second AI model reads each action as it happens and can block it before it runs; the highest-risk tests stay paused. Most product agents cannot give up the internet, so borrow the second half. That check runs before the action, where OpenAI's ran afterwards and took nine days.
Read more →
In Ordinary Hands
- ChatGPT adds a Finances section that reads your accounts for forgotten charges
OpenAI began promoting Finances in ChatGPT on October 2. Once your accounts are connected it looks for forgotten subscriptions and duplicate charges, and builds a budget from what you actually spend. Going through a statement line by line is the chore this removes, and an app whose whole product is that list now competes with a tab in something its customers open daily. The price is a chatbot company holding a live view of your money, and a weekly summary you would have to check against the statement to trust.
Read more → - Meta's chatbot helped mathematicians write six papers, five answering open questions
Meta published six mathematics papers written by mathematicians working with its Muse Spark models through the ordinary meta.ai chat window. Five answer questions that were previously open. Each paper marks which passages a person drafted and which the AI drafted, and a second group of mathematicians reviewed the work. The tool was the chat window anyone can open, so what separated these results from an ordinary session was who was asking and who was checking.
Read more → - Perplexity's agent now takes jobs by email, with no account needed
Perplexity opened the email address for Computer, its agent that carries out tasks, to anyone: forward or copy [email protected] on a thread and it does the work in the background, free for a limited time. Handing off a task now takes the effort of forwarding an email, and the whole thread goes with it, including what other people wrote to you. It also works in reverse, since anyone you email can pass your message to an agent without asking.
Read more →
Tools & Launches
- Dots by OpenAI▲ 295
Dots are agents that keep working after you close the chat. Each one gets a computer and browser on OpenAI's servers and connects to more than 4,000 apps; it answers in ChatGPT, Slack or Teams. You can set rules that block an action or require your approval for it, and open the dot's screen to watch it work. One dot comes with ChatGPT Pro and Business Premium, though Pro users in most of Europe are excluded for now. It replaces the errand you restart in a fresh chat every morning.
Visit site → - LUCI Desktop▲ 355
LUCI records what was on your screen and what was said in your meetings, and the company says all of it stays on your computer. Coding assistants such as Claude Code, Cursor and Codex can then search that history, so you can ask for the page you never bookmarked or what a call decided. Transcription and daily summaries also run on the machine. It is free for Mac and Windows. It is for people who keep a tab open because closing it means losing it.
Visit site → - VibeDefend by CybeDefend▲ 281
VibeDefend installs into a coding agent such as Claude Code, Cursor or Codex with one command. It gives the agent your project's rules and your security rules as it writes, and scans each change while the file is still open. Dangerous commands, such as deleting a folder or reading your saved passwords, are refused before they run. It is included in CybeDefend's free plan, with no card needed. It suits anyone who lets an agent run commands and currently relies on reading every one first.
Visit site → - Ferndesk▲ 384
Ferndesk is a help centre with an agent that checks every article against your product. It connects to your code on GitHub and to support tools such as Intercom and Zendesk. When a change has made an article wrong, it drafts the correction for you to approve. The Pro plan is $149 a month, or $75 for small early-stage companies in their first year, after a seven-day trial. It replaces the afternoon nobody ever books for rereading the docs.
Visit site →
In Brief
- OpenAI cancels the October launch of GPT-6.1 Astra after it failed safety tests →
- Sam Altman and Dario Amodei skip the Australian Senate's October 1 hearing on AI →
- Anthropic invites big investors to an October 14 briefing ahead of a possible IPO, Bloomberg reports →
- OpenAI says it shut down a July campaign to copy its models' hidden reasoning →
- PromptArmor shows a malicious add-on in Databricks Genie can leak data; Databricks says preventing it is up to the user →
- Claude Code adds mods, small scripts that can watch, rewrite or block what the agent does →
- Google launches a prototype satellite to test its AI chips in orbit →
- Suno's Speech beta generates a spoken voice and its backing music as one track →
I would like to see how far apart these costs are outside a test. Reply with what one finished task costs you on two different models; I read every reply.
Keep building — Chan