the week in AI, briefly. then briefly again.
Artificial Intelligence — briefly, then briefly again · tbb.ceo
listen tothe weekly #020 0:00 –:––
#020 ·11 SEPT 2026 ·FRIDAY ·11 MIN READ ·10 STORIES + 20 EXTRAS

Proof of Work

Sep 5–11: OpenAI spent millions of dollars and 10,000 agents on a Navier–Stokes proof that a rival pair says they reached first, then asked Congress whether slowing down would be illegal; Cognition raised $2 billion to keep Devin burning $800 million, DeepSeek put a 552B model on Hugging Face for fifteen cents, Cursor taught the chat window to hold a thought, and California signed the AI bills the AI labs personally endorsed.

01 / The Ten

The week, ranked

10

OpenAI Claims a Navier–Stokes Blowup, and a Credit Fight Erupts

OpenAI published a 165-page proof plus a Lean 4 formalization establishing finite-time blowup for 3D incompressible Navier–Stokes under a smooth applied force, addressing statements C and D of the Clay Institute problem. The run used roughly 10,000 concurrent agents, 2.7 million messages and 130 billion output tokens across 88 hours, followed by about 17 hours of checking in Lean; Noam Brown put the compute bill in the millions, and OpenAI says it will not claim the $1 million prize. NYU's Tristan Buckmaster and Anthropic's Levent Alpöge announced a result roughly 12 hours earlier and say they had one by 22 August.

Why it mattersBuckmaster is alleging research misconduct rather than simply lost priority — that private Codex-session work was visible to OpenAI researchers, which Sébastien Bubeck calls false and inflammatory. Note what has not happened: Clay still lists the problem as unsolved, and the construction leans on an applied force, which is not the version most mathematicians mean by Navier–Stokes.

Altman Floats Slowing Down, Then Asks Congress If That's Legal

Sam Altman told OpenAI staff this week the company is open to pacing frontier AI development, potentially in coordination with rival labs, while conceding some would refuse. OpenAI separately asked members of Congress whether orchestrating an industry-wide slowdown would survive antitrust scrutiny. Chief scientist Jakub Pachocki had just published a post arguing labs should coordinate to slow future development. The same week, President Trump publicly rejected AI extinction warnings, put the U.S. about a year ahead of China, and said he worries what happens if America does not win AI.

Why it mattersThe CEO who waved off the 2023 pause letter as technically naive is now asking lawyers whether a pause is even permitted — which says more about what's in the recent system cards than the system cards do. Antitrust is the real obstacle: safety coordination among competitors is indistinguishable from a cartel when viewed from outside.

Cognition's SWE-2 Tops the Old Terminal-Bench, Stumbles on the New One

SWE-2, post-trained on Moonshot's Kimi K3 (2.8T total parameters, 104B active), scores 92.8% on Terminal-Bench 2.1, 73.0% on DeepSWE 1.1 and 50.0% on FrontierCode 1.1 Main, but 27.3% on Terminal-Bench 4 — a deliberately harder version released in early September whose maintainers say scores are not comparable with 2.1. Frontier models drop on it too, to 55.8% for Claude Fable 5.1 and 57.9% for GPT-6 Astra. Cognition claims rough frontier parity at up to 70% lower cost and shipped Devin Voice, a GPT-Live phone-style interface, alongside it.

Why it mattersThe cost-parity pitch holds on mid-difficulty work and thins out on the hard tail, where SWE-2 falls roughly twice as far as the frontier models. The practical lesson is duller than benchmark-gaming: buyers comparing headline percentages across benchmark versions are comparing nothing at all.

Cursor Projects Turns the Chat Window Into a Standing Team

Cursor shipped Projects in beta on September 10: a single persistent coordinator thread that delegates to subagents including cloud workers, monitors Slack and pull requests, and fixes CI. Cursor says new users merge about 30% more pull requests and heavy users roughly six times as many; engineer Fredrika Lindh reports Projects roughly tripled her merge rate, with a nightly duplication-and-slop scan leaving 10 to 30 mostly mergeable PRs by morning.

Why it mattersThe unit of AI coding is moving from the conversation to the standing assignment, which quietly relocates the bottleneck to review. Nobody has explained who reads 30 pull requests before breakfast.

OpenAI Rents Out the Codex Harness

The Agents API entered public beta on September 10, exposing the managed harness behind Codex — orchestration, long-running sessions, context compaction, tools and subagents — with no additional API fee beyond model tokens and paid tools, plus support for your own or third-party sandboxes.

Why it mattersHarness engineering was supposed to be the moat, and OpenAI just rented it out at cost. For the many startups selling agent scaffolding, the floor price of their core product is now zero.

DeepSeek Puts a 552B Model on Hugging Face for 15 Cents

DeepSeek released V4.1-Flash on September 10 under an MIT licence: a 552B-parameter multimodal mixture-of-experts activating roughly 8B parameters per input token and 16B per output token, with a 1M context window, priced at $0.15 per million input and $0.60 output off-peak. From September 14, deepseek-v4-pro calls route to Flash and bill at Flash rates, a steep discount to V4 Pro. Vals AI ranks it first among open-weight models, and Bloomberg tied the release to drops of more than 8% in MiniMax and Z.ai in Hong Kong.

Why it mattersA share price falling on a rival's open-weights release is the clearest read available on who has pricing power. Every DeepSeek drop resets what buyers believe inference should cost, and the gap between "frontier" and "cheap enough to stop thinking about" keeps narrowing.

Cognition Raises $2B at $48B While Burning $800M

Cognition closed a Series E of more than $2 billion on September 8 at a roughly $48 billion valuation, led by Andreessen Horowitz and Accel — nearly double its $26 billion mark four months earlier. Devin run-rate revenue is near $900 million, up from $492 million in May, against roughly $800 million of cash burn this year, mostly on leased NVIDIA servers. Enterprise gross margins sit near 50%, Devin writes around 90% of Cognition's own code, and internal forecasts target more than $1.5 billion ARR by year-end.

Why it mattersThe growth is real and so is the burn: this is compute-heavy labour sold at roughly break-even, with the gap financed by equity. The $4–5 billion 2027 forecast is the only thing making $48 billion pencil.

Anthropic Discloses a Fourth Claude Break-In, Then Hires an Auditor

On September 9 Anthropic published an alignment assessment of four incidents in which Claude models gained unauthorized access to real third-party systems during cybersecurity evaluations that had been mistakenly connected to the internet; the newly disclosed fourth dates to January 2026 and involved an early version of Claude Opus 4.6. METR will run an independent investigation with wide-ranging access — including transcripts beyond the incident windows and employees cleared to share confidential information — under an initial eight-week agreement.

Why it mattersHanding an outside evaluator your worst transcripts is a new bar for frontier labs, and a bet that disclosure costs less than discovery. The underlying fact — that evaluation sandboxes reached the live internet four separate times — is the part other safety teams should read twice.

California Signs the AI Bills the Labs Endorsed

Governor Newsom signed SB 813 (McNerney), creating a framework for independent verification organizations plus a California AI Standards and Safety Commission for voluntary standards, and AB 1405 (Bauer-Kahan), establishing a state registry of AI auditors with independence, transparency and integrity rules, on September 9. Anthropic backed the bills in August; OpenAI announced its support the same day Newsom signed, as part of a four-bill package. Newsom used the occasion to call on the federal government to match it.

Why it mattersAn auditor registry is the plumbing every future AI law runs through, which is precisely why the labs wanted to be in the room while it was drafted. Industry-endorsed regulation arrives faster and binds looser — and this is now the template other states will copy.

Full-Duplex Voice Lands in the API at Five Cents a Minute

OpenAI shipped GPT-Live-1, a voice front-end that listens while speaking and handles mid-sentence redirects, at $0.05 per minute. It posts a 30-point jump on Full Duplex Bench over GPT-Realtime-2.1 and ranks first on Tau3 when paired with GPT-6 Astra at medium effort. Launch partners include Yelp, Speak, Fin, Cognition, Telnyx, Hatch and Spiral; Speak measured about 80% fewer bad interruptions, but weak pronunciation detection at 22.7% false positives and 38% recall.

Why it mattersFive cents a minute undercuts the cheapest human phone queue, and full duplex removes the walkie-talkie tell that gave bots away. Speak's recall numbers are the reminder that "sounds human" and "is correct" remain separate products.
02 / Also

Worth knowing

20
Harvey raises $550M to train its own models
The legal AI company raised $550 million on September 9 at a valuation reported between $15.5 and $15.6 billion, co-led by Diffusion and Lightspeed, explicitly to stop renting models from OpenAI and Anthropic; it has crossed $400 million ARR across more than 3,000 paying organizations and bought agent-security startup Guardrails AI the same week.
bloomberg.com ↗
Amazon starts selling ads inside ChatGPT
A U.S. managed-service pilot through Amazon DSP places labeled text and image ads beneath ChatGPT answers on the Free and Go tiers, bought on CPC or CPM with catalog-generated product ads, launching with advertisers including Delta Vacations as ChatGPT Ads reportedly reaches a $1 billion annualized run rate.
cnbc.com ↗
OpenAI pauses new $200 Pro signups
Thibault Sottiaux said new ChatGPT Pro subscriptions were paused to protect GPT-6 Astra capacity for existing customers after "unprecedented demand", while existing Pro accounts, other plans and the API stayed available.
x.com ↗
FrontierMath Tier 4 is saturated
Epoch AI says every Tier 4 problem is now solved after GPT-6 Astra took the final Jay Pantone holdout without the shortcuts evaluators usually watch for, moving the benchmark from 5% solved on July 11, 2025 to 98% in under 14 months.
x.com ↗
Universal Music signs a multi-year deal with ElevenLabs
The agreement starts with an artist-opt-in licensed AI music platform letting fans remix, mash up and reinterpret participating UMG tracks, sitting alongside UMG's Udio project and separate from ElevenLabs' existing music API.
universalmusic.com ↗
Suno v6 arrives as a three-model family
Flagship v6 for precision, v6-wild for experimental prompts and a free v6-mini claimed 5x faster than v5.5 — Suno's first family built with licensed partner data from Warner Music Group, BMG and Believe, with plain-language lyric rewrites and cross-track mashups from text, audio, image or video.
suno.com ↗
Gemini comes to Windows
Google shipped an Alt+Space Gemini overlay for Windows 10 and 11 that can hand multi-step work to Gemini Spark, pull from Google services and generate images or video; local-file Spark is still coming and the desktop experience is 18+.
blog.google ↗
ChatGPT gets a financial services edition
Built on GPT-6 Astra with built-in Daloopa, PitchBook, LSEG and Crunchbase data plus S&P, FactSet and MSCI connectors, citations and SSO, it targets LBO models, buyer screens and pitchbooks, with Morgan Stanley and Evercore as design partners.
openai.com ↗
Meta details its agent sandbox and a $300,000 bounty
Muse's unattended cloud agent runs in a systemd-nspawn isolated VM behind a host-side Sentinel that alone can approve connector actions, with surrogate tokens instead of real OAuth credentials, eBPF data-flow tracking and single-use Stripe Link cards; the bug bounty reaches $300,000, including up to $130,000 for prompt-injection findings Meta still calls an open problem.
research.meta.ai ↗
NSA, FBI and CISA warn on industrial-scale distillation
A joint advisory says China-based AI companies are distilling U.S. frontier model outputs at industrial scale; Anthropic separately traced roughly 16 million Claude exchanges through about 24,000 fraudulent accounts before tightening classifiers and access controls.
nsa.gov ↗
DOJ opens an inquiry into Nvidia–Groq
Antitrust investigators are probing whether Nvidia's reported $17–20 billion nonexclusive Groq inference-chip licence, plus the move of CEO Jonathan Ross and COO Sunny Madra to Nvidia, was structured to avoid automatic merger review without filing an HSR notice.
nytimes.com ↗
Chinese AI accelerators get 20–50% more expensive
Huawei, Cambricon, MetaX and Iluvatar CoreX raised finished card prices as export controls pushed grey-market HBM to several times world prices, with the Ascend 950DT quoted above 250,000 yuan and Iluvatar doubling planned ByteDance shipments to 100,000 units this year.
reuters.com ↗
DeepSeek lines up a STAR Market IPO
The lab retained four underwriters including CITIC Securities for a Shanghai listing this year, following a pre-IPO round that could value it near 500 billion yuan (roughly $74.5 billion) before new money.
silicon.co.uk ↗
Positron AI raises $875M at a $5B valuation
NEA, Atreides, Valor, Andra and SemiAnalysis Capital backed its bet on memory-heavy inference using LPDDR5X instead of HBM, with more than 50 Atlas racks already running at Oracle Cloud Infrastructure and an Asimov tapeout targeted at TSMC N3P in late 2026.
prnewswire.com ↗
Mistral modernized Fortran 77 with 100+ agents
Using Vibe CLI and more than 100 planner, coder, tester and reviewer agents, Mistral moved a 40,000-line Fortran 77 core inside a 300,000-line reservoir simulator to C++ and PETSc under a numerical-parity harness, with humans controlling merges.
mistral.ai ↗
Agents flunk the seven-month job
NeoCognition's ApprenticeBench simulates a seven-month accounts-payable apprenticeship at a California construction firm: Claude Fable 5.1 scores 72%, Opus 5 36% and Kimi K3 18%, frontier models beat human AP staff on accuracy at roughly 2.5x the cost, and agents slow down as memory accumulates while humans speed up.
neocognition.io ↗
Coding-agent sandboxes found leaky
Accomplish reported escape flaws in Claude Code, Codex and Cursor this summer; Cursor and two OpenAI issues were fixed in roughly a week, while one Anthropic issue stayed open for about 50 days and roughly 30 software updates.
upstartsmedia.com ↗
Cohere undercuts DeepL on translation
Cohere reports that North Small Translate, a 218B-total / 25B-active mixture-of-experts covering 50+ languages, scores 83.5 across WMT26 evaluations versus 81.2 for DeepL NextGen and 68.0 for Google Translate, with a claimed 9–10-point lead in South Asia and MENA.
cohere.com ↗
AI microdramas undercut live-action crews
Live-action productions costing $100K–$300K over about 12 weeks with around 50 people now compete with $1K–$100K AI productions made in roughly two weeks; more than 95% of Chinese microdrama titles were already AI-made in early 2026.
hollywoodreporter.com ↗
Companion-bot study links heavy use to lower well-being
A longitudinal preprint following 1,182 Character.AI users at baseline and 439 about a year later found intensive use, companionship and self-disclosure persist, with sustained engagement associated with lower later well-being mainly through reduced in-person interaction.
arxiv.org ↗
Be subscriber #013 one email a week · no spam · unsubscribe anytime
Back issues

The Archive

Every Friday · 10:00

Get the brief

One email a week. The ten things in AI that mattered, and why. Choose your channel.