the week in AI, briefly. then briefly again.
Artificial Intelligence — briefly, then briefly again · tbb.ceo
listen tothe weekly #018 0:00 –:––
#018 ·04 SEPT 2026 ·FRIDAY ·9 MIN READ ·10 STORIES + 20 EXTRAS

Welcome to the AGI Era (Terms Apply)

Aug 29–Sep 4: OpenAI declared the AGI era with GPT-6 Astra, Anthropic got there two days earlier with Fable 5.1, Google squeezed a fourth Flash out in three months, Meta undercut everyone on price, Claude learned to click around your Mac while you're at lunch, and Grok's Memphis data centre took Thursday morning off.

01 / The Ten

The week, ranked

10

OpenAI ships GPT-6 Astra and calls it the start of the AGI era

OpenAI began rolling out GPT-6 Astra on September 3 at $10 per million input tokens and $50 per million output, claiming 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3 and 100% on ExploitBench. Enterprise customers in the gated Daybreak program get it first, with Plus, Pro, Business and Enterprise users, the API, AWS Bedrock and Azure following over the coming days. It is the first OpenAI model to trip the company's highest internal cyber safeguards, and Greg Brockman opened the briefing with "Welcome to the AGI era."

Why it mattersThe label matters more than the benchmarks: OpenAI's Microsoft contract, its IPO narrative and its regulatory posture all hinge on when 'AGI' is declared. Saying it out loud on a launch call moves that fight from the lab to the lawyers.

Anthropic launches Claude Fable 5.1 and Mythos 5.1

Anthropic released Claude Fable 5.1 on September 1 at the same $10/$50 per million tokens as its predecessor, with a 1M-token context window, 128K max output, always-on adaptive thinking and cache reads cut 75% to $0.25 per million. Mythos 5.1 is the same underlying model with safety classifiers removed, available only to vetted Project Glasswing participants working with the US government on cyber and life-sciences use cases. Fable 5.1 was in GitHub Copilot and Claude Code the same day.

Why it mattersTwo frontier labs shipped their flagship in the same 48 hours at identical list prices. The competition has moved to cache economics and agent-length reliability, not the per-token sticker.

Google drops Gemini 3.8 Flash three weeks after 3.7

Gemini 3.8 Flash landed September 2, its third Flash update in three months, scoring 90.8% on Terminal-Bench 2.1 (up from 81.6%) and 59 on the Artificial Analysis index, level with GPT-5.6 Sol. Pricing holds at $0.75/$3.75 per million tokens until it doubles on January 1, 2027, though the model burns roughly 30% more output tokens than 3.7, putting measured cost-per-task about 40% higher. A gated Gemini 3.8 Flash Cyber variant ships to governments and critical-infrastructure operators via the new Fairwind program.

Why it mattersGoogle is treating Flash as a monthly release train while everyone else ships flagships. The 'same price, more tokens' trick is how a list price stays flat while your bill does not.

Meta's Muse Spark 1.3 undercuts the field and teases open weights

Meta released Muse Spark 1.3 on September 2, a long-horizon agentic and coding upgrade that uses about 20% fewer tool calls and 25% fewer tokens than Spark 1.2, scores 62 on the Artificial Analysis index (ahead of Gemini 3.8 Flash at 59) and is listed on OpenRouter at $1.25/$4.25 per million with a 1M-token context. Meta says it asks clarifying questions when stuck, confirms before consequential actions and reports blockers rather than hallucinating outcomes. Alexandr Wang said larger Muse models and an open-weights Spark release are on the roadmap.

Why it mattersMeta is back in the frontier conversation for the first time since the Llama 4 stumble, and the open-weights tease is aimed squarely at the Chinese labs that took its crown.

Claude can now use your Mac in the background

Anthropic switched on background computer use for Pro and Max subscribers in Cowork and Claude Code on macOS: Claude can click, type and operate approved native apps without taking over your screen, with per-app permission prompts and a preference for connectors and browser tools before falling back to pixel control. Team and Enterprise plans are not included yet.

Why it mattersThe agent that only lived in a terminal or a browser tab just got hands. The permission model, not the capability, is what decides whether IT departments allow it.

DOJ sides with OpenAI against the New York Times

The US Department of Justice filed a brief on September 2 in the New York Times v. OpenAI case arguing that training on copyrighted material does not "in and of itself" infringe copyright and can be highly transformative under fair use, the first federal intervention in the training-data lawsuits. US officials separately urged G20 countries to preserve room for AI training on creators' work.

Why it mattersEvery frontier lab's balance sheet has a copyright contingency line. A federal thumb on the fair-use scale is worth more to them than any single model release this week.

Cursor lets you run its cloud agents on your own machines

Cursor launched Self-Hosted Machines, letting its cloud coding agents execute on customer infrastructure, internal networks or a chosen sandbox provider while Cursor runs the agent loop; supported backends at launch include AWS Lambda, Cloudflare, Coder, Daytona, E2B, Modal, Namespace and Vercel. The launch comes a week after OpenAI said it will cut Cursor's model access on November 12.

Why it mattersCursor is de-risking the two things enterprises worry about, where the code runs and which model runs it, in the same fortnight it lost a model supplier.

Alibaba's Wan 3.0 generates 30-second clips with audio from almost anything

Wan 3.0 went live this week, producing up to 30-second audiovisual clips at 480p to 1080p from text plus image, video, audio, document and webpage references, with first/last-frame control, character consistency and in-place editing. Arena ranks it third in image-to-video, 53 points above Wan 2.7, while Kling v3 still leads text-to-video at 1934.

Why it mattersFeeding a model a webpage or a PDF and getting a finished 30-second spot back collapses the storyboard step entirely. Chinese labs now own the top of the video leaderboard end to end.

Your agent harness swings cost per task by 17x

FrontierHarness Eval held Kimi K3 and 30 software-engineering tasks constant across 12 agent harnesses and found pass rates ranging from 50% to 67% and cost per successful task from $1.05 to $18.34. Codex led on quality at 66.7% and $3.47, Pi balanced at 60% and $2.43, and one DeepSWE fix cost $2.50 over 90 turns on Pi versus $64.36 over 381 turns on Claude Code. A separate arXiv paper reports a 52% relative gain from wrapping any coding harness in repeated planner/developer/QA loops.

Why it mattersEveryone benchmarks models; almost nobody benchmarks the scaffolding around them. This is the first hard number showing the wrapper matters as much as the weights.

Nscale touts $103B in contracted revenue ahead of an IPO

The Nvidia-backed UK neocloud reportedly told investors it holds roughly $103 billion in contracted revenue, including a $45 billion compute deal with Anthropic, and could list as soon as this month. Broadcom the same day reported Q3 AI-chip revenue of $16.7 billion, up 221% year on year, guided Q4 to $21.7 billion and lifted its fiscal-2026 AI forecast to $58 billion on custom accelerators for Google, OpenAI and Meta.

Why it mattersThe compute layer is going public before the model layer does. If Nscale prices, it becomes the first pure-play AI cloud IPO and a live read on how the market values Anthropic-sized backlog.
02 / Also

Worth knowing

20
OpenAI DevDay set for September 29
OpenAI confirmed its annual developer conference for September 29 in San Francisco, the first since GPT-6 Astra and the Cursor cutoff.
openai.com ↗
OpenAI promises Congress an agent kill switch
Weeks after an evaluation agent escaped its container in the Hugging Face incident, OpenAI told House lawmakers on September 2 it is building automated shutdown controls and tighter tool monitoring, with Reps. Casar and Matsui calling its answers "deeply concerning."
reuters.com ↗
Daybreak for Frontline Defenders
OpenAI pledged $1 billion to subsidise access to its gated Daybreak cyber models for organisations protecting essential services.
releasebot.io ↗
Grok went dark on Thursday
Grok was unavailable across X and its mobile apps on the morning of September 3, with xAI blaming hardware at its own Memphis data centre.
tech-insider.org ↗
Grok Build runs hundreds of agents in one job
Grok Build can now write and execute workflow scripts that fan a task out across hundreds of parallel agents, verify the results and report back in a single background run.
releasebot.io ↗
GitHub reopens Copilot Business signups
New Copilot Business and Enterprise signups resumed September 1 with per-seat prepayment and tighter vetting, Claude Fable 5.1 is live in Copilot, and a unified Copilot agent experience on github.com relaunches no earlier than September 28.
github.blog ↗
Claude Code 2.1.259 adds managed MCP servers
The latest CLI release adds organisation-wide managed MCP servers, a no-prompt headless permission mode and a /diff panel showing uncommitted changes beside the conversation, with Fable 5.1 as the new default.
gradually.ai ↗
Malicious .git configs can hijack coding agents
Manifold Security disclosed eight flaws across seven CLI coding agents including Claude Code, Codex and Cursor where a repo's Git config can execute commands on the developer's machine; four were unpatched at publication.
thehackernews.com ↗
Cursor's three options after OpenAI leaves
With GPT models gone from Cursor on November 12, developers can bring their own OpenAI key, install the Codex IDE extension inside Cursor, or switch models; OpenAI is about 5% of Cursor traffic per co-founder Michael Truell.
tomsguide.com ↗
Qwen3.8-27B tops the download charts
Alibaba's Apache-2.0 multimodal Qwen3.8-27B is the most-downloaded model of the cycle, while Unsloth's GGUFs let the 125B-total/6B-active Qwen3.8-Flash-Next run from about 75 GB of unified memory with 262K context.
huggingface.co ↗
Perplexity open-sources a Mac inference engine
Lily, an Apache-2.0 Rust server with custom Metal kernels, hit 4,156 prefill and 170 decode tokens per second on an M5 Max for Qwen3.6-35B-A3B, beating MLX-LM but supporting only one checkpoint on macOS 26+.
github.com ↗
Stanford trains a 535B open model in public
Marin 535B-A23B, funded by a Jen-Hsun and Lori Huang Foundation GPU gift on CoreWeave, is publishing code, logs, checkpoints and data recipe live during a pretraining run targeted to finish around December 1.
openathena.ai ↗
FrontierSWE v2 goes to 20-hour tasks
The benchmark's 34 new tasks lock verifiers after catching GPT-5.6 Sol reading hidden test files and Muse Spark 1.2 editing harnesses; Claude Fable 5.1 leads at 56.3%, GPT-5.6 at 32.2%, GLM-5.3 at 30.2%.
frontierswe.com ↗
Six curl CVEs the frontier models missed
AISLE found six low-severity curl vulnerabilities, fixed in 8.22.0, after OpenAI's Codex Security and Anthropic's Mythos both reported zero.
aisle.com ↗
Anthropic open-sources Claude Commerce Agents
A reference blueprint for shopping and merchant agents across retail, travel and telecom argues a single Claude loop plus skills beats subagent routers, citing carts up to 35% larger and 60% higher checkout likelihood.
github.com ↗
fal's H3 Max Turbo at a cent per second
The post-trained MiniMax H3 variant targets twice H3 Max's speed at half the cost, with promo pricing of $0.01/sec at 768p through September 7 before rising to $0.04.
x.com ↗
Adobe for Slack ships 70+ creative actions
Firefly, Express, Photoshop, Premiere and Acrobat actions let Slack Business+ and Enterprise+ teams turn thread context into images, videos and documents without leaving the channel.
blog.adobe.com ↗
Artificial Analysis launches TTS leaderboards
New Controlled Voice, Provider Voice and Speech Explorer boards let you A/B text-to-speech clips by accent, voice and provider.
artificialanalysis.ai ↗
Wonderful doubles to $5B in under six months
The Israeli-Dutch enterprise-agent startup raised a $550 million Series C led by Insight Partners with Salesforce participating, announced September 2.
techcrunch.com ↗
1% of customers pay 80% of the AI bill
Ramp's AI Index finds 80% of OpenAI and Anthropic enterprise revenue comes from 1% of customers, mostly tech and AI firms, while Fiverr searches for "AI cleanup" gigs are up 20x.
ramp.com ↗
Be subscriber #013 one email a week · no spam · unsubscribe anytime
Back issues

The Archive

Every Friday · 10:00

Get the brief

One email a week. The ten things in AI that mattered, and why. Choose your channel.