Anthropic's older Opus 4.8 leads its newer flagships in enterprise usage data
Ramp’s July usage index puts Anthropic’s older Opus 4.8 at 28% of enterprise model share while the newer Opus 5 and premium-priced Fable 5 trail far behind, according to an FT report. Three separate papers landed on the same failure surface — agent memory — measuring eviction before retrieval, poisoning that survives every content screen tested, and a multi-session hygiene benchmark where the no-memory baseline passes 11.67% of tasks. The MCP maintainers published a roadmap that puts agent identity and enterprise authorization among its five priority areas.
Funding & Business #
Anthropic’s best AI model struggles to attract users as cheaper tools thrive #
Financial Times / Simon Willison
The new data point is the Ramp AI Index for July, which shows enterprise usage concentrated on the cheaper end of Anthropic’s own lineup: Opus 4.8 at 28.0%, Sonnet 4.6 at 8.3%, Fable 5 at 8.0% and Opus 4.6 at 6.9%. The FT pairs it with the revenue picture already reported this week — $65 billion annualized in July, up from $47 billion in May, 6,000 customers above $100,000 a year — to argue that the money is arriving despite, not because of, uptake on the newest models. One caveat does real work, and Willison flags it: Opus 5 shipped on 24 July, so a July-month index captures barely a week of its availability and understates it by construction. What survives the caveat is the Fable 5 number — a model available all month, sitting at 8.0% while a cheaper stablemate takes three and a half times the share.
Developer Tools #
The MCP maintainers publish a roadmap built around agent identity and transport #
Model Context Protocol
Lead maintainers David Soria Parra and Den Delimarsky set out five priority areas on 22 August, building on the 2026-07-28 specification release: agentic messaging primitives including server-initiated events and maturing the Tasks extension toward the spec; unifying transport on HTTP by running Streamable HTTP over stdio for local servers; agent identity and enterprise security via DPoP, Workload Identity Federation and standard token exchange; standardized tool-result handling plus progressive discovery for large tool catalogues; and SDK ergonomics with conformance testing. Proposals falling inside these areas get expedited review, which makes the list a scheduling commitment rather than a wish list. Two items matter most for anyone running MCP servers in production: progressive discovery is an admission that flat tool catalogues do not scale, and the DPoP/WIF work is the protocol conceding that “the client is trusted” was never a deployable assumption.
Canonical funds a three-year study of whether LLMs can turn C into safe Rust #
The Register
Canonical and UK Research and Innovation are co-funding a PhD project at the University of Bristol’s Programming Languages Research Group, supervised by Canonical VP of Engineering Jon Seager with Professor Meng Wang and Dr Cristina David, to test whether LLMs can decompose large C programs into components and rewrite them as safe, behaviourally correct Rust. The named targets are snap-confine and AppArmor. The framing is unusually honest about the state of the art: existing translators preserve C structure too literally, producing output that leans on unsafe blocks and retains C idioms, so this is scoped as three years of evidence-gathering rather than a migration plan. Treat it as a signal about timelines — a vendor with a direct commercial interest in the answer is funding a PhD to find out, not shipping a tool.
Model Releases #
An anonymous “Ox Alpha” reasoning model appears on OpenRouter #
TechCrunch
Ox Alpha went live on OpenRouter on 22 August, described as a reasoning model built for coding, sustained agentic work and production workloads, with OpenRouter saying only that it comes from “a third-party provider who has chosen to remain anonymous during this preview.” No benchmark numbers, parameter counts or evaluation results have been published, and attribution guesses have run from Z.ai to an unreleased Microsoft MAI build without converging. Stripe CEO Patrick Collison called it “very impressive,” which is currently the highest-quality public evidence available about it — a reminder that stealth previews on a router are a distribution channel, not a disclosure, and nothing here is verifiable yet.
Research & Papers #
Fable 5 closes 81.7% of the human gap on the nanoGPT optimizer speedrun #
Prime Intellect
Prime Intellect ran 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun, where an agent is given a training script and asked to make it reach the target faster. Against a starting baseline of 3,290 and a human record of 2,600, Fable 5 reached 2,726 — closing 81.7% of the gap — with Opus 5 second at 2,920 (53.6%) and Kimi K3 third at 2,930 (52.2%); individual runs consumed between 0.6 and 8.7 days of agent time. The spread is the finding: on an identical, well-specified optimization task with an executable score, the top model closes more than half again as much of the human gap as the next one down, which is a much wider capability separation than these models show on conventional coding benchmarks.
Evicting a memory’s prerequisites breaks retrieval before retrieval runs #
arXiv
The paper isolates a failure mode it calls structurally indirect prerequisite eviction: under a fixed memory budget, upstream blocks that are only weakly similar to the query get discarded, so the evidence chain is already broken by the time retrieval is asked to find it. Adding a one-hop, dependency-aware garbage collection rule (DSGC) raises full-chain retention from 0.03 to 0.90 under a lexical encoder and from 0.23 to 1.00 under a sentence encoder. The practical implication is that tuning a retriever cannot fix this class of bug, because the relevant block is no longer in the store — retention policy has to be dependency-aware, and the paper’s own robustness checks identify budget regimes where even the one-hop rule degrades.
DreamBench-SWE measures what software agents actually carry between sessions #
arXiv
DreamBench-SWE scores software agents on multi-session tasks where later work depends on non-inferable evidence from earlier sessions, graded by executable hidden oracles. In the preregistered v2.1 successor audit, which completed 360/360 work units, no external memory passed 21 of 180 (11.67%), deterministic verbatim event memory 82 of 180 (45.56%), a typed-plus-raw reference probe 83 of 180 (46.11%), and one pinned hosted Mem0 literal-storage configuration 97 of 180 (53.89%). The authors are notably disciplined about what this does not show: all three memory-bearing conditions beat no-memory after Holm correction, but the comparisons among them were nonconfirmatory, so the result establishes that having memory matters without establishing that any particular memory product is better than verbatim logging.
An index of what workers actually delegate to agents, from 53,000 skill specs #
arXiv
Rather than asking what AI could do to an occupation, this paper measures what people have already wired into workflows: it embeds roughly 53,000 agent skill specifications from the Manus Skills Marketplace, computes similarity against about 18,000 O*NET task statements, and aggregates to occupation level as an Agentic Adoption Index. Delegation concentrates in different occupations than pre-AI exposure frameworks predicted, and the index peaks below the top of the wage distribution and at the bachelor’s level, falling off at both extremes. Technical availability explains most of the variation but not the shortfall among the most educated occupations — the authors offer two candidate explanations, work that resists advance specification and professional discretion over the pace of codification, and are explicit that one cross-section cannot separate them.
Security #
Poisoning 1.2% of an agent’s memory drops accuracy from 0.85 to 0.30, and screening catches none of it #
arXiv
Using plainly worded false assertions generated in a single pass — no injection triggers, no retriever optimization, nothing an anomaly detector would find odd — the authors poisoned 1.2% of a LongMemEval corpus and watched accuracy fall from 0.850 to 0.300. A four-stage write-time screening pipeline that achieves 0.832 recall on indirect prompt injection rejected 0 of 360 poisoned memories, and provenance-weighted retrieval at its shipped weight was statistically indistinguishable from no defense at all (p=0.80). The boundary they identify is the important part: telling a false assertion from a true one generally requires grounding outside the text, so content-only screening cannot do it in principle, and a provenance weight strong enough to resist query-shaped poison also suppresses legitimate untrusted evidence — accuracy recovers to 0.700 in a mostly-benign mixed corpus but collapses to 0.0417 when the answer-bearing evidence itself arrives untrusted.
Threads to Watch #
The free lunch from the next model release has ended, and the bill is landing on harness design. The FT usage numbers show enterprises concentrating on cheaper models within a single vendor’s lineup, and Drew Breunig — quoted by Simon Willison the same day — argues the reason directly: it used to feel silly to invest in coding harnesses or context strategy because a new model would arrive and paper over the problems, and with frontier pricing where it now is, that bet has stopped paying. His cost ratios put GLM 5.2 at roughly a ninth of Fable’s price and a fifth of Opus 5’s. Prime Intellect’s speedrun points the same way from the capability side: the best model is meaningfully better at autonomous optimization, and is also the one teams are declining to route routine work to.
Agent memory has become the dominant measured failure surface, and the three papers today attack it from different ends. One shows the evidence chain breaks during eviction, before retrieval is ever consulted. One shows that 1.2% contamination is enough to halve-and-then-some a memory system’s accuracy while defeating both content screening and provenance ranking. One shows that memory-bearing configurations beat no-memory by a wide margin but cannot yet be distinguished from each other. Read together they describe a component that is now load-bearing in production agents and has neither a settled evaluation nor a working defense.
Authorization is migrating from application code into the agent protocols themselves. The MCP roadmap names DPoP, Workload Identity Federation and token exchange as a priority area, which is a standards body deciding that per-request cryptographic proof of possession belongs in the transport rather than in each server’s middleware. It arrives alongside a steady stream of papers on stateful authorization for delegated agent effects. The shape of the problem is the same in both places: approval granted at admission time stops being meaningful once the agent’s effects unfold across retries, recoveries and provider state changes.