Stripe agrees to buy AI gateway OpenRouter for more than $7 billion
Stripe agreed to acquire OpenRouter for more than $7 billion, roughly five times the AI gateway’s May valuation, according to Bloomberg. Google shut down its Imagen 4 endpoints today and Microsoft set a September retirement date for Excel’s COPILOT() function, two AI features withdrawn on the same day. Three papers landed on the same uncomfortable point from different directions: the metrics practitioners watch — offline skill-retrieval accuracy, semantic-retrieval token efficiency, and deliberation markers in a reasoning trace — are each decoupled from the outcome they are taken to predict.
Funding & Business #
Stripe agrees to buy OpenRouter for more than $7 billion #
Bloomberg / Fortune / TechCrunch
Stripe has agreed to acquire OpenRouter, the inference gateway that routes requests across more than 400 models for around 8 million developers, for over $7 billion. Neither company confirmed terms and Bloomberg notes the final price could still change; the Wall Street Journal reported the talks last month at a figure nearer $10 billion. OpenRouter was founded in 2023 by OpenSea co-founder Alex Atallah, has raised more than $150 million from CapitalG, Andreessen Horowitz and Menlo Ventures, and last raised at a reported $1.3 billion valuation in May — so the agreed price is roughly five times a three-month-old mark. What is being bought owns no model and trains nothing: it is the layer that sits between an application and every provider and knows what each request costs. That is the same substitution pressure behind the token price war reported on Saturday, priced as an asset — if a business can move a workload to whichever model is cheapest for the task, the thing that makes the move cheap captures value the labs cannot.
Developer Tools #
Google shuts down the Imagen 4 endpoints #
The three Imagen 4 GA endpoints — imagen-4.0-generate-001, imagen-4.0-ultra-generate-001 and imagen-4.0-fast-generate-001 — hit their published shutdown date on the Gemini API today, with the Gemini-native image line as the migration target. It is not a drop-in swap: the generate_images() method is gone, and image generation now runs through generate_content(), the same call used for text, so migrating means rewriting the call rather than changing a model string. The same changelog retires Gemini Robotics ER 1.5 on 31 August in favour of Robotics ER 2, and marks Gemini 3.7 Flash generally available at introductory pricing through 31 December. Google published the date well in advance, which is the point worth noting: the retirement cadence on hosted image models is now short enough that an endpoint string is a dependency with an expiry, not a constant.
Microsoft retires the COPILOT() function in Excel #
The Register
Microsoft will retire Excel’s COPILOT() worksheet function on 14 September 2026, saying it has “decided not to move forward with this feature” and directing users to the Copilot side pane, which it says offers many of the same capabilities — summarising text, classifying data, generating content and retrieving from the web. The function shipped in August 2025 and never left preview. What is actually being withdrawn is the only Copilot surface that put a model call inside the spreadsheet’s own calculation model, where it could be filled down a column, recalculated and audited like any other formula; the side pane is a chat window beside the grid, which is not the same capability restated in a different place. For anyone who built a sheet around it, there is a month to move.
Research & Papers #
Demystifying Agent Skills: Why They Work — Until They Don’t #
arXiv
Controlled experiments across several benchmarks, agent harnesses and LLMs — 8,135 normalised trial records and 238 valid labels from 240 open-coded records — yield a taxonomy of twelve skill-use modes. Procedural anchoring, where a skill turns a noisy trajectory into a stable execution path, accounts for 65.7% of the cases where skills helped, against 4.5% for explicit knowledge injection; skills beat Workflow Memory by 6.06 points in matched comparisons. So skills mostly stabilise action rather than supply missing facts. The load-bearing number for anyone maintaining a library is retrieval: as the pool grows from 5 to 100 skills, actual-use precision falls from 29.6% to 3.3%. The authors also find that exact ground-truth invocation is neither sufficient nor necessary for downstream success, and that confusable distractors hurt offline identification while leaving success rates stable — which means offline retrieval accuracy is a poor proxy for whether the library is helping at all. It lands alongside Saturday’s skill-misevolution result from the opposite side: that paper found the library accumulates unsafe entries, this one finds it stops being searchable, and both put the failure in the library rather than the skill.
Does a Language Server Save Tokens for Coding Agents? #
arXiv
The claim that LSP-based semantic retrieval is more token-efficient than grep is, the authors write, “asserted almost everywhere and measured almost nowhere”. Their five-arm ablation on Python and TypeScript repositories with Claude Opus 4.8, Sonnet 4.6 and Haiku 4.5, scored on tokens-to-success, returns a conditional and mostly negative answer. On symbol-named localisation the LSP costs 6% to 118% more tokens, and agents ignore it even when it is free, using semantic retrieval 0-6% of the time; on reference-completeness tasks they reach for it unprompted about half the time, where it buys precision but not savings and only reduces tokens for the weakest model. The sharpest result is on edits scored by real test execution: grep solves multi-file renames perfectly, a location-only LSP fails three-quarters of them by missing a call site, and even a complete, index-warmed, text-enriched LSP recovers most of the gap but cannot close it — a rename has to touch comments and strings, which semantic references exclude by construction. This is a preliminary study on two languages, but it inverts a default that a lot of agent tooling has already shipped, and the authors’ conclusion is a router keyed on task class and model capability rather than LSP-always.
Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models #
arXiv
Across 15 models and 6 text-only and vision-language benchmarks, with 15,282 annotated traces, the authors define Behavioral Lift — how much correctness changes when a behaviour is present versus absent in a trace — and find it points somewhere other than where reasoning training pushes. Thinking models strongly amplify self-correction, hypothesis testing and uncertainty acknowledgment, while the highest-lift behaviours are confidence calibration, knowledge alignment and self-awareness. Uncertainty acknowledgment is amplified 3-7x while being weakly or negatively associated with correctness; confidence calibration is among the strongest positive signals of correctness in both modalities and is barely amplified at all. The practical consequence falls on anyone reading traces as a quality signal — a judge model scoring chain of thought, a human skimming for hedging, a filter that rewards visible deliberation. Reasoning training made traces look more deliberative without moving the behaviours that track being right, so surface form is the wrong thing to reward and the paper argues for process-level objectives that score calibration and grounding instead.
Don’t Claim Benchmark-Oriented Optimization Improves General Coding Capability #
arXiv
The authors built a Django-based benchmark suite and ran foundation models against checkpoints post-trained on SWE-bench trajectories. Rankings frequently failed to generalise: the post-trained checkpoints showed little cross-task transfer, and SWE-bench optimisation produced limited or no gains on the Django tasks or on LiveCodeBench, with fine-tuning on individual Django modalities also failing to carry over. No effect sizes appear in the abstract, so the strength of the null is not checkable from it, and the recommendation — holistic evaluation for frontier models, multi-task suites for research, human-in-the-loop for narrow applications — is the familiar one. The narrow, usable claim is about inference rather than training: a SWE-bench delta is evidence about SWE-bench, and reading it as evidence about a codebase that looks nothing like SWE-bench is the step the measurement does not support.
Security #
Microsoft blames an AI-generated vulnerability backlog for a delayed Exchange update #
The Register
Exchange Server Subscription Edition Cumulative Update 1, first promised for the first half of 2026 and then pushed to the second, now has no date at all. The Exchange team attributes the slip to the volume of work generated by Microsoft’s own AI vulnerability-finding tools, describing the pipeline it has to run on each report — validating that the issue is a real security issue, reproducing it, fixing it, regression-testing after the fix, then shipping monthly — and says CU1 will not ship until the team reaches “a reasonable stable point” and gets “a month without pressing security payload”, since a major update followed immediately by security patches doubles the update work for administrators. This is a rare public accounting of where AI-assisted security work actually costs something. Discovery scaled; triage, reproduction, regression testing and release engineering did not, and the queue formed at exactly the steps nobody automated.
The AI credit resale economy #
Vectoral / Hacker News (294 points)
Vectoral, an LLM-security publication, surveys a secondary market in unused AI API credits: brokers buy startups’ unspent balances and resell access at 30-80% off list prices, with one dedicated site advertising a flat 40% off every model and a direct broker offering $100k a day in spending capacity. Distribution runs through dedicated sites, Telegram channels and subreddits, and most operators act as proxies against pooled API keys rather than handing over credentials, billing after usage thresholds. The structure is what makes this a security story rather than an arbitrage one: the provider sees a single key carrying aggregated traffic from unknown parties, the credit’s original owner keeps whatever liability attaches to what runs through it, and abuse routed this way is attributable to neither. The scale figure — “tens of millions” of credits on offer — is the author’s estimate rather than a measurement, so take it as an order of magnitude.
Model Releases #
Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things #
Simon Willison / Hacker News (516 points)
The first substantive third-party evaluation of Friday’s Apache 2.0 release, which Saturday’s digest flagged as having only Alibaba’s own numbers behind it. Willison ran a 17GB Q4_K_M build in LM Studio on a 128GB M5 Max MacBook Pro and an NVIDIA DGX Spark, getting 15-30 tokens/second against 74-184 for hosted APIs. The headline problem is a shipped default: reasoning effort is set to xhigh, and at an 8,192-token context the model spends the whole budget thinking about trivial prompts. His pelican-on-a-bicycle SVG took 21 minutes and 22,276 reasoning tokens to produce 3,223 output tokens — roughly seven reasoning tokens per output token — versus 137 seconds and 3,715 tokens with reasoning off, though the reasoned result was visibly better. Asked to “draw an svg of a circle” it deliberated over layered rings, gradients and animation and returned a beautiful animated circle he had not asked for. Vision held up, with accurate bounding boxes on two pelicans, and the model drove the Pi agent framework through documentation questions and a working transcript-conversion tool. The operative correction to Saturday’s entry is not that the benchmark scores are wrong but that the default configuration producing them is far more expensive per useful token than the model’s size suggests.
Threads to Watch #
The routing layer got priced, and it priced higher than the models it routes to. Stripe is paying more than $7 billion for a company that owns no weights, and a grey market has grown up reselling the same commodity — inference capacity — at 30-80% off through pooled keys. Both are bets on the same structural fact: with frontier models close enough in capability that workloads can move between them, the margin migrates to whoever controls the movement. The unglamorous corollary is that this layer now concentrates two things at once — spend and attribution — which is why the same shape shows up as an acquisition in one item and as an abuse channel in the other.
Two AI features were withdrawn today, and a third team cannot ship because its AI worked. Google retired the Imagen 4 endpoints, Microsoft killed Excel’s COPILOT() function before it ever reached GA, and the Exchange team has no date for a cumulative update because AI-discovered vulnerabilities filled the queue faster than humans could validate and regression-test them. None of these is a capability story, and all three are the part of the lifecycle that arrives after launch: deprecation schedules, features that do not survive preview, and operational load created downstream of a tool that worked exactly as advertised.
Three separate measurements, one finding: the visible signal is not the load-bearing one. Offline skill-retrieval accuracy degrades under distractors while task success does not, and precision collapses from 29.6% to 3.3% as a library grows — so the number you can measure cheaply says little about whether the library helps. Semantic retrieval’s token-efficiency advantage was asserted everywhere and, when measured, ran 6-118% the wrong way. Reasoning training amplifies uncertainty acknowledgment 3-7x while that behaviour is weakly or negatively associated with being correct, and barely touches the calibration that is. In each case the proxy was adopted because it was legible, and in each case measuring it directly reversed the sign.