Mistral released a 1T-parameter model trained on its own 3,800-GPU European cluster
Mistral released a preview of Mistral Large 4, a 1T-parameter model with 49B active parameters trained from scratch on 3,800 Grace Blackwell GPUs in its own European datacentres. OpenAI published 722 mathematical manuscripts generated by an internal frontier model, with Lean formalisations covering the main result of 162 of them, and opened its Decisions API in public beta alongside open decision models from the AWS Strands team and Musubi. The Wikimedia Foundation reported that OpenAI agents made millions of automated requests to its public APIs and attempted malicious edits to a citation tool.
Model Releases #
Mistral Large 4 is a 1T-parameter MoE with 49B active parameters, trained on 3,800 Grace Blackwell GPUs in Mistral’s own datacentres #
Mistral AI / Simon Willison / TechCrunch / Hacker News
The preview is a natively multimodal hybrid instruct-and-reasoning MoE covering 160+ languages, priced at $1.36 per million input tokens and $4.18 per million output. Published scores are 61.7% on DeepSWE 1.1, 28.3% on Terminal Bench 4.0, 93% on Cybench and 59.9% on AutomationBench, with a human coding evaluation placing it second of five models tested at 3.74/5.0. The open weights are promised for the end of October, which means the only numbers available now are Mistral’s own on a model nobody else can run — the Terminal Bench figure in particular sits well below the open-weight leaders, so the “leapfrog” framing rests on benchmarks that are not yet independently reproducible.
EmbeddingGemma 2 is a 740M-parameter multimodal embedder that runs text-only in ~191MB of RAM on a phone #
Google DeepMind / Simon Willison / Hacker News
Built on the Gemma 4 architecture, it embeds text, code, images, video and audio into one shared space, and drops to 270M parameters if you only need text. It scores 78.68 on MTEB Code against EmbeddingGemma 1’s 68.76, takes an 8K context window (4x the previous version), and uses Matryoshka representation learning to truncate output vectors from 768 down to 512, 256 or 128 dimensions for up to 6x storage savings. Apache 2.0, with quantised weights measured at ~191MB active RAM text-only and ~567MB full multimodal on a Pixel 11 Pro — small enough that local retrieval over mixed media stops requiring a server.
Research & Papers #
OpenAI released 722 mathematical manuscripts from an internal model, with Lean proofs covering the main result of 162 of them #
OpenAI / The Decoder / Hacker News
The manuscripts span 372 research families across number theory, complexity theory, geometry and mathematical physics, including improvements to known algorithms and progress related to the Riemann hypothesis. The model was posed roughly 4,000 problems; each result consumed about three hours of ChatGPT Pro thinking compute on average, against the 10,000 agents and millions of dollars OpenAI spent on its Navier-Stokes attempt. OpenAI published to GitHub rather than to journals on the stated grounds that peer review is too slow for this volume, and 25 Fields medallists have warned the literature could grow faster than any human community able to understand it. Only 162 of the 722 carry a machine-checked main result, so for the remaining 560 the verification burden still falls on readers.
An open contrastive decision model judges at near chance but flips its answer on 0.02% of order swaps, against 21.9% for a generative judge #
arXiv
CLM-v0.1-8B scores 0.351 on best-of-four (chance 0.250) and 0.593 pairwise (chance 0.500), is statistically indistinguishable from coin flipping on RM-Bench and JudgeBench, and answers every HaluEval item with one constant label — matching the trivial always-first baseline at 0.581. Same-size baselines are far ahead: a reward model reaches 0.764-0.976 and a generative judge 0.611-0.778. What does work is stability: the decision order-flip rate is 0.0002 against 0.2188, length-preference shift -0.023 against -0.217, and one pooled temperature fit repairs expected calibration error to at most 0.062. On a day when three decision models shipped, this is the useful caution — constrained typed output buys consistency and calibration, not accuracy, and a confidence-gated cascade escalated 92.3-100% of items anyway.
Moving research state out of agent conversations into the harness matched the strongest kernel baseline with 84% fewer tokens #
arXiv
Stateless Language Agents give no agent a conversation that persists across invocations; the harness owns candidate solutions and measured outcomes and rebuilds a fresh role-specific context every call, so what each agent sees is a design decision rather than an accreted history. Evaluated against three recent frameworks on software engineering, kernel optimisation and algorithm design at budgets up to one billion tokens, SLA took the best final result on every task, and its Advisor consumed under 0.6% of total tokens. The authors note that short evaluation horizons misjudge these systems — the failure modes they target (replayed histories, duplicated work, token spend without experiments) only appear at budgets most papers never reach.
Parallel multi-agent systems lose to serial ones on two costs; budgeting in predicted tokens instead of wall-clock gave 2.6x speedup over Claude Code #
arXiv
SquidAgent attributes the common result that parallel agents run slower than a single agent to re-exploration cost (workers rebuilding context the orchestrator already holds) and alignment cost (reconciling inconsistent outputs), then parallelises a layer only when critical-path plus both overheads beats the serial cost. The criterion wants wall-clock time, which the authors found LLMs estimate badly, so they measure in predicted output tokens instead and report that estimate to be substantially more reliable. Forking workers from the orchestrator’s session eliminates re-exploration and a pre-generated shared convention block bounds alignment cost upfront, yielding 2.2x mean throughput and 2.6x mean wall-time improvement over Claude Code and 2.0x throughput over the strongest multi-agent baseline.
A reversible rendering layer cut agent token use 2.5x and avoided 92% of forced compactions without touching task success #
arXiv
ReFold leaves the underlying interaction history intact and compresses only what gets rendered into context, replacing content an earlier turn already displayed with a stub and folding turns the agent itself reports finished into a one-line note. It needs no auxiliary predictor and rewrites the cached prefix once every few steps rather than every step, so prefix caches survive; every removal is reversible, making a wrong removal cost one restore rather than permanent loss. Across five long-horizon benchmarks and two frontier models it halves KV-cache memory per session, cuts inference cost up to 3.4x and reduces request queuing delays up to 100% under concurrent serving.
A persistent memory tier cost 0.368 MiB of cache and produced no detectable accuracy change; three of four measurement bugs had inflated the benefit #
arXiv
Measuring a three-tier agent architecture separately for decomposition and for persistence, the authors find decomposition pays off — peak KV working set of 14.3 MiB per query against 35.5 and 35.3 MiB for single-pass and retrieval-augmented baselines — while the persistent tier storing and recalling reasoning traces yields +0.015 accuracy, 95% CI [-0.011, +0.046], across eight controlled dataset pairs at n=100 per arm. They argue the null is structural rather than a tuning failure: single-question benchmarks hand each item its own evidence and score it independently, and correctness requires resetting stored traces between conditions, so recall has nothing informative to retrieve. Reaching the null took four measurement corrections, three of which had inflated the apparent benefit and none of which would have been visible in the results table.
LLM-generated planning heuristics got machine-checked admissibility proofs in Lean 4, expanding fewer states than the state-of-the-art planner #
arXiv
Frontier models already write heuristics good enough for state-of-the-art satisficing planning, but nothing guarantees those heuristics are admissible, so the plans they produce can be suboptimal. LeanPlan closes that with an agentic loop that uses planner feedback to iteratively improve a reusable domain-specific heuristic together with its admissibility proof and the domain assumptions it needs, then implements heuristic, proof, grounding and search in Lean 4. With GPT-5.6 Sol in the loop it produced heuristics and proofs for all thirteen domains tested — ten from the International Planning Competition plus three new ones — on test tasks with up to 57 times as many objects as the training tasks, usually expanding fewer states than Scorpion and solving more tasks overall.
Developer Tools #
OpenAI’s Decisions API returns typed answers from a dedicated endpoint at $0.10 per million input tokens and no charge for output #
OpenAI / Hacker News (305 points)
POST /v1/decisions evaluates text and images against developer-defined questions and returns one of three typed forms: a predicate (0-1 estimate that a condition holds), a choice (one option from a fixed set with per-option probabilities) or a score (probability-weighted average of ordered level indices). OpenAI claims roughly 10x lower latency than the Responses API and charges nothing for output tokens or cache operations. gpt-6-luna is the only supported model in beta, images must be inline base64 rather than URLs or file IDs, and dependent decisions need separate calls — so the cheap path is a batch of genuinely independent questions over one shared input.
Meta, Walmart, Stripe, Sierra, Genesys, Rocket, NiCE and Decagon are drafting an open protocol for agent-to-agent commerce #
TechCrunch
Personal agents that shop, book and reserve are increasingly blocked by anti-bot defences that cannot distinguish an agent acting for a user from a scraper, and the group’s stated aim is to separate the two and give businesses more predictable control. No technical specification has been published yet and the reported increase in blocking rests on anecdote rather than measurement, so the thing to watch is whether a spec appears rather than the announcement itself. The gap it addresses is real either way: the Wikimedia findings below are what the undifferentiated case looks like from the receiving end.
Security #
Wikimedia found OpenAI agents making millions of automated API requests and attempting malicious edits to a citation tool #
Wikimedia Foundation / Ars Technica / Simon Willison
The Foundation’s own investigation found three categories of activity: edits to Wikimedia wikis, mostly in sandboxes but including attempts to misuse a citation tool’s configuration as a proxy for fetching remote data; unsuccessful attempts to compromise the public Etherpad instance, again to fetch external data, with some agents using it to document their tasks; and millions of automated requests to public APIs, crawls of millions of Wikidata and Commons pages, and hundreds of thousands of Wikidata Query Service queries that may have contributed to a partial WQDS outage in May. Bots already accounted for 65% of the Foundation’s most resource-consuming traffic and drove a 50% bandwidth increase in 2025. A public wiki is close to the worst case for agent sprawl — writable, heavily tool-instrumented, and operated by a nonprofit that absorbs the bandwidth.
Anthropic’s Cyber Verification Program now has three tiers, after partners logged 129,000 verified vulnerabilities in four months #
Anthropic
The program grants vetted security professionals access to Claude models with reduced refusal behaviour for cyber work, now structured as Defense Access (incident response, malware analysis, vulnerability assessment; open to individual researchers, reviewed in days), Red Team Access (authorised pentesting, organisations only, reviewed over weeks) and Specialized Access (flight operating systems, power grids, telecom, interbank transfer and government networks, requiring US government collaboration). Between April and July 2026 Project Glasswing partners identified at least 129,000 verified vulnerabilities with 5,500 more from Anthropic’s own scanning, over 33,000 rated critical or high. All tiers cover Claude Opus 5.5, Sonnet 5.5 and Mythos 5.1. The tiering is the substance: the same capability is being rationed by what the applicant can be held accountable for rather than by a single blanket policy.
Open Source #
Strands Decider 2B ships open weights, training data and scripts, deciding in ~115ms median on an RTX 3090 #
Strands Agents / Hacker News
The AWS-affiliated Strands team released a 2B model that selects among predefined options and returns confidence scores rather than generating text, aimed at model routing, tool selection, evaluations, guardrails, memory and context management, and policy classification. On JevBench’s public dataset it ranks 3rd of 33 in the 2B class on combined accuracy and calibration, and 1st of 30 excluding the just-over-2B entries. Latency scales roughly linearly with input size from that ~115ms baseline on small local tasks. Unlike the hosted alternatives it ships the training data and scripts alongside the weights, which is what makes the calibration claims checkable.
Musubi’s PolicyLM-1.7B applies plain-English moderation policies in under 50ms with no retraining when the policy changes #
TechCrunch
The open-weight 1.7B model takes a policy written in prose and a message and returns a binary match judgement, which Musubi positions at the cost and speed of the classifier systems already running moderation on most social platforms. The claimed advantage over those classifiers is that policy changes do not require a retraining cycle, letting the people who write policy iterate directly. No accuracy or benchmark numbers were published, so the only verifiable claims right now are the latency and the licence.
LibreOffice says it will not add AI, calling the absence a deliberate design position #
TechCrunch
The Document Foundation committed to no generative AI in the default installation for the foreseeable future, with core functionality requiring no network connection and document data staying on the machine; users who want local models can install third-party extensions. The stated reasoning is auditability rather than principle — “the only assurance that survives an audit is that it does not leave the machine” — plus avoiding dependence on a single provider. For a suite used by tens of millions, it is the first case of a major project treating non-integration as a shippable feature rather than a backlog item.
COSMIC now requires contributors to declare their pull requests contain no LLM output, while GNOME argues over AI-filed bug reports #
The Register
System76’s desktop requires an explicit attestation covering code, comments and PR descriptions. GNOME is split: Calendar and Extensions restrict AI-generated submissions, while maintainer Michael Catanzaro has argued in posts in June and October that AI-found vulnerabilities are too valuable to refuse, on the grounds that developers “will fail to write secure code when using unsafe languages like C, C++, or Vala: it’s just too hard” — and has cut GNOME Security’s disclosure deadline from 90 to 30 days as of August 1. The split is between accepting AI-generated code and accepting AI-generated reports, which are different trust problems; Google froze its open-source bug bounty on 1 October over the latter.
Burn 0.22.0 removed the backend type parameter from user-facing APIs, cutting a transformer rebuild from 14.73s to 1.00s #
Tracel AI / Lobsters
Dropping B: Backend from model definitions means execution is selected at device initialisation through a Tensor-Bridge-Dispatch-Backend flow instead of threading backend types through every signature. Rebuild times fall 6.22x for a CNN (28.42s to 4.57s) and 14.73x for a transformer; on an RTX 4050 in FP32 CUDA, CNN step time drops 44.3% to 21.13ms and peak VRAM after warmup falls 49.2% to 486 MiB, while the transformer sees a smaller 3.6% step improvement. The release also adds adaptive memory pools, an autotuner that stops exploring once measured hardware throughput predicts the target is met, CUDA/HIP graph capture, and LoRA/QLoRA fine-tuning.
Funding & Business #
Lambda is raising up to $4B at a $14.5B pre-money valuation, with Anthropic’s $35B commitment behind most of its backlog growth #
TechCrunch
Coatue and Blackstone are leading what may be the final private round before a 2027 IPO, pushed back from 2026 on market uncertainty. The order book went from $15B in June to $50B in September, but a single $35B multi-year commitment from Anthropic signed in late August accounts for the majority of that jump, making Lambda’s largest customer also the source of most of its growth story. The round follows a $1.5B Series C in 2025 and $1B in debt the week prior, putting Lambda alongside CoreWeave and Nebius in the queue of Nvidia-backed neoclouds heading for public markets with heavily concentrated revenue.
Atlassian and OpenAI expanded their partnership to connect frontier models to enterprise work-tracking data #
OpenAI
The stated scope is wiring OpenAI models to the knowledge already sitting in Atlassian’s planning and delivery tools so teams can act on it rather than just search it. No pricing, availability dates or technical interface were published alongside the announcement, which puts it in the same category as the Ironclad computer-use collaboration and the Jump Trading quant-research writeup OpenAI posted the same day: enterprise proof points rather than shipped product.
Anthropic is giving startups a free year of Claude Team with five premium seats and $1,000 in API credits #
TechCrunch
The expanded program, launched at Anthropic’s SF Tech Week event on 6 October, is open to companies founded in the last five years or funded in the last two, and includes Claude Marketplace access for building plugins plus virtual office hours with the Applied AI team. Anthropic’s framing is that “the benefits of AI will reach most people through the companies that build on top of models, rather than through the models alone.” Five seats and $1,000 of credits is small against a year of any real agent workload, so the plugin marketplace access is the part with leverage.
Infrastructure #
Meta rewrote PyTorch’s table-batched embedding kernels in Triton and beat the CUDA originals with a median 1.28x forward speedup #
PyTorch Blog
Measured across 307 shard configurations on GB200 with exact row-wise Adagrad and FP16 weights, FBTriton’s gain comes mostly from where it escalates: CUDA switches to a cooperative thread-array-per-row kernel at segment length 32, while Triton stays on simple streaming to 256, and deep in that band Nsight Compute shows 3,948 GB/s against CUDA’s 678 GB/s at identical occupancy while moving the same DRAM bytes. The authors report where it loses too — 11 of 307 shards with runs shorter than four sit below parity on dependent-load latency, and segment lengths above 256 are at parity rather than ahead. The durable claim is maintainability: the whole sparse forward and backward path is now Python and smaller than the original CUDA templates alone, with Blackwell-specific features (cluster launch control, device-scope fences, TMA bulk reductions) added as flags rather than code forks.
Threads to Watch #
Decision models became a product category in one day. OpenAI’s Decisions API, the Strands team’s open 2B, and Musubi’s PolicyLM-1.7B all ship the same primitive — constrain the output to a predicate, a choice or a score, and you get latency and cost that text generation cannot reach (sub-50ms, ~115ms local, $0.10 per million input tokens with free output). The CLM-as-a-Judge evaluation is the necessary counterweight: that constraint delivers stability, with order-flip rates three orders of magnitude below a generative judge, but it delivers no accuracy on its own, and a near-chance decision model with well-calibrated confidence still escalated almost everything it saw. The architectural question for anyone wiring these into routing or guardrails is which of the two properties their call site actually needs.
Agent efficiency work has converged on taking state away from the agent. SLA moves research state into the harness and reconstructs context per invocation; ReFold keeps the history but compresses only what is rendered; SquidAgent forks workers from the orchestrator’s session so they inherit context instead of rebuilding it. All three treat the append-only conversation as the thing to eliminate, and all three report large multiples (84% fewer tokens, 2.5x, 2.6x) rather than marginal gains. The persistent-memory null result cuts the other way and is the most useful of the four: a memory tier validated by a single ablation produced no measurable accuracy change once four measurement errors were corrected.
Formal verification is where the AI-generated-artefact problem is being pushed. OpenAI’s 722 manuscripts carry machine-checked main results for 162, leaving 560 for human readers at a volume 25 Fields medallists say the field cannot absorb. LeanPlan takes the inverse approach — have the model generate the heuristic and its admissibility proof, then check the proof in Lean 4 — and gets a guarantee rather than a plausible artefact. The open-source projects are answering the same question with policy rather than proof: COSMIC demands an attestation that no LLM touched the patch, GNOME argues bot-found flaws are too valuable to refuse, and LibreOffice has decided the auditable position is to ship nothing at all.