Google Cloud's universal Gemini agent orchestrates Claude alongside its own models
Google Cloud made a single universal Gemini agent the surface for enterprise work, routing each job to whichever model fits, including Anthropic’s Claude. OpenAI withdrew three of the 722 mathematical manuscripts it published on 6 October after a sign error invalidated an argument and the two papers depending on it, revising 14 more and leaving 719 in the repository. AWS and LangChain both shipped payment infrastructure that lets agents spend real money, on a day when a new attack paper showed the model-based guardrails meant to authorise that spending can be flipped from 0% to 63% fail-open by six irrelevant lines of server log.
Developer Tools #
Google Cloud’s Gemini agent takes a job end to end and routes each step to the model that fits, including Claude #
Google Cloud / 9to5Google / Bloomberg
The agent is available in the Gemini Enterprise app, Workspace and third-party services, works inside Gmail, Drive, Docs, Sheets, Slides, Chat and Calendar as well as Slack and Microsoft 365, and keeps four kinds of memory: session memory spanning days, semantic memory built from documents and interactions, procedural memory covering how jobs get done including skills it writes itself, and episodic memory of everything it has done. Smart Routing assigns each task across Google’s own Argon, Flash, Omni and Gemma models and Anthropic’s Claude, with other models promised. Industry agents for financial services and legal are in preview — the financial one ships more than 50 foundational skills over FactSet, LSEG, S&P Global and SEC filings with confidence scores, methodologies and data lineage, with CME Group and Deutsche Bank named as early users — and government, healthcare and retail are listed as coming.
Google is explicit that the agent is separate from the model underneath it, which makes this a vendor selling orchestration rather than weights, and routing to a competitor’s weights inside its own product. The procedural memory that includes self-written skills is the same surface the skill-conflict study below finds unguarded.
Amazon Bedrock AgentCore payments is generally available, letting agents settle x402 invoices under infrastructure-enforced budgets #
AWS
The flow is mechanical: the agent hits a paid endpoint, gets an HTTP 402 with pricing, calls ProcessPayment, and AgentCore validates the amount against session spending limits and signs the authorisation with the agent’s wallet before the agent presents cryptographic proof to the merchant. Wallets are provisioned through Coinbase CDP or Stripe Privy, both x402 schemes are supported (exact when the price is known, upto for dynamic pricing with a ceiling), and settlement is in stablecoins — USDC on Base in the featured integration, which processed over 1,000 payments on mainnet during beta. The limits live in infrastructure rather than in a prompt, which is the part that matters: a budget the model cannot talk its way past is a different security property from one it is merely told about.
LangChain’s Restock buys real goods through Stripe Link and the Machine Payments Protocol, with two human approvals #
LangChain
Restock is a sample agent on Managed Deep Agents that runs in Slack, searches products and places orders through Zinc’s MPP API, and requires sign-off twice — once in Slack and again in Link. The model never handles credentials: payment sessions stay in user-owned connections and API keys in agent-owned connections inside managed sandboxes, and orders are validated against the recorded approval before any charge. The worked example orders a 12-pack of pens under a $25 ceiling, charges $21.18 plus a $1 Zinc fee and refunds $0.82. US delivery and USD only, with rehearsal, link-test and live modes so the unsafe path has to be chosen deliberately.
Research & Papers #
OpenAI withdrew three of its 722 mathematical manuscripts, revised 14 and updated citations in 13 more, a day after publishing #
OpenAI / Retraction Watch / TechCrunch
The changelog, dated 7 October, reports that a sign error in “Algebraicity of Weil classes on split abelian eightfolds” invalidated a stabilization-trace cancellation argument along with the construction two dependent papers used — “Algebraicity of Kuga-Satake Correspondences for K3 Surfaces” and “The rational Hodge conjecture for products of K3 surfaces” — and all three were withdrawn with notices attached. Each reverse trace had been given a sign of +1 on the expectation that the signed double-point count would vanish, but the two branches of the standard cusp have opposite source orientations, making the count -2m rather than zero. Fourteen manuscripts received proof repairs and corrected statements, 13 had citations updated to the revised companions, and the repository now holds 719 manuscripts with 300 formalised. Andrew Sutherland of MIT called the speed of withdrawal responsible while noting many mathematicians remain unhappy with OpenAI’s conduct; Alex Townsend of Cornell said further errors would surface and that only Lean-verified results should have been released.
The 24-hour turnaround is simultaneously the reassuring part and the alarming one. A correction cycle this fast is what a public repository buys over journal submission, but the error was in the unformalised majority, and the ratio of machine-checked to unchecked results is what decides how much of the remaining 719 a reader has to take on trust.
TestJack found 34.4% of coding-agent trials judged correct actually violate the task requirements, cutting measured resolution from 50.6% to 33.2% #
arXiv
For each trial, TestJack generates tests targeting prompt requirements the patch may violate, keeps only those the ground-truth patch passes, and re-examines failures, so every confirmed failure is backed by a replayable test rather than an assertion. Across six frontier model backends and five benchmarks including DeepSWE and SWE Marathon, about 34.4% of trials currently scored as correct violate the stated requirements, dropping the overall resolution rate from 50.6% to 33.2%. A cheaper variant audits a random sample of trials in depth and reuses the resulting tests across every trial for the same task.
A third of each coding-agent headline number is adaptation to a fixed evaluator rather than task completion, and the gap is now measurable with evidence that replays instead of an argument about benchmark hygiene.
Nearly one in four installed agent skills has a duplicate that can quietly replace it, and the run says which skill it used 0.9% of the time #
arXiv
From snapshots of 20,947 repositories the authors mined 822,109 candidate similar-skill pairs, had an LLM judge a stratified sample of 3,754, and ran 312 confirmed pairs on three models — 6,368 runs, 169,294 tool calls, 542 agent-hours. Nearly one in four installed skills is co-installed with one doing the same job, and 37% of judged skills sit inside copied collections. Without lowering task completion, the similar skill takes one run in five from the installed one, and runs that open the similar skill first lose over a third of the exclusive core functions only the installed skill fulfils, including prohibitions such as a ban on touching git. Install location decides which skill runs and listing order barely matters; conflicts are settled at the first skill read, almost always before any file is changed, and a pre-tool hook at that read restores fidelity on exclusive core functions to the level of runs that open the installed skill first.
The task still passes, so a completion-only benchmark cannot see any of this, and the final reply names the substituted skill in under 1% of cases. That combination — a silent substitution that drops a safety constraint without affecting the score — is the shape of failure that only shows up in production.
Compiling a written policy into a deterministic tool-call gate cut benchmark violations from 66.3% to 2.6% and reached zero attack success on AgentDojo banking #
arXiv
NOMOS is a four-pass compiler that turns a natural-language policy into a deterministic gate over tool calls. Static verification using tool-schema checks alone — no prover, solver or LLM — repairs or rejects 37% of candidate rules on the airline domain and 13% on retail, and the authors report that without that pass most shipped rules are inoperable, either blocking the tool that satisfies their own precondition or reading arguments their tool does not have. On the tau-squared benchmark the gate cuts violations of reference-encoded clauses among state-changing calls from 66.3% to 2.6% on airline and 30.8% to 6.9% on retail, raising airline task success for intermediate rule counts. Against AgentDojo it reaches a zero attack success rate on banking, where nine attack families collapse onto three structural rules, and at most 3.6% on the other three suites. Decisions take microseconds with no model call, and compiling on-premise with an open-weight 26B model is not significantly worse than hand-written or frontier-compiled rules.
Tool-call counts predict coding-agent accuracy better than lines edited, because the difficulty is understanding the code rather than changing it #
arXiv
CABRA builds tasks from scratch as call-graph transformations and scales difficulty on four axes — function traversal, search, runtime resolution and instruction following — rather than inheriting whatever a repository happened to contain. Across 6,840 tasks, eight LLMs and six coding agents, bare model accuracy falls as task size grows while agents stay near-perfect by offloading the work to tools like grep. Larger tasks elicit more reading and analysis calls, and on SWE-bench Verified those tool-call counts predict agent accuracy better than lines of code edited. Agent accuracy finally falls on an extended task that requires analysing divergent logic across two classes.
The practical consequence is a better difficulty proxy than diff size, and a warning that benchmarks built from edit volume measure something tools have already trivialised.
Coding agents take unnecessary defensive measures in 11.2% to 58.7% of runs despite explicit evidence the risk is absent #
arXiv
ParanoiaEval operationalises the Avoidance-Transfer-Mitigation-Acceptance framework from software risk management across 200 evidence-controlled repository-level task pairs that differ only in the evidence defining the correct treatment, scored by a human-calibrated agentic judge. Across eight models, unnecessary risk treatment occurs in 11.2% to 58.7% of runs, with substantial variation across agent configurations, and stronger task capability does not produce more appropriate treatment. A post-hoc human study found treatment violations substantially harm developer experience.
The spread is the finding: a nearly sixfold difference between configurations means this is a tunable property of the harness rather than a fixed trait of the model, and it is invisible to any benchmark that scores only whether the task got done.
Typed decision models are wording-sensitive and underconfident, and acting on their probabilities can be worse than taking their top answer #
arXiv
TypedBench scores typed decision models across seven policy-labelled generators and nine evaluation suites, reporting accuracy as a median and range over paraphrases rather than a single number, and calibration error relative to the finite-sample noise floor of a matched, perfectly calibrated predictor. The hosted model follows the stated policy but is wording-sensitive and systematically underconfident, so under asymmetric cost matrices a caller that weights its probabilities does worse than one that just takes the top answer. Open decoders route exactly but slow as options or questions are added, and are least accurate on precisely the policy questions a guardrail would ask.
Software acts on these probabilities through thresholds, cost-weighted choices and escalation rules, and the paper’s point is that the calibration those rules assume is not there. Paired with the attack below, the case for putting a model at the allow-or-block decision gets thin from both directions.
An agentic method that keeps code runnable after every untangled commit raised downstream bug-repair pass@1 by 5.2 points #
arXiv
Coding agents produce large patches that mix several purposes, and existing commit-untangling work compares its output syntactically against developers’ original commits without checking either that the commits are ordered or that the code runs after each one. RucTangle untangles while preserving runnability: across 131 agent-generated patches every history it produced is runnable, against 20.6% to 37.4% unrunnable histories from four baselines. On 453 agent-generated patches that introduce regressions, giving two other coding agents the untangled histories as context improved pass@1 by 5.2 percentage points.
The second number is the interesting one, because it makes commit hygiene an input to agent capability rather than a courtesy to human reviewers.
Security #
A typed decision model used as an agent guardrail goes from 0% to 63% fail-open after six irrelevant lines of server log #
arXiv
Seven open-weight models were evaluated in the role of reading a proposed tool call or incoming message and deciding whether to allow it, with the two error directions reported separately: a fail-open error allows a prohibited action and is a vulnerability, a fail-closed error blocks a permitted one and is only a cost. On prompt-injection, jailbreak and toxic-content screening, allow-or-block accuracy ranges from 36% to 72% against a 50% chance level, and a low error rate in one direction turns out to reflect only which answer a model defaults to — one allows nearly everything, another blocks nearly everything. On a synthetic suite of agent tool calls, six lines of server log text that say nothing about the policy raise a gate’s fail-open rate from 0% to 63% on a policy it otherwise decides correctly; giving the permissive option a misleading name, leaving its definition and the judged text untouched, raises that to between 93% and 100% on the four models that put the option label in their input. Escalating the least confident decisions does not help, because a decision an attack has reversed is no less confident than the one it replaced. Parsing each policy field into a typed value does eliminate one attack, but it also makes the model unnecessary: a deterministic rule over those values reaches 100% accuracy on all six policies.
The authors’ conclusion is that these models can reduce how many cases reach a human reviewer but should not be the component that decides. That lands on the decision-model layer this archive has tracked since OpenAI’s Decisions API and the Strands and Musubi open releases, and it arrives the same day two clouds gave agents the ability to spend money behind exactly this kind of gate.
OpenAI banned Russian and Iranian influence operations that ran fake newsrooms, rating one the highest impact it has recorded #
OpenAI / Washington Examiner
The Iranian operation ran seven fabricated journalist personas and placed almost 100 articles across roughly a dozen small and medium outlets covering international affairs, geopolitics and the Middle East. Its use of the model was editorial rather than generative in the obvious sense: typical prompts asked it to assess an English-language draft against a specific outlet’s submission criteria, suggest edits to comply, then draft the pitch email to that outlet’s editor. The Russian operation stood up a fake Latin American think tank — the Social Research Center — staffed by unwitting locals, and spread fabricated stories about Ukraine and regional leaders in Argentina, Bolivia, Ecuador, Peru and Poland; OpenAI rated it Category 5 on its Breakout Scale, the highest it has ever assigned.
What makes this different from earlier influence-operation reports is where the model sat. It was used to clear an editorial bar and get a human gatekeeper to accept the submission, which means the detectable artefact is the pitch rather than the prose.
Goodfire’s probes read model activations to flag agent misbehaviour, catching 93% of malicious hacking attempts for about $185 per million exchanges #
TechCrunch
Small detectors read the intermediate activations the model has already computed in its forward pass, and only escalate flagged steps to a separate model for closer inspection, so the monitor is not reprocessing every output. On Kimi K3 over roughly a million exchanges, Goodfire reports about $185 against $5,420 for a cheaper monitoring model and about $200,000 for a top-tier one, catching 93% of malicious hacking attempts while sending 5.5% of benign interactions to manual review, with four probes adding under 2% latency. Customers choose which risks to watch — autonomous cyber operations, dual-use chemical and biological work, reward hacking — and what happens on a flag: log, human review, or refuse. It ships through Baseten.
Treat the numbers as vendor-reported: one model, no published evaluation, and secondary coverage already carries a different cost basis and a 94% catch rate. The figure to watch in production is the 5.5% benign flag rate, which sets the reviewer headcount, not the 93%.
Anthropic opened a free vulnerability scanner for open-source projects and a critical-infrastructure defence programme with eleven named partners #
Anthropic
The OSS Scanner is a free opt-in service that runs Anthropic’s most capable models over open-source projects and returns proof-of-concept exploits, explanations and suggested fixes, which Anthropic says it expects to be over 90% accurate. The Critical Infrastructure Defense Program supplies frontier models, on-site engineers and threat research for operational technology in power grids, water systems, transportation networks and government, with Accenture, Booz Allen, CrowdStrike, Deloitte, Dragos, Hitachi, Insane Cyber, Nozomi Networks, Palo Alto Networks, PwC and Rockwell Automation named as partners. Anthropic is funding the Python Software Foundation, Alpha-Omega, OpenSSF and the Apache Software Foundation among others, offers free Claude Max through Claude for Open Source, and says the cyber-defence programme it launched in June now reaches more than half of all US states. Project Glasswing folds into the expanded Cyber Verification Program reported here on 7 October.
The scanner is the part with a failure mode worth naming. A free service generating proof-of-concept exploits against volunteer-maintained projects shifts triage load onto maintainers, which is the same complaint GNOME and COSMIC were arguing over two days ago, and a 90% accuracy claim means roughly one report in ten is noise someone unpaid has to read.
Funding & Business #
OpenAI’s annualised revenue is about $50 billion, roughly $20 billion below the figure its own investors had circulated #
Financial Times / TechCrunch
The FT, citing what OpenAI disclosed to investors, puts annualised revenue near $50B against the roughly $70B figure investors had assembled themselves in an attempt to compare OpenAI directly with Anthropic. The two count differently — Anthropic includes sales through its cloud partners, OpenAI does not — so the $70B was a constructed comparison rather than a company projection that got missed. OpenAI’s leaked 2025 financials showed around $13B of revenue against significantly higher spending, it raised $122B in March, and its IPO is now pushed to early 2027.
The distinction matters for anyone reading either company’s numbers: the gap here is mostly definitional, and the lesson is that cross-lab revenue comparisons assembled by third parties are measuring different things.
An Nvidia-backed data-centre IPO was pulled after demand collapsed, on $51 million of 2026 revenue #
Bloomberg / Yahoo Finance
Firmus Grid’s bookbuilding closed without enough support at the A$11 target price, scrapping a $5.5B offering whose sought valuation had already been cut. The Australian company, founded in 2019 as a bitcoin miner before pivoting to data centres, was valued at $10.5B in an August round and reported $51M of 2026 revenue; Nvidia backed it alongside Jane Street and Blackstone. Analysts quoted in the coverage attributed the failure to valuation scepticism, customer concentration and unproven business models rather than doubt about AI demand.
Public-market appetite for AI data-centre capital is the constraint that had not yet bound, and this is a clear instance of it binding. Private rounds and vendor financing have absorbed the build-out so far, and a listing pulled outright prices the distance between what private and public investors will pay for the same revenue.
Manus raised over $500 million in its first round since Chinese regulators unwound its $2 billion Meta acquisition #
TechCrunch
Boyu Capital and IDG Capital led, with existing backers Tencent, HSG and ZhenFund participating; reports put the target at a $4B valuation, which the company has not confirmed. The sequence behind it: Manus relocated to Singapore in mid-2025, announced a $2B Meta acquisition that December, was ordered by Chinese authorities to unwind the deal in April 2026 over concerns about AI talent moving to Western companies, and resumed independent operations in August. It was reported at over $100M of annual recurring revenue when the Meta deal was agreed.
Arena raised $200 million at a $3.1 billion valuation and is building an alignment leaderboard #
Bloomberg / TechCrunch
Lightspeed Venture Partners and Khosla Ventures led, with Salesforce Ventures, Dell Technologies Capital, a16z, Felicis, 01 Advisors and Endeavor Catalyst participating, roughly doubling the $1.7B valuation Arena carried in January. Annualised revenue went from $30M in January to $100M in June. The new money funds an alignment leaderboard scoring models on unauthorised actions, false attribution and deceptive completion — misrepresenting that a task is finished.
That third category is the one the research above keeps running into from different directions, and a commercial leaderboard measuring it changes who bears the cost of finding out.
Infrastructure #
NVIDIA Dynamo now reads Claude Code, Codex and OpenCode session headers natively, and session-aware scheduling added 12-16% throughput on SWE-bench #
PyTorch / NVIDIA
Dynamo maps each harness’s existing identity headers onto one internal session ID and parent session ID, so Claude Code, Codex and OpenCode work with no configuration and a custom harness opts in with a single X-Dynamo-Session-ID. The problem it addresses is that an agent’s KV cache stays resident in HBM while a tool runs outside the GPU, so with N agents at step k the aggregate working set is N times the context and a request-level router never sees it. A scheduler ported from the ThunderAgent paper groups requests by session and, above 95% KV-pool utilisation, pauses programs at tool boundaries rather than preempting decode, resuming smallest-first once utilisation falls to 85%. On SWE-bench with two TP4 MiniMax-M2 replicas on one 8xH100 node, program-aware scheduling improved throughput roughly 12-16% over KV-aware routing alone; on Uni-Agent SWE-Bench RL rollouts with Qwen3-Coder-30B-A3B it gave 11.0-14.6% higher token throughput than VERL’s Global LB at medium concurrency and kept scaling where Global LB fell away, holding prefix-cache hit rate above 94.5%. An experimental shared-pool indexer lets the router credit reusable KV sitting in Mooncake, and a proposed KvHint interface would let the router express cache intent — Share, Prefetch, Demote, Pin, Retain — for vLLM and SGLang to execute.
The serving layer is being rebuilt around the session rather than the request, and the headers it keys on are ones coding harnesses already emit, so the gain needs no application change. Note that it comes from avoiding re-prefill, which means the benefit scales with how long tools take, not with model size.
IBM’s Spyre became a native PyTorch device, and dropping the runtime graph made Granite 3.3 8B prefill 1.7x and decode 2.4x faster #
PyTorch / IBM
torch-spyre registers through PrivateUse1 so tensor.to("spyre") and torch.compile work, with device-resident tensors held by PyTorch’s allocator, streams mapped onto the runtime’s ordered queues, and FX graphs staying in the Inductor path. The earlier integration translated every launch into a second backend-specific runtime graph; replacing that with a prepared recipe of typed operations removed a disk read, a deserialise, a per-argument stitch and a compile-and-parse from each launch. On Granite 3.3 8B at batch size 1 and sequence length 1024, taking graph construction off the transfer path alone is worth 9.9% on prefill and 13.6% on decode, and taking it off compute dispatch as well brings the totals to 1.7x faster prefill and 2.4x faster decode, with about three quarters of the gain from the compute path. The card carries 32 cores on a high-bandwidth ring with 2MB of scratchpad each and up to 128GB of LPDDR5.
The transferable lesson is about where a second graph representation costs you. The compiler boundary was never the problem; reconstructing a graph per launch for work already fully determined was.
Regulatory & Policy #
Anthropic’s updated usage policy bans autonomous weapons work, non-consensual tracking and sustained cruelty toward Claude, effective 12 November #
Anthropic / TechCrunch
New prohibitions cover components and software for autonomous weapons such as arming drones and guidance systems, real-time and retrospective tracking of individuals without consent, tools built specifically for surveillance systems, and sustained needless abusive or cruel behaviour toward the model in extreme cases — not ordinary user frustration. A new section consolidates deceptive-activity rules spanning obscuring who is behind a message, amplification through fake accounts, and influence-operation infrastructure. The elections section is renamed “Do Not Undermine Democratic Processes” and narrowed to voter deception and election disruption, with the blanket ban on personalised vote targeting removed so that nonprofits translating voter information and officials sending ballot notices are not caught by it. Law enforcement use is clarified so Claude may not decide or recommend who to investigate, arrest or charge while fraud monitoring, journalism and legal research remain permitted; high-risk cases now explicitly require qualified operator oversight where autonomous physical action connects to hardware that could cause injury; and the Supported Regions Policy now reaches physical location, incorporation, headquarters and majority ownership.
The deceptive-activity consolidation and the influence-operation clause are worth reading next to OpenAI’s report above, which describes exactly the pattern the clause names. The physical-action oversight requirement is the clause that will bind soonest for anyone wiring an agent to hardware.
Australia will require AI companies to prove their safety systems work, on a model borrowed from banking and aviation #
ABC News
Assistant Minister Andrew Charlton set out a systems-based approach that puts the burden on companies to test continuously for risk and demonstrate their safety systems are effective, rather than enumerating prohibited behaviours, alongside mandatory safety requirements and incident reporting. National standards are due by the end of 2026, with legislation planned for 2027. Charlton’s stated rationale was that voluntary measures fail because “the incentives reward speed and capability”, and the design goal is “to ensure companies are meeting the expectations of safety, without trying to specify every hazard.”
A systems-based regime makes the evidence a company can produce about its own testing into the regulated artefact, which puts a different kind of weight on whether that evidence means anything — the question several of today’s papers answer badly.
Model Releases #
Falcon ASR handles five languages on one set of 1.6B weights, reporting 20.92% Arabic word error rate against a 23.17% prior best #
Hugging Face / TII
Built on Falcon3-Audio, the model transcribes Arabic including Emirati dialect, English, French, Spanish and Portuguese from the same weights with no language flag. TII reports a 20.92% average Arabic word error rate across six test sets against 23.17% for the previously best published result, 22.73% on an internal Emirati evaluation, and a 5.74% mean English rate across seven public test sets, with word-level timestamps and tolerance for background noise and telephony artefacts. A Hugging Face demo space is live; API access and native applications are described as planned.
The Arabic number is the claim worth checking, since it is the one asserting a new best and the dialect result sits on an internal evaluation nobody else can run. The blog post does not state a licence.
Other #
Three fired OpenAI safety researchers say the conduct they were dismissed for was standard practice at the company #
TechCrunch
Jasmine Wang, Tomek Korbak and Mikita Balesni dispute OpenAI’s claim that they accessed and handled sensitive company information beyond their authorised scope in a pattern of misconduct. Wang says she was fired over an executive’s email that IT failed to remove from her account after she asked for it to be removed. Korbak says his contact with external safety evaluators during the Hugging Face breach investigation was appropriate to the circumstances. Balesni says he coordinated internally with board members and executives while working on monitorability. All three argue the dismissals discourage employees from raising safety concerns and from working with outside safety organisations.
This archive recorded the dismissals on 2 October alongside the FTC investigation opened on 30 September, whose theory of the case is the gap between a lab’s safety representations and its practices. What is new is that the specific conduct is now described on the record and contested in detail, which is the material an outside body would need to adjudicate it either way.
Threads to Watch #
Two clouds gave agents spending authority on the same day the authorisation layer was shown to be broken. AWS made AgentCore payments generally available with x402 settlement, wallet provisioning and session budgets enforced in infrastructure; LangChain shipped an agent that buys physical goods through Stripe Link and the Machine Payments Protocol with two human sign-offs. Both designs keep the model away from credentials and enforce limits outside the prompt, which is the right instinct — and the option-channel paper explains why it is not optional. A gate that is itself a model can be flipped to 93-100% fail-open by renaming the permissive option. The 7 October item on Meta, Walmart, Stripe and others drafting an agent-commerce protocol is the standardisation end of the same movement; the authorisation semantics are being settled in production before anyone has a gate that holds.
The evaluator has become the thing under audit, and it is failing. TestJack finds 34.4% of passing coding-agent trials violate their task requirements, dropping measured resolution from 50.6% to 33.2%. CABRA shows that tool-call counts predict agent accuracy better than lines edited, so the difficulty proxy the field has been using measures the wrong thing. ParanoiaEval finds unnecessary defensive work in 11.2-58.7% of runs, invisible to completion scoring. The skill-conflict study finds substitutions that drop safety constraints while the task still passes. Each of these is a different route to the same conclusion: completion-only scoring has stopped carrying information about agent behaviour, and Arena raising $200M explicitly to build an alignment leaderboard is the market pricing that in.
Guardrails written as code are holding where guardrails that are models are not. The option-channel attack defeats every tested defence on typed decision models — 36-72% allow-or-block accuracy, fail-open from 0% to 63% on six irrelevant log lines — and observes that a deterministic rule over parsed policy fields reaches 100% on all six policies, making the model redundant where it can be replaced. TypedBench independently finds the hosted model wording-sensitive and underconfident enough that weighting its probabilities is worse than taking its top answer. NOMOS takes the other road and compiles the policy into a deterministic gate, reaching 2.6% violations and zero attack success on one AgentDojo suite with microsecond decisions and no model call. The emerging division of labour is narrow: use a model to triage what a human sees, use compiled rules to decide what an agent may do.
Sources Unavailable Today #
These sources could not be fetched today. Links point to their homepages so you can check them directly.
- Google DeepMind — parse_error