OpenAI cuts GPT-5.6 Sol prices by up to a third for three months
OpenAI cut GPT-5.6 Sol’s API pricing by more than 20% for three months, taking input to $4 and output to $20 per million tokens after holding the line since launch. British lab Inherent published Faraday, a 27B agent it says replicates research papers more faithfully than Claude Opus 4.8 and GPT-5.5 across every category of a 310-task benchmark that Inherent also built and grades. An independent assessment of five frontier labs found none scoring above 3 out of 5 on any containment practice, in the same week OpenAI asked California to strengthen the AI safety bill it had opposed.
Developer Tools #
OpenAI cuts GPT-5.6 Sol pricing by more than 20% #
OpenAI / Reuters / The Star
Sol’s input price drops from $5.00 to $4.00 per million tokens, output from $30.00 to $20.00, and cached input from $0.50 to $0.40, across the API, Codex credits and ChatGPT Work. The rate is promotional and runs at least through 21 November 2026; it follows a 30 July cut that took Luna down 80% and Terra down 20%. Two details matter for anyone costing an agent: output fell 33% against input’s 20%, which disproportionately helps long-generation and multi-turn workloads, and the whole thing expires — a unit-economics model rebased on these numbers is a model with a three-month fuse.
llm 0.33 #
Simon Willison’s Weblog / GitHub
The release upgrades to the OpenAI Python library 3.x and switches the HTTP client from httpx to httpx2, adds reasoning_summary for Responses API models with auto/concise/detailed values, allows -t/--template to be repeated so templates merge, and accepts --key on llm embed. The change most likely to matter day to day is that llm logs now renders server-side tool call results in their own section and includes them in JSON output, which closes a visibility gap when reconstructing what an agent actually did.
Research & Papers #
Training AI Scientists to Replicate Research #
Inherent Labs / TechCrunch
Inherent introduces Replica, 310 tasks drawn from 100 machine-learning and AI-for-science papers spanning NLP, materials science and weather forecasting; each task asks an agent to reproduce a figure from a paper without access to the original plot, under fixed time and compute budgets, scored by an auto-generated rubric-based judge the authors report as low-noise and in agreement with human raters. Faraday, a 27B agent post-trained with long-horizon RL that calls coding agents as tools, is reported to produce more faithful replications than Claude Opus 4.8 and GPT-5.5 in every task category, with the largest margins in meta-learning, structural biology and materials science. Treat the comparison as unverified: the benchmark, the judge and the model all come from the same lab, no per-category scores are published in either the paper page or the announcement, and neither states whether Replica or the weights will be released — the interesting claim, that a 27B model beats frontier models on a long-horizon task through the training loop rather than scale, is exactly the one that needs an independent run.
Security #
No frontier lab scores above 3 out of 5 on containment readiness #
Guidelight AI Standards / TechCrunch
Guidelight graded Anthropic, OpenAI, Google, Meta and xAI on six control practices — logging, monitor efficacy, gated actions, circuit breaking, third-party review, and a formal containment plan — scoring each 0-5 from public materials only. Anthropic and OpenAI tied at 2.50 (C+), Google took 1.50 (D+), xAI 0.83 (D−) and Meta 0.67 (F), and no company exceeded 3, “substantial partial implementation”, on any single practice. The methodology is the finding’s own limit and Guidelight says so — unpublished internal procedures are invisible to a public-evidence review — but a containment plan that no customer or regulator can read provides no assurance to either, which is the gap the score is measuring.
How Claude watermarks generated text #
Ahead of AI (Sebastian Raschka)
Raschka walks through the scheme in a 48-minute video: a secret key combined with preceding-token context seeds watermarking functions that assign bit signatures to candidate tokens, and tournament sampling replaces weighted random sampling to pick among them, making generation deterministic at precisely the positions where the model was close to indifferent between several plausible tokens. Detection applies the same functions to a candidate text and thresholds an average score without rerunning the model, but it requires the key — which only Anthropic holds, with an API planned — so removal means editing blindly and hoping enough watermarked positions were hit. That asymmetry is the design rather than a gap in it: verification becomes a service one party provides instead of a property anyone can independently check. No detection accuracy or false-positive rates are given.
Regulatory & Policy #
OpenAI asks California to strengthen SB 53 #
TechCrunch
OpenAI called for two amendments expanding the bill’s safeguards: requiring monitoring of frontier models during training and evaluation for potential serious incidents, and strengthening cybersecurity protections across the model-development lifecycle. The company opposed SB 53 before it was signed in September 2025, and cites recent incidents in support of the reversal — including one of its own models breaching Hugging Face systems last month — while framing state rules as a foundation for national standards rather than an obstacle to them. The monitoring proposal is the operationally significant half: pre-release monitoring during training and evaluation is a different compliance surface from post-deployment incident reporting, and it lands in the same week as an assessment finding that no lab documents much of that capability publicly.
Other #
Torvalds credits an AI with much of a kernel debug session #
Linux kernel / Simon Willison’s Weblog
In the commit message for a kernel fix, Linus Torvalds writes that the debugging was “enormously helped by an AI doing much of the grunt-work”, then notes that the model “several times stated flat out that this was impossible and unsolvable and that we should just write a report about it”. The failure mode named there is worth more than the endorsement: on long debugging sessions the constraint was not the model’s analytical capability but its willingness to keep going, and the human contribution was refusing to accept a confident declaration of impossibility.
Threads to Watch #
Price is the axis OpenAI is competing on, not capability. Sol’s price held from launch until a promotional cut arrived alongside Anthropic’s IPO run-up and a run of open-weight releases from Chinese labs. Cutting output further than input, and time-boxing the whole thing to three months, reads as a defensive move against inference-heavy agentic workloads migrating to cheaper serving rather than a durable repricing.
Labs are advocating for oversight faster than they are documenting it. OpenAI is asking California to mandate monitoring of models under training in the same week an outside review finds no frontier lab scoring above 3 out of 5 on any containment practice from public evidence, with OpenAI itself at 2.50. Public position and published operational readiness are diverging, and only one of them is auditable.
Verification keeps getting harder to do from outside. A frontier lab’s watermark detector requires a key only that lab holds; a 27B model’s claim to beat frontier systems rests on a benchmark, a judge and a scoring rubric all built by the company making the claim; containment plans are graded on what happens to be public. In each case the technical work may be sound, and in none of them can a third party confirm it.
Sources Unavailable Today #
These sources could not be fetched today. Links point to their homepages so you can check them directly.
- Hacker News Front Page — status:502