2.AUG 2026 · ISSUE 4 · 5 MIN
Today’s Lead story:
OpenAI has deployed an internal version of Astra, its designated next-major model, to autonomously solve 10 open mathematical problems that had seen no main-result progress for over a decade.
To validate the findings, OpenAI published Lean 4 formalizations of the proofs in a public repository (openai/ten-proofs), along with a research paper and an LLM-generated PDF reconstructing the proof assembly from unpublished internal reasoning traces. OpenAI reported spending under $2,000 per problem based on GPT-5.6 Sol token pricing.
The announcement highlights a shift toward automated formal proof verification, where frontier models take over heavy technical proof execution while human researchers focus on problem framing.
What changed:
Anthropic Mythos Preview for Deep Research: Anthropic revealed technical capabilities for its preview model, Mythos Preview, using it to discover novel cryptographic vulnerabilities. Anthropic logged token spending reaching $100,000 on deep research prompts specifically targeting non-trivial theoretical findings.
Automated AI Engineering & Serving Gains: Internal metrics across frontier labs reveal heavy reliance on automated AI development and customized architecture:
Anthropic now generates 80% of its internal software codebase using Claude Code.
OpenAI utilized its Sol system to cut end-to-end model serving costs by 20%.
Moonshot AI (Kimi) deployed Kimi K3 to design a custom chip optimized to serve a nano model built on its own architecture.
ChatGPT Work & Personalized Generation: OpenAI CEO Sam Altman referenced ChatGPT Work, highlighting features that ingest family calendars and user context to generate customized daily audio podcasts.
Frontier Governance & Open Weights Split: A sharp split has emerged among AI providers regarding distillation and open-weight access:
Open Weights and Distillation Support: Microsoft, NVIDIA, Amazon, Y Combinator, The Linux Foundation, and OpenAI signed a joint letter advocating for open-weight models and explicitly defending model distillation (training one model on another’s outputs) as a legitimate development technique.
Anthropic Distillation Opposition: Anthropic CEO Dario Amodei publicly opposed the letter, advocating for a policy crackdown on "industrial-scale distillation operations" due to national security and misuse risks.
Automated Research Concerns: Over 1,300 frontier AI employees (including technical leaders from OpenAI, Safe Superintelligence Inc., and Anthropic) signed the Pacing the Frontier letter, requesting government support to pace automated AI research due to risks from self-accelerating development.
Claude Shared Data Indexing: Reports surfaced that publicly shared links from Claude chats and Artifacts have been indexed by Google search engines, creating potential exposure risks for shared model outputs.
Practical details:
Model / System | Developer | Status / Availability | Notable Technical & Operational Details |
|---|---|---|---|
Astra (Internal) | OpenAI | Internal preview | Solved 10 decade-old math problems; output verified via Lean 4 ( |
Mythos Preview | Anthropic | Research preview | Specialized high-horizon reasoning; applied to advanced cryptographic research with token spend scaling up to $100,000 per run. |
Sol | OpenAI | Infrastructure / Serving | Internal serving optimization layer; reduced end-to-end model serving infrastructure costs by 20%. |
Claude Code | Anthropic | Internal deployment | Agentic coding system currently generating 80% of Anthropic’s internal codebase. |
Kimi K3 | Moonshot AI | Internal architecture | System capability demonstrated by designing hardware chips tailored to host its native nano models. |
Builder/Operator angle:
Integrate Formal Verifiers into Eval Suites: OpenAI's Astra release shows that high-level mathematical and logic reasoning outputs can be formally checked using interactive theorem provers like Lean 4. Teams building reasoning evaluation suites should pair LLM output generation with formal compiler-based verification steps to catch plausible-sounding hallucinations.
Audit Distillation Pipelines: If your startup or enterprise pipeline relies on distilling outputs from frontier models (e.g., using synthetic data from closed APIs to train smaller, specialized models), monitor policy developments closely. While OpenAI and Microsoft back distillation, regulatory pressure or terms-of-service updates driven by Anthropic's policy efforts could target commercial distillation workflows.
Account for Long-Horizon Compute Costs: Cost structures for deep reasoning tasks vary wildly based on depth. While Astra’s formal proofs cost <$2,000 per problem on GPT-5.6 Sol pricing, Anthropic’s Mythos Preview spent $100,000 on deep research runs. Operators must set strict budget caps and monitor billing alerts on open-ended research agents.
Review Data Exposure Risks: Given reports of Claude shared chats and Artifacts appearing in search engine indexes, security teams should immediately review company policies regarding public link creation on AI platforms to prevent sensitive internal prompts or code from being crawled.
Signals to watch:
Astra & Mythos Public Benchmarks: Watch for standard benchmark evaluations (e.g., SWE-bench, MATH, GPQA) and commercial API pricing for OpenAI Astra and Anthropic Mythos Preview upon public release.
Regulatory Action on Distillation: Track whether policy bodies move to define or restrict "industrial-scale distillation," which could disrupt fine-tuning workflows for open-weights developers.
Automated AI Development Guardrails: Watch for institutional responses to the Pacing the Frontier request signed by frontier researchers concerning self-improving AI research systems.
Caveats:
OpenAI did not disclose the total number of math problems Astra attempted without reaching a valid solution, making the overall failure rate for the system unclear.
Exact prompts used to generate the Astra math proofs and the reasoning traces for Mythos Preview were not fully published.
Head-to-head performance comparisons on standardized benchmarks between OpenAI Astra and Anthropic Mythos Preview are not yet available.
References to Anthropic's "Claude Opus 5" performance (e.g., managing vending machine simulations) remain limited to brief secondary reports without published technical reports.
Sources:
Open letters about AI development — https://simonwillison.net/2026/Aug/2/open-letters/
Ten advances in mathematics and theoretical computer science — https://simonwillison.net/2026/Aug/1/ten-advances-in-mathematics/
Inside the London hacker house taking a stand against founder burnout — https://techcrunch.com/2026/08/01/inside-one-london-founder-house-rewriting-the-founder-house-rules/
Sam Altman is still making the case for parenting via ChatGPT — https://techcrunch.com/2026/08/01/sam-altman-is-still-making-the-case-for-parenting-via-chatgpt/
