We Must Pace the Frontier (23 minute read)
Anthropic CEO Dario Amodei called for slowing the rate of frontier AI capability development and proposed measures including independent evaluators to verify safety commitments and incident reporting.
|
Introducing Projects (5 minute read)
Cursor Projects lets users take on larger bodies of work. It can maintain context over months of work, delegate tasks to thousands of agents, and perform recurring work without being prompted. Cursor Projects frees developers from managing agents and lets them direct the work itself by moving up a level of abstraction. The tool has been a substantial productivity multiplier at Cursor: new users merge 30% more PRs.
|
ARC-AGI-4 (2 minute read)
ARC Prize believes that open source will be the foundation for advanced AI capable of scientific innovation. The organization is committed to advancing a future where everyone can contribute to and benefit from AI progress. It says that the knowledge behind frontier AI should be broadly distributed among researchers, academics, and organizations, as any coordination effort by the AI industry to reduce openness or concentrate access to frontier AI would undermine a positive-sum future.
|
OpenAI Pushes Its IPO Beyond 2026 (2 minute read)
Sam Altman said OpenAI would not go public in 2026, arguing that current AI safety concerns made an IPO ill-advised. The company had previously filed confidentially and was reportedly considering a 2027 listing instead.
|
|
AI researchers debate how close we are to recursive self-improvement (98 minute read)
This post features a transcript of a podcast with Beren Millidge, the CTO of Zyphra, John Schulman, the chief scientist at Thinking Machines and a co-founder of OpenAI, and Charlie O'Neill, head of model training at Baseten. The episode uncovers the details of what's happening at the frontier and what comes next. A link to the full video is available.
|
GPT-6-Astra Can Do Ambitious Things (52 minute read)
Astra likely has the highest raw intelligence factor of any model. It is amazing at doing things in 3D, anything involving games, computer use, and subagent coordination. Many benchmarks show dramatic jumps from all previous models. While its performance in coding isn't a quantum leap from Sol, it is very good and makes progress over the previous model. OpenAI has already soft-announced that it has an internal model a level above Astra.
|
A cache hit is not proof that you skipped the work (12 minute read)
A cache hit can be true and still fail to prove that work was skipped. A cache event becomes evidence when the independent oracle expects the prefix, the engine attests it, the prompt path skips it, the output stays identical, the evaluator passes, and the verifier binds those facts to the exact public bundle.
|
|
Sakana: Fugu Ultra v2 (3 minute read)
Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. It uses a language model trained to route tasks across a fixed pool of open and specialized models and to recursively call instances of itself. The model prioritizes answer quality on complex multi-step reasoning, autonomous research, and full-stack software development, and does not rely on individual proprietary frontier models in its pool. It supports configurable reasoning effort, function calling, structured outputs, image and PDF input, and built-in web search.
|
Recurrent Looped Transformer (4 minute read)
The Recurrent Looped Transformer combines a causal encoder with a recurrent decoder that carries its final hidden state and layerwise sliding-window attention cache across every prompt and response token. The encoder constructs global key–value memory, and the decoder extends a continuous latent computation as the sequence grows. The design brings together latent reasoning with unbounded temporal depth, model–hardware co-design, and model–RL algorithm co-design. Realized reasoning gains, hardware efficiency, and RL scaling remain to be established.
|
SWE Benchmark (10 minute read)
The new Real-SWE benchmark tests AI models on complex tasks using private enterprise codebases, reflecting real software engineering conditions. With a maximum resolution rate of 38.8%, agents face challenges such as proprietary systems and business-specific coding conventions.
|
ToolGrad: Efficient tool-use dataset generation with textual “gradients” (3 minute read)
Google Research flipped how you make tool-use training data: ToolGrad builds a verified API chain first, then writes the user question, instead of inventing a request and hoping an agent finds a working path. That answer-first loop hit a 99.8% success rate on 16,000 real APIs, and a Gemma 3 12B model trained on just 500 of those examples matched Gemini 2.5 Pro on a tool-use test with APIs it had never seen.
|
|
The frontier now ships twice. The second copy is not for sale. (6 minute read)
Anthropic, Google, and OpenAI each shipped their best model twice this month: a public paid tier and a vetted identity-gated tier with the sharper capabilities. Public prices barely moved, but Mythos, Flash Cyber, and Astra's advanced path now ask for org IDs, government ID, or trusted-defender status instead of a bigger budget.
|
|
px0 (Website)
px0 is a read-only IDE that turns browsers into an instant verification console for whatever your agents just wrote.
|
|
Love TLDR? Tell your friends and get rewards! |
|
Share your referral link below with friends to get free TLDR swag!
|
|
|
| Track your referrals here. |
|
|
|
0 Comments