Gemini 4 Argon (9 minute read)
Google announced Gemini 4 Argon, a frontier model built for sustained reasoning across software engineering, enterprise knowledge work, and cybersecurity. Access is initially limited to trusted cyber defenders, with broader availability planned after phased safety testing.
|
Musk's SpaceXAI Considers Overhaul of Pricing for Grok, X Users (1 minute read)
SpaceX is considering an overhaul of its subscription pricing options for Grok and X. It plans to soon offer a unified subscription for the services with four pricing tiers. The options will range from a free offering with stricter usage limits for its Grok chatbot to a $100 per month Ultra tier that includes access to the Grok Bot AI agent. The $8 per month 'lite' plan will include access to Grok, a verified checkmark, and fewer ads on X.
|
Claude for Government is now generally available (3 minute read)
Claude for Government is now generally available, offering FedRAMP High authorized coding and agent capabilities to federal and state agencies. Agencies receive Anthropic commercial features with robust governance controls and customized administrative options. Early access is also available for Claude Code CLI and Claude for Microsoft 365.
|
|
We can and must solve alignment (14 minute read)
The Goodfire team identifies interpretability as the main challenge in AI alignment, highlighting its critical role in understanding and controlling AI behaviors. They stress the urgency of developing and deploying tools for detecting, debugging, and designing AI systems, focusing on controlling generalization and verifying what models learn. As AI models advance, they seek breakthroughs in interpretability to ensure safe and reliable AI alignment, advocating for a collective effort among the tech community.
|
DeepSeek Builds for Huawei Ascend (6 minute read)
DeepSeek plans to start rebuilding the software ecosystem around Nvidia's hardware so that Chinese chips become genuinely easier to use and capable of training frontier AI models. Nvidia's CUDA ecosystem provides developers with a mature set of tools for writing programs, optimizing performance, and getting thousands of GPUs to work together efficiently. DeepSeek is trying to create a layer of abstraction between models and the underlying GPU hardware that could potentially sit on top of either Nvidia GPUs or Huawei Ascend chips. If all goes well, the cost and difficulty of moving AI workloads from Nvidia to Chinese AI hardware could fall significantly.
|
|
Praxis-1 (5 minute read)
Praxis-1 is an open-weight world action model that turns Runway's video pretraining into control for real robots. The model is currently being tested with early partners, each running on their own hardware. It will be released publicly in the coming months. Clips of robots running the model are available in the post.
|
NVIDIA OpenShell Secures Autonomous AI Agents (GitHub Repo)
NVIDIA's OpenShell provides a policy-controlled runtime for autonomous agents, enforcing file, system-call, network, and credential access at the kernel level. It also uses formal verification to identify risky permissions before policy changes are applied.
|
Generalization Dynamics of LM Pre-training (4 minute read)
During pre-training, models can abruptly switch response patterns, a phenomenon termed MODE-HOPPING. Tests reveal this behavior in OLMo3-32B during arithmetic tasks, showing sudden shifts between pattern-following and correct task inference as training progresses. Controlled experiments suggest competing shallow and generalizable circuits within models, revealing that longer training doesn't always guarantee better generalization.
|
A new transformer passes its hidden state to the next token instead of recomputing it (25 minute read)
Researchers built LIFT, a transformer that feeds its internal state back into the next generation step instead of squeezing everything through the single token it outputs. Training stays parallel because the states come precomputed from an off-the-shelf model's next-token predictions, while at inference the model feeds back its own. LIFT models from 135M to 1B parameters beat token-matched standard transformers on language modeling, reasoning, and procedural tasks, so the next question is whether the gains hold at frontier scale.
|
|
Factory CEO just accused his VC board adviser of spying for Cognition (6 minute read)
Factory has fired VC Chris Degnan from his role as a board advisor due to allegations that Degnan shared confidential information with Cognition, Factory's biggest competitor. Degnan joined Cognition as its chief revenue officer two hours after the announcement. He refutes that he was fired and said that he resigned when he told Factory's CEO he was taking the job. There could be consequences for perceived board-level conflicts of interest.
|
|
Love TLDR? Tell your friends and get rewards! |
|
Share your referral link below with friends to get free TLDR swag!
|
|
|
| Track your referrals here. |
|
|
|
0 Comments