Top story
Top Story: Claude Fable 5 returns globally. Anthropic officially re-deploys Claude Fable 5 worldwide with new cybersecurity classifiers, Opus 4.8 fallback, and ongoing classifier refinement. Source
Research
Measuring the Gap Between Human and LLM Research Ideas. Empirical study comparing novelty and quality of LLM-generated vs. human research ideas. Source
Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training. Shows a single transformer layer can match full-parameter RL fine-tuning under certain conditions. Source
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?. Examines whether coding-agent benchmarks reliably measure real performance or reward overfitting. Source
Skills Are Not Islands: Measuring Dependency and Risk in Agent Skill Supply Chains. Measures interdependency and systemic risk in agent skill supply chains. Source
Tools
Kimi K2.7 Code available in GitHub Copilot. Kimi K2.7 Code becomes the first open-weight model selectable in GitHub Copilot, hosted on Azure across Pro/Pro+/Max plans. Source
Senior SWE-Bench. Snorkel releases Senior SWE-Bench, an open-source benchmark evaluating coding agents on realistic feature/bug tasks with a novel validation agent for behavioral testing. Source
Gemini Spark launches on Mac. Google's agentic assistant Gemini Spark arrives on Mac with real-time tracking and expanded app support. Source
NVIDIA Qwen3.6-35B-A3B-NVFP4. NVIDIA releases NVFP4-quantized 35B MoE variant of Qwen3.6. Source
Industry
OpenAI proposes 5% stake to Trump administration. Report via FT that OpenAI is proposing a 5% equity stake to the Trump administration. Source
NVIDIA parallel token-writing demo. NVIDIA demonstrates splitting a 30B model into two parallel token writers for faster inference. Source
DeepSeek plans agentic vulnerability model. ASPI scoop: DeepSeek job postings suggest plans for an agentic model targeting code vulnerabilities. Source
Meta looks to monetize excess AI compute. Meta explores monetizing surplus AI compute capacity, following a SpaceX-style model. Source
Community
Claude Sonnet 5 Is Not Frontier But Has Its Uses. Analytical take arguing Claude Sonnet 5 is not frontier-tier but has niche utility. Source
The Coolest Diffusion Research Isn't in LLMs. Discussion of cutting-edge diffusion model research applied to molecular AI with Genesis Molecular AI researchers. Source
Cursor's enterprise AI deployment. Inside look at how Cursor deploys AI-powered coding tools in enterprise environments. Source
Autoresearch feedback loop. Analysis of the autoresearch feedback loop mechanism enabling self-improving AI agents. Source