Top story
OpenAI launches GPT-5.6 family (Sol, Terra, Luna). Headline features include a "ultra" high-capability mode, claimed SOTA on ARC-AGI-3 for Sol, and Sol autonomously post-training its smaller Luna sibling. Source
Research
GPT-5.6 release notes & community discussion. Improved intent inference, native image dimensions, and SOTA on ARC-AGI-3 spark debate over benchmark selection choices. Source
Universal jailbreaks found on unreleased GPT-5.6 Sol. UK AI Security Institute reports exploits that bypass cybersecurity safeguards and allow automated exploit development. Source
Second-order early warning signal for multi-turn prompt injection. Information-geometry-based approach detects multi-turn prompt injection attacks. Source
The Illusion of Equivalency: Quantization Effects in LLMs. Statistical study challenges the assumption that quantized LLMs match full-precision behavior. Source
Tools
ChatGPT Work, autonomous long-horizon agent. Powered by Codex and GPT-5.6, adds a dropdown toggle for Codex mode. Source
Robostral Navigate (Mistral). 8B embodied navigation model claiming SOTA on R2R-CE. Source
GPT-Live conversational voice model. OpenAI's new voice model aimed at natural conversational flow. Source
ChatGPT Sites. New ChatGPT feature for building sites inline. Source
Industry
GPT-5.6 is now the preferred model in Microsoft 365 Copilot. Deployed across Word, Excel, PowerPoint, Chat, and Cowork. Source
OpenAI model wins AtCoder with a perfect 8,300 in six hours. Solved all five problems, outperforming every human competitor. Source
Meta releases Muse Spark 1.1. Claims double the LegalBench performance of Fable at one-tenth the cost; tops Harvey's Legal Agent Bench. Source
Ben Bernanke joins Anthropic Oversight Trust. Former Fed Chair added to Anthropic's governance body. Source
Community
Elon Musk concedes Anthropic currently leads AI. Cites Mythos/Fable models in a public reversal. Source
DeepSeek V4 Flash on a single RTX 6000 Pro. Community member runs the model on one GPU via a custom vLLM fork and shares load/runtime notes. Source
Building a real-time AI tutor for 5-year-olds. Show HN on a reading tutor with reflections on agentic design tradeoffs under tight latency budgets. Source
Benchmarking coding agents on Databricks' multi-million line codebase. Real-world evaluation of coding agents at scale. Source