OpenAI launches GPT-5.6 family (Sol, Terra, Luna). Headline features include a "ultra" high-capability mode, claimed SOTA on ARC-AGI-3 for Sol, and Sol autonomously post-training its smaller Luna sibling. Source

GPT-5.6 release notes & community discussion. Improved intent inference, native image dimensions, and SOTA on ARC-AGI-3 spark debate over benchmark selection choices. Source

Universal jailbreaks found on unreleased GPT-5.6 Sol. UK AI Security Institute reports exploits that bypass cybersecurity safeguards and allow automated exploit development. Source

Second-order early warning signal for multi-turn prompt injection. Information-geometry-based approach detects multi-turn prompt injection attacks. Source

The Illusion of Equivalency: Quantization Effects in LLMs. Statistical study challenges the assumption that quantized LLMs match full-precision behavior. Source

ChatGPT Work, autonomous long-horizon agent. Powered by Codex and GPT-5.6, adds a dropdown toggle for Codex mode. Source

Robostral Navigate (Mistral). 8B embodied navigation model claiming SOTA on R2R-CE. Source

GPT-Live conversational voice model. OpenAI's new voice model aimed at natural conversational flow. Source

ChatGPT Sites. New ChatGPT feature for building sites inline. Source

GPT-5.6 is now the preferred model in Microsoft 365 Copilot. Deployed across Word, Excel, PowerPoint, Chat, and Cowork. Source

OpenAI model wins AtCoder with a perfect 8,300 in six hours. Solved all five problems, outperforming every human competitor. Source

Meta releases Muse Spark 1.1. Claims double the LegalBench performance of Fable at one-tenth the cost; tops Harvey's Legal Agent Bench. Source

Ben Bernanke joins Anthropic Oversight Trust. Former Fed Chair added to Anthropic's governance body. Source

Elon Musk concedes Anthropic currently leads AI. Cites Mythos/Fable models in a public reversal. Source

DeepSeek V4 Flash on a single RTX 6000 Pro. Community member runs the model on one GPU via a custom vLLM fork and shares load/runtime notes. Source

Building a real-time AI tutor for 5-year-olds. Show HN on a reading tutor with reflections on agentic design tradeoffs under tight latency budgets. Source

Benchmarking coding agents on Databricks' multi-million line codebase. Real-world evaluation of coding agents at scale. Source