Top story
GPT-5.6 Sol Ultra will be in Codex. Leak/confirmation that GPT-5.6 Sol Ultra with subagent 'ultra mode' is rolling out via Codex, with enterprise users noting shift to cheaper-model guidance from OpenAI. Source
Research
Does code cleanliness affect coding agents? A controlled minimal-pair study. Controlled minimal-pair study on whether clean vs. degraded codebases affect AI coding-agent performance, with significant methodological caveats raised by the community. Source
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning. Argues monotonic inference policies are the real objective in LLM RL, critiquing training-policy optimization. Source
Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming. Unified multi-layer red-teaming framework for evaluating security of AI agents. Source
Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots. Embodied.cpp: a portable C++ inference runtime for embodied AI models targeting heterogeneous robot platforms. Source
Tools
New open model from Tencent Hy: Hy3 (295B total 21B active - apache 2.0). Tencent releases Hy3, a 295B-total/21B-active MoE language model under Apache 2.0. Source
Introducing OpenScience: A better, open-source Claude Science. SynScience launches OpenScience, an open-source, model-agnostic scientific AI agent framework with 250+ research skills and reproducible multi-agent graphs. Source
Qualcomm launches GenieX to run LLMs on their Windows Laptops. Qualcomm launches GenieX to run GGUF LLMs on Windows laptop CPU/GPU/NPU, with reported throughput for Gemma 4 26B and Qwen 3.6 27B. Source
LivePortrait distilled model that can run at 25fps in the browser. Distilled LivePortrait model that runs portrait animation at 25fps directly in the browser. Source
🤗 Kernels: Major Updates. Major updates to Hugging Face Kernels, improving custom kernel deployment for ML workloads. Source
Industry
Zuckerberg says AI agent development going slower than expected. Zuckerberg says Meta's AI agent progress is slower than hoped; HN commenters largely agree that fully autonomous coding agents still aren't viable despite increased throughput. Source
Llama-Server is Throwing Away Your Perfectly Good KV Caches, and How to Fix It. Detailed write-up of a llama-server bug that discards restored KV caches across process restarts, with root cause and fix. Source
Amazon will stop accepting new customers for Mechanical Turk. Amazon stops onboarding new MTurk customers, signaling shift in the crowdsourcing data-labeling ecosystem. Source
VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon. Lightweight detect-and-correct inference module for adaptive action horizons in robot VLAs. Source
Community
New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]. Dartmouth study claims an AI tutor produces 0.7–1.3 SD learning gains, though commenters flag methodological weaknesses and novelty/Hawthorne effects. Source
Competence Gate: gating tool-use on a small model's internal confidence signal. LoRA adapter for Qwen3.5-4B that gates tool use on internal confidence activations rather than verbalized confidence, with local inference support. Source
Does anyone have a name for that subtle "Sameness" creeping into model outputs lately?. Practitioner notices 'EchoCreep'—converging cadence, hedging, and blind spots across recent models—and hypothesizes shared synthetic-data lineage as the cause. Source
Some of the nation's rich are letting AI teach their kids. Article on affluent families using AI tools to homeschool their children. Source