Claude Sonnet 5. Anthropic launches Claude Sonnet 5, with community discussion focused on cost-per-task tradeoffs and comparisons to Opus and GLM 5.2. Source

Multi-Block Diffusion Language Models. Paper presenting a multi-block diffusion paradigm as an alternative generative approach for LLMs. Source

Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA. Analysis revealing how training on self-generated QA pairs produces hidden instability and degraded generalization. Source

TabFM: A zero-shot foundation model for tabular data. Google releases a zero-shot foundation model for tabular data, with community critique on its benchmark reporting. Source

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs. RL method using metacognitive feedback to elicit calibrated uncertainty estimates from LLMs. Source

Claude Science. Anthropic launches Claude Science, an AI workbench app for scientists with integrated tools, auditable artifacts, and flexible compute access. Source

Redeploying Fable 5. Anthropic redeploys Claude Fable 5 after the US Department of Commerce lifts export controls on Fable 5 and Mythos 5, detailing global availability and plan limits. Source

GeneBench-Pro. OpenAI introduces GeneBench-Pro, a benchmark for evaluating AI models on gene-related biological tasks. Source

NotebookLM TikTok-style clips. Google's NotebookLM adds a feature to summarize research as short-form audio/video clips. Source

Claude Code Steganographic Marking. Hacker News thread investigates reports that Claude Code embeds hidden steganographic markers in requests, with technical and ethical debate. Source

Core Dump Epidemiology: Fixing an 18-Year-Old Bug. OpenAI recounts diagnosing and patching an 18-year-old core-dump bug discovered in production infrastructure. Source

Export Controls Lifted on Fable 5 and Mythos 5. The US Department of Commerce lifts export controls on Claude Fable 5 and Mythos 5, with cybersecurity-related restrictions affecting coding tasks. Source

DeepSeek-V4-Flash MXFP4 KV Cache Behavior. LocalLLaMA users report DeepSeek-V4-Flash MXFP4 compute buffer scales roughly 3x depending on KV cache quantization type (f16 vs q8_0). Source

The Twilight of the Chatbots. Ethan Mollick reflects on the shifting chatbot landscape as agentic and multimodal systems rise. Source

Why integrations matter less for AI agents. Argues the real value of desktop AI agents is the ratio of new information surfaced, not raw app-connect count. Source

Boris Cherny on role convergence. Reflection on how engineering, product, design, and data science roles are merging into a new single discipline in the AI era. Source

Brain2Qwerty: Non-surgical brain-to-text. Meta showcases a non-surgical BCI technique moving closer to practical at-home brain-driven communication. Source