Top story
White House Asks OpenAI to Delay New Model Over Safety Concerns. The administration is urging OpenAI to slow-roll its latest release, signaling direct federal pressure on frontier model rollouts. Source
Research
Compiling Agentic Workflows into LLM Weights. A paper proposes distilling agentic workflows into static weights to reach near-frontier quality at two orders of magnitude lower cost. Source
Hallucination in World Models is Predictable and Preventable. Demonstrates that world model hallucinations follow systematic patterns and can be mitigated before deployment. Source
JetSpec: Parallel Tree Speculative Decoding. Causality-preserving parallel tree drafting achieves up to 9.64× lossless speedup and ~1000 TPS on B200 hardware. Source
The Co-Failure Ceiling on Combining Language Models. Across 67 frontier models, routing/voting/MoA gains are capped by how often member models co-fail on the same query. Source
Tools
Nemotron-TwoTower-30B-A3B-Base-BF16. NVIDIA releases a diffusion-based language model built on a Nemotron 3 Nano MoE backbone. Source
audio.cpp. A single C++/ggml runtime bundling 12 TTS models (Qwen3-TTS, PocketTTS, VeVo2) with up to 5× speedup over Python on CUDA. Source
Kuma: PyTorch to WebGPU Executables. Compiles PyTorch models into self-contained WebGPU binaries, enabling browser inference without Python. Source
Run vLLM Server on HF Jobs in One Command. Hugging Face ships one-command vLLM deployment on its managed Jobs infrastructure. Source
Industry
Z.ai Closes Frontier Gap, Plans Dual Listing. China's Z.ai (formerly Zhipu) claims GLM-5.2 nears frontier models, accelerating a dual listing amid Anthropic disruption. Source
Anthropic Accuses Alibaba of Illicit Capability Extraction. Anthropic publicly alleges Alibaba extracted its AI capabilities, escalating US–China frontier-model tensions. Source
Databricks' Former AI Chief Targets 1,000× Power Reduction. A new venture claims a path to slashing AI compute/power costs by three orders of magnitude. Source
General Intuition Raises $2.3B for Game-Trained Agents. Massive round bets that video-game environments can train real-world AI agents. Source
Community
BenchPress: LLM Benchmark Matrices are Near Rank-2. Microsoft AI Frontiers shows a handful of probes can predict full benchmark profiles before running them. Source
Linux Foundation Proposes DNS as AI Agent Identity Layer. A new initiative would repurpose DNS infrastructure to give AI agents verifiable identities. Source
2,000 Adversarial Emails vs. an AI Assistant. Large-scale red-team test of an AI email agent shows surprisingly low prompt-injection success, with methodological caveats. Source
Apple Reportedly Skips M6 for AI-Focused M7 Line. M7 Pro/Max/Ultra chips would dramatically boost memory bandwidth to enable on-device LLM inference. Source