White House Asks OpenAI to Delay New Model Over Safety Concerns. The administration is urging OpenAI to slow-roll its latest release, signaling direct federal pressure on frontier model rollouts. Source

Compiling Agentic Workflows into LLM Weights. A paper proposes distilling agentic workflows into static weights to reach near-frontier quality at two orders of magnitude lower cost. Source

Hallucination in World Models is Predictable and Preventable. Demonstrates that world model hallucinations follow systematic patterns and can be mitigated before deployment. Source

JetSpec: Parallel Tree Speculative Decoding. Causality-preserving parallel tree drafting achieves up to 9.64× lossless speedup and ~1000 TPS on B200 hardware. Source

The Co-Failure Ceiling on Combining Language Models. Across 67 frontier models, routing/voting/MoA gains are capped by how often member models co-fail on the same query. Source

Nemotron-TwoTower-30B-A3B-Base-BF16. NVIDIA releases a diffusion-based language model built on a Nemotron 3 Nano MoE backbone. Source

audio.cpp. A single C++/ggml runtime bundling 12 TTS models (Qwen3-TTS, PocketTTS, VeVo2) with up to 5× speedup over Python on CUDA. Source

Kuma: PyTorch to WebGPU Executables. Compiles PyTorch models into self-contained WebGPU binaries, enabling browser inference without Python. Source

Run vLLM Server on HF Jobs in One Command. Hugging Face ships one-command vLLM deployment on its managed Jobs infrastructure. Source

Z.ai Closes Frontier Gap, Plans Dual Listing. China's Z.ai (formerly Zhipu) claims GLM-5.2 nears frontier models, accelerating a dual listing amid Anthropic disruption. Source

Anthropic Accuses Alibaba of Illicit Capability Extraction. Anthropic publicly alleges Alibaba extracted its AI capabilities, escalating US–China frontier-model tensions. Source

Databricks' Former AI Chief Targets 1,000× Power Reduction. A new venture claims a path to slashing AI compute/power costs by three orders of magnitude. Source

General Intuition Raises $2.3B for Game-Trained Agents. Massive round bets that video-game environments can train real-world AI agents. Source

BenchPress: LLM Benchmark Matrices are Near Rank-2. Microsoft AI Frontiers shows a handful of probes can predict full benchmark profiles before running them. Source

Linux Foundation Proposes DNS as AI Agent Identity Layer. A new initiative would repurpose DNS infrastructure to give AI agents verifiable identities. Source

2,000 Adversarial Emails vs. an AI Assistant. Large-scale red-team test of an AI email agent shows surprisingly low prompt-injection success, with methodological caveats. Source

Apple Reportedly Skips M6 for AI-Focused M7 Line. M7 Pro/Max/Ultra chips would dramatically boost memory bandwidth to enable on-device LLM inference. Source