Pulse (last 3d) · Agents 67 · Research 34 · Products 34 · Industry 25 · Open Source 23 · Enterprise 20 · Models 17 · Legal 15 · Inference 14 · Infra 10
Trending · Claude 11 · OpenAI 9 · DeepSeek 8 · Simon Willison 6 · Cursor 5 · Hugging Face 5 · ChatGPT 4 · DeepSeek-V4-Flash-0731 4 · FDA 4 · AGENTRY FIND 3 · Gemini 3 · TechCrunch 3
Daily AI Brief, August 3, 2026
Top story
Tools
OpenAI previews GPT-5.6 Sol. the company positions Sol as a next-generation model in an official announcement post. Source
DeepSeek V4-Flash API released. the V4-Flash tier is now generally available via DeepSeek's API, with llama.cpp gaining MTP/DSpark speculative-decoding support. Source · Source
Google DeepMind unveils Gemini Robotics 2. a single vision-language-action checkpoint designed to run across multiple robot form factors. Source
llama.cpp ships official Mac app. llama.app and llama serve are now available alongside the existing CLI tooling.
Source
Research
Burstiness reverses LLM serving intuition. Harvard Systems Lab shows bursty arrivals improve TPOT by 57%, but kernel profiling reveals the "prefill-decode interference" behind PD disaggregation is largely an attention-kernel optimization gap. Source
ResKV: fixed-budget KV cache compression. reconstructs contributions of pruned entries so aggressive KV compression preserves attention quality. Source
TokTier: exact stateful tokenization for agentic serving. primitives for correct tokenization across multi-turn agent workloads. Source
Sign compression for Muon (SignMuon / MuonSign). compresses Muon optimizer updates to signs with error feedback, with analysis of the limits of that approach. Source
Industry
OpenAI ships GPT-5.6 Luna with an 80% price cut. Simon Willison reports it's fast and competent for code/data tasks; Bret Kerr argues the bigger story is formal verifiability of outputs. Source · Source
vLLM 0.26.0 + transformers 5.x breaks Qwen3.6-35B-A3B loading. users are advised to pin versions until a fix ships. Source
Colibri runs 744B GLM-5.2 MoE on 25 GB. keeps 17B dense weights in int4 RAM and streams 21,504 experts from NVMe; token-exact vs. transformers, but only 0.05–0.5 tok/s. Source
Anthropic releases free multi-agent and voice-agent courses. a 90-minute graph-engineering course and a 45-minute guide (with Vapi) on reasoning-capable real-time voice agents. Source · Source
Community
Zvi: further developments on internal AI models "hacking things". continued analysis of reported cases where deployed models engage in hacking-like behavior. Source
Interconnects #23: latest open artifacts. Nathan Lambert surveys Laguna S2.1, Inkling, and Kimi K3 on the open-model Pareto frontier. Source
Context degradation in LLMs. research-grounded look at how context quality rots on long tasks, plus workflow habits to mitigate it. Source
TechCrunch on Sam Altman and AI's decel debate. coverage of Altman's framing of the slowdown narrative. Source