Pulse (last 3d) · Agents 67 · Research 34 · Products 34 · Industry 25 · Open Source 23 · Enterprise 20 · Models 17 · Legal 15 · Inference 14 · Infra 10

Trending · Claude 11 · OpenAI 9 · DeepSeek 8 · Simon Willison 6 · Cursor 5 · Hugging Face 5 · ChatGPT 4 · DeepSeek-V4-Flash-0731 4 · FDA 4 · AGENTRY FIND 3 · Gemini 3 · TechCrunch 3


Alibaba/Qwen launches Qwen3.8-Max. a 2.4T-parameter MoE flagship with 1M context and native vision, marking the first open-weight release of a Max-class model alongside Qwen3.8-27B (dropping next week). Source · Source

OpenAI previews GPT-5.6 Sol. the company positions Sol as a next-generation model in an official announcement post. Source

DeepSeek V4-Flash API released. the V4-Flash tier is now generally available via DeepSeek's API, with llama.cpp gaining MTP/DSpark speculative-decoding support. Source · Source

Google DeepMind unveils Gemini Robotics 2. a single vision-language-action checkpoint designed to run across multiple robot form factors. Source

llama.cpp ships official Mac app. llama.app and llama serve are now available alongside the existing CLI tooling. Source

Burstiness reverses LLM serving intuition. Harvard Systems Lab shows bursty arrivals improve TPOT by 57%, but kernel profiling reveals the "prefill-decode interference" behind PD disaggregation is largely an attention-kernel optimization gap. Source

ResKV: fixed-budget KV cache compression. reconstructs contributions of pruned entries so aggressive KV compression preserves attention quality. Source

TokTier: exact stateful tokenization for agentic serving. primitives for correct tokenization across multi-turn agent workloads. Source

Sign compression for Muon (SignMuon / MuonSign). compresses Muon optimizer updates to signs with error feedback, with analysis of the limits of that approach. Source

OpenAI ships GPT-5.6 Luna with an 80% price cut. Simon Willison reports it's fast and competent for code/data tasks; Bret Kerr argues the bigger story is formal verifiability of outputs. Source · Source

vLLM 0.26.0 + transformers 5.x breaks Qwen3.6-35B-A3B loading. users are advised to pin versions until a fix ships. Source

Colibri runs 744B GLM-5.2 MoE on 25 GB. keeps 17B dense weights in int4 RAM and streams 21,504 experts from NVMe; token-exact vs. transformers, but only 0.05–0.5 tok/s. Source

Anthropic releases free multi-agent and voice-agent courses. a 90-minute graph-engineering course and a 45-minute guide (with Vapi) on reasoning-capable real-time voice agents. Source · Source

Zvi: further developments on internal AI models "hacking things". continued analysis of reported cases where deployed models engage in hacking-like behavior. Source

Interconnects #23: latest open artifacts. Nathan Lambert surveys Laguna S2.1, Inkling, and Kimi K3 on the open-model Pareto frontier. Source

Context degradation in LLMs. research-grounded look at how context quality rots on long tasks, plus workflow habits to mitigate it. Source

TechCrunch on Sam Altman and AI's decel debate. coverage of Altman's framing of the slowdown narrative. Source