Pulse (last 3d) · Agents 74 · Research 71 · Products 62 · Models 61 · Industry 36 · Enterprise 27 · Open Source 23 · Legal 20 · Inference 15 · Media 15

Trending · OpenAI 14 · Claude 13 · Claude Code 11 · ChatGPT 10 · Codex 10 · Anthropic 9 · Cursor 8 · Gemini 5 · GitHub 5 · Simon Willison 5 · Cloudflare 4 · FDA 4


AMD acquires Taalas to boost inference performance by etching models in silicon. AMD acquires Taalas to push hardwired model inference, etching models directly into silicon for major performance gains. The Register

Prime Agent. PrimeIntellect releases Prime Agent, an open-source agent framework built on pi with an open license. GitHub

bb. Newly-notable bb tool released. GitHub

Shieldstral. Mistral's 3B open-weight multimodal guardrail that evaluates text, images, or both from a single token output and runs locally on a 16GB GPU. Product Hunt

OpenAI updates GPT‑5.6 Sol and expands Luna access. OpenAI improves GPT-5.6 Sol with a reasoning slider and makes GPT-5.6 Luna available to free users. OpenAI

Mirendil inks $100M+ Google Cloud deal. Mirendil signs a $100M+ Google Cloud deal to scale self-improving AI infrastructure. TechCrunch

Anthropic improves Fable 5 biology safeguards. Anthropic reduces biology-related fallbacks in Fable 5 by ~85%. Anthropic

Suno to watermark AI-generated songs. Suno will watermark AI-generated songs amid ongoing lawsuits. TechCrunch

Qwen 3.8 Max tops Artificial Analysis agentic index. Qwen 3.8 Max reportedly takes the top spot ahead of Opus 5 on Artificial Analysis's agentic benchmark. Reddit/LocalLLaMA

vLLM ported to C++20. Developer ports vLLM's serving stack to C++20, producing a 66 MiB binary with token-identical output verified against vLLM. Reddit/LocalLLaMA

AV-AIVAT: 74x cheaper agent evaluation. AV-AIVAT cuts agent evaluation cost roughly 74x via anytime-valid confidence sequences in imperfect-information games. arXiv

Invisible Shortcuts in Vision Encoders. Analysis showing vision encoders can infer the source camera from image artifacts alone. Hugging Face

Zvi Mowshowitz, AI #180: No Longer In Charge. Weekly roundup covering the latest in AI developments. Substack

Inside vLLM: Anatomy of a High-Throughput Inference System. Technical retrospective arguing that web-server/GPU separation, continuous batching, and KV-caching mattered more than paged attention. aleksagordic.com

KV cache quantization benchmarks. 413-pair benchmarks on Qwen 3.6 27B and Gemma 4 31B show KVarN 6-bit beats q8_0, with tail-token precision dominating. Reddit/LocalLLaMA