Pulse (last 3d) · Research 64 · Agents 52 · Products 39 · Industry 23 · Models 19 · Open Source 14 · Inference 13 · Releases 12 · Legal 12 · Enterprise 12

Trending · Anthropic 17 · Claude Code 10 · Claude 8 · Ox Alpha 8 · OpenAI 6 · Codex 5 · Google 5 · Simon Willison 5 · TechCrunch 5 · MCP 4 · OpenRouter 4 · ChatGPT 3


GPT-5.6 lands in Kiro with better price-performance for developers. OpenAI's latest model iteration targets improved cost-efficiency for coding workflows. Source

Memory poisoning kills utility without any payload. A single non-adaptive pass of plainly-worded false statements at 1.2% of a LongMemEval corpus drops agent memory accuracy from 0.850 to 0.300, and the additive provenance fix has no usable setting. Source

SWE Refactor Bench. A new benchmark for evaluating coding agents on long-horizon, whole-repository stack migrations. Source

Agent coordination topology shapes LLM-backend traffic. First empirical characterization (SIGCOMM NAIC '26) showing multi-agent topologies produce bimodal, log-normal, decisively non-Poisson arrival processes. Source

XPerf: trace-replay benchmarking for agentic serving. UIUC + IBM's tool records an agent's full execution graph and replays it deterministically against any serving stack, making performance claims reproducible. Source

Agent Lightning v1.0. Microsoft Research's ~3,500-line disaggregated RL framework trains agents through their real production harnesses with zero harness-code changes. Source

DevRecall. A local-first, MIT-licensed searchable index of your entire work history (Git, Slack, Jira, Linear, Confluence) that doubles as an MCP server for Claude Code and Cursor. Source

Postern. A self-hosted gateway connecting Money, Health, Calendar, Mail, and Home data to any AI agent with per-sector, revocable, audited access. Source

Inferbench. An Apache 2.0 CLI tool that benchmarks local LLM inference engines on your own hardware and recommends the fastest configuration. Source

Hugging Face reportedly in talks to be acquired for $13B. A deal at this price would significantly reshape the open AI ecosystem. Source

Meta to launch consumer AI agent "Hatch" within weeks. The Information reports Hatch may cost $199.99/month and operate across DoorDash, Etsy, Reddit, Yelp, and Outlook, with the next model "Watermelon" targeted for October. Source

Thomson Reuters launches its own frontier model. The company is leveraging its world-class legal and tax data assets to build a domain-specific frontier model. Source

Mistral x HUMAIN. Mistral announces a strategic collaboration to advance sovereign AI in Saudi Arabia and the Middle East. Source

Anthropic's IPO filing to name public opposition to AI as a risk factor. Reportedly the first major AI lab IPO to formally list public opposition to AI and data centers as a written risk. Source

Chinese open models now dominate research usage. Nathan Lambert used Codex to parse 500K arXiv AI/ML papers, finding Qwen and other Chinese open models overtaking American ones in research citations. Source

A stealth AI model called "Ox Alpha" is winning over developers. A mysterious free model with an unknown creator is gaining significant developer mindshare. Source

I audited the sources my AI fact-checker was citing. An audit found roughly 1 in 18 LLM-cited sources didn't exist, exposing a common citation-hallucination failure mode. Source