Top story
DeepSeek raises $7.4B at a $60B valuation, with founder Liang Wenfeng personally investing $3B, a landmark round that cements his grip on the company as China's AI rivalry with the US intensifies. Source
Tools
Introducing Claude Tag. Anthropic launches Claude Tag, a Slack-based agentic interface where Claude joins channels, retains context, and executes delegated tasks. Source
Sakana Fugu. Sakana AI Labs releases Fugu, a full multi-agent orchestration system accessible through a single model API. Source
Qwen-AgentWorld-35B-A3B. Qwen ships a 35B MoE (3B active) trained as a language world model that simulates MCP, terminal, SWE, Android, web, and OS environments. Source
Unlimited-OCR (3.3B). Multilingual OCR model released on ModelScope under MIT for one-shot parsing of images, multi-page documents, and PDFs. Source
Research
Qwen-AgentWorld: Language World Models for General Agents. Proposes language-based world models for training and evaluating general-purpose agents. Source
DeepSWE Benchmark. Contamination-free, multi-language benchmark evaluating frontier coding agents on real pull requests. Source
IONS: External Reasoning Graph. A reasoning graph that stores claims, evidence, and reasoning paths outside the LLM to support grounded inference. Source
OpenThoughts-Agent. Data recipes for training agentic LLM models. Source
Industry
Superhuman to acquire GPTZero. The AI-text-detection tool built by Edward Tian is being acquired by email client Superhuman. Source
Google invests $75M in A24. Funding aimed at co-developing AI-powered filmmaking tools with the indie studio. Source
Chinese AI chip vendors mapping. Seven Chinese companies are already shipping H100/H200-class accelerators, most IPO'd within the last six months. Source
Chip Security Act gains industry support. Half a dozen companies back legislation requiring location-tracking mechanisms for America's most advanced AI chips. Source
Community
Russia's Social Design Agency leaks. Leaked files detail a state-affiliated effort to build fake reference platforms that contaminate AI training data and search indices. Source
1.7B fine-tune matches frontier on support. A fine-tuned 1.7B model with a 4% frontier-model fallback matched frontier performance on customer support within noise. Source
Monthly Roundup #43: June 2026. Zvi Mowshowitz's monthly synthesis of the most significant AI news and analysis. Source
Medical scribing LLM benchmark. Across 8 LLMs, hallucinations were rare but omissions remained a key failure mode. Source