OpenAI launches GPT-Live voice models inside ChatGPT - More natural, interruption-resistant conversations arrive in ChatGPT, with GPT-Live able to offload reasoning to a frontier text model in the background. Source

MCP attacks beat SOTA guardrails more than half the time - Research shows LLM agents with MCP tool access can be exploited via CVE-derived tool-call sequences that evade text-only safety guardrails across all tested model sizes. Source

LingBot-Video: sparse-MoE video diffusion (13B total, 1.4B active) - Open-sources a 13B sparse-MoE video diffusion transformer with RL post-training and an action-conditioned world-model mode for robotics. Source

Anthropic publishes GRAM: weight-level dangerous knowledge removal - A technique to surgically excise dangerous knowledge from model weights, advancing model-level unlearning research. Source

4-bit GLM-5.2 (753B MoE) on 4× DGX Spark hits 70.8% Terminal-Bench 2.1 - Quantized 4-bit run on 4× DGX Spark versus 81.0% full precision, with detailed rig, pipeline, and failure notes. Source

Scaling MoE Video Pretraining for Embodied Intelligence - Scales mixture-of-experts video pretraining specifically for embodied AI applications. Source

Is One Layer Enough? Single Transformer Layer Matches Full-Parameter RL - Shows that training a single transformer layer can match full-parameter fine-tuning in RL settings. Source

RL Post-Training Builds Compositional Reasoning Strategies - Analyzes how RL post-training induces compositional reasoning strategies in LLMs. Source

How Data Shapes RoPE Frequency Usage - Analyzes how data properties shape RoPE frequency utilization and derives conditions for length generalization. Source

xAI releases Grok 4.5 - Marketed by Elon Musk as an 'Opus-class' frontier model, trained in collaboration with Cursor for coding and agent tasks with high speed and low cost. Source

LingBot-World 2.0: one-hour interactive AI world model - Open-sources an interactive AI world model capable of continuous one-hour video generation, including model weights, code, and SGLang integration. Source

Mistral AI releases Robostral Navigate (8B) - A vision-language model for robot navigation that achieves state-of-the-art performance on the R2R-CE benchmark. Source

LangChain × NVIDIA launch NemoClaw Deep Agents Blueprint - Open, customizable enterprise AI agent architecture that cuts inference costs 10x using Nemotron 3 Ultra. Source

Microsoft releases Flint, a visualization DSL for AI agents - A JSON-based visualization language designed for LLM agents to deterministically generate charts. Source

ZML AI open-sources LLMD inference server - Streams weights directly from Hugging Face without local storage, supporting NVIDIA, AMD, and TPU hardware. Source

audio.cpp: 4 ASR models in native C++/GGML with streaming - Adds streaming support and Nemotron 3.5 ASR, Higgs Audio STT, VibeVoice ASR, and Hviske ASR models, 1.07x–2.41x faster than Python. Source

OpenAI wins AtCoder Heuristic Contest with 50 billion points - Beats the top human programmer by a wide margin, improving on its second-place 2025 finish. Source

China to allow top AI firms to buy 200,000 NVIDIA H200 GPUs - Approvals target leading domestic developers including DeepSeek, ByteDance, and Alibaba. Source

Prime Intellect raises $130M Series A - Backed by Radical Ventures, NVIDIA, Intel, and Dell to build open AI infrastructure for enterprise training and deployment. Source

Meta AI now allows deepfakes from other users' Instagram photos without consent - Raises significant privacy and consent concerns across the platform. Source

OpenAI says SWE-Bench Pro has hit a 70% noise ceiling - The benchmark's saturation is driving calls for proprietary evaluations. Source

Databricks: agent harness choice can cut task costs by 2x - Benchmarking on Databricks' multi-million line monorepo places open-source GLM 5.2 on the Pareto frontier. Source

China's MiniMax plans 2.7-trillion parameter open-source model - Codenamed M3 Pro, reportedly slated for Q3 release. Source

OpenAI announces powerful new model launches publicly July 9 - A separate frontier model reveal timed for the same day as GPT-Live's rollout. Source

Teaser for an inference-optimized model release - Upcoming model religiously optimized end-to-end for its own inference, with benchmark comparisons against frontier models. Source

Why this CEO thinks video games beat the internet as training data - Argues video games provide higher-quality training data than the open web. Source

This startup thinks robotics is about to have its ChatGPT moment - Commentary on robotics approaching a transformative inflection point similar to ChatGPT's launch. Source