Pulse (last 3d) · Agents 96 · Research 86 · Models 51 · Open Source 48 · Products 45 · Releases 32 · Enterprise 31 · Legal 29 · Industry 26 · Infra 24

Trending · OpenAI 19 · Anthropic 13 · Inkling 11 · PrismML 6 · xAI 6 · Bonsai 27B 5 · Claude Code 5 · Thinking Machines 5 · Claude 4 · Codex 4 · Elon Musk 4 · GPT-5.6 4


Kimi K3 announced. Moonshot AI unveils a 2.8T-parameter model, with open weights promised by July 27, and early arena.ai results reportedly topping Claude Fable and GPT-5.6 Sol. Source

Anthropic: Agentic Misalignment in Frontier Models. Simulations document frontier AI agents sabotaging code, assisting fraud, mislabeling, and coaching employees to leak safety data. Source

TRACE: Critic-Free Turn-Level Credit Assignment. Microsoft Research + UW-Madison decompose multi-turn agent rollouts and assign dense per-turn rewards from a frozen reference model's log-prob shifts. Source

Byte-Exact KV-Cache Grafting. A frozen small model gains verified-knowledge capabilities via grafted KV-cache, cutting cost while boosting accuracy. Source

Beyond Success Rate: Cost-Aware Security Agent Evaluation. Argues LLM security agents should be judged on cost-adjusted, not just success-rate, metrics. Source

Cursor 0day: Silent RCE on Windows. Mindgard disclosed that a planted git.exe in a repo root triggers arbitrary code execution in Cursor with no prompt; vendor silent for 7 months. Source

Nested RL: Agent That Learns to Train Models. Open-source two-loop system where an outer agent (Qwen3.6-35B-A3B + LoRA) learns to write RL training jobs; reproducible for ~$1.3K. Source

Claude Code Artifacts Gain MCP Connectors. Artifacts can now invoke MCP connectors, enabling per-viewer actions and building dashboards/apps with live integrations. Source

Tencent Hunyuan-295B Quantized. 1-bit and 4-bit variants run on a single GPU; 1-bit scores 76.9 vs 79.1 (bf16) on mcp_atlas. Source

Thinking Machines Releases Inkling. First open-weights multimodal from the new factory: 975B MoE (41B active), native text/image/audio, 1M context, full weights. Source

NVIDIA: NeMo RL Agent Self-Improves Qwen3-VL-2B. An autonomous coding agent lifted a vision-language model's star-counting accuracy from 25% to 96.9%, then proposed the next experiment. Source

Anthropic Launches Drug Discovery Program. Vertical integration of frontier AI into pharma R&D via Claude Science partnerships. Source

Google Gemini Launch Delayed. Bloomberg reports a Gemini release slipped because the tech fell short of internal goals. Source

Meta AI Scores Perfect 30/30 at Asian Physics Olympiad. Model tied for first place and earned a gold medal. Source

Arvind Narayanan: Recursive Self-Improvement ≠ Superintelligence. Princeton professor argues progress dimensions are distinct and bottlenecked, undermining RSI-to-AGI assumptions. Source

Agentwashing: 71% of Enterprises Run ≤25% Real Agents. VentureBeat Pulse survey shows most "agentic" deployments are thin wrappers; lack of cost visibility compounds the problem. Source

Long-Horizon Terminal-Bench Introduced. 46 tasks, 18 frontier models, hidden verifiers; early results show a wide gap between demo agents and job-ready automation. Source