Pulse (last 3d) · Research 127 · Agents 76 · Products 61 · Models 40 · Open Source 35 · Industry 28 · Legal 21 · Media 20 · Releases 17 · Inference 16

Trending · OpenAI 26 · Anthropic 11 · Claude Code 7 · Codex 7 · ChatGPT 6 · Hugging Face 6 · Claude 5 · Google DeepMind 5 · Meta 5 · AlphaGenome 4 · AlphaGenome Atlas 4 · Apple 4


OpenAI Launches Agents API in Public Beta. The full Codex agent harness, durable sessions, auto-compaction, tool search, subagents, and first-class MCP, is now a managed, OpenAI-operated API for building and orchestrating cloud agents. Source

Procedural Graphs: Self-Evolving Workflow Graphs for LLM Agents. Google researchers show that self-evolving graphs of (procedure, relation, procedure) triplets beat agent-memory baselines in 21 of 24 settings, and that full-context injection actively hurts performance. Source

An Open Recipe for IMO Gold. NVIDIA's Nemotron team open-sources a complete training recipe for achieving IMO gold-level olympiad math performance. Source

NCP-ArchPreview: Next Concept Prediction. Technical report proposing architectures that shift language modeling from discrete tokens to latent concept sequences. Source

MoE Models Overfit More to Repeated Data. Mixtures-of-experts overfit repeated training data more severely than dense equivalents, a timely caveat for MoE-heavy releases like today's DeepSeek launch. Source

DeepSeek V4.1 Flash. MIT-licensed 552B-class multimodal model with 1M-token context, native visual understanding, and OpenAI/Anthropic-compatible APIs, launched at a lower price than the model it claims to comprehensively surpass. Source

GPT-Live-1 in the API. OpenAI ships GPT-Live-1 for building more natural, real-time voice experiences. Source

Google Pics. Workspace-native AI image creation and editing built on Nano Banana, with object-level edits, in-image text editing/translation, and inline use in Docs and Slides. Source

Suno v6. New generation of music models in three tiers, adding section-level editing, song mashups, and text/audio/image/video starting points. Source

Anthropic's September Threat Intelligence Report. Anthropic's most detailed misuse report to date documents disrupted operations, including claims that Moonshot forwarded Kimi user requests to Claude, DeepSeek relayed exchanges, and MiniMax built a proxy network via a shell company. Source

Devin-Built GPU Sieve Factors RSA-260. Cognition's agent autonomously built a GPU lattice siever overnight and set a new RSA Factoring Challenge record at roughly 10x lower cost than the 2020 CPU-era record. Source

Anthropic Models Claude's Labor-Market Impact. New economic scenarios include an extreme case with 17.9% cognitive unemployment and labor's share of GDP falling from 60% to 45%. Source

ChatGPT for Financial Services. OpenAI pairs built-in financial data with GPT-6 Astra reasoning in a tailored ChatGPT Work offering. Source

GPT-6 Astra vs GPT-5.6 Sol on 50 Real PRs. Entelligence's benchmark finds Sol locating more confirmed bugs while Astra shows better precision and latency, with methodology open for community feedback. Source

Can Researchers Trust OpenAI with Unpublished Math?. Another case of shared ideas seemingly surfacing in later OpenAI publications without attribution sparks debate. Source

SWE-2's Terminal Bench Gap Flagged. Cognition's new coding model draws scrutiny over a 92.8% vs 27.3% gap between old and new Terminal Bench versions, cited as evidence of benchmark overfitting. Source

AI #185: Preference Cascade. Zvi Mowshowitz's weekly roundup covers the week's major AI developments with commentary. Source