Pulse (last 3d) · Research 111 · Agents 78 · Products 64 · Industry 47 · Models 39 · Open Source 30 · Media 28 · Releases 19 · Infra 19 · Legal 18

Trending · OpenAI 20 · Claude Code 13 · Claude 10 · Google 10 · ChatGPT 9 · FDA 9 · Anthropic 7 · Codex 6 · Cursor 6 · TechCrunch 6 · Meta 5 · Nvidia 5


AI agents are now being judged by what they actually change, such as the database rows they write, the money they spend and the full sequence of steps they take, rather than by what they report back. More of the improvement is coming from the code and tooling around a fixed model than from the model's weights. Headline claims, from Navier-Stokes to benchmark leaderboards, are coming out faster than anyone can check them, so validated reference artifacts such as the FDA's Model Master File are worth more now. openai.com | arxiv.org | fda.gov

  • The FDA already has a Drug Master File-style route for filing a validated model once as a regulatory reference. fda.gov
  • Argo-Bench scores data agents on what their actions do to a simulated 7.5B-row enterprise warehouse. arxiv.org
  • OpenAI says long-horizon agent safety means monitoring the whole sequence of actions, not each action alone. openai.com
  • A CMU paper uses reinforcement learning to rewrite an agent's harness code while the model's weights stay frozen. x.com
  • OpenAI claims a 10,000-agent swarm found a Navier-Stokes blowup, and two rival researchers allege priority violations. openai.com
  • Aleph Alpha's Kolibri is a 78.1B Apache 2.0 MoE trained from scratch on 24T tokens on European infrastructure. aleph-alpha.com
  • Simon Willison argues that every metered service agents can reach should have a hard budget cap by default. simonwillison.net
  • OpenAI safety resignation - a safety employee quit, saying the culture is broken and the time for trial and error is over. theverge.com | techcrunch.com | x.com | theatlantic.com
  • Claude Code mods, 48 hours in - one directory listed 26 mods within a day; a Doom mod drew 7,023 likes. rundatarun.io
  • Agents don't need memory, they need documentation - argues for enforceable docs and lint errors that explain the fix. liao.gg
  • Firetower - open-source Rust tool for running Claude Code and Codex on your own infrastructure. producthunt.com
  • Pi pod - runs the pi coding agent in isolated sandboxes on your own server. pipod.dev
  • OpenCompanion - MIT-licensed dashboard that flags agent sessions waiting for permission and answers them from a phone. producthunt.com
  • Monospace from Directus - gives agents live read-write access to legacy databases under existing permissions, without copying data. producthunt.com
  • UI Arc - MIT-licensed React motion components with llms.txt and an MCP server for coding agents. producthunt.com
  • Cue by Manus - personal agents, each with its own email, phone number, wallet and computer. producthunt.com
  • Would AI arrest Gandhi? - tests 12 models on real historical decisions. chrystianschutz.com
  • The Principles of Diffusion Models - free monograph by Lai et al. that balances rigor and intuition. reddit.com
  • Amazon drops data center NDAs - a response to community backlash over its data center deals. techcrunch.com
  • AI agents in your text messages - a roundup of agents that work inside messaging apps. techcrunch.com
  • $8M AI music streaming scam - prosecutors are asking for nearly four years in prison. lawcommentary.com
  • Capcom plans for AI co-developed games - says it is preparing for a future of making games together with AI. theverge.com

Each of these replaces the agent's own report with an outside check: the database state, the full action sequence, or a spending limit.

Scoring the consequences catches agents that report success and leave the data wrong TextQL's Argo-Bench places data agents in a simulated delivery company's ERP warehouse (235 tables, 7.5B rows) and scores 210 tasks on actions such as fraud bans and incentive budgets (preprint, unreplicated). A Microsoft post on Hugging Face describes the same failure in production, where agents claim to be done while the database disagrees. arxiv.org | huggingface.co

OpenAI moves agent safety from single actions to whole trajectories OpenAI's long-horizon safety post covers trajectory monitoring, sandbox persistence, and actions that are harmless one at a time but harmful in sequence. Per-call guardrails, which most clinical tool-use deployments rely on, cannot catch that kind of failure, so the monitor has to score the whole run. openai.com

One loose night of agent use cost more than a month of disciplined use Wagtail spent a month coding on GLM 5.3 Flash: the disciplined 1B tokens cost $68 and 4kWh, while one night of a vibe-coded MCP prototype burned 450M tokens and $150. Simon Willison draws the same conclusion and argues that every metered service agents can reach should have a hard budget cap by default. wagtail.org | simonwillison.net

Each of these improves agents by changing the code, recipes or documentation around a fixed model.

Reinforcement learning now trains one model to rewrite another agent's harness A CMU paper uses reinforcement learning to train a proposer model that edits an agent's harness code, while the solver model's weights stay unchanged (preprint, unreplicated). This turns harness self-improvement from hand tuning into a learned loop, which helps wherever the base model is locked by a vendor contract or a validation package. x.com | arxiv.org

Trillium Labs makes post-training and self-improvement research open Nathan Lambert launched Trillium Labs, a non-profit that will publish open post-training recipes and later open infrastructure for studying recursive self-improvement, reward hacking and multi-agent systems. Learned harness editing carries the same reward-hacking risk, and an open baseline gives outside teams something to test against. x.com

Every headline result in this section was self-reported or disputed when it came out.

OpenAI's Navier-Stokes claim came with a priority dispute OpenAI says a 10,000-agent swarm found a blowup in Navier-Stokes, one of the Millennium Prize problems (vendor claim). Alpöge and Buckmaster released a forced-Euler result and allege priority violations, so credit for discoveries made by agent swarms is being fought over in public before peer review. openai.com

Gemini 4 Argon tops the Vals Index Vals AI reports that Gemini 4 Argon is in the top five on 20 of 22 benchmarks and first on the Vals Index, with a large jump on its RSI index (single source). Vals is an independent evaluator, so this carries more weight than a lab's own model card. x.com

Laya now claims to beat Jev decomposition on every axis Laya's updated head-to-head claims it beats Jev on accuracy and calibration and is 7.8x faster (vendor benchmark). Calibration is the result that matters for setting confidence thresholds in a clinical workflow, and all the figures come from Laya's own site. laya.convaiinnovations.com

Regulators and sovereignty-minded buyers are getting model artifacts they can cite, audit and host themselves.

The FDA already has a master-file route for validated models An FDA document describes a Model Master File, which works like a Drug Master File: a validated model is submitted once as a regulatory reference object, and QMIN is named as the path beyond generic drugs. A precision-medicine model could then be validated once and cited by several sponsors instead of being validated again in each submission. fda.gov

Kolibri is a technically serious sovereign open-weight model Aleph Alpha released Kolibri on October 3: a 78.1B Apache 2.0 MoE with 3.46B active parameters, 384 experts and 262K context, trained from scratch on 24T tokens on German and Finnish infrastructure. It was also trained to answer "I don't know," which is worth testing on clinical questions, where confident errors cost the most. aleph-alpha.com | tej.as

22 items were left out: a running thread with no new development, a product listing whose description did not match its link, demo videos, open Reddit questions, hype threads, opinion pieces with no claim, digests, and off-topic posts.