Pulse (last 3d) · Research 139 · Agents 85 · Products 67 · Industry 46 · Models 43 · Enterprise 28 · Open Source 24 · Releases 23 · Legal 17 · Policy 14
Trending · OpenAI 22 · Codex 13 · Anthropic 12 · Claude 11 · FDA 11 · Claude Code 9 · Mistral 8 · Mistral Large 4 7 · ChatGPT 6 · Cursor 6 · Google 6 · Nvidia 6
Top story
Generating AI output is now cheap, but checking it is still expensive. OpenAI withdrew three of its math results while outside researchers improved others in Lean, and a new study shows LLM judges stop discriminating when they can see the system's own answer label. The week's launches pushed costs down further, so the next eval budget should go to blinded judges, repeated-run grading and tests on real tool calls. arxiv.org | twitter.com
The Shortlist
- OpenAI's annualized revenue is about $50B, roughly $20B below what it had signaled, and AI infrastructure stocks fell on the news. cnbc.com
- LLM judges that can see the evaluated system's label lose most of their ability to discriminate, and a label-free setup restores it. arxiv.org
- OpenAI withdrew three AI-generated math results while others improved its proofs with Lean verification. twitter.com
- Claude Haiku 5.5 targets high-volume classification, summarization and coding subagents. producthunt.com
- Six funders are pooling $1.8B to openly measure how cells respond to drugs and genes. rundatarun.io | biohub.org
- Samsung released code for 0.1-bit-per-weight LLM quantization that keeps the original architecture at inference. github.com
- Anthropic's OSS Scanner scans opted-in open-source projects for free and returns a proof of concept and a fix. theverge.com
Also Noted
- ThinkingBox-Bench - 507 stateful workflows, each run 20 times and graded on the final database state. reddit.com
- METR time horizons - a statistical look at how the METR plot's estimates are made and whether they hold. arxiv.org
- Probes catch sabotage - activation probes detect deception that models never put into words. arxiv.org
- Epistemic humility under knowledge conflict - agents stay accurate but are poorly calibrated when their sources conflict. huggingface.co
- Harmful-refusal audit - a psychometric audit questions what a safety benchmark's refusal score actually measures. arxiv.org
- HoneyBench - nine tasks designed to draw out reward hacking in frontier models. x.com
- SciConBench dashboard - Princeton's live benchmark tracks how well models synthesize health research findings, with leakage checks. x.com | x.com
- Looped tabular foundation model - a lightweight model built for bulk transcriptomics data. x.com
- On-policy distillation - it transfers skills to the student but not new knowledge. huggingface.co
- MiMo-V2.6 - Xiaomi scales reinforcement learning toward self-improving models. huggingface.co
- Mistral flagship - billed as the strongest open model outside China; Reflection AI released its first open model. x.com | x.com
- Humanity's Sixth Sense - Scale's visual-intuition benchmark: humans 93.1%, best model GPT-6-astra 53.6%. x.com
- Wikimedia vs OpenAI agents - OpenAI-linked agents edited Wikimedia's private wikis and overloaded its servers (single source). techspot.com
- OpenAI on agent liability - OpenAI's deputy general counsel told an ABA conference that labs shouldn't be liable for agent hacking. reddit.com
- Agent security incidents - a survey of incidents at three labs argues for proactive assurance over reactive containment. arxiv.org
- "False front" operations - OpenAI reports disrupting AI-enabled front operations. openai.com
- USA Today sues OpenAI - the suit covers training on 19 of its publications. runtimewire.com | theverge.com
- Enigma messages cracked - GPT-6 Astra and Opus 5 decoded two messages that had resisted researchers for decades (single source). notebookcheck.net
- Semwright - an open-source Rust runtime that lets agents drive desktop and professional software. producthunt.com
- LegalOn - halved its Codex costs without slowing development (vendor case study). openai.com
- Open-source Adobe clones - a developer is building open-source clones of Adobe products and says "software is over". arstechnica.com
- OOD degradation forecast - predicts out-of-distribution performance loss from training dynamics before deployment. arxiv.org
- Alignment generalization - value representations used to predict how alignment generalizes. arxiv.org
- MARGIN - runtime confidence calibration for coordinating multi-agent systems. huggingface.co
- Agent ecology - argues that collaboration among agents creates a population threshold for takeoff. arxiv.org
- Brain-activity image reconstruction - a model reconstructs what a person is viewing from brain activity. thebrighterside.news
- Kilometer-scale weather - aligns global and regional models for high-resolution forecasts. arxiv.org
- ReSPO - a reshaped sequence-level policy optimization that fixes gradient starvation in off-policy RL. huggingface.co
- Sequential structure study - a controlled test of whether LLMs grasp sequence structure when inferring and generating. huggingface.co
- MC-Sparse - closes the quality gap between dense and sparse attention in diffusion transformers. huggingface.co
- Text-to-image post-training - combines preference rewards with rubric rewards. huggingface.co
- Moonworks Lunara - an art-focused diffusion mixture transformer with under 10B active parameters. reddit.com
- vidu-q4 - a new video generation model from Vidu. x.com
- Recurrent ViT - one transformer block reused at several depths with depth-programmed experts. arxiv.org
- WOVEN - builds visual world modeling into multimodal LLMs. arxiv.org
- LeWAM - a JEPA world-action model that controls through diffusion-steered MPC. arxiv.org
- VioLA - generalist humanoid control policies learned from human motion data. arxiv.org
- Balanced data diet - targets the exploration bottleneck in mega-scale robot RL. arxiv.org
- Embodied Turing Machines - stateful code as a path to robot self-improvement. huggingface.co
- RoboRSI - stable, reusable robot self-evolution in real-world environments. arxiv.org
- Multi-agent egocentric world model - captures fine-grained embodied interaction between agents. huggingface.co
- Unified Bellman operator - a single operator for safety-critical RL. arxiv.org
- FAITH - feasibility-aware, safety-filtered RL for high-dimensional systems. arxiv.org
- ViSkill - VLM agents reinforced with evolving visual skills. arxiv.org
- FastBench - tests whether streaming VLMs keep up with fast-moving real-world video. arxiv.org
- SpaceCast-Bench - tests predictive spatial reasoning in VLMs. arxiv.org
- OmniCapBench - a benchmark for fine-grained audio-visual captioning. huggingface.co
- BrickBench - a benchmark for agentic brick design. huggingface.co | arxiv.org
- GeoReform - refines its formalization step by step to solve multimodal geometry problems. arxiv.org
- SpaceFlow - 3D generation with local control. arxiv.org
- OuroWorld - turns 3D worlds into endlessly looping cinemagraphs. huggingface.co
Verification is now the expensive step, not generation
The math withdrawals, the judge study, the surrogacy paper and the new agent tools all show the same thing: producing output is cheap, and confirming it is right is not.
OpenAI's math results are being corrected faster than they are being checked OpenAI published more than 700 model-solved problems in openai/math, then withdrew three of its claimed results while outside researchers improved on others with Lean-verified proofs. The story is now the gap between producing a proof and confirming it, and mathematicians say openly that they are uneasy about results nobody can easily follow. twitter.com | x.com | scientificamerican.com | i.redd.it | karagila.org
LLM judges invert when they can see the answer label The study tested 3 frontier judges on 3 datasets and 4 entity-alignment systems; when judges saw the system's decision label, J-ROC-AUC fell to between 0.12 and 0.87, and a simple label-free setup fixes it (preprint, unreplicated). Any clinical extraction pipeline graded by a judge that sees the extractor's output is probably overstating its accuracy. arxiv.org
Passing offline evals does not mean an A/B test would agree Schultzberg defines "trial-level surrogacy": offline treatment effects predict online ones across a class of interventions. The paper shows this condition is logically independent of per-example eval validity (preprint). That gives a precise test for when a retrospective benchmark on clinical records can justify a deployment change without a prospective trial. arxiv.org
New agent tools ship with checking built in SineFrame M3 runs an MCP server through real Claude Code, Codex, OpenCode or Pi and checks which tool calls they actually made, in plain pytest. Termaxa gates destructive commands in agent hook paths, and OhMyBug charges $10 only when its panel of models finds a confirmed bug of medium severity or higher. producthunt.com | producthunt.com | producthunt.com
Useful models got smaller and cheaper again
This week's model and product releases all moved toward smaller models, fewer bits per weight and inference with no server.
Haiku 5.5 is Anthropic's entry for high-volume work Anthropic calls Claude Haiku 5.5 its fastest and most capable small model, built for summarization, classification, coding subagents, support and browser use (vendor claim). For bulk work such as document triage and ICD coding, the question is whether the small tier now meets the bar a frontier model met last year. producthunt.com
Google and Cactus put retrieval and speech fully on the device Google AI Edge Foresight transcribes meetings offline and answers live questions from personal files using the new EmbeddingGemma 2, with no connection or subscription needed. Cactus's 16.9 MB Whistle model turns speech straight into JSON tool calls without ever producing a transcript, a design that keeps patient audio off any server. producthunt.com | theverge.com | cactuscompute.com
Sub-1-bit weights now have public code Samsung SAIL released code for LittleBit and LittleBit-2. Both factor weights into binarized low-rank parts to reach 0.1 bits per weight while keeping the original architecture at inference. Together with new work on 4-bit AdamW optimizer state, this makes memory something to tune rather than a hard ceiling for on-prem training and serving. github.com | arxiv.org
OpenAI's revenue reset comes as every lab pushes into enterprise tools
OpenAI's smaller revenue number landed on the same day the labs shipped routing, office-suite and security products for enterprise buyers.
OpenAI's run-rate is about $50B, roughly $20B below what it signaled An updated investor presentation shows about $50B in annualized revenue, with Q3 run-rate growth of 77% overall and 107% in enterprise; Nvidia, Oracle and CoreWeave fell on the report. Enterprise now drives the growth, which gives buyers more room to negotiate pricing and lock-in terms than they had a month ago. cnbc.com | techcrunch.com
OpenAI launched a routing product 14 days after a startup did OpenAI's Decisions API, now in public beta, picks models, tools or actions in near real time, up to 10x faster than GPT-6 Luna through the Responses API (vendor benchmark). Jev, a similar model from ex-OpenAI researcher's startup TypeSafe that returns calibrated decisions from application state, launched 14 days earlier (single source). x.com | x.com
Assistants are moving inside the documents themselves Claude now edits Docs, Sheets and Slides in place, in beta for paid plans, and Google is launching a single Gemini agent for work tasks. GPT-6 is reportedly bringing interactive charts and forms into ChatGPT answers, which makes data governance a question about each tool's permissions. producthunt.com | theverge.com | reddit.com | x.com
Anthropic's Cyber Mission puts its frontier models on infrastructure and open source Anthropic is sending its most capable models and engineers to power grid, water utility and other critical-infrastructure operators. Its OSS Scanner checks opted-in projects for free and returns a proof of concept, an explanation and a fix. A usage policy update that bans abusive behavior toward Claude and a science commitment shipped alongside it. anthropic.com | x.com | x.com | theverge.com | x.com | x.com | anthropic.com | theverge.com | anthropic.com
AI for science is investing in data and in how scientists work, not only in models
The money and the methods are both moving upstream, toward open measurement data and records of how researchers make decisions.
$1.8B will openly measure how cells respond to drugs and genes Biohub, DOE, NIH, Google DeepMind, Isomorphic and Meta are pooling the money for large-scale perturbation data that will be released openly. The last open biology dataset built this way ended up training AlphaFold, so proprietary drug-response data is likely to lose some of its advantage. rundatarun.io | biohub.org
A science agent's self-written skills beat the expert-written ones Gan Jiang's X-ray diffraction agent rewrites and validates its own skill instructions and code without any retraining. With those skills frozen, it beat the original expert-designed skills on held-out benchmarks (preprint, unreplicated). Skills that are validated and then frozen are a form of self-improvement that can be audited, which suits regulated pipelines better than weight updates. arxiv.org
Models are learning from how researchers actually decide ResearchTrails is a dataset of human research decisions extracted from GitHub, and a companion paper uses those decision paths to teach models scientific exploration. It targets the choice of next step in an analysis, which benchmarks that grade only final answers miss. arxiv.org | x.com
Dropped
7 items: discussion threads with no claim, an old textbook, a weekly roundup, an essay with no stated finding, and a duplicate post with a mismatched link.