Pulse (last 3d) · Research 131 · Agents 78 · Products 60 · Models 44 · Industry 42 · Open Source 30 · Releases 25 · Enterprise 25 · Legal 23 · Policy 18

Trending · OpenAI 25 · Claude Code 12 · Codex 12 · Anthropic 11 · ChatGPT 11 · Claude 10 · FDA 10 · Mistral 8 · Mistral Large 4 8 · Cursor 7 · GitHub 7 · Simon Willison 6


Capability got cheap and plentiful on the same day: Claude Haiku 5.5 costs $0.10 per million input tokens, a 1T-parameter Mistral model gets open weights this month, and a 27B open medical model beats GPT-6 Astra on HealthBench. What stays scarce sits around the model: checked output (mathematicians now urge a stop to work with OpenAI over unchecked proofs), open biological training data (a $1.8B coalition formed to build it), and a settled protocol for agents that act for a user. The small-model routing tier and the open-weight fallback both deserve a new bake-off before November, and the lasting advantage is in data and verification. anthropic.com | terrytao.wordpress.com

  • Claude Haiku 5.5 costs about 75% less than Haiku 4.5 and adds an effort setting. anthropic.com
  • Mistral Large 4 is a 1T-total, 49B-active open-weight MoE with 1M context, weights due by end of October. producthunt.com
  • HuatuoGPT-3, a 27B open medical model, scores 70.1 on HealthBench, above GPT-6 Astra. arxiv.org
  • A $1.8B coalition of Biohub, DOE, NIH, Meta, DeepMind and Isomorphic will build the largest open AI-ready biology dataset. x.com
  • GeneICL, a 4.2M-parameter tabular model with a transcriptomics prior, wins on clinical outcomes. arxiv.org
  • Tao reposted a mathematicians' statement urging a stop to OpenAI collaboration after its dump of unverified proofs. terrytao.wordpress.com
  • Meta and Sierra's Personal Agent Protocol signs Shopify, Stripe and Walmart against Visa's rival protocol. sierra.ai
  • Nano Banana 2.1 - Google image model with mask editing, better text rendering and Search grounding up to 4K. producthunt.com
  • Reflection's first model - Axios reports it and other Western open-weight models arrive this month (single source). x.com
  • DeepSeek nears a $12B round - backed by Tencent and CATL, ahead of an early-2027 IPO (single source). qz.com
  • OpenRouter usage and spend split apart - the models with the most tokens are not the ones earning the most spend. reddit.com
  • Spend versus traffic - 29 of the top 50 AI vendors by spend appear in neither a16z traffic ranking. theneuron.ai
  • Polars 2.0 - streaming is the default for lazy queries, with disk spill, first-class SQL and a Map type. producthunt.com
  • TidesDB for MySQL - a new storage engine optimized for writes and space. tidesdb.com
  • openTPU - agent-designed open accelerator that runs Qwen3 and Gemma 4 on a $300-class FPGA. github.com
  • Databench by Alkera - open-source shared notebooks with parallel agents and results traced to data and code. producthunt.com
  • Docker Agent - open-source framework for defining collaborating agents without code. github.com
  • Figma Agent - canvas agent that works with real components, pulls MCP context and runs team skills. producthunt.com
  • Manus 2.0 Video Editor - combines AI video generation with editing workflows. producthunt.com
  • Pegasus 1.6 - TwelveLabs egocentric video model that outputs labeled, timestamped robot training data. producthunt.com
  • vidu-q4 - Vidu released a new video model (single source). x.com
  • DevAlly AI Agent - records user journeys and audits each step against WCAG criteria. producthunt.com
  • Reika - coding-agent CLI built for local models first. producthunt.com
  • Thalia - free Mac app for Meta's Muse Code CLI with read-only plan mode and snapshot review. producthunt.com
  • GSS - 3D language in CSS syntax that compiles to a single raymarching shader. producthunt.com
  • An Application in Lisp You Grow by Talking to It - live-image Lisp development driven by LLM conversation. ghuntley.com
  • TUI-to-UI model - a 1.26M-parameter model turns htop, vim and emacs output into UI components. reddit.com
  • Codex 28-day sprint - OpenAI promises one shipped improvement a day or a full usage reset for everyone. thenewstack.io
  • Anthropic startup program - a free Claude Team year, five premium seats, $1,000 in credits and perks. techcrunch.com | x.com
  • Atlassian adopts GPT-6 Astra - GPT-6 Astra and GPT-5.6 run across Rovo, Jira and Confluence. openai.com
  • Nous Research at $1.5B - confirms the valuation and launches agents for business users. techcrunch.com
  • Lambda raises $4B - at a $14.5B pre-money valuation, its last private round before an IPO. wsj.com
  • Nvidia nears $6T - the stock is back at a record high. bloomberg.com
  • Microsoft's Nvidia AI PCs - revamped Windows 11, Surface Laptop Ultra, and a $5,999 Surface RTX Spark Dev Box. techcrunch.com | theverge.com | theverge.com
  • Boston Dynamics CEO - former Amazon Alexa chief Rohit Prasad takes over. bloomberg.com
  • Anduril's week - Arsenal-2, NGC2 and a $6.6B shipbuilding investment. tectonicdefense.com
  • Australia's AI safety plan - AI companies would have to prove their safety systems work. abc.net.au
  • ChatGPT for Teens - Common Sense Media rates it an "unacceptable risk." theverge.com
  • Generative protein design for DNA editing - generative AI designed better DNA-editing proteins. thebrighterside.news
  • Before They Can Solve - predicts a base model's coding-agent performance before post-training. arxiv.org
  • On-policy distillation - three papers on collapse modes, multi-teacher composition and adaptive self-teaching. huggingface.co | arxiv.org | arxiv.org
  • Decoupling exploration from optimization in RLVR - separates the two in RL with verifiable rewards. arxiv.org
  • DLoop - looped speculative decoding for faster inference. huggingface.co
  • Forget-only unlearning needs memorization - such methods depend on memorizing retained data, which matters for patient deletion requests. arxiv.org
  • EngramEdit - updates model knowledge through a separate conditional memory. arxiv.org
  • Retrieval instructions and evidence routing - how instructions affect embeddings, plus adaptive context routing (RECAST). arxiv.org | arxiv.org
  • PHRBench - tests how LLMs reason after they have hallucinated. arxiv.org
  • Validity Without Ground Truth - uses stated-preference economics to validate LLM evaluations. arxiv.org
  • SkillForge - evolves agent skills and agents together through skill lifecycles. huggingface.co
  • RunningTab - agents act on workspaces through tabs held on the environment side. huggingface.co | arxiv.org
  • A Society of Researchers - institutional designs for populations of autonomous research agents. arxiv.org
  • SciExam for ENSO - tests whether agents can build El Niño climate models. arxiv.org
  • SWE-Game - benchmark for coding agents that build games to a user's spec. huggingface.co
  • Robot world-action models - RoboJEPA, UniWAM, Long-WAM, compositional WAMs and the RobotWorld benchmark. arxiv.org | huggingface.co | arxiv.org | huggingface.co | huggingface.co
  • Language sensitivity in VLA models - instruction wording changes robot behavior, and rephrasing reduces the effect. arxiv.org
  • Nvidia's DreamDojo ICML spotlight disputed - small gains over Cosmos 2.5 despite 44k hours of human data. reddit.com
  • JPEG XL in Chrome - Chrome ships JPEG XL support. developer.chrome.com

Today's releases moved price and openness more than peak capability, and the narrow models posted the frontier-beating scores.

Haiku 5.5 cuts small-model cost by three quarters and adds an effort setting Claude Haiku 5.5 costs $0.10/$0.50 per million tokens in/out for prompts under 100k tokens and $0.50/$2.50 above, about 75% below Haiku 4.5, and it is the first Haiku with adjustable effort (vendor benchmark). Prompts over 100k tokens cost five times more per token, so the chunking of long clinical records now sets the bill. anthropic.com | x.com

Mistral Large 4 makes a 1T-parameter model something an organization can host itself The natively multimodal MoE runs 49B active of 1T total parameters with a 1M-token context and was trained in Mistral's own European datacenters. The preview API is live now and weights follow by end of October (press release). Mistral calls it the strongest open-weight model from the US or Europe, which suits patient data that cannot leave owned infrastructure. producthunt.com | x.com

A 27B open medical model beats GPT-6 Astra on HealthBench HuatuoGPT-3 scores 70.1 total on HealthBench, above GPT-6 Astra (preprint, unreplicated). A 27B model that can run on owned hardware and beats a frontier generalist on the main clinical benchmark weakens the case for sending PHI to a hosted API for medical QA. arxiv.org

Perplexity's decision model is a classifier head, not a text generator pplx-decider-v1.1-27b raises the Jev Decision Index from 56.4 to 61.56, above Jev's own 57.9 on the same Qwen3.8-27B backbone (vendor benchmark). The checkpoint outputs a classification from a readout head instead of text, a design that fits triage and eligibility decisions better than prompting a chat model for a label. huggingface.co

Biology AI money went to shared datasets and data-aware priors rather than to bigger models.

A $1.8B coalition will build the largest open AI-ready biology dataset Biohub, DOE, NIH, Meta, Google DeepMind and Isomorphic are launching the effort (single source), and Google is separately putting millions into Biohub's virtual-cell program. Open cell data pooled at this scale pushes proprietary advantage toward linked clinical outcomes, which no public consortium holds. x.com | theverge.com

A 4.2M-parameter tabular model wins on clinical outcomes by using a transcriptomics prior ETH Zurich's GeneICL builds a transcriptomics-aware prior into pretraining and posts the first credible clinical-outcome win for tabular foundation models (preprint, unreplicated). At this size it runs anywhere, and tables of expression data plus outcomes are exactly the shape of a precision-medicine cohort. arxiv.org

This week platforms gave agents access to files, documents and checkout, while the standards for authorizing them stayed split.

Commerce agents now have two competing protocols with the same backers Meta and Sierra published the Personal Agent Protocol with Shopify, Stripe and Walmart signed on. Shopify and Stripe also still back Visa's rival Trusted Agent Protocol. With merchants hedging across both, agent identity and consent stay unsettled for at least one more standards cycle. sierra.ai | thenextweb.com

Assistants moved into the operating system and the office suite Copilot gets more control over Windows and user files, Claude now works in Google Docs, Sheets and Slides, and GPT-6 ships with an Intelligent UI that puts charts and buttons in replies. Agents now reach the documents that hold regulated data, so audit and permission scoping move to the OS and suite layer. theverge.com | claude.com | openai.com | openai.com | theverge.com

Claude Code's next-prompt prediction is accused of collecting preference data Claude Code now pre-fills the prompt box with a predicted next message. An Oct 6 post argues that a sent prediction becomes a positive label and an edited one becomes an RLHF preference pair; Anthropic denies this (single source). The dispute tests whether coding-assistant telemetry counts as training data under enterprise terms. zohaib.cc

The bottleneck in AI research output is now human verification, and the people who do it are pushing back.

Mathematicians move from criticism to a call for boycott over OpenAI's proof releases OpenAI's Oct 6 release added hundreds more math results. Terence Tao welcomed AI's role but criticized dumping unverified proofs, then reposted an Association for Human Mathematics statement urging mathematicians to stop working with OpenAI. The Institute for Advanced Study formed an independent math-and-AI advisory group, making unreviewed volume a reputational cost. scientificamerican.com | mathstodon.xyz | terrytao.wordpress.com | agmai.org

Epoch AI finds LLMs still well short of human researchers on innovation Epoch AI's analysis puts LLMs well behind human researchers at producing new ideas (single source). Together with the math backlash, it shows AI's research output growing faster than both its originality and the capacity to review it. i.redd.it

16 items dropped: narrow robotics, graphics and optimization-theory papers, opinion pieces, and anecdotes that make no claim anyone can act on.