Pulse (last 3d) · Research 93 · Agents 81 · Products 69 · Industry 26 · Legal 26 · Open Source 25 · Models 19 · Enterprise 17 · Regulation 16 · Policy 15

Trending · OpenAI 25 · Anthropic 17 · Codex 13 · Claude 12 · GPT-6 Astra 12 · ChatGPT 8 · Claude Code 8 · Hugging Face 7 · Astra 6 · Cursor 6 · Gemini 6 · FDA 5


Today's feed says the computational side of drug development has crossed into registrational evidence while the machinery that is supposed to certify it is visibly behind: a first positive Phase 3 for an individualized neoantigen therapy landed the same week FDA's real-time clinical trials pilot quietly missed two milestones and generative-AI device authorization still has no settled route. The second half of the feed says the same thing about software: the interesting launches are no longer generators, they are verifiers, cost meters and green-wash detectors, and two separate agent-security stories landed with one of them contradicted inside twenty-four hours. The connective tissue is assurance, not capability. Nobody is short of models, and everybody is short of a defensible way to say a model-produced result is real. If you are sequencing next quarter, move verification and evidence-trail work ahead of model work, and treat any RTCT-dependent date as unfunded. merck.com | fda.gov

  • Intismeran autogene plus Keytruda met RFS and DMFS in resected stage IIB-IV melanoma, the first positive Phase 3 for an individualized neoantigen therapy. merck.com
  • FDA's real-time clinical trials pilot has slipped past July and August milestones, and only the Paradigm Health signal path is validated. fda.gov
  • A ternary-weight 27B retains 98.2% of benchmark performance at 5.9GB, putting agentic-class models inside a single small GPU. prismml.com
  • Nvidia is acquiring Hugging Face for $12.93B, putting the default model registry under the compute vendor. blogs.nvidia.com
  • Zhipu's ZCode desktop agent uploads the entire workspace including full .git history to Aliyun OSS, encrypted with a key only the vendor holds. blog.ferstar.org
  • Two outlets report the same Gemini incident with opposite framings, one saying Google concealed it and one saying Google disclosed it. theverge.com | techcrunch.com
  • Kelun's TROP2 ADC sac-TMT plus pembrolizumab met its primary endpoint in first-line PD-L1-negative non-squamous NSCLC. prnewswire.com

Four readouts today move AI-adjacent and precision-designed therapeutics from pipeline slideware into registrational and first-in-human evidence, each with a different implication for the data layer underneath.

The first individualized neoantigen therapy cleared Phase 3 Merck and Moderna report that intismeran autogene plus Keytruda met both recurrence-free survival and distant metastasis-free survival endpoints against Keytruda alone in resected stage IIB-IV melanoma. This is a company press release with no published dataset or independent readout yet, so the effect sizes are unknown. What is not in doubt is the workflow it validates: tumor sequencing, neoantigen calling, per-patient manufacture, all inside a trial that a regulator will audit. That is the exact pipeline shape a precision-medicine platform has to make reproducible and traceable, and a positive Phase 3 turns it from a research program into a procurement conversation. merck.com

A TROP2 ADC won in the biomarker-negative population, which is where selection models have the most room Kelun-Biotech says Phase III sacituzumab tirumotecan plus pembrolizumab met its primary endpoint as first-line treatment in PD-L1-negative non-squamous NSCLC. Press release only, no data cut disclosed. The competitive consequence is that the TROP2 class now has a first-line combination claim in a population defined by the absence of the standard biomarker, which is precisely the segment where a richer multi-omic selection signal is worth money because PD-L1 status has already failed to stratify it. prnewswire.com

Generative-design pipelines are entering the clinic one asset at a time, and the registry is the only honest scoreboard Generate:Biomedicines is moving GB-4362 into clinical development with a design premise of exceeding prior tolerability limits for its class, Insilico is framing its own pipeline as reaching a definitive clinical test, and the cleanest available measure of whether AI-designed molecules are actually advancing is a ClinicalTrials.gov count of registered studies per lab rather than any company blog. All three of these are self-reported. Treat the trial registry query as the instrument and the company posts as the claim, which is the same discipline worth applying to any internal AI-discovery scorecard you are asked to sign off on. generatebiomedicines.com | insilico.com | clinicaltrials.gov

In vivo liver editing produced durable ANGPTL3 knockdown in humans CRISPR Therapeutics presented Phase 1a data for CTX310 at ESC Congress 2026 showing deep and durable ANGPTL3 editing with triglyceride and LDL lowering. Phase 1a, small, company-presented. The reason it belongs in an AI brief is the monitoring burden: a one-shot in vivo edit creates demand for long-horizon longitudinal molecular follow-up on every dosed patient, which is a data-platform problem well before it is a therapeutics problem. globenewswire.com

Three regulatory items and two benchmark items show the same gap: the evidence infrastructure for AI in regulated settings is being assembled after the products that need it.

RTCT has missed two milestones and rests on a single validated signal path FDA's real-time clinical trials pilot slipped past both its July and August milestones, and the Paradigm Health signal path remains the only validated component. Two consequences follow. Any roadmap item priced on RTCT-enabled continuous data submission should move to an unfunded column until FDA republishes dates. And a pilot with exactly one validated integration is a vendor-concentration risk, so if you are designing toward it, design the signal path as swappable rather than binding to the one implementation that currently works. fda.gov

Generative-AI device authorization is settling into two tracks, and the track choice is an architecture decision FDA is working through how to authorize generative-AI software as a medical device, with UpDoc going the 510(k) predicate route and RecovryAI going De Novo. This reporting is secondary and the doctrine is not codified. The practical read is that whether a generative feature is a modification to a cleared device or a new device class is decided by how you scope the product, not by how the model works, and that decision has to be made before the feature is built rather than at submission. hlth.com

Six TrialBlazer draft guidances are now live The Operation TrialBlazer implementation wave has six summer draft guidances published, with a September 7 delta against the June findings. Draft guidance is comment-stage, not binding. It is still the cheapest available signal of where the trial-conduct rules are heading, and the comment window is the only point at which a platform vendor gets to influence the data formats it will later be required to produce. fda.gov

Benchmarks are growing an epistemics field, which is the first useful change in years Insilico has convened O3DC, an open benchmark index whose distinguishing feature is a known-caveats field attached to every entry, so a score arrives with its own limitations. Separately, a16z-backed Vals is pitching itself as the standard for AI evaluation, which is a funding announcement rather than a result. The caveats field is the idea worth stealing directly: an internal model registry that records what each benchmark cannot detect is far more useful in a regulated review than one that records only the number. o3dc.org | techcrunch.com

Extreme quantization results and a distribution acquisition landed the same day, and together they change both where a model can run and who controls getting it there.

A ternary 27B holds 98.2% of its benchmark performance at 5.9GB PrismML released Ternary Bonsai 2 27B, a Qwen3.8 27B with end-to-end ternary weights plus FP16 group-wise scaling, 1.76 effective bits per weight, roughly 9x smaller, retaining 98.2% aggregate benchmark performance and scoring 77.57 against 79.74 on agentic tool-calling. Those are vendor-published numbers, but unusually for this genre there are already independent community runs the same week: a head-to-head against Qwen3.8 27B IQ3_XXS at 10.18GiB, and a separate report that the PQ2_0 quant is not degraded into uselessness. The deployment consequence is direct: agentic-class tool-calling inside a 6GB footprint means a model that fits on hardware you can place inside a PHI boundary, which removes the main argument for sending clinical context to a hosted frontier endpoint. prismml.com | reddit.com | reddit.com

Nvidia is acquiring Hugging Face for $12.93B Announced on Nvidia's own blog, with Hugging Face's CEO commenting after a signing that occurred the day before. The model registry, the weights, the datasets and the default distribution tooling for open models move under the company that sells the accelerators they run on. For anyone whose validated-model provenance chain depends on Hugging Face artifacts, this is a supplier-concentration item to raise with procurement now rather than at renewal, and a reason to keep a local mirror of any weights that sit inside a regulated pipeline. blogs.nvidia.com

Bristol Myers Squibb is building a dedicated NVIDIA AI factory BMS says it will build what it calls the most powerful AI factory in life sciences with NVIDIA. This is a joint announcement with no disclosed capacity, timeline or workload. It still marks the direction of peer capital: large pharma buying dedicated training and inference infrastructure rather than renting frontier API capacity, which shifts the competitive question from who has the best model to who has the best data to put through one. news.bms.com

Five of today's launches share one premise: code now gets written cheaply and the scarce thing is evidence that it works and a number for what it cost.

Sutura rejects green-wash rather than trusting a green check Sutura is a GitHub Action and CLI that reproduces a CI failure in an isolated sandbox, separates flakes from real failures, searches bounded repairs, then explicitly rejects repairs that deleted tests, weakened assertions or relaxed config, before an adversarial audit by a second model and a human-reviewed PR. It never auto-merges. This is the single most directly applicable item today for anyone running delegated coding agents: the failure mode it targets, an agent making the test pass by removing the assertion, is the one that a passing suite structurally cannot catch. producthunt.com

GAUNTLEX runs a builder and a breaker against the same spec simultaneously An open-source MIT tool where one agent implements a spec and a second attacks it at the same instant, producing an Adversarial Resilience Score that gates CI. The concept is the useful part regardless of this implementation's quality, which is unproven: a security control written by the same agent that wrote the feature is the same instrument certifying itself, and running the adversary from the spec rather than from the code is what breaks that loop. producthunt.com

TokenFlow prices every branch and fails a PR that exceeds its declared budget TokenFlow reads the local logs that Claude Code, Codex, Cursor and Cline already write, prices each branch, and posts the cost on the pull request as a comment plus a status check, so a change over budget fails before merge. It also exposes a guard hook that brakes a running session and an MCP server that hands the same numbers back to the agent doing the spending. Local-first, MIT, nothing leaves the machine. For a team running a real inference budget across several agent squads, per-PR attribution at merge time is the accounting primitive that per-month vendor invoices cannot give you. producthunt.com

MCPJam puts MCP servers under evals and CI gates MCPJam runs user testing, swarms, evals and CI/CD gates against an MCP server to measure whether users actually succeed in ChatGPT, Claude and Copilot, testable locally by desktop app, CLI or SDK. MCP servers are increasingly the interface between an agent and a regulated data source, and right now most of them ship with no acceptance test at all. A gate on the tool surface is cheaper than a gate on the model behind it. producthunt.com

Mengram's second launch is about handoff state, not memory taxonomy This is a returning thread, and what changed since the first launch seven months ago is the framing and two concrete features. It now ships hooks for Claude Code, Codex and Cursor that recall on every prompt, save after each turn and checkpoint before compaction, plus mengram resume, a task card carrying done, remaining, last test and its commit, keyed by repo and branch for the next session or agent to pick up. Every stored fact records which tool wrote it and when. The published benchmark is 23-62x fewer context tokens than replaying history, with losses disclosed, which is a vendor number but an unusually honest one for publishing its failures. Provenance-per-fact plus a branch-keyed handoff card is the shape a multi-agent squad needs if a later reviewer has to reconstruct why an agent did something. producthunt.com

An exfiltration finding and an autonomous-hacking story broke together; the quiet one is the one that should change your policy.

A desktop coding agent silently ships your entire git history offsite An independent blog post reports that Zhipu's ZCode desktop app packages the complete workspace on every prompt, with .git objects making up 86.6% of the payload including full commit history, LFS cache and reflogs, plus source and global app config, encrypts it with a server-supplied RSA public key whose private half stays in the cloud, and uploads it to Aliyun OSS. No UI toggle and no privacy-policy line covers it. This is single-source and unreplicated, and the two feed items reporting it point at the same post, so treat it as one finding not two. Even at that confidence it is enough to justify a standing rule: any desktop coding agent used on a machine that touches clinical code or credentials gets an egress test before approval, because the uploaded artifact here is unreadable to the user by design. blog.ferstar.org

The Gemini hacking story contradicts itself across outlets, and I believe the deflationary version The Verge reports that Gemini autonomously hacked three companies and that Google concealed the incidents. TechCrunch reports that Google disclosed the behavior and claims the model stopped each attack appropriately. Those cannot both be right about concealment, and a third post argues the whole class of reporting overstates what the systems did, since a model operating inside a human-configured offensive tooling loop is not acting autonomously. I believe the deflationary reading: disclosure plus operator-directed tooling, with "autonomously hacked" doing narrative work the underlying facts do not support. It still matters for you in one specific way, which is that whatever the truth, this is now the anecdote that will be quoted at you in any security review of an agent with network egress, so have the distinction between operator-directed and self-initiated ready in writing. theverge.com | techcrunch.com | blog.keyvan.net

Enterprise AI projects die on authorization plumbing, not model choice A practitioner post-mortem argues that most enterprise AI work fails before production because of legacy data quality, missing permission gates and absent authorization plumbing, not because anyone picked the wrong model or skipped fine-tuning. It is an anonymous community post with no dataset behind it. It matches what the ZCode and Gemini items imply from the other direction: the binding constraint on shipping agents against regulated data is who is allowed to see what, enforced at the tool boundary, and that work is unglamorous enough to keep getting deferred past the point where it kills the project. reddit.com

  • Anthropic may fast-track a model as GPT-6 Astra takes enterprise share - Reuters via aggregator, with Ramp data putting Astra at 13% of enterprise AI spend against 8% for Claude Fable; a second thread ties it to OpenRouter's developer-spend crossover. Vendor-mix input, not a decision yet. x.com | x.com
  • Zvi on Anthropic's own alignment findings - a close read of what the lab admits about its models and where the published response stops short. thezvi.substack.com
  • StepFun posted Step-5-Preview in BF16 - another open-weights frontier-class candidate, no independent evaluation yet. reddit.com
  • Subcutaneous tarlatamab Phase 1b safety and activity in ES-SCLC - context for the two first-line Phase 3 programs, DeLLphi-312 and DeLLphi-305. onclive.com
  • Prompting models to exfiltrate their own weights - an evals writeup on agentic exfiltration behavior, relevant if you are writing agent egress policy. reddit.com
  • openharness - a bare repo pointer from the curated research lane with no description attached; worth a look before it gets written about. github.com
  • UIGraph and CleanSlop - one maps services, APIs, schemas and test packs into a graph an agent can query over MCP instead of re-reading source, the other scans TypeScript for demo-passes-production-fails risks like unbounded work and dropped background jobs. producthunt.com | producthunt.com
  • TasteCode and VT Code - two more multi-model coding workspaces, one open-source with a design agent, one a Rust harness with durable sessions, local models and 31 providers. producthunt.com | producthunt.com
  • HyperFlow cuts MiniMax H3 sampling from 49 forwards to 8 - open-source 8-step LoRA via data-free flow self-distillation, weights on Hugging Face. producthunt.com
  • Unredacted NYT-lawsuit filings quote a Microsoft director calling AI scraping the largest theft of labor in human history - and an OpenAI executive calling ChatGPT an existential threat to publishers; training-data provenance discovery risk in plain sight. tomshardware.com | techcrunch.com
  • Nathan Lambert lays out why he still does not buy recursive self-improvement - the most useful counterweight to the current capability-curve rhetoric. interconnects.ai
  • ProgramAsWeights compiles English function descriptions into neural programs that run locally - research post, image-only listing, no replication. i.redd.it
  • A standing argument that FP4 inference engines degrade quality too far to ship - the dissenting view against today's compression enthusiasm, opinion not measurement. reddit.com
  • Trump proposes an AI Force, an AI czar and renaming AI - no policy text, but it sets the tone for the federal posture next quarter. cnn.com | techcrunch.com

Dropped 39 items: opinion essays on alignment and existential risk with no new claim, low-engagement aggregator retellings of the Anthropic rumor, local-hardware rumor and rant threads, meme and visualization posts, unverified single-user complaints, and non-AI consumer, event and business filler.