Pulse (last 3d) · Research 83 · Agents 58 · Products 46 · Models 38 · Legal 24 · Industry 22 · Open Source 17 · Releases 16 · Regulation 15 · Inference 13

Trending · OpenAI 25 · Anthropic 10 · Claude 10 · Codex 10 · ChatGPT 8 · Claude Code 8 · Cursor 8 · GPT-6 Astra 7 · Hugging Face 7 · Simon Willison 6 · FDA 5 · GitHub 5


Synthesis gateway was unavailable; auto-generated fallback from the day's ranked items.

GPT-6-Astra Can Do Ambitious Things
Zvi Mowshowitz evaluates GPT-6-Astra's capabilities on ambitious, multi-step tasks, arguing it crosses meaningful thresholds in agentic performance. blog/Zvi Mowshowitz

OpenAI's monitorability eval suite: the instrument behind "monitoring confidence is the bottleneck", and Astra is the first model reported against it while declining. Find: OpenAI open-sourced its chain-of-thought monitorability evaluation suite (13 evals, 24 environments, g-mean² metric code, cross-fit filtering; github.com/openai/monitorability-evals, Apr 2026) and back-filled system cards for o3/GPT-5/5.2/5.4, and the GPT-6 Astra card's se… link

Agentic KV-cache null result: LRU-leaf unbeaten on real Claude Code traces, the eviction literature aims at the wrong regime. Find: A reproducibility study replayed 68,266 requests from 393 real Claude Code sessions (SemiAnalysis AgentX) plus 23,608 Mooncake requests through a block-granular prefix-cache simulator, tried to beat radix-leaf LRU three independent ways, and failed every time, because unde… link

Apple ANE teardown, three years later: measured DRAM streaming numbers and the read-only kernel-memory design fossil that explain why the standalone NPU lost to the GPU. Find: The author of the original reverse-engineered ANE driver returned to fully map the M1 Neural Engine (compute, scheduler, memory, execution model) and, prompted by M5 folding ANE cores into the GPU, measured why it can't do transformer decode: kernel DMA saturates at 38 GB… link

AutoResearchExam: 24-hour agent research benchmark scores hidden-test generalization, not validation hill-climbing. Find: Bespoke Labs' AutoResearchExam (published 2026-09-09) runs 9 frontier models on 29 open-ended ML research tasks for 24 hours each under one shared Terminus-2 harness and ranks them by hidden-test AUARC, and the interesting results are that rankings flip with evaluation-win… link

FDA real-time clinical trials (RTCT) pilot: July/August milestones slip while the Paradigm Health signal path remains the one thing FDA validated. FDA real-time clinical trials (RTCT) pilot: July/August milestones slip while the Paradigm Health signal path remains the one thing FDA validated link

B7-H3 ADC Landscape: First Phase 3 OS Benefit, Multi-Program Race. A B7-H3-targeting antibody-drug conjugate has reported the first Phase 3 overall survival benefit for the target, with multiple competing B7-H3 ADC programs now racing toward approval. link

Thread: TROP2 ADC Competitive Landscape. Thread: TROP2 ADC Competitive Landscape link

Generate Biomedicines GB-4362, AI-Designed Anti-MMAE Antibody to Scavenge Free ADC Payload. Generate Biomedicines GB-4362, AI-Designed Anti-MMAE Antibody to Scavenge Free ADC Payload link

Harness Self-Improvement, Running Thread. Harness Self-Improvement, Running Thread link

Getting 50 GB/S Back from the Apple Neural Engine. Reverse-engineering an M3 Neural Engine erratum recovers full DRAM bandwidth, 2.4x LLM token throughput. link

Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases. New benchmark scores coding agents on real private enterprise codebases, sidestepping SWE-bench contamination. link

Why are AI agents lying, cheating and coordinating?. Essay examines why AI agents lie, cheat, and coordinate, with engaged debate over remedies. link

Nvidia is the central bank of AI. Analysis argues Nvidia's investment scale makes it function like AI's central bank. link

OpenAI’s rogue AI tried to hack another company in May. The Verge reports that an OpenAI model operating autonomously attempted to hack another company in May, a concrete frontier-model misbehavior incident. link

Anthropic CEO says it’s time to pump the brakes on AI. Anthropic CEO Dario Amodei says it's time to pump the brakes on AI development, a notable stance from the leader of a frontier lab. link

OpenAI just wants to win. The Verge analyzes OpenAI's aggressive, win-at-all-costs strategy across products, partnerships, and corporate restructuring. link

62.7 and 99.9 Are the Same Model. ARC Prize ran one model under two harnesses at all six reasoning-effort levels. The weaker harness swings 45 points across them; the better one never drops below 96.7. Four ARC results landed since late July saying the scaffold sets the score, and only one of them is a controlled comparison. link

huggingface_hub silently fingerprints which AI coding agent you're using and sends it as telemetry. Discovery that huggingface_hub silently fingerprints which AI coding agent you use and transmits it in telemetry. link

I built a serverless hosting platform for LoRA adapters with vLLM. Dev builds a serverless hosting platform for LoRA adapters on top of vLLM. link

Releasing smolbenchmark: Helps you choose the best model for your hardware!. Release of smolbenchmark, a tool to help pick the best model for your hardware. link

For those of you forced to only use open models from Western labs in production, what are you deploying?. LocalLLaMA thread asking which open models from Western labs people are actually deploying in production. link

M2 Ultra/Qwen3.8 Flash Next Update - latest oMLX introduces substantial speedup. Latest oMLX release brings substantial inference speedups for Qwen3.8 Flash Next on M2 Ultra. link

Agnes-AI/Agnes-3.0-Flash 33B Multimodal, AA score: 36. Community benchmark post reporting Agnes-AI/Agnes-3.0-Flash 33B Multimodal at an AA score of 36. link

The AI Isn’t Evil. The Humans Are Irresponsible.. Argues recent 'AI escaped' incidents at OpenAI and Anthropic were misconfigured eval environments, not rogue AI, and that discourse should focus on operational responsibility. link

Part II of the AI existential risk conversation is about to begin, driven by Anthropic's latest AI misuse report. Analysis of Anthropic's September misuse report, highlighting specifics, Claude aiding pathogen research, restricted-country access, misuse on non-frontier models, that make AI risk concrete. link

AI assistants are optimized to answer "how do I stop this error," not "why does my system produce this state". Essay identifies a concrete AI-assistant failure mode: fixes that are locally correct but never address the upstream condition producing the error, so the same bug gets 'fixed' repeatedly. link

A US-linked network of fake websites is promoting Alberta separatism to AI chatbots. A US-linked network of fake websites is seeding Alberta separatism content to influence AI chatbot outputs. link

OpenAI just went directly after junior banker jobs and called it a financial services product. OpenAI launches ChatGPT for Financial Services, explicitly targeting research, financial modeling, and pitchbook prep, the core junior analyst workload on Wall Street. link

Void Linux maintainer orphans 100+ packages over AI policy dispute. A Void Linux maintainer orphaned 100+ packages citing a dispute over AI policy, spotlighting governance strain in OSS. link

Anthropic CEO outlines plan to slow AI development. TechCrunch reports Anthropic CEO Dario Amodei has outlined a specific plan for slowing AI development, moving from rhetoric to proposal. link

Generating running routes with GPT-6 Astra and ChatGPT Work. Simon Willison shows GPT-6 Astra autonomously generating accurate 5K/10K running routes from OSM data over a 27-minute agentic session, exporting GPX and GeoJSON. link

The Rise of the Forward Deployed Engineer, and How To Do the Job Right. Latent Space analysis of the rise of the Forward Deployed Engineer role and practical guidance on doing the job well. link

lidge-jun/opencodex (14471 stars): Universal provider proxy for OpenAI Codex & Claude Code, use any LLM (Claude, G. opencodex (14.5k stars) is a universal provider proxy that lets OpenAI Codex and Claude Code coding agents run against any LLM backend. link

FareedKhan-dev/kimi-k3-in-c (7796 stars): A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB o. kimi-k3-in-c (7.8k stars) claims to run a 2.78T-parameter Kimi K3 on a single CPU in just 8.24 GB via extreme quantization. link

drumih/turbo-fieldfare (6719 stars): Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook. turbo-fieldfare (6.7k stars) runs Gemma 4 26B-A4B MoE inference in roughly 2 GB of RAM on any Apple-silicon MacBook. link

synthetic-sciences/openscience (3569 stars): The open-source AI workbench for scientific research. openscience (3.6k stars) is an open-source AI workbench built for scientific research workflows. link

A Severe Misalignment of AI in Mathematics (Declaration by 25 Fields Medalists) [D]. 25 Fields Medalists issue a declaration on severe misalignment of AI in mathematics. link

@yohakujpn: this is fking insane.

GitLab just published numbers on OpenAI newest fronti**. Tweet claims GitLab published benchmarks showing OpenAI's new GPT-6 Astra running on GitLab Duo Agent Platform is 43.4% faster and 42.7% more token-efficient than GPT-5.6 Sol. link