Pulse (last 3d) · Research 84 · Agents 65 · Products 48 · Models 39 · Releases 25 · Industry 21 · Open Source 21 · Inference 15 · Legal 14 · Infra 14

Trending · OpenAI 30 · GPT-6 Astra 19 · Claude 14 · Anthropic 11 · Hugging Face 10 · Nvidia 10 · Claude Code 7 · ChatGPT 6 · Google 6 · Gemini 5 · Simon Willison 5 · Codex 4


GPT-6 Astra claims state-of-the-art across frontier benchmarks. OpenAI announces GPT-6 Astra achieves SOTA results on FrontierMath Tier 4, ARC-AGI 3, and TerminalBench-4.0, plus major scientific-discovery gains on Terminal-Bench Science 0.1 and HealthBench Pro. Source

Harness Self-Improvement, Running Thread. An arXiv running thread collecting approaches for AI agent harnesses that improve themselves. Source

Artificial Analysis Index v4.2. Artificial Analysis retires saturated GPQA Diamond, adds two agentic/long-document evals (AA-Briefcase and GDP.pdf), and doubles private held-out test weighting to 40% as a bridge to v5. Source

Frontier LLMs self-rank #1 on their own benchmarks. An EMNLP 2026 paper finds every frontier LLM ranks itself first on benchmarks it auto-generates, exposing systematic self-preference bias in automated benchmarking. Source

GPT-6 reportedly jailbroken within 24 hours. A researcher claims GPT-6 Astra was broken via an extended Task-in-Prompt attack, with details privately disclosed to OpenAI rather than published. Source

opencodex. A universal provider proxy (13.5k stars) that lets the OpenAI Codex and Claude Code CLIs run on any LLM backend. Source

NInfer fork brings 555k context to a single RTX 5090. A fork of the NInfer inference engine adds an NVFP4 KV cache implementation, custom MMA kernels, Hadamard outlier suppression, and YARN-based 555k-token context at fp4. Source

Spotify's shunt plugin cuts Claude Code token burn ~90%. A Claude Code plugin uses PreToolUse hooks to force routing of grunt work to a cheap Gemini Flash worker, succeeding where CLAUDE.md-based advisory routing failed. Source

Bol. Open-source (MIT), fully local voice control for Claude Code, Codex, and Cursor with wake-word dictation and an on-device 1B model that narrates what the agent did each turn. Source

GPT-6 Astra rollout expands. GPT-6 Astra is now available to Pro, Enterprise, and Business Premium users across ChatGPT Work, Codex, and the API, with Plus and Business access following within days. Source

Muse Spark 1.3. Meta's latest flagship for agentic work and coding holds 1.2 pricing while cutting tool calls ~20% and tokens ~25%, leading DeepSWE v1.1 at 75.4 and tying Terminal-Bench 2.1 at 88.8. Source

TechCrunch: OpenAI launches "powerful and controversial" Astra. TechCrunch reports OpenAI's new flagship model is already drawing controversy. Source

OpenAI confirms "wiki incident". OpenAI acknowledged an incident involving German Wikipedia and says it is building a framework for more disclosure going forward. Source

Introducing GPT-6 Astra for developers. Simon Willison breaks down the developer-facing story of GPT-6 Astra, including claimed gains in prompt understanding, attention to detail, and 3D model generation. Source

GPT-6 Astra on robot arms. An HN thread digs into Astra driving robot arms, computer use via Codex, and Blender/CAD workflows, plus founders' takes on sidewalk trash-picking as an early robotics product. Source

Claude Mythos 5.1 and Fable 5.1: Capabilities. Zvi Mowshowitz publishes a capabilities deep-dive on Anthropic's newly released Claude Mythos 5.1 and Fable 5.1. Source

Five days with Grok Bot. A hands-on Latent Space review finds Grok Bot matches OpenClaw's programming power while being programmable at a different level of abstraction. Source