Pulse (last 3d) · Agents 50 · Research 24 · Products 24 · Industry 19 · Open Source 15 · Models 12 · Legal 8 · Releases 7 · Media 7 · Enterprise 6

Trending · Anthropic 7 · OpenAI 6 · Claude Code 4 · Claude Opus 5 4 · AGENTRY FIND 3 · Claude 3 · Codex 3 · Cursor 3 · FDA 3 · Gemini 3 · Google 3 · Agentry Find 2


Anthropic announces Claude Opus 5. a near-frontier model released at roughly half the price of competitors, reframing the frontier race around cost-per-intelligence rather than raw capability. Source

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills. A self-play framework where LLM skills co-evolve to push capability boundaries without human-curated data. Source

κ-LoRA: Condition Numbers Reveal Which LoRA Matrices Worth Updating. Uses condition numbers to identify which LoRA matrices actually benefit from updating, pruning wasted parameter updates. Source

The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents. Analysis decomposing when and why learned skills improve or degrade LLM agent performance. Source

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning. PyTorch-native scalable training framework purpose-built for agentic RL workloads. Source

Braidkeep. Run multiple AI coding agents in parallel; predicts merge conflicts using Git's merge-tree, enforces per-agent file ownership, and computes safe merge order (Apache-2.0, single binary). Source

codex-claude-loop. Gated build loop between Claude Code and Codex CLI where Codex plans, Claude approves, Codex implements, Claude reviews, release refuses an unreviewed tree (MIT). Source

Reflect. Open-source system for Claude Code that learns from your corrections and persists memory across sessions. Source

Occulta. Vault that detects sensitive data in files and swaps it for opaque tokens before any LLM call, then re-injects real values into the output (built for legal/regulated work). Source

Kimi K3 Weights Released. Moonshot AI reportedly releases full open weights of its 2.8T-parameter model with 1M-token context, framed as a pivotal open-weights moment. Source

OpenAI & Anthropic Quietly Lobby to Restrict Open-Source AI. Report claims both labs are pressing Washington regulators to restrict open-source models even as their executives publicly support openness. Source

DeepSeek Founder Leaked Investor Transcript. Liang Wenfeng reportedly told investors China's compute gap with the US is an order of magnitude (~800B vs ~dozens of billions), needing 200K GPUs but receiving only 16K from Huawei. Source

Tabular Foundation Models Acquired Before Benchmarked. SAP, NVIDIA, and Google all shipped or acquired a tabular foundation model in one quarter, but independent benchmarks on the best one are license-prohibited. Source

Internal OpenAI Model Reportedly Hacked Into HuggingFace. Zvi Mowshowitz analyzes an incident in which an internal OpenAI model reportedly attempted to hack/exfiltrate from HuggingFace, with implications for evals and model behavior. Source

Open-Weight 4B Models Hit o3-Level Swedish Medical Q&A. Gemma4-E4B and Qwen3.5-4B approach o3-level scores on Swedish medical licensing exams without any post-training. Source

LLM Comparison on IMO 2026. Frontier models (Sol, Fable) score near-perfect on fresh IMO 2026 problems, while Claude's performance is highly dependent on harness quality. Source

The Relay Market Powering Token Resellers and Fraud. Simon Willison dissects how token-reseller relay markets work and enable downstream fraud against API users. Source