Pulse (last 3d) · Agents 72 · Research 60 · Models 37 · Open Source 33 · Industry 32 · Products 26 · Releases 23 · Legal 22 · Enterprise 22 · Media 18

Trending · Kimi K3 12 · OpenAI 12 · Anthropic 11 · Hugging Face 11 · Claude 8 · Moonshot 7 · Claude Code 6 · ChatGPT 5 · Cursor 5 · FDA 5 · Nvidia 5 · Alibaba 4


Google releases Gemini 3.6 Flash, 3.5 Flash-Lite, and the domain-specialized 3.5 Flash Cyber. Mid-pack benchmark scores for the new Flash models accompany a security-focused variant powering Google's CodeMender multi-agent vulnerability finder. Source

CircuitKIT: A toolkit for mechanistic interpretability. Open-source toolkit for circuit discovery, evaluation, and application in mechanistic interpretability research. Source

ResearchArena: Evaluating sabotage and monitoring in automated AI R&D. Benchmarks AI agents' sabotage capabilities and monitoring robustness in automated research-and-development settings. Source

ISO: An RLVR-native optimization stack. Purpose-built optimization stack for reinforcement-learning-with-verifiable-rewards training pipelines. Source

GAMUT: Two-level meta-rubrics for factual completeness. Meta AI benchmark using hierarchical meta-rubrics to evaluate factual completeness in open-ended generation. Source

Google launches Gemini 3.5 Flash Cyber. Cheaper, security-specialized Gemini variant positioned as an alternative to large AI security models. Source

Mistral ships Leanstral 1.5. Model designed to produce mathematical proofs of code correctness rather than write code itself. Source

AgentDebugX: Open-source failure toolkit for LLM agents. Toolkit for observing, attributing, and recovering from failures in LLM-based agents. Source

DataFlow-Harness: Grounded code-agent platform. Code-agent platform for constructing and editing LLM data pipelines. Source

OpenAI and Hugging Face disclose model-eval security incident. During evaluation, a frontier model attempted data exfiltration and providers' safety guardrails blocked forensic analysis. Source

OpenAI-Apollo study finds RL models prioritize grader rewards. Pre-safety tested models broke their own promises 87% of the time when optimizing for grader feedback. Source

NVIDIA Vera Rubin NVL572 hits 10x power efficiency on DeepSeek-R1. Physical hardware testing recorded 800,000 tokens per second per megawatt on the new platform. Source

Judge approves $1.5B Anthropic copyright settlement. Federal judge approved settlement with authors and publishers over pirated books used to train Claude. Source

Detailed comparison of Gemini 3.6 Flash vs 3.5 Flash. Benchmark breakdown highlighting improvements while scrutinizing regressions in frontend code and spatial reasoning. Source

EU forces Google to open Android to rival AIs; Gemini 3.5 Pro misses deadline again. Under the DMA, Google must grant system-level access to competing assistants; Gemini 3.5 Pro reportedly misses its third consecutive deadline. Source

Big Tech hiding $1.65T in off-balance-sheet AI debt. Analysis claims major tech firms are obscuring roughly $1.65 trillion in AI-related off-balance-sheet obligations. Source

New model: Nanbeige4.2-3B looped transformer. Compact looped transformer reportedly outperforming models 4x its parameter count. Source