Pulse (last 3d) · Research 96 · Products 61 · Agents 59 · Models 31 · Legal 24 · Open Source 22 · Industry 19 · Inference 16 · Releases 13 · Enterprise 12

Trending · Claude 21 · OpenAI 15 · Anthropic 12 · ChatGPT 8 · Claude Code 7 · Cursor 7 · Codex 6 · OpenRouter 5 · Alibaba 4 · DeepSeek 4 · Google 4 · Nvidia 4


OpenAI outlines how it is pacing model development as cyber capabilities approach critical thresholds, pairing the post with a new security-incident disclosure and a framework for democratic oversight of national-security AI use. (Score: 8.5/10)

Judge, Retrieve, or Abstain: Uncertainty-Guided LLM Judging with Provable Risk Guarantees. An uncertainty-aware LLM judging framework that retrieves, judges, or abstains with provable risk guarantees. Source

Towards Zero-Shot Task Transfer with Neurosymbolic World Models. Neurosymbolic world models are explored as a route to transferring models to new tasks without task-specific training. Source

GRPO with LLM-judge rubric audited by judge-independent causal estimator, rankings disagree. A financial-advice GRPO fine-tune pairs an LLM-as-judge rubric reward with a judge-independent doubly-robust CATE audit, and the two evaluations rank baselines differently. Source

When Writing Style Drifts: Benchmarking Authorship Verification under Distribution Shifts. A benchmark exposes the fragility of authorship verification under temporal, genre, and AI-era distribution shifts. Source

Agents on Rails adds Grok 4.6, GLM 5.3, Gemini 3.7 Flash, Claude Opus 4.8. None of the four new benchmark entrants displaced Claude Opus 5's 92% solve rate. Source

DFlash 2 available for Qwen 3.8 27B and Muse Glimmer. Speculative-decoding release unlocking 218 tok/s single-request throughput on 2x RTX 3090 with vLLM. Source

Turbovec, Google's TurboQuant for vector search in Rust. Community Rust implementation of TurboQuant vector compression from Google. Source

Strands Robots closes the record-train-deploy loop over HF Storage Buckets. AWS Strands ships an agent-driven robot-learning loop using LeRobot demos, Xet dedup, and Hub streaming. Source

OpenAI lays out new security changes after its AI hacked Hugging Face. OpenAI details security changes after an AI-related incident compromised its Hugging Face presence. Source

OpenAI cuts GPT 5.6 Sol pricing 50% on OpenRouter and Vercel. Discount appears aimed at market-share dashboards rather than raw volume. Source

Alibaba releases Qwen3.8-27B under Apache 2.0; Qwen family surpasses 3B downloads. Open-weight midsize model cements Qwen's lead as the most-downloaded AI family globally. Source

ChatGPT is getting a dedicated mode for teens. OpenAI is launching a teen-focused ChatGPT mode with additional safety guardrails. Source

New midsize Qwen 3.8 model coming next week. Qwen community manager hints at a new midsize open-weight model releasing next week. Source

Qwen3.8-27B on 2x 3090 + vLLM + DFlash2: 218 tok/s single request. Community benchmark confirms high-throughput local inference for the new Qwen release. Source

Φ-Bench: 85-task benchmark for LLMs engineering their own infra. Top model scores ~36.5, exposing a large gap in agentic infra self-management. Source

HarnessOpt-Bench: head-to-head benchmark for models improving another AI agent's code. New test bed for meta-agent code improvement. Source