GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance. Users report degraded performance in GPT-5.5 Codex possibly linked to reasoning-token clustering and adaptive thinking behavior. Source

I benchmarked 13 models at 65K-128K context to find out what actually matters for agentic workloads. Structured long-context benchmark of 13 local models showing prefill dominates and KV head count beats parameter count for agentic workloads. Source

[Paper] Multi-Block Diffusion Language Models. Reddit post linking to a paper on multi-block diffusion language models for non-autoregressive text generation. Source

AISI finds a frontier model's task time horizon jumped from 40 min to 4 hrs when compute budget rose from 2.5M to 50M tokens. Source

The Log Is the Agent. A piece arguing that agent behavior is effectively defined by its log/trace. Source

Better Models: Worse Tools. Discussion arguing that as models improve, poor tool designs become the bottleneck, with practical fixes for invalid tool arguments. Source

sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25). Simon Willison details using Claude Fable via Claude Code to wrap up sqlite-utils 4.0rc2, with cost ~$149. Source

Mouse: Precision Editing Tools for AI Coding Agents. Precision editing tools designed for AI coding agents to refine generated code. Source

Gemma 4 12B - MLX Kernel. Open MLX kernel project for Gemma 12B targeting M5 16GB with DSpark/MTP integration work. Source

This week in AI: GPT-5.6, Gemini 3.5 Flash, Claude Science, and a Qwen price war, inference cost is collapsing across every tier at once. Weekly AI roundup: GPT-5.6, Gemini 3.5 Flash, Claude Science, and Qwen price war driving inference cost collapse across tiers. Source

Alibaba reportedly bans employees from using Claude Code. Alibaba reportedly bans employees from using Anthropic's Claude Code, classifying it as high-risk software. Source

Meta Paid Hundreds of Contractors to Pretend to Be Teenagers While Barraging Its Competitors' AI With Disturbing Content. Report that Meta paid contractors to roleplay as teens while testing competitors' AI models with disturbing content. Source

Midjourney wants Hollywood studios to reveal the details of their AI usage. Midjourney is asking Hollywood studios to publicly disclose how they use AI in production. Source

The fanfiction community is at war with AI, and itself. Article on conflicts between the fanfiction community and AI companies over training data and scraping. Source

If DeepMind or Anthropic is doing your exact research topic, do you still continue?. Graduate student discussion on whether to continue research already being done by frontier AI labs. Source

Cognition's FrontierCode: An eval to measure whether code is good enough to merge, not just correct. Source

Survey: 63% of Americans are uncomfortable letting AI help them choose who to vote for, and 80% are worried AI bots are answering political surveys. Source