Models
Frontier models with sustained coverage here. One-off mentions live on their analysis pages.
claude 12
- Fable's Guardrails Are Blocking the Security Researchers Who Want to Use It
- Cyber agents are constrained by permissions, audit, and accountability
- Project Glasswing is about cyber operations, not offense demos
gpt 10
- Ads and finance push ChatGPT's trust stack into view
- ChatGPT commercialization is a context-boundary problem
- ChatGPT's Dreaming moves context engineering into the product default
gemini 5
- Gemini 3.5 Live Translate: Real-Time Voice Translation Leaves the Demo Reel
- Gemini’s real Apple win is developer distribution, not just Siri
- Apple hid Gemini inside Private Cloud, and rewrote who gets credit for Siri
deepseek 4
- Are Local Models Good Enough Yet: Two Camps Measuring Two Different Things
- DeepSeek V4 Moves 1M Context Into the Cost-Structure Era
- DeepSeek V4: Open Weights Finally Lead on the Efficiency Frontier, Not the Leaderboard
gemma 4
- Are Local Models Good Enough Yet: Two Camps Measuring Two Different Things
- DiffusionGemma: Text Diffusion Finally Reaches Mainstream Open Source
- Gemma 4 12B Drops the Multimodal Encoder: Google's Bet on a Unified Token Space
mai 4
- Microsoft's MAI-Thinking-1: The Logic Here Is Control, Not Catching Up to GPT
- MAI-Code-1-Flash Matters Because Microsoft Put Its Own Model Near Copilot's Default Path
- Frontier Tuning Turns Enterprise Tuning Paths Into Microsoft Platform Assets
mimo 4
- Xiaomi MiMoCode: Open-sourcing the Claude Code Playbook for Free
- MiMo UltraSpeed's Value Is the Real-Time Interaction Cost Curve
- MiMo UltraSpeed Pulls 1T Models Toward Real-Time Agents, But Not as a General Entry Point
qwen 4
- Are Local Models Good Enough Yet: Two Camps Measuring Two Different Things
- Qwen3.7-Max Is an Agent Foundation
- Qwen3.7-Max: Alibaba's Advantage Is the Enterprise Agent Stack, Not a Single Benchmark
cosmos 3
kimi 3
minimax 3
codex 1
fable 1
glm 1
glm-5-2 1
holo3.1 1
llama 1
mellum2 1
mythos 1