17 vendors. 25+ models. 1 month. I don't feel the difference anymore.
Let me just list what shipped in the past month. Not an exhaustive list — only the notable ones:
Anthropic: Claude Sonnet 5 (Jun 30), Claude Opus 5 + Opus 5 Fast (Jul 24)
OpenAI: GPT-5.6 Sol / Terra / Luna (Jul 9), then an 80% price cut on GPT-5.6 (Aug)
Google: Gemini 3.6 Flash + 3.5 Flash Lite + 3.5 Flash Cyber (Jul 21), Gemini Robotics 2 (Aug)
xAI: Grok 4.5 (Jul 8)
Meta: Muse Spark 1.1 (Jul 16)
DeepSeek: V4 Flash (Jul 31 / Aug 1)
Alibaba: Qwen3.7 Flash (Jul 27), Qwen 3.8-Max — 2.4T params (Aug 3)
Moonshot: Kimi K3 — 2.8T params, largest open weights ever (Jul 16)
ByteDance: Seedance 2.5 (Aug 3)
MiniMax: H3 multimodal (Aug 3)
Huawei: openPangu-2.0-Pro — 505B total / 18B active, open source (Aug 1)
Meituan: LongCat 2.0 (Jul 20)
NVIDIA: Nemotron TwoTower (Jul 1)
poolside: Laguna S 2.1 (Jul 21)
thinkingmachines: Inkling (Jul 17), Inkling Small (Jul 30)
Ant Group: Ling-3.0-flash (Jul 23)
Kuaishou: KAT-Coder-Air V2.5 / Pro V2.5 (Jul 10)
SenseTime: SenseNova U1.5 (Aug 3)
That's 17 vendors, 25+ models, in roughly one month. A new model roughly every 36 hours.
A year ago, every time a major model dropped — especially from OpenAI or Anthropic — I'd drop what I was doing and go test it. Run my usual prompts. See if the reasoning got sharper, if the code got better, if it could finally handle some edge case that used to trip it up. It felt like opening a gift.
Now? I see the announcement and scroll past. Not out of cynicism — out of saturation. There's just no time to process one before the next three land.
And here's the thing: in my actual day-to-day work, I can no longer feel the difference. I swap the model, run the same tasks, and… it's fine. A little faster maybe. A little cheaper. But the output quality, the reasoning depth, the "wow" factor — it's been flattening for a while. The gap between "last month's best" and "this month's best" has shrunk to the point where the swap doesn't justify the eval effort.
I'm not saying the models aren't improving. They are. But the improvement is happening in dimensions that matter more to benchmark leaderboards and pricing charts than to someone like me who just needs to get work done.
Replies