GROK 4.6 · NOW LIVE ON HAPPYCAPY

Humans Can't Tell Grok 4.6 From Grok 4.5 Anymore — So We Let an AI Grade the Exam

Every model release follows the same script: benchmarks up a few percent, ELO up a few dozen points, the timeline cheers for a day. Then ask anyone cheering what the new version actually does differently in daily use. They can't tell you. Not because they're lazy — because the gap has become too fine for human eyes to judge.

So when Grok 4.6 landed on Happycapy this week, we skipped the human judges entirely. An AI set the exam: three mathematical visualization problems — Fourier series, the Lorenz attractor, and the Kakeya needle problem that reached the top of mathematics' honors this year. In two fresh sessions, Grok 4.6 and Grok 4.5 received the identical questions, each answer a single runnable web page. The grader was Fable-5, a third model living in the same model picker, which ran all six answer files itself. Humans did exactly one job in this pipeline: invigilation.

Final score: 90 to 70. The winner also used 15% fewer output tokens (11,702 vs 13,799). If you don't trust an AI's grading — all six raw answer files are attached below, unedited, and the code runs on click. Grade them yourself. See you in the comments.

EXAM NOTICE

Grok 4.6 moved into Happycapy this week. House rules: new students take a placement exam. But we don't test benchmarks — it publishes those itself. We set three problems about the beauty of mathematics, seated it next to its senior, Grok 4.5, and asked Fable-5 from next door to grade.

This exam doesn't measure who computes faster. It measures: handed a piece of deep mathematics, can you make the beauty visible — and explain it like a person?

— The Registrar (a capybara)

Q IDraw a capybara with a hundred circles(Fourier series · 30 pts)

Fourier told the world that any closed curve can be decomposed into a chorus of circular motions. Candidates must chain rotating circles tip-to-tail so the outermost point traces a capybara's profile out of thin air — and explain, in plain words, why circles can draw anything.

Q IDraw a capybara with a hundred circles(Fourier series · 30 pts)

Three differential equations; two trajectories whose starting points differ by 0.0001. For a few seconds they shadow each other — then part ways without warning. That is chaos, and it is why no weather forecast survives two weeks. Candidates must bring this butterfly to life on screen.

Q IIWhy the butterfly flaps its wings(Lorenz attractor · 30 pts)

Turn a needle around inside a flat region: how little area must it sweep? Intuition says half a disc. Mathematics says: as little as you like. This near-absurd fact is the Kakeya problem; its three-dimensional cousin was cracked and honored at mathematics' highest podium in 2026. Candidates must paint the absurdity for ordinary eyes.

Conclusion: 

This comparison was run on (Capy): Grok 4.6 beat Grok 4.5 90:70 across all three problems while producing 15% fewer output tokens (11,702 vs 13,799). Grok 4.6's advantage concentrated in readability and aesthetic judgment; Grok 4.5 retained bright spots in instrumentation and single-panel visualization. Both candidates and the grading model, Fable-5, are available in Capy's model picker — one subscription, no separate API keys. All six raw answer files are published with this article and are runnable.

Don't trust the grading? The exam room is open to everyone. Both candidates, the grader, and twenty-plus other models share one model picker on Happycapy — one subscription covers them all, no API keys to manage. Set your own exam and let them answer. Connect your Notion, GitHub, or your own computer, and they'll turn the answers into finished work.

12 views

Add a comment

Replies

Be the first to comment