Sponsored by

Chip stocks had an ugly Friday: the Philadelphia Semiconductor Index closed into a bear market, and Apple briefly overtook Nvidia as the world's most valuable company — roughly $4.88 trillion to Nvidia's $4.84. Analysts kept pointing at one thing: a leaderboard.

On July 16, Moonshot AI released Kimi K3 — 2.8 trillion parameters, open weights. It debuted at #1 on LMArena's Frontend Code Arena at 1,679 Elo, ahead of Claude Fable 5 and GPT-5.6 Sol; its predecessor sat at #18. The full weights are promised by July 27 — a promise on the calendar, not a file you can download.

Leaderboards are a number someone else calculated. So we sent identical prompts to three models, ran what came back, and scored it. One disclosure up front: our Claude column is Claude Opus 4.8, not Anthropic's newer Fable 5.

Test 1 - the creative one. Build an animated particle flow field, no libraries. Claude Opus 4.8 and GPT-5.6 Sol delivered on attempt one. Kimi K3's first two runs burned their reasoning budget and returned nothing — the output we scored was its third attempt. That reliability gap is part of the finding; the Elo score doesn't show it.

Test 2 - where we proved nothing. We wrote a spec where all 256 cells of a pixel-art image are fixed by a formula. One correct answer. All three scored 100% — a floor check, not a ranking. All three clear the bar; we draw no ranking from it.

Test 3 - six pixels. So we made it harder: render the Mandelbrot set across 160,000 pixels with an exact escape condition. Now the model has to implement math, not copy rules.

Claude Opus 4.8: perfect. GPT-5.6 Sol: perfect. Kimi K3: 159,994 out of 160,000.

Six wrong pixels — 99.9963%. The signature matches a loop that skips the final escape check, so six points that should be white came out black, in mirror-symmetric pairs — a real mathematical boundary, not noise. You have to hunt at the fourth decimal place to find daylight between an open model and the closed leaders — and it's one prompt, one run, one edge case a differently worded spec wouldn't have caught.

Test 4 - what image models can't count. We described a famous painting in precise text and asked two image models to recreate it photorealistically. Both nailed the atmosphere. The one machine-checkable requirement split them: one returned an ~3:4 portrait as asked; the other returned a square. Models follow atmosphere instructions convincingly; the only way to catch them is to count something.

One admission — our first version of that test asked for the wrong number of flowers. We had built the test wrong, so we rebuilt it; the scores here are from the corrected test. Most benchmarks fail the same way: the test was wrong, not the model.

What it costs. K3 runs $3/$15 per million tokens against $5/$25 for Claude Opus 4.8 and $5/$30 for GPT-5.6 Sol. Cost per task from our runs: $0.94 for Kimi K3, $1.04 for GPT-5.6 Sol, $1.80 for Claude Opus 4.8 — single-run figures that exclude K3's two failed attempts.

Can you run it yourself? Almost certainly not. Moonshot's own serving guidance calls for a "supernode" of 64 or more accelerators, and the weights alone are roughly 1.4 TB — data-center hardware, not a workstation purchase. What open weights buy you is provider competition, no lock-in, and distillation into smaller models that do run on normal hardware.

Be skeptical of single benchmarks, including ours. K3 tops frontend coding but sits at #3 among distinct models on Artificial Analysis's Intelligence Index and, as of this writing, #8 on the general Text Arena. Moonshot roughly tripled its input price (about $1 to $3 per million tokens) and nearly quadrupled output ($4 to $15, per Artificial Analysis) versus K2.6. On Artificial Analysis's AA-Omniscience test it answers more questions correctly than its predecessor (46% versus 33%) — but when it's wrong, it's now likelier to answer confidently than admit it doesn't know: that hallucination rate rose from 39% to 51%.

The bottom line: on these four tasks, one run each, an open model priced near zero matched two expensive closed models almost pixel for pixel — when it answered at all. Six pixels out of 160,000 were all that separated them.

— AI Super Simplified

The Smartest Way To Get 37 Claude Prompts Free

Subscribe to The AI Report, the free 5-minute daily AI brief that 400,000+ business leaders at IBM, AWS and Microsoft read every morning, and we'll send our 37 Claude Prompts pack free in your welcome email.

Most prompt lists are sorted by function. This one's built around the 8 situations every manager actually faces: team comms, hiring, meetings, strategy, reporting and more. Find your moment, copy the prompt, fill the brackets, run it in Claude. Managers are saving 10+ hours a week.

Join the free daily newsletter and get the prompt pack free, instantly, today. Subscribe and grab both

Keep Reading