News

Kimi K3 tops frontend code rankings but trails badly on hard math

Moonshot's Kimi K3 sits at #1 on the Code Arena Frontend leaderboard, but scores only ~39% on FrontierMath Tier 4 versus ~90% for top OpenAI and Anthropic models.

Moonshot’s Kimi K3 landed at the top of the Code Arena Frontend leaderboard with an Elo score of 1,679, beating Claude Fable 5 (1,631) and GPT-5.6 Sol (1,618). According to The Decoder, it’s the first time a Chinese model has claimed that particular top spot. For frontend work, that’s the headline.

The math story is different. According to data from Epoch AI cited by The Decoder, K3 scores roughly 39% on FrontierMath Tier 4, the benchmark’s hardest expert-level tasks. OpenAI and Anthropic models score close to 90% on the same tier in some cases. That’s not a narrow gap; it’s a different category of capability.

The practical read: if your work is frontend code, UI components, web apps, that kind of thing, K3 looks like a serious option once the weights drop on July 27. If your work involves complex math, symbolic reasoning, or anything that would land on FrontierMath’s hardest tier, the current open-weight field (K3 included) isn’t there yet. One benchmark doesn’t prove a model, but the split between “best frontend coder” and “can’t touch the hard math” is specific enough to be useful. Worth keeping the weights drop date in the calendar and running your own tests.

Source: The Decoder ↗