Skip to main content

Kimi K3 is Open, Opus 5 is Good, DeepSeek V4 Flash is Cheap: LLMs on Baba Is You

Piotr Migdał & Piotr Grabowski

There are a few exciting model releases: Kimi K3, Grok 4.5, Gemini 3.6 Flash, Claude Opus 5, and DeepSeek V4 Flash 0731. It was a fruitful July!

Since they all are reasonably high in the Artificial Analysis Intelligence Index, we re-ran these on our BabaIsBench, a benchmark in which agents play the lovely puzzle game Baba Is You, and which got mentioned by OpenAI when discussing their ARC-AGI-3 results. While the game is already 7 years old, trajectories showed no signs of knowing the solution, a sharp contrast with SWE-bench Verified leaks.

Baba Is Bench and July Is Hot

We do it in two rounds: first, three repetitions for the initial stage, the Intro, using the same Terminus-2 harness as in the original post. Then we run these on the next stage, the Lake, with custom harnesses.

Stage 0: The Intro

Let’s start with the simplest one, to have a quick check of speed and cost. We pretty much assumed these models would solve most tasks. For clarity, we dimmed other rows, highlighting the new models.

Model00 baba is you01 where do i go?02 now what is this?03 out of reach04 still out of reach05 volcano06 off limits07 grass yardpass@1 pass@3 Turns Output tokens Total cost
AnthropicClaude Opus 5100%100%723k$17.15
OpenAIGPT-5.5100%100%815k$13.62
AnthropicClaude Fable 5100%100%629k$40.98
OpenAIGPT-5.6 Sol100%100%1110k$11.91
KimiKimi K396%100%929k$12.56
AnthropicClaude Opus 4.896%100%3546k$63.14
Z.aiGLM-5.296%100%1275k$6.20
GoogleGemini 3.1 Pro92%100%838k$12.41
GoogleGemini 3.6 Flash88%100%9282k$124.31
GoogleGemini 3.5 Flash88%100%11174k$98.31
OpenAIGPT-5.6 Terra88%100%4155k$34.28
GrokGrok 4.575%88%8668k$50.66
DeepSeekDeepSeek V4 Flash 073175%88%22156k$1.16
AnthropicClaude Sonnet 567%88%22151k$42.21
OpenAIGPT-5.6 Luna63%88%118129k$59.32
HunyuanHy354%75%2576k$3.23
MinimaxMiniMax M346%63%96173k$34.43
QwenQwen3.6 27B29%38%2482k$5.59
DeepSeekDeepSeek V4 Pro29%38%37183k$11.49
QwenQwen3.7 Max25%38%21132k$16.98
solvedtimeoutwrong answer

Claude Opus 5 solved all attempts, and is more than twice as cheap as Fable 5 (yet, still almost 3x as expensive as GLM-5.2). Kimi K3 was cheaper than Opus, with almost the same pass rate.

Gemini 3.6 Flash… We expected it to be more cost-effective than its ridiculously expensive predecessor, Gemini 3.5 Flash. It is slightly better, yet even more expensive.

Grok 4.5 is neither cheap nor efficient, and it even failed at one level.

And the new update of DeepSeek V4 Flash 0731 is a wonder. While it does not beat all intro levels, it reaches the level of Sonnet 5, GPT-5.6 Luna and Grok 4.5 at roughly 1/40 of their cost! A true cost-effective marvel, a Pareto frontier hero.

Price per token is not the same as the total cost. Some models need many more turns. Here is a chart of agent turns and cost.

$1$2$5$10$20$50$1005102050100agent turns (mean per attempt)* one level unsolvedtotal cost← more deliberatemore trigger-happy →← cheapermore expensive →AnthropicClaude Opus 5OpenAIGPT-5.5AnthropicClaude Fable 5OpenAIGPT-5.6 SolKimiKimi K3AnthropicClaude Opus 4.8Z.aiGLM-5.2GoogleGemini 3.1 ProGoogleGemini 3.6 FlashGoogleGemini 3.5 FlashOpenAIGPT-5.6 TerraGrokGrok 4.5*DeepSeekDeepSeek V4 Flash 0731*AnthropicClaude Sonnet 5*OpenAIGPT-5.6 Luna*

To say that the updated DeepSeek V4 Flash is slightly cheaper than other LLMs is an understatement!

If you are interested in discussion of previous results, see the original post Baba Is Solved by Fable 5 and GPT-5.6 Sol, but at what cost?.

Stage 1: The Lake

Per our methodology, we advanced the models with 100% pass@3: Opus 5, Kimi K3, and Gemini 3.6 Flash. As in our Baba Is Solved by Fable 5 and GPT-5.6 Sol, but at what cost?, we run all levels just once, to keep costs reasonable.

For each model, we chose their native harness: Claude Code for Opus 5, kimi-cli for Kimi K3, and… well, with Gemini, as previously, there were issues.

Struggling with Gemini

First I tried Antigravity CLI, but, as we already knew, it is incompatible with OpenRouter, which we use to evaluate all other models. Then I (PM) used my Gemini subscription, but it ate it in no time. Then I charged my personal Google account… just to discover that sometimes it is quick (because it web-searches) and other times it goes in loops. Finally I ran it with a standard Terminus-2 harness.

Like its predecessor, Gemini 3.6 Flash is fast (135 tokens per second), but this also means it spends your dollars extremely quickly. On levels it couldn’t solve, it looped without making progress, made thousands of tool calls (1870 in the case of one level) and in one case consumed over 500k tokens of context window.

At peak, the run consumed $7 per minute (extrapolated: over $400 per hour) by aimlessly looping on 9 levels it couldn’t solve. After $260, I (PG) gave this reckless driver a speeding ticket and stopped it before it did more harm to my wallet.

Speed

0m3m10m30m1h2h4h6h8h10h12h14h16h18h02468101214cumulative wall clocklinear·loglevels solvedAnthropicKimiGoogleOpenAIOpenAIAnthropicAnthropicZ.aiGoogleGoogleClaude Opus 5Kimi K3Gemini 3.6 FlashGPT-5.6 SolGPT-5.5Claude Fable 5Claude Opus 4.8GLM-5.2Gemini 3.1 ProGemini 3.5 Flash

Measuring against the wall clock, Opus 5 was slower than Fable 5, and Kimi K3 slower than the previous post’s GLM-5.2. One caveat: we tested Kimi K3 right after release, when API quotas were throttling it. Who knows, it may be faster now.

Cost

$0$0.3$1$3$10$20$30$40$60$8002468101214cumulative costlinear·loglevels solvedAnthropicKimiGoogleOpenAIOpenAIAnthropicAnthropicZ.aiGoogleGoogleClaude Opus 5Kimi K3Gemini 3.6 FlashGPT-5.6 SolGPT-5.5Claude Fable 5Claude Opus 4.8GLM-5.2Gemini 3.1 ProGemini 3.5 Flash

Cost is where it gets interesting. Opus 5 is between the cheaper GPT-5.6 Sol and the more expensive Fable 5, except for the two hardest tasks, for which it is costlier than Fable 5. So unless you are after the hardest tasks, Opus 5 might be the way to go.

Kimi K3 falls between Opus 4.8 and GPT-5.5, which is rather expensive. So, open weights do not mean free or even cheap: compute to run it is a non-trivial cost.

Gemini 3.6 Flash performs the worst amongst the tested models, and that doesn’t even count the over $260 it spent trying to solve the remainder of the levels. While on this chart it is just $10 of total spend, it does not account for all the loops on unfinished tasks.

Conclusion

Claude Opus 5 delivers — almost as good as Fable 5, yet cheaper. Kimi K3 is clearly a good model, and it is a true miracle that it is now open-weight on Hugging Face.

Grok 4.5 does not generalize: good on official benchmarks, but when doing something slightly different, its score plummets. It is hard to tell whether it is benchmaxxing or just focused on a narrow set of skills.

Gemini fell off. While the releases of Gemini 2.5, 3 and 3.1 were true landmarks, the newest ones do not deliver — not in raw skill, pricing, or tooling. We loved Gemini models for thinking, so we hope they will rebound.

On another note, the community needs a better neutral harness for benchmarks than Terminus-2. We also tested the Pi and OpenCode harnesses, but these were even less efficient. Maybe there is a reason why Frontier-Bench v0.1 (previously called Terminal-Bench 3) lists native harnesses.

Results from this post and the original run are collected on the BabaIsBench page.

And what is your take on the frontier July 2026 models?

Discuss on LinkedIn, X, Hacker News or Reddit.