SEELE02-pro leads GameDevBench at 65.8% Pass@1, ahead of Codex with GPT-5.6 Sol at 62.46%.
RESULTS
Ranked by Pass@1 on GameDevBench. Each score reflects one final effective attempt per task, evaluated by hidden Godot validation scripts.
LEADERBOARD
SEELE02, GPT-5.6 Sol, and Claude Opus 4.8 use complete 333-task validation runs. Earlier external rows follow the GameDevBench README-style leaderboard reference.
| Rank | Model | Harness | Feedback | Pass | Fail | Total | Pass@1 |
|---|---|---|---|---|---|---|---|
| 1 | SEELE02-pro | SeeleAgent | Final effective run | 219 | 114 | 333 | 65.8% |
| 2 | GPT-5.6 Sol | Codex | Screenshot + Video | 208 | 125 | 333 | 62.46% |
| 3 | SEELE02-flash | SeeleAgent | Final effective run | 193 | 140 | 333 | 58.0% |
| 4 | Claude Opus 4.8 | Claude Code | Video | 186 | 147 | 333 | 55.86% |
| 5 | gemini-3-pro-preview | Gemini CLI | Screenshot + Video | — | — | — | 53.8% |
| 6 | gpt-5.4 | Codex | Screenshot + Video | — | — | — | 52.0% |
| 7 | gemini-3-flash-preview | Gemini CLI | Video | — | — | — | 46.9% |
| 8 | gpt-5.4-mini | Codex | Video | — | — | — | 43.2% |
| 9 | gpt-5.4-mini | OpenHands | Baseline | — | — | — | 38.4% |
| 10 | claude-sonnet-4-5 | Claude Code | Screenshot + Video | — | — | — | 34.8% |
| 11 | gemini-3-flash-preview | OpenHands | Screenshot + Video | — | — | — | 31.8% |
| 12 | kimi-k2.5 | OpenHands | Screenshot + Video | — | — | — | 20.7% |
| 13 | claude-haiku-4-5 | Claude Code | Video | — | — | — | 18.6% |
| 14 | claude-haiku-4-5 | OpenHands | Screenshot + Video | — | — | — | 17.7% |
| 15 | qwen3.5-397b | OpenHands | Baseline | — | — | — | 5.4% |
KEY FINDINGS
The headline takeaways from the leaderboard above — every figure traces back to a row in the Pass@1 table.
SEELE02-pro leads GameDevBench at 65.8% Pass@1, ahead of Codex with GPT-5.6 Sol at 62.46%.
SEELE02-pro ranks first and SEELE02-flash ranks third, with GPT-5.6 Sol in second and Claude Opus 4.8 in fourth.
Even the efficiency-tier SEELE02-flash clears gpt-5.4 on Codex (52.0%) and gemini-3-pro-preview (53.8%).
Every pass is confirmed by hidden Godot validation scripts. Model self-reporting is never counted as success.
WHAT IS GAMEDEVBENCH?
GameDevBench evaluates agents on real Godot projects derived from web and video tutorials. Tasks require edits across scenes, scripts, UI, physics, shaders, TileSets, particles, resources, and runtime behavior. A submission passes only when Godot validation scripts say it passes.
Official categories cover 2D Graphics & Animation, 3D Graphics & Animation, User Interface, and Gameplay Logic.
Pass@1 counts one final effective attempt per task. Model self-reporting is not used as evidence of success.
EXAMPLE TASKS
These examples are compressed from real GameDevBench task_config instructions, preserving the actual task substance while making them readable on the page.
Update projectile behavior so enemy hits advance a QuestManager kill step, remove the projectile on impact, and clean it up after it travels off screen.
Clone a SafePlatform into an UnsafePlatform with an AnimationPlayer that shifts warning colors and disables Area2D processing during the red danger window.
Create a viewport-filling Control scene with launch, pause, and restart panels, exact node names, offsets, labels, buttons, font overrides, CanvasLayer, and script wiring.
Update a Godot shader with border-smoothing uniforms and wire the required ShaderMaterial parameters on the MeshInstance3D while preserving its texture input.
Modify TileSet metadata so six waterfall atlas tiles render above the player with a higher z-index and half-opacity modulation, without changing the map structure.
Place six tight semi-transparent Polygon2D circles over yellow star coins in a platformer image, minimizing spill outside each coin.
SOURCES
SeeleAgent builds 2D and 3D games — mobile and multiplayer — with export to Unity 6, Three.js, and WebGL, and distribution to Steam, App Store, and Google Play. Open the workspace and ship your first game, video, or scene tonight.