Fixture control
Pin revision, Unreal version, plugins, assets, platform, build mode, prompt, and tools.
Qwen3.8 × Unreal evaluation lab
A useful benchmark fixes the revision, task, evidence, budget, scorer, and recovery path. It measures whether a change works—not whether the answer sounds senior.
No official public Qwen3.8 Unreal benchmark is established. Build a local suite that scores completion, first-pass compile rate, invented API rate, regressions, evidence quality, human review time, and packaged behavior under matched conditions.

Snapshot date: July 20, 2026. Re-check time-sensitive preview, availability, plan, pricing, and open-weight statements before acting.
Pin revision, Unreal version, plugins, assets, platform, build mode, prompt, and tools.
Score compile, runtime, automation, package, correctness, latency, cost, review, and recovery.
Hide model identity when practical and use one rubric for model and human attempts.
Keep first failures, rejected diffs, logs, and rollback results; silent retries distort reliability.
Choose a compile fix, Blueprint review, runtime bug, asset/config issue, and packaged-build check.
Document expected owner, acceptance path, prohibited changes, and minimum proof.
Use the same context, tools, time, Credits ceiling, and stopping rules.
Apply rubric, review blindly, repeat unstable tasks, and publish limits.
Repair one deterministic compile failure. Explain the first error, propose the smallest diff, and stop if evidence is missing.
Diagnose a client/server mismatch and name the authority boundary plus one reproducible correction test.
Review before/after captures, gameplay steps, and logs; return a failure matrix and rollback trigger.
Separate missing assets from module/config errors and propose one diagnostic change at a time.
Task IDs, revisions, inputs, tools, budgets, weights, and exclusions.
Prompts, responses, diffs, commands, timings, costs, failures, and notes.
Completion, compile, runtime, package, invention, regression, review-time, and reproducibility.
Approved uses, prohibited uses, monitoring, rerun date, and rollback owner.
SEELE AI can generate a native Unreal 5 game, preview it in-browser, optimize and package it, and provide a downloadable game or packaged build for external publishing or paid Seele games. Sales are not guaranteed.
The cited launch sources do not provide a public Unreal-specific benchmark.
Use task correctness proven by build, runtime, and package evidence.
Start with five representative failures, then expand across main systems and regressions.
Provide only relevant, authorized context and measure retrieval separately.
No. Repeat tasks, preserve failures, test recovery, and measure reviewer effort.
It can provide a browser prototype target while native scoring stays in Unreal.
Generate a native Unreal 5 game in SEELE AI, review the browser preview, then optimize, package, download, or publish the game. Keep licensing, target-platform testing, and third-party model claims as separate verification gates.
Paid download is optional for eligible SEELE AI outputs. Availability, demand, pricing, and revenue are not guaranteed.