Expert verdicts on what AI builds.

AI has made it easy to generate worlds, games, and systems. It still can't tell you whether what it built is any good.

Quality doesn't show up in a unit test. It only registers in expert judgment.

We build instrumented worlds where models create and experts judge — so what AI builds isn't just functional, but good.

Private evals. Preference datasets.
A public benchmark.

Every score traces back to a human who played the build.

Ship things people love?
Judge what AI builds next.

Get paid to decide what AI gets right and wrong.

The playtest is the eval.