HyperAIHyperAI

Command Palette

Search for a command to run...

RECURSIVE GAME CREATOR: AN AGENTIC PRODUCT-LEVEL EXPERIENCE-ORIENTED GAME HARNESS

Jiajun Chen Haoyu Wu Mingda Jia Xihui Liu

Abstract

Recent game design agents have made substantial progress in generating playable games. However, program correctness does not ensure an enjoyable experience for players. We present Recursive Game Creator, an experience-oriented harness to advance agentic game development from rough game prototypes into entertaining games. Recursive Game Creator organizes recursive development around four components: Designer, Builder, Player, and Reviewer. The Designer translates user instructions and Reviewer’s feedback into detailed plans. The Builder turns these plans into candidate games. The coding-native Player creates and executes reusable policies through programmatic interfaces to efficiently collect diverse gameplay trajectories, mitigating evaluation bias caused by slow GUI-based collection. The Reviewer uses carefully designed trajectory-based metrics to induce player preferences, integrating with visual evidence and explicit textual preferences to evaluate games against game-specific criteria. Finally, the Reviewer accepts the better version and provides improvement reviews for the next round, closing the recursive loop. Our method achieves stateof-the-art overall performance of 77.89 on GameCraft-Bench. On GameASG-Bench, it achieves a strict task success rate of 53.2%, a 34.1% improvement over the same-model baseline, and the highest mean runtime-check pass rate at 93.4% among compared methods. A user study shows longer playtime and higher ratings. Code is coming soon.

One-sentence Summary

Researchers from HKU MMLab at The University of Hong Kong and Shenzhen Loop Area Institute propose Recursive Game Creator, an experience-oriented agentic harness that iteratively refines games through Designer, Builder, coding-native Player, and Reviewer components, achieving a state-of-the-art overall score of 77.8977.8977.89 on GameCraft-Bench and, on GameASG-Bench, a strict task success rate of 53.2%53.2\%53.2% with the highest mean runtime-check pass rate of 93.4%93.4\%93.4% among compared methods.

Key Contributions

  • The paper introduces Recursive Game Creator, an experience-oriented recursive harness that coordinates a Designer, Builder, Player, and Reviewer to evolve rough game prototypes into more entertaining games.
  • The coding-native Player creates and executes reusable policies through programmatic interfaces, enabling efficient collection of diverse gameplay trajectories and reducing evaluation bias caused by slow GUI-based play.
  • The Reviewer uses trajectory-based metrics to induce player preferences, integrates visual evidence and explicit textual preferences, and selects the better version while producing improvement reviews for the next round; evaluations show a 77.89 overall score on GameCraft-Bench, a 53.2% strict task success rate on GameASG-Bench with a 34.1% improvement over the same-model baseline, and the highest mean runtime-check pass rate of 93.4%.

Introduction

The authors address a key gap in agentic game development: coding agents can produce runnable games, but successful execution does not measure whether the gameplay is engaging, well-paced, coherent, or responsive. Prior iterative workflows often rely on GUI-driven testing, which can miss structured state, legal actions, events, and progress signals, while visual decision-making adds latency and limits frequent, fine-grained evaluation. The authors introduce Recursive Game Creator, an experience-oriented recursive harness that simulates a game studio with Designer, Builder, Coding-Native Player, and Experience-Oriented Reviewer roles. It collects programmatic gameplay trajectories, performs game-specific experience review, and converts preference-informed feedback into multi-round game refinement, reaching state-of-the-art results on GameCraft-Bench and GameASG-Bench.

Method

Recursive Game Creator: A Closed-Loop Experience Refinement Framework

The authors introduce Recursive Game Creator as a closed-loop system that connects game construction, policy-based play, experience assessment, and iterative refinement. The loop consists of four cooperating roles: the Designer, Builder, Player, and Reviewer. The Designer converts user requirements and prior evaluation feedback into a structured plan. The Builder implements that plan as a playable candidate. The Player executes code-based policies against the candidate to collect behavioral trajectories and visual records. The Reviewer then evaluates those trajectories and observations against experience criteria, and its findings guide the next design plan.

Designer and Builder

The Designer expands the user brief into a concrete specification covering the core play loop, mechanics, progression and difficulty, visual direction, and production priorities. It combines the current game state and recent feedback with the user requirements to produce a plan of the form

Pt=Design(U,Gt,Ht−1)=(ht,Δt,At,Xt),P _ {t} = \mathrm{Design} (U, G _ {t}, \mathcal {H} _ {t - 1}) = (h _ {t}, \Delta_ {t}, A _ {t}, X _ {t}),Pt​=Design(U,Gt​,Ht−1​)=(ht​,Δt​,At​,Xt​),

where UUU is the user brief, GtG _ {t}Gt​ is the current game, and Ht−1\mathcal {H} _ {t - 1}Ht−1​ is recent feedback when available. The hypothesis hth _ {t}ht​ states the intended experiential improvement and its rationale. The change specification Δt\Delta _ {t}Δt​ describes concrete modifications and strengths to preserve. The acceptance goals AtA _ {t}At​ define target scenes, player inputs, and observable outcomes. The asset requests XtX _ {t}Xt​ list required production resources. This representation turns experience goals into implementation decisions and testable objectives.

The Builder inherits the Designer's plan and selects implementation tools according to the game mechanics, interaction requirements, visual direction, and asset needs. Implementation proceeds from a playable draft to an integrated candidate:

Zt=Draft⁡(Gt,Δt),Xt∗=Generate⁡(Xt;Zt),Ct=Integrate⁡(Zt,Xt∗),Z _ {t} = \operatorname{Draft} (G _ {t}, \Delta_ {t}), \qquad X _ {t} ^ {*} = \operatorname{Generate} (X _ {t}; Z _ {t}), \qquad C _ {t} = \operatorname{Integrate} (Z _ {t}, X _ {t} ^ {*}),Zt​=Draft(Gt​,Δt​),Xt∗​=Generate(Xt​;Zt​),Ct​=Integrate(Zt​,Xt∗​),

where ZtZ _ {t}Zt​ is the draft, XtX _ {t}Xt​ denotes requested assets, Xt∗X _ {t} ^ {*}Xt∗​ contains assets actually generated by the framework, and CtC _ {t}Ct​ is the integrated candidate. During drafting, the Builder implements the core loop and planned changes, using placeholders when needed. Generated assets are then integrated and checked in gameplay scenes, including their interaction with collision, animation, and interaction regions. Targeted self-checks and configured project checks complement the broader play evidence collected later by the Player. The workflow explicitly links asset production to in-game validation because generating an asset alone does not establish an improvement to play.

Coding-Native Player

The Player constructs executable gameplay policies and separates policy generation from repeated interaction. Each policy reads the game's exposed state and chooses valid actions, allowing state-based control without requiring a language-model response or visual interpretation at every action. For a candidate version CtC _ {t}Ct​ at round ttt, a policy πt,k\pi _ {t, k}πt,k​ produces a trajectory τt,k\tau _ {t, k}τt,k​ through

τt,k=Rollout(Ct,πt,k),Et={(τt,k,Vt,k)}k=1K,\tau_ {t, k} = \text {Rollout} (C _ {t}, \pi_ {t, k}), \quad \mathcal {E} _ {t} = \{(\tau_ {t, k}, V _ {t, k}) \} _ {k = 1} ^ {K},τt,k​=Rollout(Ct​,πt,k​),Et​={(τt,k​,Vt,k​)}k=1K​,

where Vt,kV _ {t, k}Vt,k​ denotes visual records when available, and the evidence set Et\mathcal {E} _ {t}Et​ contains records from KKK rollouts.

Policy construction begins by reading the state schema, available actions, and testing objectives. The Player writes a policy that observes states, chooses actions, and repeats until the task ends or the rollout budget is reached. Policies run through the game's programmatic or command-line interface, and rollout requests are dispatched as execution jobs so that repeated trials and separate game instances can be used. The Player inspects outcomes and execution failures, revising policy code when necessary. A conversational response is not used as a substitute for executing the rollout.

Policies vary in strategy, skill, exploration, and risk tolerance. One policy may favor fast progress while another explores optional content. The resulting trajectories record exposed states, actions, events, progress, elapsed time, and termination outcomes. Parallel or repeated execution supplies multiple attempts and failure cases under known policy settings. Policy and observation settings are recorded with each trajectory so that comparisons across game versions remain controlled.

The trajectories support estimates of completion, difficulty, explored content, and points of failure or stagnation. Direct state and event access avoids reconstructing these variables from every rendered frame. Lightweight execution and parallel sampling make repeated trials practical, and diverse strategies can expose failures that a single strategy misses. The total testing cost includes policy generation, repair, execution, and review. Proxy strategies provide varied test behavior, but the authors note that they do not establish representation of the full human player population. Human trajectories and explicit feedback supply additional evidence about individual preferences.

Experience-Oriented Reviewer

The Reviewer combines behavioral trajectories, visual observations, and user input into an experience assessment and actionable revision feedback. Its inputs are tied to the evaluated game version and policy or player session, and the assessment preserves the distinction between observed behavior, inferred preference, and explicit user requests.

The Player supplies many trajectories from policies with different skills and strategies. The Reviewer uses this evidence to estimate success rates and examine where play ends or gets stuck. The objective is not the highest possible success rate but a suitable difficulty level for the intended players. Success that comes too easily may offer little challenge, while repeated failure may cause frustration. Comparing outcomes across policies helps assess whether progress is achievable and whether stronger play is rewarded. For version comparisons, policy settings and observation access must remain consistent.

The same trajectories record playtime. The Reviewer examines run times together with progress, explored content, and reasons for stopping. This helps distinguish a long run with varied activity from one spent stuck at the same point. Success rate and playtime support a broader assessment of difficulty and playability, but both remain measures of the tested policies. Proxy playtime alone does not show human interest or enjoyment.

The Reviewer also uses a game-specific rubric. It starts from shared assessment dimensions such as control responsiveness, content richness, narrative or progression pacing, visual coherence, and support for exploration or replay. It then instantiates these dimensions as criteria suited to the game's genre and intended experience. For example, a racing game may require clear track guidance, responsive steering, and visible collision feedback, while a visual novel may require coherent dialogue and choices with discernible consequences. The Reviewer checks consistency between rules, visual cues, and observed play, and explains strengths, weaknesses, and tradeoffs rather than averaging unrelated criteria.

Visual evidence is handled separately from the fast control loop. The Reviewer obtains sampled screenshots and available recordings from rendered gameplay and examines them alongside play reports without seeing source code or version order. Trajectories show where a policy succeeds, fails, or stops progressing, while rendered views show what a player could see at those moments. The Reviewer uses this evidence to assess text readability, visual style, and action feedback. Each observation is tied to its game version, policy, and replay, and the Reviewer separates what happened from why it may have happened. A failed turn could reflect poor controls, an unclear cue, or a weak policy. When the cause is uncertain, the feedback specifies what further evidence would help resolve it.

The Reviewer translates trajectory patterns, visual findings, and explicit user input into a structured preference record. The record describes the desired experience, supporting evidence, whether a preference is inferred or user-stated, and the corresponding revision priority. For instance, a tendency to explore optional areas may suggest interest in discovery, while a user request for less demanding combat supplies an explicit difficulty target. These observations inform design hypotheses rather than establishing preferences from playtime alone.

For version comparison and retention, the Reviewer compares a candidate with the retained game using the same shared signals and game-specific criteria. It checks whether the changes address the earlier revision goals and introduces new problems. The output is an A/B preference, a tie, or an unavailable judgment, with reasons and evidence paths. The feedback states the problem, its possible cause, its effect on play, and a revision goal with a follow-up check. The harness then maps the Reviewer's judgment back to the game versions and selects which artifact to retain.

Experiment

The evaluation covers GameCraft-Bench for complete Godot games, GameASG-Bench for requirement compliance, qualitative multi-round analyses, player rollout coverage, and a user study on experience-oriented customization. Recursive Game Creator improves overall game quality across three refinement rounds, outperforming the same-model baseline and the strongest reported baseline, with gains in all five quality categories. Qualitative examples show revisions expanding content and clarifying game state, available actions, and feedback across action, visual novel, and arena games. The user study reports clear improvements in controls, playability, depth, and art, along with increased average playtime.

The recursive approach improves overall GameCraft-Bench score from its first to third refinement round, moving from slightly above the same-model baseline to clearly ahead of all reported baselines. All five evaluated categories improve, with action games showing the largest relative gain but requiring more rounds to surpass the baseline action score. The iterative baseline also improves with additional rounds, but its final score remains below the recursive method's final result. Recursive refinement improves overall quality across rounds, and all five game categories gain over the three rounds. Action shows the largest category improvement, though its early rounds stay below the same-model baseline before surpassing it in the final round. The final overall score exceeds the strongest reported baseline, while the multi-round baseline also improves with additional refinement rounds.

GameASG-Bench separates source-code checks from browser-based gameplay checks, and L1 source-code pass rates stay high across systems while L2 gameplay pass rates vary more. Recursive Game Creator completes 25 of 47 tasks, slightly below the top baseline on task success but above all reported configurations on mean L2 and extended-feature pass rates. Lower-ranked baselines show progressively weaker L2 performance, particularly on required and extended gameplay checks. Recursive Game Creator achieves the highest mean L2 and L2 P2 pass rates among reported configurations, with task success just behind the leading baseline. Across baselines, task success falls substantially as L2 P1 and L2 P2 pass rates decline, while L1 source-code scores remain consistently high. Startup and interface checks are generally strong, with the top task-success baseline and Recursive Game Creator both reaching 99.0% on L2 P0.

Across refinement rounds, expert ratings for controls, playability, depth, and art all increased, with controls and art moving from low to high scores. Average playtime also rose from under two minutes to over twelve minutes, indicating more substantial playable content. The improvements span interaction, content, and presentation. Controls, playability, and art showed the largest gains across rounds, while depth improved more modestly. Average playtime rose from under two minutes to over twelve minutes, paralleling the rating improvements.

The experiments evaluate recursive refinement across three complementary settings. In GameCraft-Bench, recursive refinement improves overall score across rounds and surpasses all reported baselines, with all five game categories improving and action games showing the largest late gains. In GameASG-Bench, source-code checks remain strong across systems while gameplay checks vary more; Recursive Game Creator trails the top baseline slightly on task success but leads on L2 and extended-feature pass rates, with startup and interface checks generally robust. Expert ratings further show recursive refinement improving controls, playability, depth, and art, while average playtime rises from under two minutes to over twelve minutes, indicating more substantial playable content.


Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp