WorldMark: A Unified Benchmark Suite for Interactive Video World Models

Xiaojie Xu1,2 Zhengyuan Lin1,2 Kang He1,3 Yukang Feng1,3 Xiaofeng Mao1 Yuanyang Yin1,3 Yongtao Ge1,† Kaipeng Zhang1,3,†

1Alaya Lab    2The University of Tokyo    3Shanghai Innovation Institute

Corresponding authors: yongtao.ge@shanda.com, kaipeng.zhang@shanda.com

Alaya Lab The University of Tokyo Shanghai Innovation Institute
Paper Code & Data World Model Arena
WorldMark overview: per-model adapters translate one shared action vocabulary into each model's native control format; a round-trip probe compares outbound and return views at equal accumulated motion.

An interactive world model is not only a video generator: it is an environment. You press a key, and the world is supposed to move that way, keep moving while you hold it, stop when you let go, and still be there when you turn back.

WorldMark is a benchmark that measures exactly that.

It drives ten heterogeneous models—caption-, pose-, keyboard-, and trajectory-controlled—from one shared WASD-style action vocabulary over 500 standardized cases, and scores the result with nine deterministic metrics covering action dynamics, world memory, and visual quality.

Action Dynamics Does the world go where it is told—and when?

Command Forth & Back  W→S  40s  ·  Look for whether the camera actually reverses when the command flips, or keeps advancing through both halves.

High Matrix-Game 3.0
Low Yume 1.5

Command Pan L & R  L→R  40s  ·  Look for whether the motion stays on the yaw axis, or leaks into strafing and forward drift.

High AlayaWorld
Low Yume 1.5

Command Strafe Right  D  20s  ·  Look for whether motion holds steady, or stutters, stalls, and decays to a standstill while the command still runs.

High Matrix-Game 2.0
Low SANA-WM

Command Forth & Back  W→S  40s  ·  Look for how soon the new command takes effect at the 20s switch—a step, or a ramp that arrives seconds late.

High Matrix-Game 2.0
Low SANA-WM

World Memory Does the world it generated stay the same?

Command Forward  W  20s  ·  Look for flicker, mutation, or an outright cut between adjacent moments.

High HY-GameCraft 1.0
Low LingBot-World

Command Pan L & R  L→R  40s  ·  Look for whether turning back returns you to the view you left, or to somewhere else entirely.

High Lyra 2.0
Low Matrix-Game 2.0

Command Pan L & R  L→R  40s  ·  Look for whether the scene layout survives the full sweep, or its structures rearrange as the camera passes—invisible in any single pair of frames.

High AlayaWorld
Low LingBot-World

Visual Quality How does each frame look?

Command Zigzag  A→D→A  60s  ·  Look for per-frame fidelity: sharpness and detail against haze, clutter, and low-level distortion.

High AlayaWorld
Low HY-GameCraft 1.0

Command Reverse + Pan  S→L  40s  ·  Look for aesthetic appeal: composition, colour, and lighting rather than pixel-level fidelity.

High Yume 1.5
Low SANA-WM

Benchmark Suite

Unified Action Interface

Models differ in what they accept—key presses, keyboard–mouse vectors, pose strings, camera trajectories, or directional keywords in a caption—and in how they internalize it: Plücker ray maps, PRoPE, cross-attention modules, or AdaLN-injected poses. A per-model action-mapping adapter translates the shared vocabulary into the native format along three paths:

Native keyboardModels with a keyboard interface take the primitives directly.
Matrix-Game 3.0, DreamX-World, …
Caption injectionText-driven models take motion types and durations written into the caption.
Yume 1.5, HY-GameCraft 1.0, …
Trajectory conversionCamera-controlled models take a per-frame unified trajectory integrated from the primitives.
HY-World 1.5, Lyra 2.0, …
WASD LR

Six primitives—forward, backward, strafe left/right, yaw left/right—each parameterized by duration, compose all 15 trajectories.

Image Suite — 50 scenes × 2 viewpoints

25 photorealistic and 25 stylized references drawn from the WorldScore pool of 1000 per domain, deduplicated to preserve its semantic diversity at a fortieth of its size. Each first-person reference is paired with a synthesized third-person view, giving 100 test images across three scene categories (Nature, City, Indoor) and styles from oil painting and Ukiyo-e to Minecraft.

All 100 test images: photorealistic above, stylized below, each scene as a first-/third-person pair
The complete Image Suite: all 100 test images, photorealistic above and stylized below. Each adjacent pair is one scene—the first-person reference on the left, its third-person counterpart on the right: the same camera with a character composited into frame, so a pair agrees everywhere except around the added figure.

Action Suite — 15 trajectories, 3 functional tiers

Easy · 20s, single-segmentTests basic compliance—all a model without command switching can complete.
Medium · 40s, two-segmentAdds exactly one switch, without which there is no transient; provides the first round trips.
Hard · 60s, three-segmentPatrol routes and 360° rotations run past most models' conditioning window, where revisit and global memory fail.

For each image, a VLM selects the five most plausible sequences from the library, identifying physical constraints—lateral obstacles in a corridor make prolonged strafing implausible—so models are not penalized for refusing to walk through a wall. Across 100 images this yields 500 evaluation cases.

The 15 standardized action sequences
The 15 standardized action sequences, from elementary translations and rotations to combined and cyclic trajectories.
Context-aware action selection via VLM reasoning
Context-aware action selection. A VLM analyzes the initial image to identify physical constraints and selects plausible action sequences from the predefined library.

Evaluation Dimensions — 9 metrics, 3 categories

Action Dynamics

Does motion go the right way, only that way, soon enough—and does it hold?

  • Direction Accuracytrans / rot
    Does the world move the commanded way?
  • Direction Puritytrans / rot
    Does it move only that way, without off-axis leakage?
  • Response Latencytrans / rot
    How long until a new command takes effect?
  • Motion Stabilitytrans / rot
    Does motion hold, or stutter, stall, and decay?

World Memory

Does the world stay the same—at three separate timescales?

  • Local Memory
    No flicker, mutation, or hard cuts between adjacent moments.
  • Global Memory
    Geometry stays self-consistent as drift accumulates over the full video.
  • Revisit Memory
    Round-trip probes: turn away and back—is the street still there? Frames paired by equal accumulated motion, not equal time.

Visual Quality

Does everything else reduce to appearance? (It does not.)

  • Perceptual Quality
    Frame fidelity from Q-Align logits—deterministic, no sampling.
  • Aesthetic Quality
    Aesthetic appeal from the same deterministic reader.

Quantitative Results

All metrics normalized to [0,100], higher is better. Bold = best, underline = second best. 125 videos per model per split.

Model Action Dynamics World Memory Visual Quality
Direction
Accuracy
Direction
Purity
Motion
Stability
Response
Latency
Local Global Revisit Perceptual Aesthetic
transrot transrot transrot transrot
AlayaWorld78.6792.3172.5590.3268.0598.0875.4896.7692.5375.7793.4385.9458.50
DreamX-World96.9176.6078.1870.3865.4127.7196.1097.0093.2934.0479.8971.1651.98
HY-GameCraft 1.089.072.7474.8263.0766.7613.7090.9751.9991.3136.5474.2064.5554.38
HY-World 1.592.5585.0581.1880.3131.6686.3985.4485.1394.1758.8887.5276.1556.59
LingBot-World89.1487.3275.4383.6479.1494.2686.5699.2092.2636.8579.3075.6556.49
Lyra 2.098.1685.4486.3286.1887.2180.3590.6286.1096.9165.3493.2684.0357.66
Matrix-Game 2.095.1280.6680.5775.9779.8978.8598.2386.4791.3140.5672.2271.1551.00
Matrix-Game 3.098.3778.3084.8373.3876.3996.7491.8396.3991.5250.6080.3869.6252.86
SANA-WM88.7483.0865.5178.5546.7983.3976.8397.4191.7053.4582.0576.7653.16
Yume 1.551.8559.4775.1652.5355.3579.4540.6677.1199.3655.2580.9687.2464.72
Model Action Dynamics World Memory Visual Quality
Direction
Accuracy
Direction
Purity
Motion
Stability
Response
Latency
Local Global Revisit Perceptual Aesthetic
transrot transrot transrot transrot
AlayaWorld71.8981.1775.7879.4867.1386.2955.5586.6987.6567.8289.9484.0163.83
DreamX-World96.5774.6277.6569.5657.2036.4394.0396.0290.8730.7176.8966.3850.37
HY-GameCraft 1.091.772.8372.6355.4649.2415.8589.3851.9784.9229.9469.5259.3450.07
HY-World 1.587.9185.0970.4177.3045.3785.4881.1581.9087.9752.8087.5773.8360.86
LingBot-World89.9883.2072.3179.8172.2290.2589.6799.3183.4330.8178.1370.4260.85
Lyra 2.095.4476.0080.0775.5669.2170.8887.0383.6283.5347.3687.1573.9157.21
Matrix-Game 2.092.1777.3474.1071.9167.0978.7294.5889.2682.8938.1370.0369.4249.40
Matrix-Game 3.097.1377.0880.5871.2159.5090.7687.7595.5479.0844.9177.1069.3552.75
SANA-WM86.2779.6263.1375.1153.4689.9682.2094.4278.1947.1777.5071.1250.98
Yume 1.554.2451.7772.9244.0956.7970.8034.5880.0096.3643.3476.9679.6861.29
Model Action Dynamics World Memory Visual Quality
Direction
Accuracy
Direction
Purity
Motion
Stability
Response
Latency
Local Global Revisit Perceptual Aesthetic
transrot transrot transrot transrot
DreamX-World97.2975.6476.4869.3061.9226.4995.9897.1593.6431.1679.4471.6052.87
HY-World 1.590.4982.1281.3477.7541.9287.8380.6984.8793.9255.9787.5576.5759.11
LingBot-World74.2255.7869.8562.8066.7664.4281.1574.3989.8947.0881.6876.6758.74
Matrix-Game 2.091.3446.5076.4939.8277.5975.2298.1187.7178.7631.6875.6467.2345.78
SANA-WM88.8983.1064.6078.3346.0382.5276.8497.3591.3252.2582.6275.4452.72
Model Action Dynamics World Memory Visual Quality
Direction
Accuracy
Direction
Purity
Motion
Stability
Response
Latency
Local Global Revisit Perceptual Aesthetic
transrot transrot transrot transrot
DreamX-World96.3975.0375.9369.7453.1536.3395.4594.5890.1229.9477.5966.2250.76
HY-World 1.593.0279.0772.0571.6550.1585.4679.2583.0389.2648.0485.4675.9264.05
LingBot-World69.9330.9065.9743.4959.6732.7478.1052.1787.5947.7176.8777.5268.03
Matrix-Game 2.093.9442.4276.9537.4974.0367.4686.7379.9072.9731.2875.2667.2945.98
SANA-WM87.4380.8463.9276.7954.2189.3079.7893.9175.7846.1776.3370.5951.63
Radar capability profile of each model over thirteen columns, Real vs Stylized overlaid
Model capability profile. First-Person Real (solid) against Stylized (dashed).

Key Findings

01Direction accuracy hides the differences that matter. Eight of ten models exceed 88 in translational direction accuracy—a solved problem, apparently. The other columns disagree: motion stability spans 31.7–87.2 and latency 40.7–98.2. Models indistinguishable by direction differ several-fold in how they get there.
02The fastest responders are the least stable. DreamX-World has the highest response latency score in the suite yet nearly the lowest rotational stability—a ~50-point gap. The trade-off recurs across fast responders; no single action metric can capture it.
03Command axes fail independently. HY-GameCraft 1.0 follows translation well (89.1) but is unresponsive to rotation (2.7). Milder ~20-point asymmetries appear in other models, and AlayaWorld inverts the pattern, with rotation ahead of translation.
04Memory needs its three timescales. Local memory is saturated (σ = 2.6) while global memory spreads over 34–76 (σ = 13.1). An averaged score would follow the saturated component—concealing that the smoothest frame-to-frame model yields the least coherent geometry.
05Third-person viewpoints attack rotation. Switching to third-person costs 13.9 points of rotational direction accuracy on average, against 4.0 translationally—and unevenly: Matrix-Game 2.0 loses 34.2 while SANA-WM is unaffected. Camera control around a visible character is a distinct capability.
06Stylization attacks memory, not action. From Real to Stylized, all ten models lose local and global memory, while action dynamics barely moves. Command following is largely style-agnostic; keeping a stylized world self-consistent is not.
07Visual quality is not a proxy for controllability. Perceptual quality correlates positively with all three memory metrics (ρ = 0.76–0.79) yet negatively with translational action dynamics, most strongly with response latency (ρ = −0.82). Yume 1.5 is the clearest case: first in perceptual quality and local memory, last in translational direction accuracy and latency.
08Difficulty breaks rotation, not translation. From Easy to Hard, rotational direction accuracy falls for eight of ten models while translational accuracy does not move (88.2 → 88.0); global memory loses 35% and is the only metric that collapses with difficulty. This also fixes 60s as the suite's end: a longer tier would rank models by the speed of their disintegration rather than expose a new failure.

World Model Arena

Beyond offline metrics: pit leading world models against each other in side-by-side battles and watch the live leaderboard.

Enter the Arena →

Citation

@article{xu2026worldmark,
  title={WorldMark: A Unified Benchmark Suite for Interactive Video World Models},
  author={Xu, Xiaojie and Lin, Zhengyuan and He, Kang and Feng, Yukang and Mao, Xiaofeng and Yin, Yuanyang and Ge, Yongtao and Zhang, Kaipeng},
  journal={arXiv preprint arXiv:2604.21686},
  year={2026}
}