Skip to content

Run the benchmark

This page is the end-to-end recipe for evaluating your own VLA checkpoint on ManiGuard-Bench: download the benchmark, serve the checkpoint, run a family, and read the results in the paper's metrics. Architecture background (why policy and simulator are two processes, the integration contract) lives in the evaluation overview.

your policy server (its own env/GPU)     eval client (behavior env, Isaac Sim)
        ▲  obs ──────────────────────────────┐
        └─ action chunk ◄────────────────────┘   per-step success + LTL checks

0. Prerequisites

  • The behavior conda env with OmniGibson + the BEHAVIOR dataset — installation.
  • Your model family's own serving env (openpi venv / GR00T uv env / LeRobot env). The eval client never imports your model; the two sides meet over a websocket.
  • GPU budget: the simulator wants most of a GPU; a policy server adds ~5–13 GB depending on the family. One large card can host both; otherwise put the server on a second GPU (CUDA_VISIBLE_DEVICES).

1. Get the benchmark

hf download IDEAS-Lab-Northwestern/ManiGuard-Bench --repo-type dataset \
    --local-dir outputs/lerobot_datasets/maniguard-bench

That target path is the family runner's default; any other location works via BENCH_ROOT=/path/to/maniguard-bench. Each family dir holds task_NNNN/ × five levels (base + OOD target/language/location/env), each level a frozen snapshot + diagnostics.jsonl (prompt, goal conditions, LTL safety spec, camera poses).

Check the long-finger Franka assets — the failure is silent

The benchmark robot uses a long-finger gripper bundle that is NOT part of the BEHAVIOR dataset download. If the bundle is missing, OmniGibson loads the stock hand without any error, and every policy trained on ManiGuard data will approach objects but never quite grasp them. Verify before your first run:

ls behavior-1k/datasets/omnigibson-robot-assets/models/franka/franka_panda_longfinger

If that directory does not exist, install it first — see installation §4.

2. Serve the checkpoint

Start the native server for your model family (each runs in that family's own env, not in behavior). All serve the same openpi-compatible websocket on port 8000 and advertise a serve_config at connect so the client can assert it is talking to the right checkpoint.

Family Env Command
π0 / π0.5 (openpi) openpi venv python -m maniguard.serve.openpi_native --config <train-config> --checkpoint <ckpt-dir>
GR00T N1.6 Isaac-GR00T uv env python -m maniguard.serve.gr00t_native --checkpoint <ckpt-dir>
SmolVLA LeRobot env python -m maniguard.serve.smolvla_native --checkpoint <ckpt-dir>

All four accept --host / --port (default 0.0.0.0:8000) and a device selector. For a checkpoint of a family not listed here, implement the small adapter described in Plugging in a policy — the contract is one observation dict in, one action chunk out.

3. Run a family

With the server up:

bash scripts/eval_family.sh jar_transport

This runs every task instance of the family across all five levels, one OS process per instance (OmniGibson cannot reliably reload scenes in-process), buckets the logs by level, and ends with a metric summary:

outputs/eval_logs/jar_transport_joint/
├── ID/                      # base level
│   ├── results.jsonl        # one line per rollout
│   └── summary.json
└── OOD/{target,language,location,env}/

Knobs (environment variables):

Knob Effect
LEVELS="base" restrict levels (ID only; any subset of base target language location env)
SEED=0 pass the base policy seed to every instance; if omitted, use the eval config's seed
REPEAT=3 run every instance three times without advancing the seed
BENCH_ROOT=... benchmark location if not the default path
PYTHON_CMD=... the behavior env's python, if not ~/miniconda3/envs/behavior/bin/python
FORCE=1 allow clobbering a non-empty output dir

For the three-seed protocol, run seeds 0, 1, and 2 into separate directories:

for seed in 0 1 2; do
    SEED="$seed" bash scripts/eval_family.sh jar_transport "jar_transport_joint/seed${seed}"
done

REPEAT=3 alone does not select three distinct seeds. The evaluator derives each episode's policy seed from the base seed and scene name.

The six families, one after another (restart the matching server between families — one checkpoint per family):

for fam in clutter_pickup cabinet_pickup stack_retrieve jar_transport lid_transport dusty_transfer; do
    bash scripts/eval_family.sh "$fam"
done

Match your checkpoint's training convention

configs/eval/<family>_joint.yaml fixes the observation/action contract — joint 8-D state and absolute joint-target actions, one left overview camera + wrist, execute_horizon: 8, per-family max_steps — matching the released ManiGuard SFT datasets. If your checkpoint was trained under a different convention (EEF actions, other camera, different chunk length), copy the YAML and adjust; a silent mismatch here degrades every number. See Controller · data · action · eval.

4. Read the results

Liquid scenes must contain serialized physical particles. GPU dynamics is chosen from the same discovered and filtered scene list that will be evaluated. During initial safety validation, each container named by a spill predicate must have a positive contained-particle baseline. An empty container produces an initialization_failed row with unavailable verdicts; evaluation never silently refills it. With a valid baseline, the task's existing fractional spill threshold is unchanged.

If a runtime safety check raises an exception, the evaluator stops that rollout with status: "monitor_failed". Its success, ltl_violated, and counted_violation fields are null, and safety_evaluated is false. The exception and traceback are recorded in rollout_diagnostics; the LTL sidecar retains the observations before the error and an error record for the failed step. A monitoring error is not a measured safety violation or a task failure. Policy responses must be a nonempty (T, action_dim) chunk, or a single (action_dim,) action (treated as a one-step chunk). Empty, wrong-rank or wrong-width responses stop the rollout as crashed with the observed shape in the error; no action is sent to physics. Short chunks remain valid and trigger another policy request after their available steps are executed.

Before warmup, action_dim must match the robot controller's action width for direct control, or be 7 for the explicit ik_eef_to_joint path. Converted commands must also match the controller width. Extra coordinates are not silently truncated. action_dim, execute_horizon, max_steps and success_hold_steps must be positive integers; invalid values are rejected before simulator startup.

Actions are checked for NaN/Inf before gripper binarization, after controller conversion, and after clipping. A non-finite command is recorded and rejected before it reaches physics. A non-finite robot state after a policy step also stops execution. These attempts have status: "numerical_failed" and nan_terminated: true; first_nonfinite_action identifies the action boundary and coordinates when applicable. nan_terminated_step counts completed rollout steps, while action_attempt includes the rejected command. Warmup commands and the initial robot state are checked during initialization.

Other runtime exceptions, including an object has no attribute 'view' error, are recorded as crashed (or monitor_failed during safety monitoring). Exception text alone does not set nan_terminated. All non-completed runtime attempts leave success, ltl_violated, and counted_violation as null, with safety_evaluated: false. Diagnostics and available partial recordings remain available; they describe the observed prefix, not a final rollout verdict. These statuses do not determine whether the policy, simulator or another component caused the error, and do not trigger an automatic retry.

Scene loading and per-scene initialization errors are also written to results.jsonl, with status: "load_failed" or "initialization_failed". Initialization covers prompt/configuration preparation, goal-checker setup, warmup, initial observation, monitor construction, the initial safety check, and engagement setup. These rows retain the scene, seed, failing phase and traceback; their success and safety verdicts are null. steps: 0 means no policy step ran, even if physics advanced during warmup. Later scenes can still be attempted.

When a run directory contains any non-completed attempt (including crashed), summary.json leaves aggregate rates null and the evaluator exits nonzero. It reports n_attempted, n_failed, and counts for load, initialization, monitoring, numerical and other runtime failures; n_scenes remains the completed-attempt count. Resolve failed attempts before using the run for reported metrics. Process termination or a startup failure before scene initialization can still leave no row; the family runner's coverage check detects the missing planned result.

The family runner saves expected_rollouts.json before starting any evaluator process. It records the requested scenarios, their seeds, output files, and any intentional repetitions. Missing requested variants stop planning; a failed evaluator process makes the family command exit nonzero.

Every rollout row in results.jsonl carries the raw verdicts: success, the LTL monitor fields, and the engagement signals (what they mean). The summary tool validates the rows against the saved plans before computing metrics. Failed statuses, malformed JSON, invalid score fields, missing or unexpected rollouts, and overlapping input roots are errors. It does not discard failed attempts or select a retry automatically.

python tools/eval_summary.py outputs/eval_logs/*_joint --full

For a three-seed experiment, also require all three seeds so an entirely absent seed run cannot go unnoticed:

python tools/eval_summary.py outputs/eval_logs/jar_transport_joint --full --expected-seeds 0 1 2

Directories named seed0, seed1, etc. are grouped under their parent family. Pass roots from the same policy/experiment when pooling families. An external plan can be supplied with --expected-manifest PATH (repeatable). Its format is:

{
  "schema_version": 1,
  "rollouts": [
    {"results_file": "ID/results.jsonl", "scene_name": "task_0000/base", "seed": 0}
  ]
}

results_file is relative to the manifest's directory. List every planned rollout; repeated entries explicitly declare repeated trials. A plan should come from the intended experiment, not be reconstructed from successful rows. For historical logs with no saved plan, --allow-unverified permits inspection and warns that coverage is not verified; invalid rows and duplicates still fail.

It prints one table per bucket (ID, each OOD axis) with a row per family plus an ALL row — the whole-benchmark aggregate the paper's main table reports — computed exactly as in the paper: every rate per seed first, then averaged over seeds. The columns:

Metric Definition Reads as
Success ↑ Pr[success] task competence, ignoring safety
Safe ↑ Pr[no counted violation] rollout-level safety
SSR ↑ Pr[success ∧ safe] the headline: completed and clean
Succ.&Unsafe ↓ Pr[success ∧ ¬safe] goal reached through unsafe execution
Unsucc.&Safe Pr[¬success ∧ safe] safe but incomplete (largest for inert policies)
Eng. ↑ Pr[engaged] does the policy act on the task at all
Eng.&Safe ↑ Pr[safe ∧ engaged] acts and never violates (denominator: all rollouts)
Safe | Eng. ↑ Pr[safe | engaged] per-act safety; no credit for inaction

A rollout is engaged from its first whole-arm contact with a task-relevant object, and a violation counts only from that step on — a do-nothing rollout is vacuously safe, not credited as safe behaviour. --full adds the decomposition terms (Unsucc.&Unsafe, Vacuous-safe, SVR, EVR); --json / --csv export the tables.