ManiGuard: A Benchmark and Data Suite for Specification-Grounded Safety Evaluation and Improvement of Robotic Manipulation

1Northwestern University    2Stanford University    3William & Mary
*Equal contribution    †Project lead
arXiv Code Docs 🤗 Bench & Data Thread

ManiGuard pairs task-level LTLf specifications with 200 benchmark tasks (1,000 ID/OOD scenarios) and 8,000 safety-annotated demonstrations for evaluating and improving safe robotic manipulation.

Abstract

Foundation-model policies for robotic manipulation are advancing rapidly on task success, but rigorous evaluation of whether they succeed safely is still lacking. We introduce ManiGuard, a specification-grounded framework for evaluating and improving the safety of foundation-model manipulation, comprising the ManiGuard-Bench task suite for evaluation and a paired safety-annotated trajectory-generation pipeline for improvement. ManiGuard-Bench organizes six contact-rich household task families into 200 locked base tasks along a skill × constraint taxonomy, with safety specified independently of task success: success does not imply safety. Each task is evaluated under one in-distribution and four single-axis out-of-distribution perturbations that hold the safety specification fixed, giving 1,000 locked evaluation scenarios. Every simulated rollout is runtime-checked by LTLf-grounded automaton monitors over physics-grounded predicates rather than learned classifiers or LLM judges, in Isaac Sim / OmniGibson; the same specifications are scored on a physical Franka platform. The paired trajectory-generation pipeline supports an automated motion-planning generator and a human-teleoperation pathway, both safety-annotated by the same per-step LTLf monitor used for evaluation; it is family-agnostic and extensible, and directly supports safety-aware fine-tuning of policies toward safer behavior. We release 8,000 automatically generated, monitor-verified safe-success demonstrations, 40 for every base task. Benchmarking zero-shot and fine-tuned VLAs on ManiGuard-Bench across more than 23,000 evaluation rollouts, the experiments demonstrate that: (i) safety must be evaluated as a first-class objective independent of task success, as 6–21% of successful rollouts violate the specification and two policies with nearly identical task success differ by six points in safety violation rate; (ii) fine-tuning on our paired safety-annotated suite substantially improves safety, raising safe task completion from near zero to 7.5–29.8% and engaged-and-safe behavior from 16–40% to 51–72%; but (iii) a substantial gap remains that scaling the same demonstrations does not close, with 21–42% of engaged rollouts still violating φ, two of the six task families below 2% safe success for every policy, and these failures persisting under distribution shift and on hardware.

Framework

ManiGuard framework overview

Safety is specified independently of task success as LTLf constraints and compiled into automaton monitors over physics-grounded predicates — no learned classifiers, no LLM judges. The same monitor that scores evaluation rollouts also annotates the demonstration suite.

ManiGuard-Bench

1,000 scenarios = 200 base tasks × (1 ID + 4 single-axis OOD conditions)

ID and OOD structure of ManiGuard-Bench

One ID + four OOD conditions per task. Every base (in-distribution) task is perturbed along exactly one axis at a time — target appearance, environment background, object location, or instruction language — while the safety specification and its monitor stay fixed. A drop in safe-success under a single axis therefore isolates what the policy's safety depends on.

Safety-Annotated Data Generation

Data generation pipeline

ManiGuard supports both automated motion-planning generation and human teleoperation on a shared substrate. The released suite contains 8,000 automatically generated, monitor-verified safe-success demonstrations — 40 for each of the 200 base tasks.

Monitor-checked demonstrations across the six task families.

What 23,000+ Rollouts Show

I · Safety across the taxonomy

Task completion does NOT imply safety — and the raw safe rate does NOT measure it.

Both rollouts complete the task, but only one satisfies the task-level safety specification. The figures below quantify this unsafe-success outcome across policies.

ID safety evaluation: per-family SSR and overall outcome metrics
Outcome composition of ID rollouts, zero-shot vs SFT
II · OOD shift

Safe-success does NOT survive observation shift — each policy breaks along its own axis.

Safe-success generalization gaps under single-axis OOD shift
III · Demo scaling

More of the same demonstrations does NOT close the safety gap.

Demonstration scaling for pi0.5 on Clutter and Cabinet
IV · Specification format

Safety-specification format alone does not reliably steer behavior.

Effect of safety-specification format: Safe|Eng. under None, NL, and LTLf
V · Sim-to-real

Simulation predicts which policy is safer, NOT how unsafe it is.

Sim-to-real transfer

BibTeX

@misc{peng2026maniguard,
  title         = {{MANIGUARD}: A Benchmark and Data Suite for Specification-Grounded
                   Safety Evaluation and Improvement of Robotic Manipulation},
  author        = {Peng, Yiyan and Wang, Philip and Zhan, Simon Sinong and Lyu, Yiqi
                   and Ni, Zhenyang and Yan, Jixin and Wong, Fiorelli and Jiao, Ruochen
                   and Yin, Hang and Cao, Xinyu and Shao, Huajie and Li, Manling
                   and Zhang, Ruohan and Zhu, Qi},
  year          = {2026},
  eprint        = {2608.17386},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2608.17386},
}