SmolVLA SFT¶
SFT of HuggingFace LeRobot's SmolVLA on the same ManiGuard joint LeRobot v2.1 datasets used for openpi and GR00T — the dataset is model-agnostic; SmolVLA only differs in how the shared cameras/state/action are presented to it.
SmolVLA is LeRobot-native: it is fine-tuned with the upstream lerobot-train
CLI. The dataset uses standard observation/action keys; the training launcher
maps its two camera names to the pretrained policy's expected keys. Saved
processors carry this mapping and the dataset normalization statistics into
evaluation. The model pads state/action vectors to its internal width.
Why SmolVLA needs a data-prep (rename) step¶
The ManiGuard datagen export uses flat, non-standard keys (image_left,
wrist_image, state, actions, …) and ships 5 camera streams. openpi and
GR00T consume those flat keys through an indirection layer (openpi's
RepackTransform, GR00T's modality.json original_key). lerobot-train has no
such indirection — it classifies features purely by standard key prefix. So a
one-time prep step creates a 2-camera, standard-keyed v3.0 copy of the dataset:
| datagen (source) | → SmolVLA (standard) | note |
|---|---|---|
image_<external_cam> |
observation.images.top |
chosen overview (default left) |
wrist_image |
observation.images.wrist |
wrist |
state (8-D) |
observation.state |
absolute joint |
actions (8-D) |
action |
absolute joint target |
image_opposite / image_right / image_left_shoulder |
— | dropped |
actions_commanded |
— | dropped |
The preparation script copies the selected videos without re-encoding, renames Parquet columns and statistics without changing state/action values, and converts the copy from v2.1 to v3.0 with one episode per video file. The five-camera release remains v2.1. Choose a fresh output directory; existing output and converter intermediate directories are not overwritten.
Embodiment & schema¶
- Base model:
lerobot/smolvla_base(SmolVLM2 backbone + flow-matching action expert) - State / action: absolute joint 8-D (7 arm joints + 1 gripper), padded to SmolVLA's internal width; the model outputs absolute joint targets fed straight to a
JointControllerat eval (no delta transform — see end to end) - Cameras (2):
observation.images.top(overview) +observation.images.wrist - Tuning: SmolVLA default — vision encoder frozen, action expert trained (no LoRA, same "freeze VLM" strategy as the GR00T N1.6 path)
Tooling¶
| Purpose | Path |
|---|---|
| Embodiment contract (flat → standard key map, dims) | maniguard/smolvla_sft/embodiment.py |
| Dataset prep (5-cam flat → 2-cam standard-keyed copy) | tools/smolvla_sft/prepare_dataset.py |
SFT launcher (wraps lerobot-train) |
tools/smolvla_sft/run_sft.sh |
| Push checkpoint + card to HF | tools/smolvla_sft/push_to_hf.py |
| End-to-end 6-family driver | tools/smolvla_sft/run_all.sh |
Runtime setup¶
Use Python 3.12 and the pinned upstream LeRobot source (package version 0.5.1):
1396b9fab7aecddd10006c33c47a487ffdcb54b4.
The supplied runtime patch bounds video-decoder caching and applies the configured
tokenizer length when loading pretrained processors.
From the ManiGuard repository root, in a dedicated training environment:
MANIGUARD_ROOT="$(pwd)"
git clone https://github.com/huggingface/lerobot.git /path/to/lerobot
git -C /path/to/lerobot checkout 1396b9fab7aecddd10006c33c47a487ffdcb54b4
git -C /path/to/lerobot apply "$MANIGUARD_ROOT/tools/smolvla_sft/lerobot-runtime.patch"
python -m pip install -e '/path/to/lerobot[smolvla]'
FFmpeg must be available for dataset conversion. The default training video
backend is torchcodec; set VIDEO_BACKEND=pyav to use the PyAV backend.
Running¶
Prepare a training copy in the LeRobot environment:
python tools/smolvla_sft/prepare_dataset.py \
--src /path/to/datagen-clutter-v1-joint-5cam \
--out /path/to/clutter-smolvla-v3 \
--repo-id maniguard/clutter --external-cam left
Then launch training with the chosen schedule:
bash tools/smolvla_sft/run_sft.sh \
--dataset /path/to/clutter-smolvla-v3 --repo-id maniguard/clutter \
--output /path/to/checkpoints --steps 20000 --batch 64 --gpus 1 --exp-name clutter
--batch is per GPU. --gpus defaults to 1; larger values launch Accelerate DDP.
The launcher maps observation.images.top and observation.images.wrist to the
base model's camera1 and camera2 keys and saves that mapping with the processors.
State/action vectors remain eight-dimensional; state padding occurs inside the
model. BASE_MODEL can name a Hub model or a local checkpoint directory.
The public launcher uses HF_TOKEN for base-model access and WANDB_API_KEY for
its online training logs. Training saves locally with automatic Hub publication
disabled; publishing checkpoints is a separate operation.