Nvidia · Hugging Face
LeRobot v0.6.0: Imagine, Evaluate, Improve
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
★ Tier-1 Source
LeRobot v0.6.0 introduces world model policies (VLA-JEPA, FastWAM, LingBot-VA) that learn to imagine the future, a wave of new VLAs (GR00T N1.7, MolmoAct2, EO-1, EVO1, Multitask DiT), and a new reward models API (Robometer, TOPReward).
Key facts
- LeRobot v0.6.0 introduces world model policies (VLA-JEPA, FastWAM, LingBot-VA) that learn to imagine the future, a wave of new VLAs (GR00T N1.7, MolmoAct2, EO-1, EVO1, Multitask DiT), and a new
- Inference fits in ~12 GB at bf16, and LoRA fine-tuning fits on a single 24 GB GPU
- EO-1, a VLA pretrained upstream on interleaved vision-text-action data, joins LeRobot: a Qwen2.5-VL-3B backbone with a flow-matching action head, contributed by one of the paper's own authors
- The Multitask Diffusion Transformer policy brings the TRI Large Behavior Models recipe to LeRobot: a ~450M-parameter diffusion transformer conditioned on CLIP vision and language embeddings, so one
Summary
LeRobot v0.6.0: Imagine, Evaluate, Improve TL;DR Table of contents World models: policies that imagine VLA-JEPA LingBot-VA FastWAM VLAs: the model zoo keeps growing GR00T N1.7 MolmoAct2 EO-1 Multitask DiT EVO1 Reward models: knowing when your robot succeeds Robometer TOPReward Datasets: faster loading, richer data Your codec, your rules Depth support, end to end Language annotations at scale Up to 2x faster data loading Benchmarks: one CLI to evaluate them all Training & inference lerobot-rollout: deployment gets its own CLI FSDP: train models bigger than your GPU Cloud training with HF Jobs Codebase: leaner and cleaner Community & ecosystem Final thoughts. Your codec, your rules. Depth support, end to end.