Streaming 100 TB+ Robot Datasets

Affiliation

Hugging Face

Published

2026-07-23

Table of Contents

Training as datasets grow

Training a robot to grasp an object means learning from many attempts, viewpoints and situations. One robot with three cameras at 20 fps produces 216,000 images per hour. Across a fleet, demonstrations can grow to hundreds of terabytes. Now that robotics is scaling, how can we efficiently train on all this data?

LeRobot’s new episode-based streaming loader is designed for workloads of 100 TB and beyond. It reads selected episodes directly from a Hugging Face dataset or Bucket and prepares their samples for training. This means the dataset can keep growing without requiring a complete local video copy.

What a streaming loader needs

Avoiding downloading is only useful if training still works well. We need varied batches, and we need them ready when the model asks for them. That gives us four requirements:

The design at a glance

The idea is to work with a few episodes at a time. StreamingLeRobotDataset divides complete episodes between training processes, reads their table rows, and keeps a bounded pool of compressed video bytes in RAM. From this pool, it builds samples around selected frames, which we call anchors. Each assigned anchor is visited once per streaming epoch.

An MP4 index tells us where to find the video bytes. We fetch them ahead of time, mix samples from the active episodes, and decode images while the model trains on earlier batches. This lets us spread network requests over many samples while still giving the model varied data. Let’s look at how those pieces fit together.

Remote data, local training
Remote storage
Hugging FaceDataset repository or Bucket
Full dataset MP4 video · Parquet tables
HTTP
range reads
Selected
episode data
Your training machineOne GPU shown
IndexLocate episode bytes
SamplerChoose & shuffle samples
RAM

Episode pool

ABC

Compressed video
+ table rows

CPU workers

Prepare

Decode video
Build ready batches

GPU

Train

Use batches
in planned order

Prepare the next batch while this one trains
Only selected episode data crosses the network. Each GPU has its own episode pool and sample order.

Finding the data for a training sample

To build a training sample, we need to bring two kinds of data together. LeRobot stores camera observations in compressed video, while state, actions, timestamps and language live in tables. A sample might combine past camera frames with future actions, so the loader reads both from the same episode and lines them up in time.

Why read more than one frame?

Fetching one frame at a time sounds efficient. But to decode it, we usually need to start from an earlier keyframe and work forward. If we fetch nearby frames separately, we can end up reading and decoding the same data again. Reading an episode’s compressed video range lets us reuse those bytes across many samples.

Each network request also takes time, roughly request latency + bytes / bandwidth. By reading an episode at once, we spread that waiting time over many samples. The video stays compressed in memory until we need its images.

Finding an episode in a file

A training sample uses one or more frames and their associated data. Reading by episode helps us prepare many such samples efficiently. However, an episode does not always have its own file: one Parquet or MP4 file can contain several episodes. We use the dataset’s metadata to find where the selected episode begins and ends.

Two formats, one episode
Storage format
Choose an episode
01 · Remote file

Episode rows

One file, several episodes
02 · Locate

Row-group statistics

03 · Local RAM

Episode A's rows

Aligned groups shown; mixed groups may require broader reads.

Choose a format and an episode. Follow the highlighted data from remote storage into RAM.

For tables, Parquet’s statistics help us skip unrelated row groups. We read the non-video columns and keep the selected episode’s rows.

For video, an MP4 sidecar index locates frames and keyframes. We fetch the episode’s bytes from a usable keyframe and add a header, creating a self-contained mini-MP4 in RAM. The decoder reads it like a local video.

With both parts available, we assemble camera history and future actions, padding requests at episode boundaries. Asking for more future actions reads table values without decoding extra images.

Preparing the video index

We need an index that matches the videos. On first use, LeRobot checks its local cache, then looks for a published sidecar. If neither is valid, it builds one locally. A lock prevents competing builds, and only a complete, validated index is installed.

To share the index between training processes, LeRobot converts it once into a read-only, memory-mapped file. Processes using the same cache share its data in memory and prepare entries only for their episodes’ video files.

To keep the index in sync, we pin a dataset version for repositories and check content hashes for Hugging Face Buckets. Keep the source unchanged during training.

Turning episodes into varied batches

Now that we can read an episode, how should we choose its samples? Nearby frames show similar scenes, while shuffling the entire dataset scatters reads across many files. We balance these needs by sampling from a bounded pool of active episodes, reusing their video bytes in memory.

Each sample starts from an anchor frame, with history and future data selected around it. We shuffle each episode’s anchors and use each once, then replace the finished episode. Windows can overlap, so the once-per-epoch guarantee applies to anchors.

By default, an episode’s chance of being chosen follows its remaining anchors. With 3, 2 and 4 left, the probabilities are 3/9, 2/9 and 4/9. Every remaining anchor in the pool has an equal chance; other episodes enter as space becomes available.

Sample, finish, replace
Illustrative epoch · 6 episodes · 18 anchorsSample. Finish. Replace.ready to sample
Active pool3 / 3 slots
Available anchor Already sampled
Waiting to enter
Sample order0 / 18 anchors
Each draw uses one anchor. A finished episode frees a slot.
Follow 18 anchors through a three-slot pool. Each label appears once in the sample order.

You can also choose shuffled round-robin, taking one anchor from each active episode in a shuffled order per round. Episodes appear more evenly in batches. Both strategies cover every anchor, but neither gives a uniform dataset-wide shuffle or has established convergence benefits. The measurements below use the default.

Preparing ahead and delivering in order

While the model trains on one batch, workers fetch and decode upcoming samples. Ready samples wait in a bounded queue and leave in planned order. A later sample waits its turn even if it finishes first, so slow reads delay the sequence without skipping samples.

Multi-worker support lets --num_workers set several DataLoader processes per rank. Each owns separate episodes and a pool, sharing the rank’s video-byte budget. More processes can use more CPU cores, but add memory overhead. Each batch draws from one worker’s pool, making pool size important for variety.

Balancing memory and throughput

How much data should we keep ready? We hold compressed episode video in RAM and decode images as needed. Reusing those bytes supports shuffled samples and overlapping time windows within a fixed video budget.

What limits speed?

To make training faster, look at where it waits: fetching, sample preparation or model compute. The slowest sets the pace. If decoding is slower than fetching, more network connections will not help.

Compare these stages in training samples per second. For the network, divide bandwidth by average bytes fetched per delivered sample, including cameras, tables and prefetch. Try changing each stage’s capacity below to see which sets the limit.

Find the bottleneck
Illustrative inputs, not measured network capacity · one training process
Try a limit:
1 · Fetch data 256 samples/s
2 · Prepare samples 200 samples/s
3 · Train the model 120 samples/s
Best possible training speed ≤ 120 samples/s

Training the model is the slowest step. Speeding up another step will not help.

Fetching data: 64 × 1024 ÷ 256 = 256 samples/s.

An illustrative upper bound, not a benchmark. Real throughput can be lower.

A queue of ready samples gives us a little breathing room when a read takes longer than usual. At 200 samples/s, 20 ready samples cover 100 ms. That helps with short delays, but the queue will eventually empty if the loader consistently produces samples slower than the model uses them.

What should you adjust?

Once we know what sets the pace, we can adjust three controls:

Does it keep up with training?

The practical question is whether the loader can keep a training job fed for hours. We tested this by training π0.5 for eight hours on four H100 GPUs, streaming from RealOmni, which contains about 10.7 TB of source data. The model used recent movement to predict future end-effector positions, so the test exercised the loader as part of a real training loop.

During the run, the model consumed more than a million samples at an average of 42.5 samples/s per node. It spent just 0.52% of step time waiting for data. Training continued through regular checkpoints, and restart checks returned the expected next samples. The loader was supplying data quickly enough for this model to keep training.

To see how much data the loader could prepare, three-minute multi-worker tests used two public datasets on an 8-vCPU machine. With four workers, the two-camera SO-101 pick-and-place dataset reached 607 samples/s, up from 283 with one worker. The three-camera R3 dataset reached 1,130–1,354 samples/s, around 2.3–2.5× faster in matched runs.

What the training run shows is that we could train for hours while reading the data remotely, without first downloading the video dataset.

Start streaming

Ready to put your robot data to work? Choose a dataset and a policy, then add --dataset.streaming=true to your training command. LeRobot will prepare the samples as training needs them, so you can start without downloading all the videos first:

lerobot-train \
  --policy.path=<policy-repo> \
  --dataset.repo_id=<owner>/<dataset> \
  --dataset.streaming=true \
  --output_dir=outputs/streaming

If your data lives in a Bucket, use its name for --dataset.repo_id=<owner>/<bucket> and add --dataset.repo_type=bucket. The Bucket should contain the dataset’s metadata as well as its data files.

Want each active episode to contribute a sample before revisiting any of them? Add --dataset.streaming_sampling_strategy=round_robin to try shuffled round-robin sampling.

Try streaming on your next training run and share what you learn with the LeRobot community. What will you train next?

Streaming is available on LeRobot’s main branch. See the streaming documentation to get started.