A robot arm that learns by watching.
Guided through one task 160 times, then tested on a spot it had never seen.
- 160 demonstrations
- 2 policies
- 18/20 and 16/20 trials succeeded
- under $300
The 3D arm on this page is a reconstruction from the SO-101’s URDF. The real footage is below.
The setup: a sub-$300 arm.
The SO-101 is an open-source arm with 3D-printed links and six Feetech STS3215 servos. A person drives one arm by hand and a second arm copies it, and that’s how every demonstration was recorded.
≈ 1/100the cost of the research arms behind published tactile-VLA work
One job, and one rule.
The job: pick up a short piece of foam pool noodle and drop it in a yellow cup. A run counts only if the noodle ends up in the cup.
The noodle starts on a 3×3 grid of table positions, A1 to C3. Eight cells were used to teach the arm. The center, B2, was left empty on purpose.
Why leave one out?
An arm that only memorized its demonstrations can look perfect on the cells it practiced. B2 shows whether it learned the job itself.
How it learned: 160 demonstrations, one hand at a time.
Each one is a person driving the leader arm while the follower copies it, then resetting the scene by hand. Eight cells, exactly 20 demonstrations each, with the noodle rotated through ±90° in 10° steps.
Scroll to replay demonstrations 1–160 in recording order.
Episode 0 / 160
No episodes recorded yet
Replayed in the protocol’s recording order. Dots, path and tethers are illustrative, not recorded positions.
- episodes
- 160
- frames
- 67,693
- fps
- 30
B2 was never recorded.
Two cameras. At the grasp, one goes blind.
Each policy sees an overhead and a wrist camera, plus its own six joint positions. At the grasp the gripper hides the noodle from overhead, so the wrist view keeps it visible.
Two policies, the same 160 episodes.
About 52M parameters
- ResNet18 vision backbone
- Predicts 100 actions at a time
- 42,308 training steps
- Learning rate 1e-5, constant
About 450M parameters
- Pretrained vision-language model, vision encoder frozen
- Predicts 50 actions at a time
- 20,000 training steps
- Learning rate 1e-4, warmup then decay
Both trained on one RTX 3060; both run on Apple Silicon. The trail follows the reconstruction’s scripted path, sized to each policy’s chunk length (100 vs 50). It illustrates chunking; it is not policy output.
It never saw B2. It picked the noodle up anyway.
B2 appears in none of the 160 training episodes. So doing the job there shows the policy learned the task itself, rather than memorizing the moves it was shown. Both policies were tried there several times, and every attempt succeeded.
Four real runs, all successful
Filmed end to end, cell called out before each run: two in B2, two in cells the arm trained on. These are not the scored trials.
held outACT in B2 held outSmolVLA-450M in B2 trainedACT in A2 trainedSmolVLA-450M in C3
18 of 20. 16 of 20.
20 autonomous trials per policy, spread across the whole grid, B2 included. A trial counts only if the noodle ends up in the cup.
-
ACT
18 of 20
050100%90% estimate95% interval 68–99%
-
SmolVLA-450M
16 of 20
050100%80% estimate95% interval 56–94%
In plain words: 18 of 20 is consistent with a true success rate anywhere from 68% to 99%, and 16 of 20 with 56% to 94% (Clopper–Pearson intervals). They overlap heavily, so both work and 20 trials can’t say which is better.
Run on the physical arm and scored by hand. No per-trial log exists, so the dots show counts, not order.
What this doesn’t show.
- Nothing about touch. Both policies are vision-only; no tactile data was recorded, trained on, or evaluated.
- Not a ranking of ACT against SmolVLA. Separating 90% from 80% would take on the order of a hundred trials each.
- One task, one object, one lighting setup, and a held-out cell surrounded by trained ones. Nothing here tests reaching outside the grid.
What’s next: touch. Not built yet.
Touch is Phase 3, and it isn’t in either policy. Both are the vision-only baseline it has to beat. Every STS3215 servo reports its own load, which makes it a candidate touch signal with no added hardware.
- Shoulder pan92
- Shoulder lift152
- Elbow flex112
- Wrist flex48
- Wrist roll52
- Gripper56
Try it yourself: pick a cell.
Choose where the noodle starts and watch the arm do the job. Tap any cell, B2 included, click a tile in the scene, or click anywhere on the table to put the noodle there.
Choose a cell.
This is a kinematic simulation built from the SO-101’s URDF, not the trained policy. The real runs are above.
What I built.
- Hardware. An SO-101 leader and follower arm, with overhead and wrist cameras.
- Data. 160 teleoperated demonstrations, with one grid cell held out on purpose.
- Training. ACT and SmolVLA-450M on one RTX 3060, with LeRobot.
- Evaluation. 20 autonomous trials per policy on the real arm, plus the held-out cell.
- Tooling. Recording, training and evaluation scripts, and a test suite CI runs on every pull request.
It builds on the open-source SO-101 design, a pretrained SmolVLA and the LeRobot pipeline.