TactileVLA-Edge

Loading

3D: kinematic reconstruction, not policy output

TactileVLA-Edge

A robot arm that learns by watching.

Guided through one task 160 times, then tested on a spot it had never seen.

  • 160 demonstrations
  • 2 policies
  • 18/20 and 16/20 trials succeeded
  • under $300
How it works GitHub
Real footage: the arm working on its own

The 3D arm on this page is a reconstruction from the SO-101’s URDF. The real footage is below.

The setup: a sub-$300 arm.

The SO-101 is an open-source arm with 3D-printed links and six Feetech STS3215 servos. A person drives one arm by hand and a second arm copies it, and that’s how every demonstration was recorded.

≈ 1/100the cost of the research arms behind published tactile-VLA work

One job, and one rule.

The job: pick up a short piece of foam pool noodle and drop it in a yellow cup. A run counts only if the noodle ends up in the cup.

The noodle starts on a 3×3 grid of table positions, A1 to C3. Eight cells were used to teach the arm. The center, B2, was left empty on purpose.

Why leave one out?

An arm that only memorized its demonstrations can look perfect on the cells it practiced. B2 shows whether it learned the job itself.

How it learned: 160 demonstrations, one hand at a time.

Each one is a person driving the leader arm while the follower copies it, then resetting the scene by hand. Eight cells, exactly 20 demonstrations each, with the noodle rotated through ±90° in 10° steps.

Scroll to replay demonstrations 1–160 in recording order.

Episode 0 / 160

No episodes recorded yet

Replayed in the protocol’s recording order. Dots, path and tethers are illustrative, not recorded positions.

episodes
160
frames
67,693
fps
30

B2 was never recorded.

Time-lapse of recording the demonstrations

Two cameras. At the grasp, one goes blind.

Each policy sees an overhead and a wrist camera, plus its own six joint positions. At the grasp the gripper hides the noodle from overhead, so the wrist view keeps it visible.

overhead 800×600 wrist 640×480
The actual observation space, from episode 0 of the training set.

Two policies, the same 160 episodes.

About 52M parameters

  • ResNet18 vision backbone
  • Predicts 100 actions at a time
  • 42,308 training steps
  • Learning rate 1e-5, constant

Both trained on one RTX 3060; both run on Apple Silicon. The trail follows the reconstruction’s scripted path, sized to each policy’s chunk length (100 vs 50). It illustrates chunking; it is not policy output.

It never saw B2. It picked the noodle up anyway.

B2 appears in none of the 160 training episodes. So doing the job there shows the policy learned the task itself, rather than memorizing the moves it was shown. Both policies were tried there several times, and every attempt succeeded.

Four real runs, all successful

Filmed end to end, cell called out before each run: two in B2, two in cells the arm trained on. These are not the scored trials.

  • held outACT in B2
  • held outSmolVLA-450M in B2
  • trainedACT in A2
  • trainedSmolVLA-450M in C3

18 of 20. 16 of 20.

20 autonomous trials per policy, spread across the whole grid, B2 included. A trial counts only if the noodle ends up in the cup.

  1. ACT

    18 of 20

  2. SmolVLA-450M

    16 of 20

In plain words: 18 of 20 is consistent with a true success rate anywhere from 68% to 99%, and 16 of 20 with 56% to 94% (Clopper–Pearson intervals). They overlap heavily, so both work and 20 trials can’t say which is better.

Run on the physical arm and scored by hand. No per-trial log exists, so the dots show counts, not order.

What this doesn’t show.

  • Nothing about touch. Both policies are vision-only; no tactile data was recorded, trained on, or evaluated.
  • Not a ranking of ACT against SmolVLA. Separating 90% from 80% would take on the order of a hundred trials each.
  • One task, one object, one lighting setup, and a held-out cell surrounded by trained ones. Nothing here tests reaching outside the grid.

What’s next: touch. Not built yet.

Touch is Phase 3, and it isn’t in either policy. Both are the vision-only baseline it has to beat. Every STS3215 servo reports its own load, which makes it a candidate touch signal with no added hardware.

Peak servo load under hand pressure, bench test (raw register units)
  • Shoulder pan92
  • Shoulder lift152
  • Elbow flex112
  • Wrist flex48
  • Wrist roll52
  • Gripper56

Single presses; values include pose-dependent static torque, so joints aren’t directly comparable. The glow on the 3D servos is drawn from these numbers. It isn’t live sensing.

Try it yourself: pick a cell.

Choose where the noodle starts and watch the arm do the job. Tap any cell, B2 included, click a tile in the scene, or click anywhere on the table to put the noodle there.

Choose a cell.

This is a kinematic simulation built from the SO-101’s URDF, not the trained policy. The real runs are above.

What I built.

  • Hardware. An SO-101 leader and follower arm, with overhead and wrist cameras.
  • Data. 160 teleoperated demonstrations, with one grid cell held out on purpose.
  • Training. ACT and SmolVLA-450M on one RTX 3060, with LeRobot.
  • Evaluation. 20 autonomous trials per policy on the real arm, plus the held-out cell.
  • Tooling. Recording, training and evaluation scripts, and a test suite CI runs on every pull request.

It builds on the open-source SO-101 design, a pretrained SmolVLA and the LeRobot pipeline.