Open hardware · Apache-2.0

TactileVLA-Edge

A vision-language-action stack for contact-rich robot manipulation on sub-$300 hardware, working toward on-device tactile inference.

Phase 2 complete — vision-only baseline Phase 3 tactile — not built Interactive 3D demo Source on GitHub
The interactive 3D demo: a reconstructed SO-101 arm above the 3×3 task grid

See it in 3D

A guided tour of the whole project: the setup, the held-out test, the real runs and the numbers. Then send the arm to any cell yourself, B2 included.

The 3D arm is a kinematic reconstruction. The videos are the real runs.

Open the interactive demo
The SO-101 arm grasping a foam pool noodle and dropping it into a cup
ACT in cell B2 — approach, grasp, transport, release. No training episode ever placed the object here.

Autonomous trials

ACT52M params · 42,308 steps 18 / 20
SmolVLA-450Mfrozen vision encoder · 20,000 steps 16 / 20

20 trials cannot separate these two — the confidence intervals overlap heavily. What they establish is that both work.

The claim

Tactile VLA models exist, but the published work runs on ~$30k research arms. This project aims to be a complete, reproducible, open-source tactile-VLA reference stack that runs on hardware costing under $300 — the arm, the touch sensing, the dataset format, the fusion method, and the measured results, in one place.

That is roughly a 100× reduction in hardware cost. Reading touch off the servos the arm already has is a large part of why: it keeps the touch channel at zero marginal cost.

To be precise about the claim: this is not the first tactile-VLA. Tactile-VLA, TacVLA, and TacFiLM came first. The intended contribution is the integration at the low-cost end — and it does not exist yet. What exists today is the baseline it has to beat.

Collecting the data

Recording sessiontime-lapse

160 episodes, teleoperated one at a time

Every episode is a human driving the follower arm through the task on the leader arm, then resetting the scene by hand: reposition the noodle into the next cell, rotate it to the next yaw in the sweep, start the next take.

Four rounds walking the perimeter A1→B1→C1→C2→C3→B3→A3→A2, five episodes per cell per round. The walk order reverses on rounds 2 and 4 so cell identity does not alias with fatigue and drift within a round.

This is the part no architecture choice can substitute for, and it is most of the wall-clock cost of the project.

Watch the runs

Four trials filmed end to end, each narrated with the workspace cell called out before the run. All four succeeded. B2 is the held-out cell — it appears in none of the 160 training episodes.

ACTB2 · held out
SmolVLA-450MB2 · held out
ACTA2 · trained
SmolVLA-450MC3 · trained

These are demonstrations, not the scored trials — the 20-trial runs were tallied by hand without video.

Dataset & protocol

160 demonstrations, stratified by design

Teleoperated episodes spread over a 3×3 grid projected onto the table by homography from four clicked corners. Eight cells were recorded at exactly 20 episodes each. Within every cell the object's yaw sweeps the full ±90° range in 10° steps, so no cell is a single memorised pose.

The centre cell, B2, was never recorded — it exists only to be tested against.

Episodes160
Frames67,693
Rate30 fps
Cameras800×600 overhead, 640×480 wrist
FormatLeRobot v3.0
Workspace grid showing eight cells at 20 episodes each and cell B2 held out
Rendered from the recording session's own saved geometry, so the figure cannot drift from what was recorded. The grid is a perspective quad, not a square — outer cells cover more table than inner ones.

Generalization

Held-out cell · zero training episodes

Both architectures succeeded in a cell they were never shown

This is the result worth reporting, more than either success rate. Rather than matching the nearest trained trajectory, both policies interpolated into an unseen region of the workspace — on a 160-episode dataset.

The grid overlay on screen during the recording session, with B2 marked hold
The live overlay mid-session, B2 marked (hold), with the finished dataset's 160 episodes / 67,693 frames visible in the terminal behind it.

What the policy sees

Side-by-side overhead and wrist camera streams from a training episode
Overhead (left) and wrist (right), from episode 0 of the training set — the actual observation space, not a staged shot.

Why two cameras

The overhead camera loses the object behind the gripper at exactly the moment contact matters. The wrist camera is what keeps the grasp legible.

That occlusion is also the argument for touch: a channel that still reports when vision is blocked. Which is Phase 3, and is not built.

Hardware

LayerComponentState
ArmSO-101 leader + follower, Feetech STS3215working
VisionOverhead 800×600 + wrist 640×480, 30 fpsworking
PolicyACT and SmolVLA-450M via LeRobotworking
ComputeRTX 3060 training; Apple Silicon / MPS inferenceworking
TouchServo load (Present_Load, STS3215 reg. 60)not integrated
FusionFiLM conditioning, token-concat firstnot built

Touch reads load off servos the arm already has, so it needs no additional hardware. This replaced an earlier plan to build an AnySkin magnetic sensor; that sensor was never built.

What this does not establish

Full protocol, training configuration, and limits: docs/results.md.