LIBERO-Spatial
“quickly pick up the black bowl next to the ramekin and slowly place it on the plate”
Vision-Language-Action (VLA) models can perform manipulation tasks from language, but task completion does not necessarily imply following instructions about execution speed. We introduce TempoBridge, a lightweight framework that makes tempo information in a frozen VLA actionable. TempoBridge extracts tempo cues from contextual VLM representations, aligns them with task progress through a causal phase router, and modulates nominal motion commands during execution. It requires no external language model, additional speed-conditioned robot demonstrations, or tempo-specific policy fine-tuning. Across 40 LIBERO tasks, TempoBridge improves Tempo Success Rate from 52.6% to 89.7% under quickly/slowly instructions, with an overall Task Success Rate of 92.9%. It also maintains near-baseline task success under original instructions and generalizes to unseen tempo expressions. Experiments on an xArm6 further demonstrate that changing the language instruction produces corresponding changes in physical execution speed.
Once per instruction, a prototype-based readout extracts a tempo sequence from the internal language features of the frozen π₀.₅ policy. The sequence is cached and reused throughout execution.
For two-entry tempo sequences, a causal router uses visual features reused from the policy and robot-state history to select the active event. It permits one forward transition from Event 1 to Event 2; single-entry sequences do not require routing.
In LIBERO, the selected tempo scales nominal Cartesian translation by 1.3 for FAST, 0.7 for SLOW, or 1.0 for NORMAL. Rotation and gripper commands are retained, and the entire base policy stays frozen.
The readout is constructed from text examples containing quickly and slowly. The phase router is trained on existing LIBERO-90 demonstrations with phase-transition annotations, using tasks disjoint from all 40 evaluation tasks. Neither component requires tempo-conditioned robot demonstrations or tempo-specific updates to the base policy.
Language-conditioned tempo control across 40 LIBERO tasks.
We evaluate TempoBridge across LIBERO-Spatial, Object, Goal, and Long. Under canonical quickly/slowly instructions, TempoBridge achieves 89.7% Tempo Success Rate, compared with 52.6% for π0.5. Task Success Rate is 92.9%, compared with 93.6% for the base policy.
Joint Task & Tempo Success Rate increases from 48.8% to 82.0%, showing that the tempo improvement is retained when task completion and tempo compliance are evaluated together.
Task Success Rate measures completion over all rollouts. Tempo Success Rate is conditional on successful, tempo-evaluable rollouts: the mean TCP speed for FAST must exceed that for SLOW. Speeds are averaged over the central 50% of each evaluation segment. Two-event boundaries come from simulator-state predicates independently of the router; single-event comparisons use matched initial states and seeds. Task & Tempo Success Rate reports joint success over all rollouts.
With the 40 original benchmark instructions, task success remains near baseline: 94.3% versus 95.0%. The readout returns NORMAL for 38 of 40 instructions (95%), leaving nominal translation unchanged for those instructions.
Choose a method and task to compare both tempo instructions side by side.
All executions succeed. Each baseline matches its TempoBridge counterpart in seed and initial state; the two tempo instructions use different seeds or initial states.
LIBERO-Spatial
“quickly pick up the black bowl next to the ramekin and slowly place it on the plate”
LIBERO-Spatial
“slowly pick up the black bowl next to the ramekin and quickly place it on the plate”
Measured TCP speed
All examples share a TCP speed color scale of 0.03–0.25 m/s, from blue (slow) to red (fast). Speeds outside this range use the endpoint colors. Green marks the start; white marks the end.
New tempo expressions, with the same fixed readout.
We evaluate rapidly, swiftly, carefully, and at a slower pace, none of which is used for readout construction or configuration selection. Across 3,200 rollouts per method, TempoBridge achieves 80.1% tempo success, compared with 53.5% for π0.5, with 92.3% task success versus 95.1% for the baseline. Joint Task & Tempo Success Rate rises from 50.6% to 72.9%.
The readout recognizes rapidly, swiftly, and at a slower pace with accuracies of 98.3%, 98.3%, and 95.0%, respectively. Recognition of carefully is lower at 57.9%, highlighting the challenge of expressions whose relationship to speed is less direct.
Component-wise diagnostics under canonical and unseen instructions.
Readout accuracy measures exact agreement of the full tempo sequence over unique instructions. Balanced phase agreement gives equal weight to the two reference events and includes rollouts where the router never transitions. Evaluation boundaries are defined independently of the router.
The median FAST-to-SLOW speed ratio is computed over the same task-successful comparisons used for Tempo Success Rate. Ratios of 1.58× and 1.52× quantify the realized speed separation beyond the binary FAST > SLOW criterion.
We evaluate three manipulation tasks using an xArm6 robot with external and wrist-mounted Azure Kinect cameras:
We compare the task-adapted π0.5 and TempoBridge under FAST → SLOW and SLOW → FAST instructions. The words quickly and slowly modify the pickup and placement or dropping clauses, and are swapped between the two conditions. Both methods use the same task-specific checkpoints without tempo-specific policy fine-tuning. TempoBridge completes 60 of 67 attempts (89.6%), compared with 60 of 66 (90.9%) for π0.5. We analyze 10 task-successful rollouts per task, method, and tempo ordering, for 120 successful rollouts in total.
Segment boundaries run from execution start to gripper close, and from gripper close to release, independently of the router. TCP speed is the duration-weighted mean over the central 50% of each segment. In real-world control, the tempo gain scales joint displacements and motion limits.
Across all three tasks, TempoBridge produces a higher median speed for each segment when quickly is requested than when slowly is requested. Swapping the two expressions reverses the median segment-speed ordering, while the base policy responds less consistently. These results demonstrate language-conditioned changes in physical execution speed without additional speed-conditioned demonstrations or tempo-specific policy fine-tuning.
Coming soon.