From Language to Motion: Phase-Aware Tempo Control for Vision-Language-Action Policies

TempoBridge concept: language tempo interpretation and phase-aware motion execution
Figure 1. TempoBridge reads tempo cues from a frozen VLA and applies the requested tempo as the task progresses, without additional speed-conditioned demonstrations or tempo-specific policy fine-tuning.

Abstract

Vision-Language-Action (VLA) models can perform manipulation tasks from language, but task completion does not necessarily imply following instructions about execution speed. We introduce TempoBridge, a lightweight framework that makes tempo information in a frozen VLA actionable. TempoBridge extracts tempo cues from contextual VLM representations, aligns them with task progress through a causal phase router, and modulates nominal motion commands during execution. It requires no external language model, additional speed-conditioned robot demonstrations, or tempo-specific policy fine-tuning. Across 40 LIBERO tasks, TempoBridge improves Tempo Success Rate from 52.6% to 89.7% under quickly/slowly instructions, with an overall Task Success Rate of 92.9%. It also maintains near-baseline task success under original instructions and generalizes to unseen tempo expressions. Experiments on an xArm6 further demonstrate that changing the language instruction produces corresponding changes in physical execution speed.

Method Overview

TempoBridge architecture with a prototype-based tempo readout, causal phase router, and frozen pi0.5 policy
Figure 2. TempoBridge connects a cached tempo sequence to causal phase routing and execution-time action modulation, while keeping the base VLA frozen.

Prototype-Based Tempo Readout

Once per instruction, a prototype-based readout extracts a tempo sequence from the internal language features of the frozen π₀.₅ policy. The sequence is cached and reused throughout execution.

Causal Phase Routing

For two-entry tempo sequences, a causal router uses visual features reused from the policy and robot-state history to select the active event. It permits one forward transition from Event 1 to Event 2; single-entry sequences do not require routing.

Tempo-Conditioned Motion

In LIBERO, the selected tempo scales nominal Cartesian translation by 1.3 for FAST, 0.7 for SLOW, or 1.0 for NORMAL. Rotation and gripper commands are retained, and the entire base policy stays frozen.

The readout is constructed from text examples containing quickly and slowly. The phase router is trained on existing LIBERO-90 demonstrations with phase-transition annotations, using tasks disjoint from all 40 evaluation tasks. Neither component requires tempo-conditioned robot demonstrations or tempo-specific updates to the base policy.

Simulation Results

Language-conditioned tempo control across 40 LIBERO tasks.

LIBERO executions under canonical and unseen tempo expressions, with TCP trajectories colored by measured speed
Figure 3. The left two examples use quickly and slowly; the right three use unseen tempo expressions. TCP trajectories are colored from slower motion in blue to faster motion in red.

We evaluate TempoBridge across LIBERO-Spatial, Object, Goal, and Long. Under canonical quickly/slowly instructions, TempoBridge achieves 89.7% Tempo Success Rate, compared with 52.6% for π0.5. Task Success Rate is 92.9%, compared with 93.6% for the base policy.

Joint Task & Tempo Success Rate increases from 48.8% to 82.0%, showing that the tempo improvement is retained when task completion and tempo compliance are evaluated together.

Task Success Rate measures completion over all rollouts. Tempo Success Rate is conditional on successful, tempo-evaluable rollouts: the mean TCP speed for FAST must exceed that for SLOW. Speeds are averaged over the central 50% of each evaluation segment. Two-event boundaries come from simulator-state predicates independently of the router; single-event comparisons use matched initial states and seeds. Task & Tempo Success Rate reports joint success over all rollouts.

Table 1. Canonical instructions. Overall Task / Tempo / Task & Tempo success: pi0.5 93.6 / 52.6 / 48.8%; TempoBridge 92.9 / 89.7 / 82.0%.
Table 1. Canonical tempo instructions: 800 rollouts per method across 40 LIBERO tasks. Click the table to enlarge.

Task Competence without Tempo Cues

With the 40 original benchmark instructions, task success remains near baseline: 94.3% versus 95.0%. The readout returns NORMAL for 38 of 40 instructions (95%), leaving nominal translation unchanged for those instructions.

Table 2. Task success without tempo cues. Spatial / Object / Goal / Long / Overall: pi0.5 95.0 / 97.0 / 95.0 / 93.0 / 95.0%; TempoBridge 98.0 / 96.0 / 95.0 / 88.0 / 94.3%.
Table 2. Original benchmark instructions: 400 rollouts per method, without added tempo cues.

See TempoBridge in Action

Choose a method and task to compare both tempo instructions side by side.

All executions succeed. Each baseline matches its TempoBridge counterpart in seed and initial state; the two tempo instructions use different seeds or initial states.

LIBERO-Spatial

“quickly pick up the black bowl next to the ramekin and slowly place it on the plate”

Robot execution · original speed
TCP speed trajectory: quickly pick up the black bowl next to the ramekin and slowly place it on the plate
TCP trajectory · click to enlarge

LIBERO-Spatial

“slowly pick up the black bowl next to the ramekin and quickly place it on the plate”

Robot execution · original speed
TCP speed trajectory: slowly pick up the black bowl next to the ramekin and quickly place it on the plate
TCP trajectory · click to enlarge

All examples share a TCP speed color scale of 0.03–0.25 m/s, from blue (slow) to red (fast). Speeds outside this range use the endpoint colors. Green marks the start; white marks the end.

Generalization to Unseen Tempo Expressions

New tempo expressions, with the same fixed readout.

We evaluate rapidly, swiftly, carefully, and at a slower pace, none of which is used for readout construction or configuration selection. Across 3,200 rollouts per method, TempoBridge achieves 80.1% tempo success, compared with 53.5% for π0.5, with 92.3% task success versus 95.1% for the baseline. Joint Task & Tempo Success Rate rises from 50.6% to 72.9%.

The readout recognizes rapidly, swiftly, and at a slower pace with accuracies of 98.3%, 98.3%, and 95.0%, respectively. Recognition of carefully is lower at 57.9%, highlighting the challenge of expressions whose relationship to speed is less direct.

Table 3. Unseen expressions. Overall Task / Tempo / Task & Tempo success: pi0.5 95.1 / 53.5 / 50.6%; TempoBridge 92.3 / 80.1 / 72.9%.
Table 3. Unseen tempo expressions: 3,200 rollouts per method. Click the table to enlarge.
Unseen cue recognition: rapidly FAST 98.3%; swiftly FAST 98.3%; carefully SLOW 57.9%; at a slower pace SLOW 95.0%.
Table 4. Cue-specific recognition for unseen tempo expressions. Click to enlarge.

From Tempo Readout to Physical Motion

Component-wise diagnostics under canonical and unseen instructions.

Canonical: readout accuracy 97.5%, balanced phase agreement 89.7%, median speed ratio 1.58 times. Unseen: 78.4%, 90.4%, 1.52 times.
Table 5. Readout accuracy, balanced phase agreement, and realized speed separation. Click to enlarge.

Readout accuracy measures exact agreement of the full tempo sequence over unique instructions. Balanced phase agreement gives equal weight to the two reference events and includes rollouts where the router never transitions. Evaluation boundaries are defined independently of the router.

The median FAST-to-SLOW speed ratio is computed over the same task-successful comparisons used for Tempo Success Rate. Ratios of 1.58× and 1.52× quantify the realized speed separation beyond the binary FAST > SLOW criterion.

Real-Robot Experiments

xArm6 experimental setup with external and wrist-mounted Azure Kinect cameras and three manipulation tasks
Figure 4. Real-world setup with an xArm6 robot, external and wrist-mounted Azure Kinect cameras, and three manipulation tasks.

We evaluate three manipulation tasks using an xArm6 robot with external and wrist-mounted Azure Kinect cameras:

  • Task 1: Pick up the white mug and place it on the green coaster.
  • Task 2: Pick up the red apple and place it in the pink basket.
  • Task 3: Pick up the red block and drop it into the yellow box.

We compare the task-adapted π0.5 and TempoBridge under FAST → SLOW and SLOW → FAST instructions. The words quickly and slowly modify the pickup and placement or dropping clauses, and are swapped between the two conditions. Both methods use the same task-specific checkpoints without tempo-specific policy fine-tuning. TempoBridge completes 60 of 67 attempts (89.6%), compared with 60 of 66 (90.9%) for π0.5. We analyze 10 task-successful rollouts per task, method, and tempo ordering, for 120 successful rollouts in total.

Measured TCP speed for three tasks: pi0.5 above and TempoBridge below, comparing FAST-to-SLOW and SLOW-to-FAST instructions
Figure 5. Columns correspond to tasks, with π0.5 in the top row and TempoBridge in the bottom row. Thin lines connect the two segment mean speeds within each rollout. Bold lines and error bars indicate medians and interquartile ranges over 10 successful rollouts per condition.

Segment boundaries run from execution start to gripper close, and from gripper close to release, independently of the router. TCP speed is the duration-weighted mean over the central 50% of each segment. In real-world control, the tempo gain scales joint displacements and motion limits.

Across all three tasks, TempoBridge produces a higher median speed for each segment when quickly is requested than when slowly is requested. Swapping the two expressions reverses the median segment-speed ordering, while the base policy responds less consistently. These results demonstrate language-conditioned changes in physical execution speed without additional speed-conditioned demonstrations or tempo-specific policy fine-tuning.

BibTeX

Coming soon.