TOAST: Stochastic Robot Action Tokenization
for Autoregressive Vision-Language-Action Models

AIST

TL;DR: TOAST samples alternative tokenizations of the same robot action sequence during training, which makes autoregressive VLA policies more data-efficient without any additional demonstrations.

Abstract

Autoregressive Vision-Language-Action (VLA) models represent continuous robot actions as discrete token sequences. Frequency-based tokenizers such as FAST compactly encode action trajectories, but assign a single deterministic tokenization to each action sequence, although many token sequences decode to exactly the same motion. We propose TOkenization of Action Sequences with STochastic Sampling (TOAST), which samples alternative tokenizations of the same quantized action sequence during policy training. This diversifies the discrete supervision while preserving the underlying robot action, and requires no additional demonstrations. On LIBERO, TOAST consistently improves over its deterministic counterpart, with larger gains as training data decreases (+6.8 points with 1/16 of the data). Across four real-robot manipulation tasks, TOAST improves the mean success rate by 15.8 points over the deterministic counterpart.

Method

TOAST tokenization pipeline

A continuous action chunk is quantized (e.g., by DCT) and flattened in dimension-major order. TOAST builds a unigram language model over such action sequences, computes probabilities over the n-best token sequences for each chunk, and samples one of them as the training target. All sampled token sequences decode to the same action, so the policy architecture and next-token objective stay unchanged, and sampling is used only during training.

Real-Robot Experiments

We evaluate TOAST on four real-robot manipulation tasks with a Franka Research 3 (7-DoF joint-angle actions + gripper, 20 Hz), equipped with a wrist camera (RealSense D405), two scene cameras (RealSense D435), and a Robotiq 2F-85 gripper. Demonstrations are collected by teleoperation with GELLO.

TaskDescription#Demos
Table BussingSort trash into the bin and dishes into the dish rack200
Grocery BaggingPut grocery items into a paper bag200
Breakfast SetupMove a toast onto a plate and a cup onto a coaster100
Drawer StowingOpen a drawer, put an object inside, and close the drawer60

All policies use PaliGemma-3B as the backbone and differ only in the action tokenizer. Each policy is evaluated with 30 rollouts per task over three random seeds. TOAST achieves the highest mean success rate (50.8%), outperforming TOAST (deterministic) by 15.8 points and FAST+ by 20.0 points.

Real-robot results

Rollout Videos

Policy rollouts of TOAST on each real-robot task, played at 1x speed. Select a task, then choose a successful or failed rollout.

BibTeX citation

@misc{shirai2026toast,
author = "Keisuke Shirai, Tomohiro Motoda, Hanbit Oh, Ryoichi Nakajo, Roman Mykhailyshyn, Ryo Hanai, Shotaro Miwa, Yukiyasu Domae",
title = "TOAST: Stochastic Robot Action Tokenization for Autoregressive Vision-Language-Action Models",
year = "2026",
}