Stanford researchers have developed Transformer Transformer, a unified model that generates complete robot designs — including every link, joint, motor, and inertial property — optimized for specific manipulation tasks.

The system works by analyzing a demonstration of desired end-effector motion and producing a tailored robot embodiment. When tested on cloth flinging using an ALOHA bimanual platform, the generated design reduced tracking error by 73% and maximum joint speed by 30% compared to the original robot.

The breakthrough centers on RoboTokens, a unified tokenization system that represents robot embodiments, states, and actions in a single format. This allows one model to work across different robot types — from wheeled bimanual systems to quadrupeds and humanoids — without requiring separate adapters for each embodiment.

How the unified architecture works

The model uses a diffusion transformer trained on RoboTokens with a DDIM noise schedule. By changing which tokens are masked during training, the same network can serve multiple functions: unconditional robot generation, cross-embodiment control, and embodiment optimization.

Rather than training on specific reward functions, Transformer Transformer learns as a dynamics model. At inference time, it converts reward-agnostic predictions into reward-specific value predictions through a process called Dynamics Self-Guidance, which steers embodiment diffusion toward optimal designs.

The researchers tokenized 11 robots from the MuJoCo Menagerie, spanning masses from 0.65 kg to 67.5 kg and 6 to 35 active joints. Each robot compressed into just 28-101 RoboTokens — up to 110 times more compact than traditional MJCF text representations.

Experiments across three design spaces demonstrated zero-shot optimization of unseen rewards and trajectories, outperforming evolutionary baselines in both performance and runtime. The system can generate embodiments, control them, and validate designs using the same unified architecture.

The research, led by Huy Ha, Karen Liu, and Shuran Song, will be presented at the Conference on Robot Learning 2026. The team has released code and demonstration videos showing the system's ability to co-design robots for specific manipulation tasks.