SAIL improves robot trajectory generation success from 25% to 73% through test-time scaling

A new inference-time search method boosted simulated robot task success from 25% to 73% by testing and refining up to 45 candidate trajectories without any model retraining.

SAIL improves robot trajectory generation success from 25% to 73% through test-time scaling

Researchers from the University of Tokyo and Sakana AI have developed a method that lets vision-language models (VLMs) generate more reliable robot trajectories by spending extra computation at inference time. The approach, called SAIL, increased the rate of finding a successful trajectory in simulation from 25% with a single generation to 73% when the system was allowed to test and refine up to 45 candidates, without ever updating the model's weights. The work has been accepted to IROS 2026.

Why single-shot trajectory generation falls short

Foundation models trained on images, video, and robotics data can already produce end-effector poses and gripper states from camera input and a few demonstrations. But a single prediction is often unreliable. Small errors in motion targets can cause failure, and executing a trajectory without evaluation leaves no chance to catch mistakes before the robot moves. The core question SAIL addresses is whether the same test-time scaling trade-off seen in language models-more inference computation for better outputs-extends to physical action.

How SAIL refines trajectories through search

SAIL uses Monte Carlo Tree Search (MCTS) to combine refinement of promising trajectories with exploration of alternatives. A policy VLM generates candidate trajectories from a few successful demonstrations provided in context. Each trajectory is then executed in simulation, and a separate evaluation VLM estimates task completion from the resulting video. Completion scores are aligned with trajectory waypoints and returned as step-level feedback, telling the policy VLM which segments to preserve and which to revise.

The revised trajectory becomes a child node in the search tree, evaluated again in simulation to guide further exploration. Candidate testing stays entirely in simulation; only the final selected trajectory is sent to the physical robot. All experiments used Gemini Robotics-ER 1.5 as both the policy and evaluation VLM, with no weight updates.

Search budget drives success rates across six tasks

Across six manipulation tasks in the ALOHA simulator-covering objects like a banana, pen, bowl, drawer, laptop, and marker-success rates climbed steadily with the number of search nodes. A single generated trajectory succeeded 25% of the time on average. Breadth-first search with 15 nodes reached 51%, while depth-first search with the same budget reached 37%.

SAIL outperformed both. With a budget of 6 nodes, the average success rate hit 55%. At 15 nodes it reached 65%, at 30 nodes 71%, and at 45 nodes 73%. The drawer and marker tasks proved hardest; the banana and bowl tasks reached 95% and 100% respectively at higher budgets. These gains came purely from additional inference computation-no fine-tuning, no extra training data.

Physical robot trials and training data generation

On a physical block placement task using a LeRobot SO-101 arm, SAIL succeeded in five of six trials with a budget of 15 candidate trajectories. The team reconstructed the scene from color and depth observations, searched for a trajectory in simulation, and executed it open-loop on the robot. The single failure was attributed to pose estimation errors and contact dynamics differences between simulation and the physical setup.

The researchers also showed that successful search trajectories can serve as training data for an imitation policy. This approach likewise succeeded in five of six trials and reduced execution time compared with running the full MCTS search at deployment. The method has clear limitations: trajectories run open-loop without visual feedback during motion, simulation-based search requires additional computation, and real-world validation remains limited to one task.

Why this matters for operations and product development

SAIL demonstrates a practical path to more reliable robot control without the cost and complexity of collecting large demonstration datasets or fine-tuning foundation models. For teams building robotic systems, the implication is that inference-time computation can be treated as a tunable budget-spend more search nodes to raise success rates on harder tasks, or spend fewer when speed matters. The approach also turns successful search trajectories into training data, creating a flywheel where test-time effort pays dividends for future deployments. Professionals working with AI Agent Courses or Generative AI Courses may recognize the parallel to how chain-of-thought and self-consistency methods improved language model outputs-here, the same principle applies to physical action.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)

US energy department allocates $5.25 billion to upgrade grid for AI datacenters

Related AI News for Science and Research

Related AI News for Product Development Professionals