Skip to content
Glossary

Vision language action model

Physical AI

A model that uses visual observations and language instructions to produce actions.

In robotics, a VLA connects what the robot sees and is asked to do with commands it can execute. Training examples need consistent relationships between the instruction, observations and actions. Action format and timing depend on the robot and dataset.

In practice

An instruction to place an object should correspond to the recorded demonstration and its action sequence.

Further reading