Vision language action model
Physical AI
A model that uses visual observations and language instructions to produce actions.
In robotics, a VLA connects what the robot sees and is asked to do with commands it can execute. Training examples need consistent relationships between the instruction, observations and actions. Action format and timing depend on the robot and dataset.
In practice
An instruction to place an object should correspond to the recorded demonstration and its action sequence.