Skip to content
Glossary

Training, validation and test split

Data quality

Separating data used to fit a model, choose its settings and estimate its final performance.

The split should reflect the generalization being tested. Nearby frames or repeated captures of the same episode may be strongly related. Keep appropriate groups together, and decide whether the test should hold out episodes, objects, environments or collection sessions.

In practice

Every frame from one robot episode stays in the same partition, instead of placing nearly identical frames in training and test data.

Further reading