Glossary
The language of data quality,
human judgment and physical AI.
83 of 83 terms
- Acceptance criteriaHuman judgment The requirements an item or dataset must meet to be accepted for its intended use.
- Action chunkingPhysical AI Predicting a sequence of future actions in one policy prediction.
- Action conditioned predictionPhysical AI Predicting future observations or states while taking proposed actions into account.
- Action spacePhysical AI The set of actions available to an agent and the representation used to express them.
- Active learningData quality A method that selects useful examples for labeling instead of treating every unlabeled item as equally valuable.
- AI judgeHuman judgment An AI system configured to evaluate an item against specified criteria.
- Annotation agreementHuman judgment A measure of how consistently different reviewers label the same examples under the same instructions.
- Annotation confidenceHuman judgment An indication of how certain a person or system is about an annotation.
- Annotation consistencyData quality The degree to which the same labeling rules are applied across examples, reviewers and time.
- Annotation error analysisData quality The examination of labeling mistakes to understand their patterns and causes.
- Annotation guidelinesHuman judgment Instructions that explain how to label data, including definitions, boundaries, examples and exceptions.
- Annotation metadataData quality Information about how an annotation was created and reviewed, beyond the label itself.
- Annotation pipelineData quality The workflow that turns source data into labels ready for their intended use.
- Audit trailData quality A record of the actions and decisions that produced a result.
- Balanced datasetData quality A dataset whose classes have similar numbers of examples, or follow a deliberately chosen balance.
- Behavior cloningPhysical AI A form of imitation learning that trains a policy to predict demonstrated actions from observations.
- Blind expert spot checkHuman judgment A review of selected production items against expert answers that are hidden from the automated judges.
- Calibration set for AI judgesHuman judgment Expert reviewed examples used to compare judges before selecting them for a review task.
- Computer visionPhysical AI Methods that extract information from images, video or other visual measurements.
- Concept driftMeasurement A change in the relationship between inputs and the outcomes a model is expected to predict.
- Confidence intervalMeasurement An interval calculated from sampled data to express uncertainty in an estimated population quantity.
- Consensus scoreMeasurement A measure of agreement among reviewers on an item or set of items.
- Consensus thresholdHuman judgment The required level of agreement before a review result is treated as sufficiently supported.
- Data annotationData quality The addition of labels or other structured descriptions to raw data.
- Data curationData quality The selection, organization and maintenance of data for a particular purpose.
- Data leakageData quality Information entering training or model selection that should be unavailable for the evaluation being performed.
- Data lineageData quality A record of where data came from and how it changed before reaching its current form.
- Data miningData quality The search for patterns or useful information within a larger body of data.
- Data pipelineData quality A sequence of processes that moves data from its source to a usable output.
- Data qualityData quality How well data meets the requirements of its intended use.
- Data validationData quality Checking whether data meets specified requirements before it is accepted for a use.
- Dataset coverageData quality How well a dataset represents the tasks, conditions and variations relevant to its intended use.
- Domain shiftMeasurement A difference between the conditions represented in training or evaluation data and those encountered elsewhere.
- Egocentric videoPhysical AI Video recorded from the perspective of a person or agent participating in a scene.
- EmbodimentPhysical AI The physical body, sensors and controls through which an agent perceives and acts.
- Evaluation metricsMeasurement Measures used to assess a model, dataset or review process against a defined objective.
- Expert reference setHuman judgment Examples with documented expert judgments that provide a reference for evaluating other reviewers or systems.
- False negativeMeasurement A missed positive case: the system predicts negative when the reference says positive.
- False positiveMeasurement A positive prediction for a case that is negative according to the reference.
- Ground truthHuman judgment The reference answer used to train or evaluate a system for a particular task.
- Human escalationHuman judgment Routing an item to human reviewers when automated review cannot resolve it under the agreed rules.
- Human in the loopHuman judgment A workflow in which people contribute judgments or decisions to an automated process.
- Image annotationPhysical AI Labels that describe an image or elements within it, such as objects, regions and landmarks.
- Imbalanced datasetData quality A dataset in which some classes or conditions occur much more often than others.
- Imitation learningPhysical AI Learning how to act from demonstrations of a task.
- Instance segmentationPhysical AI Identifying individual objects and assigning pixels to each separate instance.
- Intersection over UnionMeasurement A measure of overlap between two regions: their intersection divided by their union.
- Judge alignmentMeasurement How closely a judge's evaluations match expert reference judgments under a particular rubric.
- KeypointsPhysical AI Specific locations used to represent visual features or meaningful landmarks on an object or body.
- Label noiseData quality Errors or inconsistencies in the labels associated with data.
- Long tail dataData quality Examples of less frequent conditions or events within the distribution a system needs to handle.
- Model driftMeasurement A change in a model's performance or behavior over time relative to the task it is meant to perform.
- Multimodal learningPhysical AI Learning from more than one kind of input, such as video, language, audio or robot state.
- Object detectionPhysical AI Identifying objects in visual data and locating where they appear.
- Observation and action alignmentData quality The correct temporal and contextual pairing of observations with the actions they describe or inform.
- OcclusionPhysical AI A condition in which part or all of an object is hidden from a sensor's view.
- Out of distribution detectionMeasurement Identifying inputs that differ from the data a system was developed or evaluated on.
- Physical AIPhysical AI AI systems that interpret physical environments and support decisions or actions within them.
- PrecisionMeasurement The fraction of positive predictions that are correct against a reference.
- Proof of QualityHuman judgment Sapien's approach to reviewing data against customer criteria, with expert judgments used to evaluate AI judges.
- Quality assurance in annotationData quality The practices used to make annotation quality repeatable across a project.
- Quality scoreMeasurement A measure of how well an item or dataset meets defined quality criteria.
- RecallMeasurement The fraction of actual positive cases that a system correctly identifies.
- Reinforcement learningPhysical AI Learning a policy to maximize expected cumulative reward through interaction data.
- Review coverageMeasurement The extent of the data and conditions examined by a review process.
- Review panelHuman judgment The set of reviewers assigned to evaluate work under a shared decision rule.
- Robot episodePhysical AI A bounded sequence of observations, actions and other records from a robot's interaction with its environment.
- RoboticsPhysical AI The design and operation of machines that sense, compute and act in physical environments.
- RubricHuman judgment A structured set of criteria and scoring rules for evaluating a piece of work or data.
- Semantic segmentationPhysical AI Assigning a category to each pixel in an image.
- Sensor fusionPhysical AI Combining measurements from different sensors to estimate an environment or system state.
- Sensor synchronizationPhysical AI Aligning measurements from different sensors to a consistent time reference.
- Simulation to reality transferPhysical AI Applying a model or policy learned in simulation to a physical environment.
- Synthetic dataData quality Data created through simulation, procedural generation or generative models rather than directly recorded from the target setting.
- Task success labelPhysical AI An annotation that records whether an attempt achieved a defined task outcome.
- TeleoperationPhysical AI Control of a machine by a human operator through an interface rather than direct physical contact.
- Temporal consistencyPhysical AI Consistency in identities, labels and events across a sequence over time.
- Training dataData quality Examples used to fit a model's parameters or behavior during learning.
- Training, validation and test splitData quality Separating data used to fit a model, choose its settings and estimate its final performance.
- TrajectoryPhysical AI An ordered sequence of states or observations and actions from an agent’s interaction with its environment.
- UncertaintyMeasurement A lack of sufficient knowledge or evidence to make a dependable judgment.
- Vision language action modelPhysical AI A model that uses visual observations and language instructions to produce actions.
- World modelPhysical AI A learned representation of an environment that can help predict how it changes over time.