How Proof of Quality Brings Measurable Evaluation to City Intelligence

September 25, 2026
Moritz Hain
Marketing Coordinator

PoQ is an evaluation system for AI work. In city intelligence, it measures data derived from sensors and the recommendations built on it against the standard the user sets, then records the result.

City intelligence combines sensor and camera networks with models that turn readings into counts, incidents, conditions, and recommendations. Collection and processing already operate at machine scale. The difficult stage comes after processing, when someone has to decide whether a reading meets the standard required for the decision attached to it. That judgment often has less structure than the systems around it, so evaluation becomes the slow stage in an otherwise automated pipeline.

TL;DR

  1. City sensor systems report collection rates, processing speeds, and model outputs, while the judgment between a reading and a decision often lacks an equivalent measure.
  2. Field performance also differs from controlled performance, so an accuracy figure may describe a different question from the one a city needs answered.
  3. A useful evaluation separates deterministic checks from contextual judgment, measures failures against a written rubric, and records the strength of reviewer agreement.
  4. PoQ does this with a Quality Score, a Consensus Score, and a Proof Report the user keeps.

City intelligence has a review queue

Cameras on works vehicles, air-quality sensors, and intersection systems all produce enormous amounts of data. Models turn those readings into detections or recommendations that feed operational decisions.

The quality question appears between those two stages. Across pothole detection, air quality monitoring, and traffic incident classification, a plausible datapoint may still fail the standard attached to the downstream decision. The system therefore needs a way to distinguish acceptable work from work that needs deeper judgment.

Public evidence shows how much time this stage absorbs. Toronto’s Auditor General reported an average of 48 days from red light camera capture to screening in 2025 across about 2.2 million penalties. The technology collected the evidence far faster than the process judged it. The same structural constraint appears whenever automated collection grows faster than evaluation capacity.

A sensor reading passes through a model to a reading, then judgment, then a city decision. Collection and processing run at machine scale; judgment is the stage that needs an explicit measure.
Figure 1 · Collection and processing report a rate. Judgment is the stage without one.

Street conditions change what accuracy means

Sensor performance changes under operating conditions. A 2021 study of PurpleAir PM2.5 sensors found that, on average, raw readings overestimated concentrations by about 40 percent across most of the United States before manual correction.

Accuracy therefore has to be read against its definition and use. A strong result under one metric may still leave the downstream decision unanswered.

The 2024 New York City Comptroller audit shows the distinction. A gunshot detection system met its 90 percent contractual performance target in most boroughs, while confirmed shootings represented 8 percent to 20 percent of alerts. The contract metric measured one stage of performance; the operational decision depended on another.

The same principle applies to infrastructure data. A rubric specific to a city may define what an acceptable reading, detection, or recommendation could mean under the conditions in which the system operates. Evaluation then measures the work against that definition instead of treating a general accuracy figure as a substitute.

Make judgment measurable

Some parts of quality have a known answer. Deterministic checks cover properties such as timestamp validity, image focus, or counts that can be compared with a manual sample. Other parts depend on interpretation, including whether a detection represents the intended object or whether a recommendation follows from the readings underneath it.

The rubric separates those cases. Automated checks handle criteria with explicit tests, while qualified experts judge work that depends on context or expertise. The acceptance threshold then becomes measurable: how often does the system accept a datapoint that qualified experts would reject?

Sapien spent two years and 195 million tasks learning this problem in practice. That work showed that reliable judgment needs structure.

How Proof of Quality evaluates the work

PoQ uses five steps: Standard, Route, Judge, Consensus, and Record.

The user starts by defining quality through a rubric, including evaluation criteria and acceptance thresholds. A drafting agent helps author the rubric, while the user owns the standard. This ties the evaluation to the intended use.

Routing applies the appropriate depth of evaluation to each datapoint. Agents handle volume where automation operates reliably and surface uncertainty where deeper review adds value. Qualified experts evaluate work that depends on context, expertise, or consequence. Reviewers score independently so one person’s judgment never determines the result.

Consensus combines those assessments into two measures. The Quality Score shows how the work measured against the rubric. The Consensus Score shows how strongly the reviewers agreed. A unanimous panel and a divided panel therefore carry different information even when their median rating is similar.

Every evaluation ends in a Proof Report. The report records the standard applied, the work evaluated, the scores, and how the conclusion was reached. It also carries a signature that allows a recipient to confirm that the record is authentic and unchanged.

One datapoint moves through Standard, Route, Judge, Consensus and Record, producing a Quality Score, a Consensus Score and a Proof Report.
Figure 2 · Standard, Route, Judge, Consensus, Record, and what the user is left holding.

Improvement comes before the record

PoQ first shows where work falls below the user’s standard. The city or vendor then decides what to recollect, recalibrate, change, or discard. PoQ reads work in order to evaluate it. After an evaluation, only the Proof Report is retained.

The report preserves the evidence behind that evaluation. The public Proof Report example ties each result to the work and rubric that produced it, records reviewer scoring, and preserves the sequence through timestamps and a signed payload. Routing applies deeper attention where the evaluation requires it.

This record gives the result a stable reference. The Quality Score remains relative to the rubric the user authored, while the Consensus Score records the strength of reviewer agreement.

Evaluation speed affects decision time

Collection scales through cameras, vehicles, and connected sensors. Processing scales through software. Evaluation has to scale with both if the resulting work is going to reach a decision on useful terms.

PoQ concentrates deeper review on datapoints that require it. Automation handles explicit checks, uncertain work moves to qualified experts, and consensus resolves cases where a single judgment carries too much variance.

A pilot starts with one queue of completed AI work and a rubric that defines the acceptance threshold. The evaluation method, expected effort, and cost are set up front, and the first Proof Report lands while the decision it informs is still open.

As AI does more work, quality depends on making judgment scale with it.

Sapien is starting to test this approach with city intelligence teams

If sensor-derived data feeds a decision in your system, we would like to compare how that judgment is handled today and which parts can be measured more clearly.

FAQ

What is city intelligence?

City intelligence combines sensor or camera networks with software that turns readings into conditions, detections, recommendations, or actions. The AI work includes the data going in, the reasoning and processing in the middle, the outputs coming back, and the decisions built on those outputs.

Why does sensor accuracy differ between a laboratory and the street?

Operating conditions change the relationship between a sensor reading and the condition it represents. Evaluation in the deployment environment shows how the system performs against the standard used for the decision.

Does PoQ decide whether sensor data is objectively correct?

PoQ measures work against the user’s rubric. Deterministic checks apply where a source, specification, test, or known answer exists. Qualified experts and consensus handle cases where quality depends on context or expertise.

Who are the reviewers?

Reviewers are qualified experts selected for the work being evaluated. Agents handle checks where automation is reliable, while experts concentrate on work that requires context or expertise. Reviewers score independently, and consensus combines their assessments.

What does a Proof Report contain?

A Proof Report records the standard applied, the work evaluated, the resulting scores, and how the conclusion was reached. The report belongs to the user and carries a signature that allows a recipient to confirm that it is authentic and unchanged.

What do the two scores mean?

The Quality Score shows how the work measured against the user’s rubric. The Consensus Score shows how strongly the reviewers agreed. Together they separate the rating from the strength of agreement behind it.

References

  1. Toronto Auditor General, Audit of the City’s Administrative Penalty System for Parking and Red Light Camera Violations (2026).
  2. Barkjohn, Gantt and Clements, Atmospheric Measurement Techniques (2021).
  3. South Coast AQMD, AQ-SPEC sensor evaluation results.
  4. US EPA, Air Sensor Toolbox.
  5. New York City Comptroller, audit of the NYPD ShotSpotter agreement (2024).
  6. City of Chicago Office of Inspector General, ShotSpotter analysis (2021).