← All posts
research · benchmark · CCTV

StateWatch: a benchmark for persistent-state verification on fixed cameras

Existing vision-language benchmarks do not measure what a monitored camera needs: is a described state true in a zone, under clauses, with calibrated confidence. StateWatch will.

By Oussama Messai Published 5 October 2026 5 min read Lire cet article en français
A sequence of warehouse camera frames minutes apart, with a state appearing and clearing

Benchmarks shape what models get good at. The public vision-language benchmarks reward captioning, visual question answering on web photos, document reading and chart understanding. None of them asks the question a monitored camera asks all day: does this described state hold, in this zone of this fixed camera, under the clauses the operator wrote, with a confidence one can act on and few false alarms. StateWatch is the benchmark we are building to measure that. This post describes its design before any number exists, so the design cannot be fitted to the numbers.1

What is missing from existing benchmarks

Three things. First, persistence: a state that matters on a site is one that lasts, and the question “did it appear, persist, clear” cannot be asked of a single photo. Second, clauses: “ignore the goods in the racks”, “only if it touches the floor” are part of the task, and a benchmark without them measures the wrong skill. Third, the operating point: a site cares about false alarms at high recall and about whether the confidence means anything, not about accuracy averaged over easy items.

Warehouse CCTV data to build on barely exists. PhysicalAI SmartSpaces is the only public, commercially usable, multi-camera warehouse source, and it was made for tracking: nothing in it is blocked, obstructed, saturated or missing. LOCO has the objects but not the cameras. Everything else is campus, traffic, retail or crime footage. Real warehouse CCTV at scale exists only on customer sites and leaves them only under a data agreement, blurred, with the works council informed.1

Two tracks

SW-FrameSW-Clip
InputOne frame, plus an optional earlier frame of the same cameraFour to eight frames, minutes apart
QuestionA clause-bearing question and a zoneA state to track over the sequence
OutputA JSON verdict with points in the image and a short grounded descriptionThe state’s appearance, persistence and clearing, with times
What it measuresDecision quality under clauses, on hard negativesTemporal consistency and time to detect

Every item carries a zone, a clause, a condition tag (day, night infrared, compression level, occlusion) and, where it applies, a counterfactual twin: the same frame with the clause or the zone changed, so that clause compliance can be scored on pairs rather than guessed from averages.

Metrics

They are decision metrics, chosen to match how a site experiences a model:

  • balanced accuracy, so that the negative-heavy distribution does not hide the misses;
  • false alarms per 1,000 negatives at 90% recall, the operating point of an alerting system;
  • expected calibration error of the confidence field, which gates alerts;
  • clause-compliance rate on counterfactual pairs;
  • consistency: agreement of verdicts across frames minutes apart with no state change;
  • time to detect on the clip track;
  • a faithfulness score for the short descriptions, judged against the ground-truth claims.

Baselines and leaderboard

Zero-shot: Gemini 3.x Flash and Pro, GPT-5.x, Claude, Qwen3.5 4B and 9B, the Qwen3.8 sizes, Qwen3.6-35B-A3B, InternVL3, Gemma 3, each with and without reasoning, with and without an evidence-first output schema. Fine-tuned: the Camly models by training stage. The leaderboard is split by size class (up to 4B, up to 10B, up to 35B, larger, closed API) and by cost, because a 4B model that is close to a frontier API at a hundredth of the price is a result in itself.1

Splits, licence, hosting

SetItemsComposition
Public development split3,000CC BY 4.0; built on SmartSpaces and staged scenes; no pilot footage
Hidden test2,000Real staged positives, pilot footage under agreement; refreshed yearly

The public split is for development and for anyone who wants to reproduce the numbers. The hidden test is scored by submission, so that nobody, including us, tunes on it. The harness is written for lmms-eval or VLMEvalKit and hosted in the EU. Pilot footage enters the hidden test only under a data agreement, blurred, with the site’s works council informed; it never enters the public split.

When it becomes public

Not before months seven to nine of the plan, and only after the model work is done. If the Camly 9B model beats every API on at least two of the three claims (false alarms on hard negatives, clause compliance, consistency and calibration) on the hidden split, StateWatch v1 is released with a human agreement study (300 items, three raters), a datasheet and a paper. If it does not, the benchmark is still published, with the APIs on top of the leaderboard. That is an honest and useful result too: a public measure of what deployed camera monitoring actually needs.1

If you run warehouse cameras and would like staged scenes from your site in the hidden test, or you work on vision-language models and want to be on the first leaderboard, we would like to hear from you before the development split is frozen.

Sources

  1. Camly AI, Beating frontier models with a small VLM: research plan, October 2026.
  2. NVIDIA PhysicalAI SmartSpaces dataset; LOCO, Logistics Objects in Context (dataset pages, by name).