Compositing states with exact labels: training data without a single customer frame
To train a model to verify warehouse states, we paste the states into real footage, systematically, so every example comes with ground truth known by construction. Here is how, and what can go wrong.
A model that verifies states on fixed cameras needs examples of those states: an exit blocked, a fire aisle narrowed, combustibles beside a charger. Nobody will stage ten thousand of them on a working site, and customer footage is not training data. Our demo already pastes a pallet into real warehouse footage. Done systematically, with masks, geometry and artefact control, that becomes a data engine whose labels are exact because they were constructed. This post describes the engine, the reason exact labels matter, and the main way it could fool us.1
The principle
Start from a real frame of a fixed camera where the state is false: the exit is clear. Composite the state into it: a pallet, a cage, a pile of cardboard, taken from a library of segmented objects. The engine knows everything about what it did, so every example carries ground truth that nobody had to annotate:
- the state (blocked or not) and the clause under which it holds;
- the object mask, at pixel level;
- the zone, in the camera’s own coordinates;
- for sequences, the moment the state appeared and the moment it cleared;
- the untouched frame, which is a hard negative for free.
Counterfactual pairs come out by construction too: the same frame with the pallet placed just outside the zone, or raised onto the rack so that it no longer touches the floor, gives the item on which “only if it touches the floor” flips the answer.
The components
| Step | What it does |
|---|---|
| Floor plane | Estimated per camera from four clicks or from the SmartSpaces calibration |
| Scale | From the floor plane, so a pallet at the far end is the right size |
| Shadow | A drop shadow from the scene’s dominant light |
| Occlusion | Order from existing masks: a forklift in front hides the pasted object |
| Colour and noise | Matched to the frame |
| Conditions | Augmentation: day, night infrared, compression level, rain on the dock |
| Time | Temporal composition over frames minutes apart: appear, persist, clear |
Why exact labels matter
Supervised fine-tuning tolerates some label noise. Reinforcement fine-tuning with verifiable rewards does not: the reward has to be computed from a truth the model cannot argue with. “Did the verdict match the state, did the points fall inside the mask, did the description mention the right object and nothing else” are checkable questions only if the state, the mask and the object are known exactly. Human annotation of CCTV at this volume would be slow, expensive and, on thirty-pixel objects, not exact. Construction is.
The main risk: learning the paste
A model trained on composited positives can learn “there is a pasted object here” instead of “the exit is blocked”. It would score well on composited test items and fail on the first real pallet. This is the risk that decides whether the engine is useful, so it gets several controls, all measured:
- artefact control in the compositor: the same JPEG re-encoding applied to the whole frame, no sharp edges, a randomised blend;
- an adversarial artefact detector, trained to tell pasted from real, which must be at chance on the training set; pastes it can detect are rejected;
- at least 20% of positives are real: staged on the test bench, filmed by real cameras, later from pilots;
- the hidden test split contains real positives only, so any gap between composited and real accuracy is visible, not hidden in an average.
The hypothesis we commit to is H5: a model trained on composited positives only, tested on real staged ones, should lose under 5 points of balanced accuracy.1 If it loses more, the engine is improved before the model is.
Where the objects come from
SA-1B and SA-V provide a library of segmented objects and video masks to composite; the licence for products derived from them is to be checked before anything ships.2 COCO, LVIS, POPE and RefCOCO are not used for training at all: the vision encoders of the candidate models have already seen far more. They serve as regression tests in the release gate, so that a fine-tune does not lose general grounding while it learns warehouses.3
Privacy, by construction too
No customer frame is used for training. The base footage is public (SmartSpaces) or staged on our own bench. Pilot footage leaves a site only under a data agreement, blurred, with the works council informed, and it goes to the hidden test, not to training. The model is trained on states, not on people, and the engine makes that a property of the data rather than a promise.
Compute
The DGX Spark, with 128 GB of unified memory, covers the compositing and augmentation, the artefact detector, and LoRA supervised fine-tuning of the 4B and 9B models; inference numbers measured on it are the product’s numbers, since it is the sovereign edition’s own target. The GRPO runs, several GPU-days each, go to Inria clusters. Nothing in the plan needs more than one node.1
If you have a test bench, a few cameras and the patience to stage a blocked exit twenty times under different lights, you can contribute real positives to the hidden test. That is the part no engine can replace.
Sources
- Camly AI, Beating frontier models with a small VLM: research plan, October 2026.
- Meta AI, Segment Anything 1 Billion (SA-1B) and SA-V. ai.meta.com
- COCO, Common Objects in Context. cocodataset.org