Why a 9B model can beat frontier APIs at verifying warehouse states
A small vision-language model will not out-describe Gemini or GPT on arbitrary images. It does not need to. On one narrow question, asked on fixed cameras, it can win where it matters.
Camly’s cameras do not ask a model to describe the world. They ask one narrow question, over and over: is this state true, in this zone, on this fixed camera, under these clauses, and why. We think a model of 4 to 27 billion parameters, trained on exactly that distribution, can beat a frontier API on the things that decide whether a deployment works. This post explains why we believe it, how we intend to measure it, and what we have not measured yet.1
The question is narrow, and that is the point
“Describe this image” is an open task. Frontier models are very good at it, and a small model will not catch them there. “Is emergency exit B clear, ignoring the goods stored in the rack beside it, counting a pallet only if it touches the floor?” is a closed task with a yes-or-no answer, a confidence, and a short justification grounded in the picture. It is asked thousands of times a day on the same cameras, under the same clauses.
On that task, four things matter in deployment, and none of them is “general image understanding”:
- false alarms on hard negatives, because a site stops reading alerts after a week of noise;
- obedience to exclusion clauses, because every zone comes with “ignore this” rules written by the operator;
- consistency over time, because the same scene minutes apart must get the same verdict, or episodes flicker;
- cost per frame, because periodic checks on every camera of a site add up.
Why it is winnable
Distribution. Web photography is eye-level, centred, well lit. CCTV is high-angle, wide, compressed, often monochrome infrared at night, with the object of interest at thirty pixels. Frontier models are strong but were not trained on this. A small model trained on it is.
Clauses. “Ignore the goods in the racks” and “a pallet is in the zone only if it touches the floor” are domain knowledge. A prompt can carry them, but not reliably: the API reads the clause, it was never trained on the pair of frames where the clause flips the answer. A model trained on counterfactual pairs, same frame, clause changed, learns the clause as a rule rather than as a hint.
Calibration. The confidence field gates alerts. Frontier models emit confident numbers that are not calibrated. A trained model can be rewarded for calibration directly.
Cost. A 9B model on an on-site GPU is, by our earlier inference study, two orders of magnitude cheaper per frame than a cloud API, and the images never leave the site.
Four claims, and how each will be measured
| Claim | Metric | Why a small model can win |
|---|---|---|
| C1 Fewer false alarms on hard negatives at equal recall | False alarms per 1,000 negative frames at 90% recall, broken down by condition | Hard negatives (goods beside the exit, a person walking through, a shadow) make up half the training distribution, with exact labels |
| C2 Clause compliance and zone conditioning | Share of counterfactual pairs answered correctly | Pairs are generated by construction; the model is trained on the pair, the API only reads the clause |
| C3 Consistency and calibration over time | Agreement across frames minutes apart with no state change; expected calibration error; time to detect on clips | Consistency is rewarded directly during reinforcement fine-tuning; an API is stateless and uncalibrated by design |
| C4 (supporting) Cost | Tokens, seconds and euros per 1,000 frames, on a site GPU and on the APIs | A 9B model on-site is about a hundred times cheaper per frame |
The targets we have set ourselves, as hypotheses: false alarms on hard negatives at least halved against the best API at 90% recall (H1), and description faithfulness within 3 points of the best API while being more than a hundred times cheaper (H7).1
What COCO and SA-1B are for
Not pretraining. The vision encoders of the candidate models have already seen billions of image-text pairs; another pass over 118,000 COCO photos changes nothing, and the domain is wrong anyway.2 COCO, LVIS, POPE and RefCOCO serve as regression tests in the release gate: a fine-tune must not lose general grounding. SA-1B and SA-V serve as a library of segmented objects to composite into real warehouse footage; the licence for derived products is to be checked.3
The model ladder
| Model | Role | Footprint (bf16 LoRA) |
|---|---|---|
| Qwen3.5-4B | Edge candidate, cheapest site GPU, first fine-tuning target | about 10 GB |
| Qwen3.5-9B | The product model | about 22 GB |
| Qwen3.8 family | Newer generation: run the whole ladder zero-shot first; if a size under 10B exists, it replaces Qwen3.5 as the base. Sizes and licence to be checked on release | to be checked |
| Qwen3.8-27B | Upper bound for on-premise; teacher for distillation | about 56 GB, fits a DGX Spark for inference and LoRA |
All of these are public weights from the Qwen family.4
What we have not done yet
Nothing above is a result. It is a plan with targets, written down before the first measurement so that the measurement cannot bend it. The order is product first: a frozen evaluation set and a zero-shot ladder in the first two months, which alone may change which model ships; supervised fine-tuning and a release gate in months three and four; reinforcement fine-tuning after that. A paper comes only if the 9B model beats every API on at least two of C1 to C3 on a hidden test split with real positives. If it does not, the benchmark is still published, with the APIs on top, and that is a useful result too.
If you run a warehouse and want your cameras in the evaluation, or you work on vision-language models and want to compare on the same items, the next two posts describe the benchmark and the data engine.
Sources
- Camly AI, Beating frontier models with a small VLM: research plan, October 2026.
- COCO, Common Objects in Context. cocodataset.org
- Meta AI, Segment Anything 1 Billion (SA-1B). ai.meta.com
- Qwen model family on Hugging Face. huggingface.co/Qwen