← All posts
research · VLM · benchmark

Why a 9B model can beat frontier APIs at verifying warehouse states

A small vision-language model will not out-describe Gemini or GPT on arbitrary images. It does not need to. On one narrow question, asked on fixed cameras, it can win where it matters.

By Oussama Messai Published 1 October 2026 5 min read Lire cet article en français
A warehouse camera frame with a zone outline and a verdict, next to a model size ladder

Camly’s cameras do not ask a model to describe the world. They ask one narrow question, over and over: is this state true, in this zone, on this fixed camera, under these clauses, and why. We think a model of 4 to 27 billion parameters, trained on exactly that distribution, can beat a frontier API on the things that decide whether a deployment works. This post explains why we believe it, how we intend to measure it, and what we have not measured yet.1

The question is narrow, and that is the point

“Describe this image” is an open task. Frontier models are very good at it, and a small model will not catch them there. “Is emergency exit B clear, ignoring the goods stored in the rack beside it, counting a pallet only if it touches the floor?” is a closed task with a yes-or-no answer, a confidence, and a short justification grounded in the picture. It is asked thousands of times a day on the same cameras, under the same clauses.

On that task, four things matter in deployment, and none of them is “general image understanding”:

  • false alarms on hard negatives, because a site stops reading alerts after a week of noise;
  • obedience to exclusion clauses, because every zone comes with “ignore this” rules written by the operator;
  • consistency over time, because the same scene minutes apart must get the same verdict, or episodes flicker;
  • cost per frame, because periodic checks on every camera of a site add up.

Why it is winnable

Distribution. Web photography is eye-level, centred, well lit. CCTV is high-angle, wide, compressed, often monochrome infrared at night, with the object of interest at thirty pixels. Frontier models are strong but were not trained on this. A small model trained on it is.

Clauses. “Ignore the goods in the racks” and “a pallet is in the zone only if it touches the floor” are domain knowledge. A prompt can carry them, but not reliably: the API reads the clause, it was never trained on the pair of frames where the clause flips the answer. A model trained on counterfactual pairs, same frame, clause changed, learns the clause as a rule rather than as a hint.

Calibration. The confidence field gates alerts. Frontier models emit confident numbers that are not calibrated. A trained model can be rewarded for calibration directly.

Cost. A 9B model on an on-site GPU is, by our earlier inference study, two orders of magnitude cheaper per frame than a cloud API, and the images never leave the site.

Four claims, and how each will be measured

ClaimMetricWhy a small model can win
C1 Fewer false alarms on hard negatives at equal recallFalse alarms per 1,000 negative frames at 90% recall, broken down by conditionHard negatives (goods beside the exit, a person walking through, a shadow) make up half the training distribution, with exact labels
C2 Clause compliance and zone conditioningShare of counterfactual pairs answered correctlyPairs are generated by construction; the model is trained on the pair, the API only reads the clause
C3 Consistency and calibration over timeAgreement across frames minutes apart with no state change; expected calibration error; time to detect on clipsConsistency is rewarded directly during reinforcement fine-tuning; an API is stateless and uncalibrated by design
C4 (supporting) CostTokens, seconds and euros per 1,000 frames, on a site GPU and on the APIsA 9B model on-site is about a hundred times cheaper per frame

The targets we have set ourselves, as hypotheses: false alarms on hard negatives at least halved against the best API at 90% recall (H1), and description faithfulness within 3 points of the best API while being more than a hundred times cheaper (H7).1

What COCO and SA-1B are for

Not pretraining. The vision encoders of the candidate models have already seen billions of image-text pairs; another pass over 118,000 COCO photos changes nothing, and the domain is wrong anyway.2 COCO, LVIS, POPE and RefCOCO serve as regression tests in the release gate: a fine-tune must not lose general grounding. SA-1B and SA-V serve as a library of segmented objects to composite into real warehouse footage; the licence for derived products is to be checked.3

The model ladder

ModelRoleFootprint (bf16 LoRA)
Qwen3.5-4BEdge candidate, cheapest site GPU, first fine-tuning targetabout 10 GB
Qwen3.5-9BThe product modelabout 22 GB
Qwen3.8 familyNewer generation: run the whole ladder zero-shot first; if a size under 10B exists, it replaces Qwen3.5 as the base. Sizes and licence to be checked on releaseto be checked
Qwen3.8-27BUpper bound for on-premise; teacher for distillationabout 56 GB, fits a DGX Spark for inference and LoRA

All of these are public weights from the Qwen family.4

What we have not done yet

Nothing above is a result. It is a plan with targets, written down before the first measurement so that the measurement cannot bend it. The order is product first: a frozen evaluation set and a zero-shot ladder in the first two months, which alone may change which model ships; supervised fine-tuning and a release gate in months three and four; reinforcement fine-tuning after that. A paper comes only if the 9B model beats every API on at least two of C1 to C3 on a hidden test split with real positives. If it does not, the benchmark is still published, with the APIs on top, and that is a useful result too.

If you run a warehouse and want your cameras in the evaluation, or you work on vision-language models and want to compare on the same items, the next two posts describe the benchmark and the data engine.

Sources

  1. Camly AI, Beating frontier models with a small VLM: research plan, October 2026.
  2. COCO, Common Objects in Context. cocodataset.org
  3. Meta AI, Segment Anything 1 Billion (SA-1B). ai.meta.com
  4. Qwen model family on Hugging Face. huggingface.co/Qwen