Can software sort vehicle photos without a big AI project?
A car rental business had 67,000 handover photos already tagged by staff. Reusing those tags sorted new photos with no training and no new hardware.
Your staff photograph every vehicle at handover and again on return: the odometer, the fuel gauge, each side of the car, the signed checklist, any damage. Each photo is supposed to be tagged with what it shows. Many arrive untagged, and somebody has to go through them by hand. That is skilled people doing filing, and the record they are creating is the one you will reach for when a customer disputes a scratch.
The usual answer is an AI project: collect labelled examples, rent expensive hardware, train a model over a couple of weeks. You do not want that for a tagging problem.
For this car rental business there was a cheaper route, because of two facts. There were already around 67,000 photos tagged by staff, a by-product of years of ordinary work. And the only server available had eight processor cores, no graphics hardware, and was already running the live application and two databases. Training anything on it would have used every core for hours.
What was actually going on
The photos had been labelled for operational reasons and never thought of as training material. Most businesses with a sorting problem are sitting on the same thing: support tickets with categories, documents with types, transactions with codes. Before commissioning a model, ask what has already been labelled by the ordinary course of business.
What we changed
We used a general-purpose image model off the shelf, unchanged, to turn each of the 67,000 photos into a set of numbers describing what it looks like, once. For each of 24 photo types we averaged those numbers into one typical example. A new photo is compared with the 24 and given the closest match. Nothing is trained, and adding a new type later takes seconds.
It runs away from the live server. It reads two light queries over a secure connection and fetches pictures straight from storage, so the production server is never in the picture path. It checks the server is not busy before it works and limits how fast it writes. It only ever touches photos that are still untagged, so it cannot overwrite a tag a person made. That last rule is worth keeping on its own: a system that suggests must never be able to overwrite a decision a person already made.
We measured it on tagged production photos that were kept out of the building of the typical examples. Across everything, the match was right 70.6% of the time. That is not good enough: three wrong tags in ten makes more work than it saves.
But the method also says how sure it is, as the gap between its best match and its second best. A large gap means the photo sits clearly nearer one type. A small gap means it is nearly a coin toss between two. Accepting only confident answers:
- gap above 0.015: 86.9% right, covering 55% of photos
- gap above 0.025: 89.6% right
- gap above 0.04: 96.1% right, covering 30% of photos
So you choose the accuracy the operation needs and accept less coverage in return. At the tightest setting it correctly tags three photos in ten and leaves the rest to a person, at an accuracy where nobody has to check its work.
What it did not fix
Some categories are simply weak. Odometer and fuel readings, the signature and the checklist were near 100%, because they look alike every time. Damage close-ups were 36%. Damage is not a visual category in any coherent sense: a scratched bumper and a cracked windscreen have nothing in common except the intention. The decision recorded was to leave damage photos to a person, not to keep trying to improve it.
The other finding was that the method recognises which view of the car a photo shows but often cannot tell left from right, a known weakness of this kind of model. The design absorbs it without special code. Left and right sit close together, so a photo that could be either gets a small gap, and the confidence setting already sends it to a person.
And the result is not clever. It handles part of a tedious job at an accuracy where nobody checks it, and declines the rest. We have no figure for the hours the manual tagging used to take, so none is claimed.
The pattern, for anyone with a sorting job your staff dislike
Ask what has already been labelled by the people doing the work. Ask for a measured accuracy table, not a claim. Ask that the software reports how sure it is, and that the unsure cases go to a person. And ask that nothing it does can overwrite what a person decided.
What to settle up front is not accuracy but what happens to the cases the system is not confident about, and who decides where that line sits. In a small operation, this is usually what AI in production looks like: a measured cut in volume, with the remainder routed to the people who were doing all of it before.
Where this ends up
The photographs are the handover and return record Sazinga Rentals keeps against a booking, which is why the tagging is worth reducing at all, and why the one rule that cannot bend is that it never overwrites what a person wrote.