The pricier image model labelled the wrong picture perfectly
We asked two AI image models for a pizza with three of eight slices left. The dearer one drew a perfect 3/8 label over a whole pizza. A parent would approve it.
You want pictures in your product, and you ask which AI image model is the cheapest that is still accurate. The obvious answer is to pay more for the better one. We tested that, and the dearer model produced the worse result.
The product is a children’s learning app with a library of lessons and no pictures in any of them. The test request was exact: a pizza with eight equal slices, three remaining, labelled 3/8. The cheaper model drew a whole, photoreal pizza with no label. The one that cost about four and a half times as much drew a clean, correct “3/8” above a whole pizza.
The second is the worse failure, because it looks finished. A parent reviewing it would approve it, and a child counting the slices would learn the wrong answer. The cost of a wrong picture in a lesson is not the price of the picture. It is a mistake that has been signed off.
What was actually going on
An image model is asked to do two jobs at once: decide what to draw, and draw it. Deciding is language, and drawing is geometry. Models are good at the look of a picture and poor at the count of things in it. The label came out right because text is easy to make look right. The slices came out wrong because nothing was counting them.
We also checked what was available instead of reading it from web pages, which had it wrong. The region this app runs in offers no image models of any kind. The one other region we checked with a text-to-image model offered a single legacy one, closed to new customers and eight days from retirement. The live options were in a different region again.
Then we looked at the library itself, which changed the plan. It held 59 lessons and 975 questions, with no images anywhere. The four drawing types we had first built served about three lessons of the 59. What the library needed most was pictures for vocabulary, in 18 French lessons with 340 questions, then charts and timelines, then abacus beads for a subject about a device it never shows.
What we changed
We split the job. A language model reads the lesson line and writes a short specification, such as eight slices with five eaten. Ordinary code then draws from that specification. The model never touches a pixel, so it cannot draw seven slices on an eight-slice pizza. The same idea made the 3D shapes: we wrote the geometry by hand, because a generated box cannot promise to be four by three by two, and a child counting unit cubes in a wrong picture learns the wrong answer.
Testing it live found four things no reading would have. The image provider rejects prompts that are not in English, which would have failed every French picture, and it does so with a success code and no image, so it looked like an outage. The model sometimes returned an empty plan for an obviously illustrable lesson, and needed one retry; it then passed five times out of five. And a French example in our own prompt leaked into the description of an English lesson.
The cost, priced before building: about 70 French nouns at three cents each is roughly two dollars, once. Everything else costs a fraction of a cent per drawing. Illustrating the whole existing library came to under five dollars.
What it did not fix
We priced the French vocabulary pictures. The log does not record that they were generated, and they still need a real image model, so accuracy on a photographic picture is avoided, not solved.
The hand-built shapes cover seven solids. Anything outside them needs another renderer. And the comparison used a single prompt. One pizza is a fair warning, not a survey of every model.
There was also a mistake of ours. The early image tests ran against the wrong cloud account on the laptop, a client’s, because it was the first set of working credentials. About 17 cents landed on the wrong bill. It was small, and it was ours.
The pattern, for anyone buying AI-made pictures or documents
A result that looks finished is not evidence that it is correct. Ask what in the process counts, checks or measures, and if the answer is nothing, the model’s confidence is all you have.
Split the work where you can. Let the model decide and describe, and let something that cannot invent do the drawing, the adding or the lookup. Then test with a request whose right answer you already know, such as a count, and test the model you did not expect to fail.
Where this ends up
The general pattern here, a model that decides and deterministic code that does, is how we build custom systems that have to be right rather than merely plausible, as described under custom software development.