I have never enjoyed paying for a calorie tracker. It always felt like renting a calculator. So when the renewal notice landed I cancelled it and decided that for two weeks I would photograph every meal and ask ChatGPT to do the arithmetic instead.
To make it a fair test rather than a vibe, I weighed everything first and logged it against USDA composition data before I uploaded a single photo. That way I had a real number to compare each estimate against, and ChatGPT never saw it.
Here is the short version. It is much better at this than it has any right to be, for about three days. Then the thing that breaks is not the accuracy.
Breakfast: it got surprisingly close
Egg whites and steamed spinach, which is what I eat most mornings because it requires no decisions. Flat plate, even lighting, nothing hidden under anything.
My weighed log came to 212 calories. ChatGPT estimated 190 to 240, and said it was assuming no added fat, which was correct.
That is a good answer. If I had not weighed it I would have taken the top of the range and moved on, and I would have been about 13% over — which, for a meal I eat five times a week, is a rounding error I could live with.
Lunch: still good, and it explained itself
Chicken breast and roasted vegetables in a shallow container, photographed at an angle rather than straight down.
Weighed log: 418 calories. ChatGPT: 380 to 450, with a note that the oil on the vegetables was the biggest uncertainty and that it had assumed about a teaspoon.
That reasoning is the part no app gives you. A tracker hands you a number and stares. ChatGPT tells you the number is mostly a bet about the oil, which is genuinely how a cook should think about it. For learning why a dish costs what it costs, this is the best tool I have used.
Dinner: the overhead shot broke it
Beef and rice in a deep bowl, photographed from directly above, which is how everyone photographs a bowl.
Weighed log: 540 calories. ChatGPT: 700 to 820.
It could not see depth. From overhead a bowl with two inches of rice looks identical to a bowl with four, and it guessed high — about a third high. This is not a ChatGPT failing so much as a photograph failing, and it is the single most common way people take pictures of food.
Three meals in, my honest scorecard was: impressive on flat food, unreliable on anything with volume, and useful in a way an app isn’t when you want to understand a dish.
If the experiment had stopped there I would have written a mildly positive piece. It ran another eleven days.
The problem that actually ended it
Ask twice, get two answers. On day six I re-uploaded the dinner photo from day two, out of curiosity. It came back 610 to 700 — a different answer to the same picture, several hundred calories from the first attempt, with nothing changed except how I happened to word the request.
That is the finding that matters, and it is easy to miss if you only test for a day. When you are tracking, you are not measuring a meal. You are measuring change over weeks. A number that drifts based on your phrasing means your week-over-week comparison is partly measuring your prompt style, and you cannot tell which part.
A database does not have this problem. A chicken breast is the same chicken breast on Tuesday as it was on Friday.
It cannot scan a barcode. I eat packaged food like everyone else — yoghurt, bread, a tin of chickpeas, the protein bar in my bag. All of it has exact figures printed on the side and a code designed to be scanned. That is a two-second job in any tracker and a typing exercise here.
It does not remember your day. No running total, no “you have 400 left”, no yesterday to compare against. I kept the tally in a notes app, which is to say I hand-built the worst food diary ever made. By day nine I had started skipping entries, and skipping entries is how tracking dies.
There is nothing to look at. No ring closing, no week as a bar chart, no weight line against intake. Just paragraphs of text. I had not appreciated how much of the habit is carried by that feedback until I removed it, and two weeks of prose was enough to bore me out of my own experiment.
And barely any nutrients. Calories and a rough macro split. Nothing on fibre, sodium, iron — nothing you would actually change a shopping list over.
The maths I did not see coming
The whole premise was that I resented paying for a tracker. So it is worth writing down what the alternative costs.
| Per year | |
|---|---|
| ChatGPT Plus | $240 |
| MyFitnessPal Premium | $79.99 |
| MacroFactor | $71.99 |
| Cronometer Gold | $54.99 |
| Lose It! Premium | $39.99 |
| PlateLens Premium | $34.99 |
I was avoiding a $35-a-year subscription by leaning on a $240-a-year one. If you already pay for ChatGPT, the marginal cost is zero and that is a fair argument. As a strategy for not paying for things, it is the most expensive route on the table — and PlateLens has a free plan that never expires, with three photo scans a day plus unlimited manual and barcode logging, so the genuinely free option was never the chatbot either.
What the purpose-built tools do better
Two weeks of doing it the hard way made the differences sharper than I would have described them beforehand.
They are more precise, and somebody outside the company has checked. PlateLens’s calorie error was measured at ±1.1% by the Dietary Assessment Initiative across 180 weighed reference meals, then reproduced independently by the open-source Foodvision Bench project on its own 231-meal set. It is the only figure in this category a second lab has replicated — a low bar that almost nobody clears. No general chatbot has an equivalent number, and cannot, because it is not performing the same operation.
They are easier to follow. The day adds itself up. Yesterday is still there. The target moves when your weight moves. None of it asks you to remember anything, and the whole game is whether you are still logging in week six.
They are visual. The week as a chart, calories in against calories out, a weight trend with the daily noise smoothed off. It sounds trivial next to accuracy and it is not — it is most of what keeps people going.
And they do what a chatbot structurally cannot. Barcode scanning. 82+ nutrients per entry instead of one figure. Voice logging when your hands are covered in flour. A full web app on the same diary for when you would rather work at a desk. Exporting your entire history as a file when you want to leave. There is no diary inside a chat window, so none of these can exist there.
Where I landed
I did not renew the app I cancelled. I moved to PlateLens, for the unglamorous reason that it was the only one that did not make me choose: photograph what you cooked, scan or search what came in a packet, and neither path is the consolation prize. The deep bowl that defeated ChatGPT is still the hardest case for any camera — that is physics, not software — but there I can correct the portion in two taps instead of re-arguing with a chatbot.
The part of the experiment I liked, I kept. PlateLens runs a read-only MCP server you can connect to an AI assistant on any account, free tier included. The assistant reads your real diary — meals, trends, weight history — and answers questions about it. It cannot log or change anything, which is exactly right.
So I still ask an AI about my food. It just reads the tracker first instead of guessing from a photograph.
Should you try it?
If you want to understand what is driving the calories in an unfamiliar dish, or you are staring at takeout with no nutrition information anywhere, spend ten minutes with ChatGPT and a photo. It is a good teacher and it will show its working in a way no app does.
If you want to track — across weeks, with numbers you can compare to each other — you need something that remembers, stays consistent, and has a real database underneath. That is not what a chatbot is, and it is faintly ridiculous that we all had to try it to find out.