A 25-meal benchmark shows calorie estimation doesn't favor the biggest model
A test on 25 Nutrition5k meals compares seven multimodal LLMs on calorie estimation under a 20% error threshold. Open-weights Muse Spark 1.3 hits 48%, while Qwen 3.8 27B stops at 16%. The ranking doesn't track model size—a useful signal for anyone ev...