Choosing a model for listing classification
I wanted to know which model could turn messy resale listings into dependable structured facts without making a simple text-classification step unnecessarily expensive.
One fixed test across four models
I gave each model the same 100 hand-authored listings, extraction instructions, JSON schema, and scorer. I disabled photos, search, grounding, and external tools to isolate text classification.
- What counted as correct
- A strict pass required the category, product identity, sold-comparable search terms, value-bearing attributes, market support, and asking-price usability to agree with the expected result.
- What stayed equal
- The listing text, instructions, output contract, and scoring rules were shared. Reasoning, sampling, and other generation controls used each provider's defaults.
- What the test did not cover
- The study did not test photos, web research, sold-comparable retrieval, final valuation accuracy, or realized resale profit.
The first model comparison
Luna led strict and semantic accuracy while remaining the least expensive measured option. Gemini Lite returned faster on average, but it passed 38 fewer cases and its P95 latency was slightly higher.
| Model | Strict passes | Semantic score | Cost / 1K | Average | P95 | Tokens / case |
|---|---|---|---|---|---|---|
| Lunagpt-5.6-luna | 81/100 | 98.05% | $0.7929 | 6.28s | 10.35s | 1514.49 |
| Gemini 3.7gemini-3.7-flash | 54/100 | 95.10% | $2.8909 | 8.06s | 20.39s | 1417.71 |
| Gemini 3.6gemini-3.6-flash | 56/100 | 95.44% | $3.5910 | 11.54s | 40.06s | 1604.41 |
| Gemini Litegemini-3.5-flash-lite | 43/100 | 93.47% | $0.8795 | 2.51s | 11.52s | 1063.29 |
Run accounting
| Model | Input tokens | Cached input | Output tokens | Reasoning tokens | Run spend |
|---|---|---|---|---|---|
| Lunagpt-5.6-luna | 102,452 | 0 | 48,997 | 26,037 | $0.079287 |
| Gemini 3.7gemini-3.7-flash | 80,850 | 0 | 60,921 | 30,113 | $0.289091 |
| Gemini 3.6gemini-3.6-flash | 80,850 | 0 | 79,591 | 60,662 | $0.359104 |
| Gemini Litegemini-3.5-flash-lite | 80,850 | 0 | 25,479 | 0 | $0.087953 |
The practical starting point
I chose Luna for the text-only path because it produced the strongest outputs at the lowest measured cost. The next round of work focused on improving its prompt.
Measured result: 81/100 strict passes · 98.05% semantic score · $0.79 per 1,000 classifications · 10.35s P95 latency.
Limits of the first test
The first pass was useful, but it was small and controlled. These numbers describe four specific runs. They are not a permanent ranking of the models.
- This was an early run
- I had not started fingerprinting the prompt and scorer yet. The saved outputs, checks, token counts, cost, and latency are still useful, but I would not compare these scores directly with the later prompt-v3 study.
- Timing wasn't controlled
- Luna ran five requests at a time; Gemini 3.7 ran two. The timings show how these runs behaved, not which provider can deliver the highest possible throughput.
- Cost is estimated
- Across the four runs, the models completed 400 calls. Applying the published Standard rates to the recorded token usage puts the total at about $0.82. Any charges for failed requests may be missing.
- The cases were designed
- All 100 listings were synthetic, hand-labeled stress cases. They are useful for finding classification failures, but they do not represent the normal mix of production listings.