zlm-v2-iab-classify-edge-enriched: multilingual IAB classification at 4–9× lower latency
ZeroGPU's multilingual, enrichment-aware classification model was evaluated against gpt-5.4-nano on 5,000 real production prompts — 3,000 of them non-English across 31 languages — using GPT-5.5-generated gold labels over the IAB Content 2.2 and Audience 1.1 taxonomies. Zero request errors.
01How this benchmark was measured
Both models saw the same 5,000 prompts under identical conditions, scored against a single frozen gold set and timed from the same co-located worker.
Sampled from live orchestration prompt logs (taskType=iab_classify), deduplicated,
median ~62 characters of short ad-copy and headline text. 3,000 of the rows are non-English,
spanning 31 languages.
Gold labels were produced by gpt-5.5, constrained to the exact IAB Content 2.2 (704 labels) and Audience 1.1 (1,567 labels) label space and validated locally. Gold reflects an LLM's judgement rather than certified human truth, so figures are best read as relative comparisons on identical data.
All requests were issued from a Fly worker in region iad, co-located with the
classification API, so client network latency is negligible (wall − server ≈ 10–70 ms).
Accuracy is measured over all 5,000 rows; latency separately at concurrency = 1 after warm-up.
02IAB classification quality
Top-1 = the model's first label is among the gold labels. F1@5 / precision@5 / recall@5 = set overlap over the returned lists. Tier-1 top-1 = the first tier-1 category matches gold. v2 lands within 11 points of gpt-5.4-nano on content top-1 and clearly ahead on audience targeting.
| Metric | zlm-v2-iab-classify-edge-enriched | ZLM v2 | gpt-5.4-nano |
|---|---|---|---|
| Content top-1 | 0.657 | 0.771 | |
| Content F1@5 | 0.438 | 0.510 | |
| Content precision@5 | 0.408 | 0.475 | |
| Content recall@5 | 0.494 | 0.567 | |
| Tier-1 top-1 | 0.574 | 0.674 | |
| Audience F1@5 | 0.292 | 0.164 |
Audience F1@5 is 78% higher than gpt-5.4-nano (0.292 vs 0.164), the metric that drives audience segmentation and targeting quality.
03English vs non-English
The multilingual path holds up on foreign-language traffic: v2 scores 0.647 top-1 on non-English content versus 0.671 on English — a 2.4-point spread.
| Metric | Set | ZLM v2 | gpt-5.4-nano |
|---|---|---|---|
| Content top-1 | English | 0.671 | 0.750 |
| Content top-1 | Non-English | 0.647 | 0.785 |
| Content F1@5 | English | 0.451 | 0.475 |
| Content F1@5 | Non-English | 0.429 | 0.533 |
| Audience F1@5 | English | 0.303 | 0.171 |
| Audience F1@5 | Non-English | 0.285 | 0.160 |
Content top-1 by language
| Language | ZLM v2 top-1 | ZLM v2 | gpt-5.4-nano |
|---|---|---|---|
| English | 0.671 | 0.750 | |
| French | 0.602 | 0.759 | |
| German | 0.619 | 0.757 | |
| Italian | 0.844 | 0.908 | |
| Spanish | 0.587 | 0.741 | |
| Hebrew | 0.549 | 0.750 | |
| Dutch | 0.578 | 0.656 | |
| Portuguese | 0.586 | 0.741 |
04Accuracy by content area
Each prompt is bucketed by its gold top-1 tier-1 category. Buckets with at least 25 prompts, sorted by volume. n = prompts in the bucket. Green bars mark the areas where v2 leads on top-1.
| Category | n | ZLM v2 top-1 | ZLM v2 | GPT top-1 | ZLM F1@5 | GPT F1@5 |
|---|---|---|---|---|---|---|
| Real Estate | 592 | 0.836 | 0.921 | 0.655 | 0.730 | |
| Medical Health | 503 | 0.738 | 0.841 | 0.440 | 0.529 | |
| Business and Finance | 396 | 0.646 | 0.795 | 0.398 | 0.477 | |
| Personal Finance | 358 | 0.682 | 0.679 | 0.461 | 0.472 | |
| Home & Garden | 279 | 0.642 | 0.788 | 0.461 | 0.469 | |
| Pop Culture | 252 | 0.436 | 0.631 | 0.337 | 0.436 | |
| Healthy Living | 241 | 0.635 | 0.813 | 0.457 | 0.553 | |
| News and Politics | 225 | 0.707 | 0.822 | 0.391 | 0.474 | |
| Style & Fashion | 216 | 0.556 | 0.616 | 0.399 | 0.476 | |
| Automotive | 211 | 0.621 | 0.739 | 0.407 | 0.475 | |
| Technology & Computing | 210 | 0.767 | 0.819 | 0.524 | 0.493 | |
| Travel | 198 | 0.702 | 0.828 | 0.426 | 0.491 | |
| Shopping | 136 | 0.507 | 0.691 | 0.397 | 0.509 | |
| Sports | 126 | 0.786 | 0.794 | 0.412 | 0.500 | |
| Family and Relationships | 114 | 0.614 | 0.772 | 0.458 | 0.476 | |
| Pets | 111 | 0.757 | 0.739 | 0.482 | 0.400 | |
| Food & Drink | 108 | 0.750 | 0.824 | 0.506 | 0.546 | |
| Education | 97 | 0.660 | 0.835 | 0.402 | 0.503 | |
| Hobbies & Interests | 75 | 0.560 | 0.640 | 0.343 | 0.481 | |
| Events and Attractions | 59 | 0.610 | 0.678 | 0.364 | 0.439 | |
| Careers | 51 | 0.392 | 0.765 | 0.302 | 0.469 | |
| Science | 40 | 0.725 | 0.800 | 0.474 | 0.506 | |
| Television | 39 | 0.538 | 0.692 | 0.265 | 0.549 | |
| Books and Literature | 37 | 0.460 | 0.757 | 0.309 | 0.482 | |
| Movies | 36 | 0.861 | 0.833 | 0.364 | 0.566 | |
| Fine Art | 29 | 0.517 | 0.655 | 0.301 | 0.493 |
05Latency
Measured warm, one request at a time (concurrency = 1) — the latency a single
production caller experiences. ZLM figures are the server-reported processing_time_ms;
gpt-5.4-nano has no server timing, so its client wall-clock from the same co-located worker is used.
| Model | Set | p50 | p90 | p95 | mean |
|---|---|---|---|---|---|
| ZLM v2 edge-enriched | English | 238 ms | 348 ms | 370 ms | 243 ms |
| ZLM v2 edge-enriched | Non-English | 515 ms | 856 ms | 1061 ms | 548 ms |
| gpt-5.4-nano | English | 2232 ms | 3169 ms | 3505 ms | 3381 ms |
| gpt-5.4-nano | Non-English | 2173 ms | 3107 ms | 3656 ms | 2510 ms |
Median speed-up
The bottom line
- 4–9× faster at the median than gpt-5.4-nano, on both English and non-English traffic.
- Audience F1@5 0.292 vs 0.164 — materially stronger audience signal extraction.
- Content top-1 0.657 vs 0.771 — within 11 points of gpt-5.4-nano on the same gold labels.
- Consistent across languages: 0.647 top-1 on non-English vs 0.671 on English, across 31 languages.
- Leads gpt-5.4-nano on top-1 in Personal Finance, Pets and Movies, and on F1@5 in Technology & Computing and Pets.
Notes & limitations
| Test set | 5,000 real production prompts · 3,000 non-English across 31 languages · median ~62 chars |
| Taxonomies | IAB Content 2.2 (704 labels) · IAB Audience 1.1 (1,567 labels) |
| Gold labels | gpt-5.5, constrained to the exact label space, validated locally |
| ZeroGPU model | zlm-v2-iab-classify-edge-enriched — multilingual, content + audience |
| Reference model | gpt-5.4-nano |
| Environment | Fly worker, region iad, co-located with the classification API |
| Latency protocol | warm, concurrency = 1, split by English / non-English |
| Date | 2026-08-03 |