Benchmarks & Research
Head-to-head evaluations of ZeroGPU Nano Language Models, and production research on the compute edge network that runs them.
Model benchmarks
Nano Language Models evaluated head-to-head against frontier models.
Moderation vs OpenAI omni-moderation
0.899 vs 0.853 binary F1, wins on 9 of 13 harm categories, and 1.2-1.8x faster p50 latency on production-range inputs.
View benchmarkMultilingual IAB Classify with Enrichment vs GPT-5.4 Nano
5,000 production prompts across 31 languages: 4-9x lower median latency, 0.292 audience F1@5 vs 0.164, and content top-1 within 11 points of GPT-5.4 Nano.
View benchmarkIAB Classify vs GPT-5.4 Nano
Head-to-head evaluation across 10,000 samples: 66.0% win rate over GPT-5.4 Nano on accuracy, speed, cost, and reliability.
View benchmarkDomain Classify vs GPT-5.4 Nano
Held-out benchmark mapping domains to IAB categories: more accurate and roughly 50x faster than GPT-5.4 Nano.
View benchmark