Customer case study · B2B outbound
Limelight scales AI enrichment to millions of rows with ZeroGPU
How a B2B outbound agency processed 11.6 billion tokens across nearly 2.9 million prospect enrichments while dramatically reducing the cost of AI inference.
Customer
Limelight
Cold email that books qualified meetings for SaaS founders.
Limelight runs outbound end to end for B2B SaaS founders. They build the prospect list, write the copy, and run the campaigns that put qualified sales conversations on the calendar.
thelimelightagency.comSix weeks on ZeroGPU · Jul 29 - Sep 6, 2026
The challenge
At millions of rows, enrichment becomes a cost problem
For Limelight, list quality is the product.
Every prospect has to be classified and qualified before an email goes out. But some enrichment jobs can involve millions of prospects, with thousands of tokens of website content processed for every row.
That means the cost of general-purpose AI models can quickly become the constraint on how large a list Limelight can economically build.
“Flow before was using Claude Haiku, then I switched to Gemini Flash-Lite. But it was still quite expensive because some enrichments can be millions of rows.”
Before ZeroGPU
Scrape
Prospect website · ~3,800 input tokens per row
Classify
General-purpose LLM determines company type
Extract
Pull client-specific signals from website content
Qualify
Build the final outbound prospect list
The constraint wasn’t demand.
It was how many prospects Limelight could economically enrich.
The solution
Use the right open-weight model for the workload
Limelight connected to ZeroGPU through its OpenAI-compatible API and used multiple open-weight models across its enrichment, research, and outbound workflows.
This allowed Limelight to evaluate and use different models while keeping the application integration simple.
For its largest production enrichment workload, gpt-oss-120b processed scraped website content and performed classification and client-specific signal extraction at massive scale.
Prospect domain
Website content
ZeroGPU API
Open-weight model
Classification + signal extraction
Qualified outbound list
Models used
Different models for different outbound workloads
Limelight used multiple open-weight models through the same ZeroGPU API while optimizing its enrichment, research, and personalization workflows.
Qwen3
Personalise at list-wide volume without frontier pricing
DeepSeek
Fit an entire website into one pass for richer research
gpt-oss-120b
Reasoning for personalisation and multi-step qualification
The results
Nearly 2.9 million enrichments. Zero failed requests.
Limelight went from signup to production in approximately 34 hours.
The workload could also be extremely bursty. On its busiest day, Limelight processed more than 1.14 million enrichment requests through ZeroGPU.
Across nearly 2.9 million production calls, ZeroGPU recorded zero failed requests.
The economics
The same job at a fraction of the cost
| Stack | Total cost | Cost / 1K rows |
|---|---|---|
| Claude Haiku 4.5 | $14,299 | $4.97 |
| Gemini 3.5 Flash-Lite | $4,956 | $1.72 |
| ZeroGPU · gpt-oss-120b | $1,560* | $0.54* |
Same 10.97B input tokens and 667M output tokens.
*ZeroGPU reflects actual cash charged during the measured period, including account credits. Claude Haiku 4.5 and Gemini 3.5 Flash-Lite are calculated using published list prices for the identical token workload. ZeroGPU usage at standard consumption value was approximately $2,002.
The workload
One inference turns website content into qualified prospect data
For the featured production workflow, gpt-oss-120b reads scraped website content and performs two critical jobs:
- 1.Classifies the company into a specific type
- 2.Extracts client-specific signals from the homepage
The workflow was validated against a labelled set of approximately 100 to 150 rows and tuned to roughly 90% precision and recall before being scaled into production.
Website text
~3,813 input tokens / row
gpt-oss-120b
Company classification + client-specific signals
Qualified prospect
Why ZeroGPU
Not every task needs a frontier model
High-volume AI pipelines contain many repeatable workloads such as classification, extraction, enrichment, personalization, summarization, and signal detection.
ZeroGPU provides specialized small and open-weight models through one OpenAI-compatible inference API, allowing companies to use the right model for each workload instead of automatically sending everything to a frontier model.
Right-sized models
Use the model that clears the accuracy bar for the task, from purpose-built small models to larger open-weight reasoning models.
One API, no GPU ops
OpenAI-compatible endpoints without having to provision clusters, manage autoscaling, or pay for idle GPU capacity.
Built for bursty workloads
Scale from small model evaluations to millions of requests without rebuilding the application around dedicated infrastructure.
What comes next
Right-size even more of the enrichment pipeline
The large production workload featured here runs classification and extraction through a general-purpose open-weight reasoning model.
Over time, deterministic parts of that workflow can move to smaller purpose-built ZeroGPU models, reducing inference costs even further.
The purpose-built models below are potential next-step optimizations.
They are not all models currently used by Limelight.
Domain
Classify
Extract
Scrub PII
Personalise
zlm-v1-iab-domain-classifier
Classify directly from a domain
gliner2-base-v1
Extract structured fields from website text
deberta-v3-small
Score ICP fit or classify against custom labels
gliner-multi-pii-v1
Detect and redact PII
Qwen3
Personalise at list-wide volume without frontier pricing
DeepSeek
Fit an entire website into one pass for richer research
gpt-oss-120b
Reasoning for personalisation and multi-step qualification
The takeaway
AI enrichment becomes a unit-economics problem at scale
At a few thousand calls, model pricing can feel insignificant.
At millions of enrichment rows and billions of tokens, model selection becomes core infrastructure.
Limelight shows how teams can use ZeroGPU to evaluate open-weight models, move high-volume workloads away from expensive frontier models, and scale without managing their own GPU infrastructure.
Run workloads on ZeroGPU
Have millions of AI tasks that don’t need a frontier model?
Run specialized small and open-weight models through one OpenAI-compatible API with ZeroGPU.