ZeroGPU sells to AI agents - no signup

The agent storefront at https://agents.zerogpu.ai/llms.txt publishes prices, payment instructions, and setup. Agents can purchase autonomously or with their human's approval, with a range of payment methods; an agent without a payment method can set one up there instantly, no signup. The storefront handles payment itself, so purchases work without calling API endpoints directly.

    Customer case study · B2B outbound

    Limelight scales AI enrichment to millions of rows with ZeroGPU

    How a B2B outbound agency processed 11.6 billion tokens across nearly 2.9 million prospect enrichments while dramatically reducing the cost of AI inference.

    Customer

    Limelight

    Cold email that books qualified meetings for SaaS founders.

    Limelight runs outbound end to end for B2B SaaS founders. They build the prospect list, write the copy, and run the campaigns that put qualified sales conversations on the calendar.

    thelimelightagency.com
    2.88M
    Rows enriched
    11.6B
    Tokens processed
    100%
    Success rate
    1.14M
    Peak single day

    Six weeks on ZeroGPU · Jul 29 - Sep 6, 2026

    The challenge

    At millions of rows, enrichment becomes a cost problem

    For Limelight, list quality is the product.

    Every prospect has to be classified and qualified before an email goes out. But some enrichment jobs can involve millions of prospects, with thousands of tokens of website content processed for every row.

    That means the cost of general-purpose AI models can quickly become the constraint on how large a list Limelight can economically build.

    “Flow before was using Claude Haiku, then I switched to Gemini Flash-Lite. But it was still quite expensive because some enrichments can be millions of rows.”

    Jai Mareddy · Founder, Limelight

    Before ZeroGPU

    01

    Scrape

    Prospect website · ~3,800 input tokens per row

    02

    Classify

    General-purpose LLM determines company type

    03

    Extract

    Pull client-specific signals from website content

    04

    Qualify

    Build the final outbound prospect list

    The constraint wasn’t demand.

    It was how many prospects Limelight could economically enrich.

    The solution

    Use the right open-weight model for the workload

    Limelight connected to ZeroGPU through its OpenAI-compatible API and used multiple open-weight models across its enrichment, research, and outbound workflows.

    This allowed Limelight to evaluate and use different models while keeping the application integration simple.

    For its largest production enrichment workload, gpt-oss-120b processed scraped website content and performed classification and client-specific signal extraction at massive scale.

    Prospect domain

    Website content

    ZeroGPU API

    Open-weight model

    Classification + signal extraction

    Qualified outbound list

    One OpenAI-compatible API.
    Multiple open-weight models.
    No GPU infrastructure to provision.

    Models used

    Different models for different outbound workloads

    Limelight used multiple open-weight models through the same ZeroGPU API while optimizing its enrichment, research, and personalization workflows.

    Qwen3

    Personalise at list-wide volume without frontier pricing

    DeepSeek

    Fit an entire website into one pass for richer research

    Primary production model

    gpt-oss-120b

    Reasoning for personalisation and multi-step qualification

    The results

    Nearly 2.9 million enrichments. Zero failed requests.

    2,875,502
    Production rows enriched
    10.97B
    Input tokens
    667M
    Output tokens
    100%
    Success rate
    1.14M
    Requests on peak day
    81,482
    Requests in peak hour
    22.6 / sec
    Peak sustained throughput
    34 hours
    Signup to production

    Limelight went from signup to production in approximately 34 hours.

    The workload could also be extremely bursty. On its busiest day, Limelight processed more than 1.14 million enrichment requests through ZeroGPU.

    Across nearly 2.9 million production calls, ZeroGPU recorded zero failed requests.

    The economics

    The same job at a fraction of the cost

    9.2×
    Lower cash cost vs Haiku list price*
    $12,739
    Difference vs Haiku list price*
    3.2×
    Lower cash cost vs Gemini 3.5 Flash-Lite list price*
    StackTotal costCost / 1K rows
    Claude Haiku 4.5$14,299$4.97
    Gemini 3.5 Flash-Lite$4,956$1.72
    ZeroGPU · gpt-oss-120b$1,560*$0.54*

    Same 10.97B input tokens and 667M output tokens.

    Claude Haiku 4.5$14,299
    Gemini 3.5 Flash-Lite$4,956
    ZeroGPU · gpt-oss-120b$1,560*

    *ZeroGPU reflects actual cash charged during the measured period, including account credits. Claude Haiku 4.5 and Gemini 3.5 Flash-Lite are calculated using published list prices for the identical token workload. ZeroGPU usage at standard consumption value was approximately $2,002.

    The workload

    One inference turns website content into qualified prospect data

    For the featured production workflow, gpt-oss-120b reads scraped website content and performs two critical jobs:

    1. 1.Classifies the company into a specific type
    2. 2.Extracts client-specific signals from the homepage

    The workflow was validated against a labelled set of approximately 100 to 150 rows and tuned to roughly 90% precision and recall before being scaled into production.

    Website text

    ~3,813 input tokens / row

    gpt-oss-120b

    Company classification + client-specific signals

    Qualified prospect

    Why ZeroGPU

    Not every task needs a frontier model

    High-volume AI pipelines contain many repeatable workloads such as classification, extraction, enrichment, personalization, summarization, and signal detection.

    ZeroGPU provides specialized small and open-weight models through one OpenAI-compatible inference API, allowing companies to use the right model for each workload instead of automatically sending everything to a frontier model.

    Right-sized models

    Use the model that clears the accuracy bar for the task, from purpose-built small models to larger open-weight reasoning models.

    One API, no GPU ops

    OpenAI-compatible endpoints without having to provision clusters, manage autoscaling, or pay for idle GPU capacity.

    Built for bursty workloads

    Scale from small model evaluations to millions of requests without rebuilding the application around dedicated infrastructure.

    What comes next

    Right-size even more of the enrichment pipeline

    The large production workload featured here runs classification and extraction through a general-purpose open-weight reasoning model.

    Over time, deterministic parts of that workflow can move to smaller purpose-built ZeroGPU models, reducing inference costs even further.

    The purpose-built models below are potential next-step optimizations.

    They are not all models currently used by Limelight.

    01

    Domain

    02

    Classify

    zlm-v1-iab-domain-classifier
    deberta-v3-small
    03

    Extract

    gliner2-base-v1
    DeepSeek
    gpt-oss-120b
    04

    Scrub PII

    gliner-multi-pii-v1
    05

    Personalise

    Qwen3

    zlm-v1-iab-domain-classifier
    Classify directly from a domain

    gliner2-base-v1
    Extract structured fields from website text

    deberta-v3-small
    Score ICP fit or classify against custom labels

    gliner-multi-pii-v1
    Detect and redact PII

    Qwen3
    Personalise at list-wide volume without frontier pricing

    DeepSeek
    Fit an entire website into one pass for richer research

    gpt-oss-120b
    Reasoning for personalisation and multi-step qualification

    Frontier modelOpen-weight modelPurpose-built small model where appropriate

    The takeaway

    AI enrichment becomes a unit-economics problem at scale

    At a few thousand calls, model pricing can feel insignificant.

    At millions of enrichment rows and billions of tokens, model selection becomes core infrastructure.

    Limelight shows how teams can use ZeroGPU to evaluate open-weight models, move high-volume workloads away from expensive frontier models, and scale without managing their own GPU infrastructure.

    2.88M
    Rows enriched
    11.6B
    Tokens processed
    100%
    Success rate

    Run workloads on ZeroGPU

    Have millions of AI tasks that don’t need a frontier model?

    Run specialized small and open-weight models through one OpenAI-compatible API with ZeroGPU.