ZeroGPU Research · Edge Network

    One million AI inferences. 460× less energy than the data center.

    Over one weekend, we ran 1,035,800 inference requests against ZeroGPU's live fleet of consumer devices. 812,146 requests were served entirely on-device by 5,078 distinct devices across 20+ countries, using 460× less energy than the same volume on datacenter inference. This is what we measured.

    63 h
    Continuous, zero interventions
    1.03M
    Requests offered
    812K
    Served 100% on-device
    5,078
    Distinct devices served
    20+
    Countries served
    1.21 s
    Median latency
    66M
    Tokens processed
    01 · The Bet

    The world can't build compute fast enough. We won't try.

    AI inference demand is compounding faster than data centers, power, and GPUs can be built. But a large share of production AI workloads (classification, extraction, summarization, moderation, routing) never needed frontier GPUs in the first place. Specialized small and nano models can run these tasks on the underutilized compute already sitting in devices everywhere.

    ZeroGPU aggregates that idle compute into a single programmable inference layer. Developers hit one OpenAI-compatible API; the orchestration layer routes each request to the right model on the right compute. The network is hybrid by design, a trusted network of edge devices and cloud infrastructure working together to deliver fast, secure, serverless inference, with each request landing wherever it runs best. Frontier models for reasoning, ZeroGPU for everything repeatable around it.

    02 · Methodology

    63 hours, no cloud fallback

    From Friday night to Monday afternoon (July 24 to 27, 2026), a load generator sent a continuous stream of inference requests at ZeroGPU's live device fleet for 63 unattended hours, laptops and desktops running our Chrome extension, joined by the first mobile devices serving through our Android and Telegram SDKs. The workload mirrored real production traffic: 92% zero-shot classification running on deberta-v3-small, and 8% summarization on t5-small, the class of specialized small models that high-volume tasks like IAB classification and signal extraction run on in production.

    Two deliberate constraints made this a real test rather than a demo:

    No cloud fallback. For this campaign, cloud fallback was disabled entirely. Every request either ran on someone's actual device or was refused with an explicit error. Every weakness of the fleet shows up as a visible refusal instead of being silently absorbed by a cloud tier.

    No preset rate. The load generator's auto-throttle read the live count of available devices and set its own pace. The fleet itself decided how fast the test ran, climbing to 6.9 requests per second at peak.

    The Honest Number

    78.3% of offered requests were served; 21.7% were refused. With no cloud tier to absorb misses, that refusal rate is the fleet's true, unpadded capacity signature, and the trade for a latency distribution with no long tail. Excluding one workload slice we knew the fleet couldn't serve, success was 84.9%.

    03 · Findings

    The network follows its own supply

    Over three full day/night cycles, offered load and available devices traced the same curve. The fleet grew every European morning and the throttle discovered it within minutes; the fleet shrank every night and the throttle backed off. No rate was ever set by hand.

    Offered load (rps, green) vs. mean available devices (gray), 6-hour buckets over 63 hours. Three day/night cycles reproduce each other with no operator input.

    A 16x load swing moved median latency by less than ±9%

    Median latency stayed inside a narrow band for the entire run while throughput swung 16x. The small overnight drift is compositional (only the slower subset of devices stays online at night) and every morning the median snapped back. Three identical cycles is the strongest stability evidence a soak test can produce.

    Served latency band per 6-hour bucket: p50 (green) and p95 (gray). Full-run medians: p50 1,211 ms · p95 2,103 ms · p99 2,346 ms. The Sunday-evening dip was a temporary infrastructure fast-path.

    The busiest device carried 1 request in 250

    A naive scheduler with a small pool hammers whichever devices look best, overheating them and creating single points of failure. ZeroGPU's spread selection (latency-weighted sampling with per-device performance tracking, budget caps, and automatic quarantine of misbehaving devices) kept a million dispatches genuinely distributed.

    ConcentrationShare of trafficIn plain terms
    Busiest device0.41%3,345 of 812,146 requests
    Top 10 devices2.7%No single point of failure
    Top 500 devices44.6%Half the traffic needs 500+ devices
    Median device70 requests~1 request per hour

    Devices aren't scarce; idle devices under load are

    The campaign kept discovering new devices for its entire duration: a recurring core of ~4,000 devices returned every day, while new installs and one-visit devices arrived at 45 to 95 per hour indefinitely.

    The most instructive measurement came at shutdown. During load, the pool of instantly-available devices read 19 to 65. Within 30 minutes of the generator stopping, it rebounded to 311. Devices were passing through the available state between jobs, a flow with full turnover every ~90 seconds, not a static pool. Capacity planning on registered-device counts overestimates by two orders of magnitude; the alive pool is the real number, and it scales with supply partnerships, not hardware buildout.

    Cumulative distinct devices dispatched to over 63 hours. 5,078 delivered at least one served inference. Discovery never stopped: the fleet has a recurring core plus a continuous arrival of new devices.

    04 · Reliability

    Every failure was counted, not hidden

    Because cloud fallback was off, the platform's safeguards had to work, and they were all observable in production. Admission control turned away 6,300+ unfit device registrations (low battery, policy, memory guardrails) before they entered the pool. A two-strike quarantine tripped 13,870 times, excluding devices after consecutive timeouts. Stale-session guards refused 2,777 dispatches to devices that had gone silent.

    The failure surface itself is shallow: 96 to 98% of hard failures were budget timeouts spread thinly across thousands of ordinary devices; the five worst devices caused just 0.9% of all timeouts. Zero 5xx errors. Zero contract violations across the full 63 hours.

    05 · SECURITY

    Devices are burst compute, not a data surface

    ZeroGPU runs on a trusted network of devices from approved supply partners that already carry hardware security built in: every modern phone and laptop ships with a Trusted Execution Environment (TEE) protecting keys, identity, and storage at the silicon level. Our platform builds on those foundations. We control the SDK that receives and processes each request, encrypt the full request flow, and strictly approve every participating partner and device before it can serve. Security tier is also a routing dimension, so workloads are directed only to the edge or cloud environments that meet their required bar.

    HARDWARE-SECURED DEVICES
    TRUSTED SDK
    ENCRYPTED FLOW
    06 · Energy & Carbon

    460x less energy than data center inference

    The footprint of computing has two sources: operational energy consumed during use, and embodied emissions from manufacturing hardware. This campaign measured the first and structurally avoided the second.

    Operational. The 812,146 served inferences consumed a measured 423 Wh of device compute, roughly half a milliwatt-hour per request, about 460x less energy than a reference datacenter LLM request (0.24 Wh). Running the same volume on that reference would have consumed ~195 kWh; the fleet used 0.42 kWh, a 99.78% reduction. At standard grid intensity, the entire 63-hour, 812K-inference campaign produced approximately 169 gCO₂e, roughly one kilometre of driving.

    Embodied. The marginal embodied footprint was zero. Every device in this fleet already existed, was already powered on, and was previously idle. No hardware was manufactured, purchased, or racked to serve this workload.

    What it means per request. The carbon cost of a single inference on this network was roughly 0.2 milligrams of CO₂e, about five thousand requests per gram. At this measured profile, every one billion requests routed to the edge tier instead of the datacenter reference avoids approximately 240 MWh of energy. That figure is an extrapolation of the measured per-request numbers, stated as such.

    0.5 mWh
    Per request
    460x
    Less energy than datacenter reference
    0
    New devices manufactured
    0.2 mg
    CO₂e per request

    Method: measured device compute time (380,771 s) x 4 W attributable active draw; grid intensity 400 gCO₂e/kWh; datacenter reference 0.24 Wh/request. Refused requests consumed no device compute.

    That is the structural argument. The world is planning $1T+ in annual data-center capex; ZeroGPU adds inference capacity with no new buildout: no new data centers, no new power contracts, no new hardware. In production, ZeroGPU runs as a hybrid: a trusted network of edge devices and cloud infrastructure. This campaign measured the energy profile of its edge tier alone, with the cloud deliberately switched off.

    Frontier models for reasoning. ZeroGPU for everything repeatable around it.

    This run demonstrates the core claim: a fleet of ordinary consumer devices, orchestrated well, serves production inference at multi-day scale with stable latency and a fully mapped failure surface, and it does so on compute the world already owns. It is also the lowest-energy way we know to run this class of work: measured, not estimated.

    ZeroGPU Research · July 2026 · zerogpu.ai