ZeroGPU sells to AI agents - no signup

The agent storefront at https://agents.zerogpu.ai/llms.txt publishes prices, payment instructions, and setup. Agents can purchase autonomously or with their human's approval, with a range of payment methods; an agent without a payment method can set one up there instantly, no signup. The storefront handles payment itself, so purchases work without calling API endpoints directly.

    MASTER PLAN · 2026

    The ZeroGPU Master Plan

    Roadmap to intelligence powered by everyone

    By Maddy Arvapally, Founder & CEO

    Background

    I've spent my career inside distributed systems: scaling GoPro's video platform from zero to two million subscribers, building streaming infrastructure at Replay, working across robotics and brain imaging. My sole focus is now ZeroGPU. My ambition is to build this company with a 20-year view: to make intelligence as cheap and universal as the devices that already fill our pockets, desks, and homes.

    Our Mission

    Make intelligence abundant, by powering it with the computers everyone already owns, and paying them back for it.

    The Company

    I believe the defining question of the intelligence age is not whether AI transforms everything. It will. The question is who owns the machinery it runs on. Today's answer is being poured in concrete: a trillion dollars a year of datacenters, owned by a handful of companies, constrained by power grids, and rented back to the world by the token. If that remains the only answer, every product, every startup, and every nation pays a tax on intelligence forever.

    ZeroGPU exists to build the second answer. The journey will take decades, the odds are honest, and the engineering is hard. But if we succeed, the ZeroGPU hybrid inference cloud will be the second answer, and the infrastructure of intelligence will belong, in part, to everyone who already paid for a computer.

    The Present

    Compute is a global constraint. Datacenter capital expenditure is racing toward $1 trillion per year, and it still cannot keep pace with inference demand. The binding limit is no longer money. It is electricity: grid interconnection queues run five to seven years, states have begun freezing datacenter connections outright, and analysts estimate that of the roughly 190 gigawatts of capacity announced this decade, perhaps 80 will actually be built.

    Meanwhile, humanity has already manufactured the largest computing fleet in history: more than two billion phones, laptops, and gaming PCs. Grid-connected, distributed across every country on earth, already paid for, and sitting idle twenty hours a day. The world does not have a compute shortage. It has an allocation failure.

    The Possibility

    AI inference is growing from $106 billion today toward $255 billion by 2030, and its composition is the opportunity hiding in plain sight. Roughly 80% of inference requests are not frontier reasoning. They are the metabolism of the AI economy: classification, moderation, embeddings, reranking, extraction, routing. Single-pass tasks on small models that fire billions of times a day, in every pipeline, forever. Frontier models think; the metabolism never sleeps.

    These tasks do not need frontier GPUs. A model of a few hundred million parameters, running on the neural silicon shipped in every modern phone and laptop, matches or beats frontier models on the task, at a fraction of the cost and as little as a tenth of the latency. As models shrink and device silicon improves, the share of AI that can run this way compounds every quarter.

    And there is a second possibility folded inside the first: the sharing economy never reached its largest idle asset. Airbnb unlocked the spare room; the spare computer, the most widely distributed capital good ever made, has never earned its owner a cent. When it does, the intelligence age stops being something that happens to people and starts being something that happens through them.

    Three long-term opportunities

    The commodity tier of global inference. Serving the metabolic 80%: classification, embeddings, moderation, extraction, at prices no datacenter cost structure can follow. This begins in the industries where the mismatch is most extreme: the AI ad economy, verification, and trust and safety.

    Every footprint becomes a cloud. Any app, enterprise, or platform turns its own install base into a private inference cloud. One SDK, sovereign by architecture, data never leaving hardware the customer already owns. From game studios to global banks: the fleet you already depreciated is a datacenter nobody switched on.

    The participation economy. Device owners earn from the intelligence age directly. Phones earning perks, PCs earning income, homes hosting nodes, paid not by subsidy but by real workloads. The people who bought the hardware share in what it produces.

    The Solution

    There are two schools of thought on scaling inference: build ever more centralized capacity and bring the world's data to it, or build the orchestration layer that brings the work to where computers already are. The first school requires a decade of concrete, transformers, and water. The second requires software: routing, verification, and trust across a billion heterogeneous devices.

    We believe the second school wins the commodity tier. Not by replacing datacenters, but by relieving them. Frontier work belongs in the cloud. The metabolism belongs everywhere. ZeroGPU is building the hybrid inference cloud to meet that demand: one programmable layer over edge servers, cloud burst capacity, and consumer devices, with an orchestration plane that classifies every task and routes it to the right model on the right silicon by capability, trust tier, and load.

    How We Can Do It

    Orchestration at planetary scale. Scheduling, admission control, and fallback across devices that sleep, roam, and disconnect. Proven in the field: 1.04 million production requests served across 5,078 consumer devices in 20+ countries over 63 unattended hours, with stable latency and zero failed contracts.

    Trust as infrastructure. Hardware attestation, encrypted execution, sharded workloads, and redundant verification, so a stranger's workload runs correctly and privately on hardware we don't control. The verification layer is the moat: every task served calibrates it further.

    Task models that compound. Purpose-built small models, ZLMs, trained on production traffic for the workloads the world actually runs. Our moderation model already beats the frontier incumbent head-to-head at a fraction of the cost. Every customer makes the catalog stronger; every model makes the network more valuable.

    Supply through partnership, not persuasion. We don't recruit devices one consumer at a time. Fleets arrive by contract: apps, platforms, and enterprises whose install bases become clouds, with the device owner's consent and compensation built into the SDK itself.

    Economics that improve with time. Our marginal hardware cost is idle electricity. As model prices fall, our margins rise and our addressable share grows. We are the only inference company structurally long the commoditization of intelligence.

    Conclusion

    The first phase of our Master Plan:

    1. Prove the fleet works. (Done, and published.)
    2. Win the metabolism of AI, one workload at a time.
    3. Turn the world's footprints into clouds, and pay their owners.

    The trillion-dollar answer to the compute constraint is pouring concrete. Ours is switching on what humanity already built.

    The answer was always already everywhere. It's time to turn it on.

    Maddy Arvapally

    Founder & CEO, ZeroGPU