qwen3-30b-a3b-fp8
Text Generationby Qwen
qwen3-30b-a3b-fp8 is Qwen’s open-weight mixture-of-experts model designed for efficient reasoning, coding, multilingual text generation, and agentic workflows. It activates only 3B of its 30B parameters per token, delivering strong model quality with lower latency and inference cost. It supports function calling, streaming, and multilingual workloads, making it well suited for agents, code assistance, automation, structured data tasks, and high-volume production applications. On ZeroGPU, it runs across the hybrid inference cloud to optimize speed and cost for production workloads.
Specifications
Parameters
30B
Context length
32,768 tokens
Task
Text Generation
Provider
Qwen
Billing
Per token
API
OpenAI-compatible
Pricing
Input
$0.10
per 1M tokens
Output
$0.45
per 1M tokens