llama-3.1-8b-instruct-fast
Summarizationby Meta
llama-3.1-8b-instruct-fast is Meta’s Llama 3.1 8B Instruct model optimized by ZeroGPU for fast, cost-efficient summarization and text processing at scale. Its 128K context window makes it well suited for long documents, transcripts, articles, email threads, and conversations that need to be processed in a single pass. The model supports multilingual workloads and is a strong fit for summarization pipelines, content processing, and agent workflows where low latency and predictable inference cost matter. On ZeroGPU, it runs across the hybrid inference cloud to optimize speed and cost for high-volume production workloads.
Specifications
Parameters
8B
Context length
131,072 tokens
Task
Summarization
Provider
Meta
Billing
Per token
API
OpenAI-compatible
Pricing
Input
$0.15
per 1M tokens
Output
$0.28
per 1M tokens