Fireworks AI
Fireworks AI
Inference, fine-tuning and dedicated deployments for open models, billed by tokens and GPU hours.

Platforms
Overview
Fireworks AI is an inference platform for open models. Its pricing page — "pricing to seamlessly scale from idea to enterprise" — is split three ways: serverless inference billed per token, training billed per training token, and on-demand dedicated GPUs billed per hour. Training on models up to 16B runs $0.50 per million training tokens for LoRA SFT and $1.00 for full-parameter SFT. On-demand H100/H200 is $8 an hour and B200 is $13. New accounts get $1 in free credits.
Main features
- Serverless inference billed per token
- Fine-tuning (LoRA and full parameter)
- On-demand dedicated GPUs
- $1 in free credits to start
- Custom enterprise terms
Pricing
Training (up to 16B, LoRA SFT)
$0.50 / 1M training tokens
Published price (USD) · Source
- Up to 16B: LoRA SFT $0.50 / Full Param SFT $1.00
- 16.1B-80B: LoRA SFT $3.00 / Full Param SFT $6.00
- 80B-300B: $6.00-$24.00
- >300B: $10.00-$40.00
- All per 1M training tokens
Last verified: September 2026
On-demand deployments
$8 / GPU hour
Published price (USD) · Source
- H100/H200: $8.00/hour
- B200: $13.00/hour
- B300/GB300: $15.00-$20.00/hour
Last verified: September 2026
Serverless inference (embeddings example)
$0.10 / 1M input tokens
Published price (USD) · Source
- Qwen3 8B embeddings: $0.1 per 1M input tokens
- Embeddings up to 150M params: $0.008
- Embeddings 150M-350M params: $0.016
- Get started with $1 in free credits
Last verified: September 2026
Enterprise
Contact sales
- Faster speeds and lower costs
- Higher rate limits
- Contact us for a quote
Fireworks AI publishes one price worldwide and bills in USD. There is no separate price in your local currency.
Screenshots

Related AI tools
Together AI
Together AI is an AI infrastructure and API service useful for AI development and inference.
Replicate
Replicate is an AI infrastructure and API service useful for AI development and prototyping.
Groq
An inference cloud built on its own LPU processor, operating dedicated capacity at scale.
OpenRouter
OpenRouter is an AI infrastructure and API service useful for AI development and model comparison.