Transparent Pricing
High-performance inference at competitive prices. Pay only for what you use with no hidden fees or commitments.
Transparent Pricing
High-performance inference at competitive prices. Pay only for what you use with no hidden fees or commitments.

Serverless Pricing
Flexible token pricing, high usage limits, and postpaid billing—plus $1 in free credits to get you started!
DeepSeek
DeepSeek released the first open‑weight model and has gained global attention for building highly capable, cost‑efficient LLMs. Models such as DeepSeek‑V3.2 and DeepSeek‑R1 are competitive with top international models, delivering remarkable performance in reasoning, coding, and mathematical problem‑solving.
DeepSeek released the first open‑weight model and has gained global attention for building highly capable, cost‑efficient LLMs. Models such as DeepSeek‑V3.2 and DeepSeek‑R1 are competitive with top international models, delivering remarkable performance in reasoning, coding, and mathematical problem‑solving.
Model Name
Context Length
Context Length
Input
Cached Input
Output
Actions
Model Name
DeepSeek-V4-Flash-Vision-Exp
1049K
Context Length
Input (/M Tokens)
$
0.44
Cached Input (/M Tokens)
$
0.028
$
1.32
Output (/M Tokens)
Model Name
DeepSeek-V4-Pro-0813
1049K
Context Length
Input (/M Tokens)
$
1.32
Cached Input (/M Tokens)
$
0.044
$
3.96
Output (/M Tokens)
Model Name
DeepSeek-V4-Flash-0731
1049K
Context Length
Input (/M Tokens)
$
0.22
Cached Input (/M Tokens)
$
0.014
$
0.66
Output (/M Tokens)
Model Name
DeepSeek-V4-Pro
1049K
Context Length
Input (/M Tokens)
$
1.50162
Cached Input (/M Tokens)
$
0.135
$
3.135
Output (/M Tokens)
Model Name
DeepSeek-V4-Flash
1049K
Context Length
Input (/M Tokens)
$
0.13
Cached Input (/M Tokens)
$
0.028
$
0.28
Output (/M Tokens)
Prices shown are per 1 million tokens.

Qwen
Open-source AI model family built by Alibaba Cloud, ranging from sub-1B to 480B+ parameters, designed to scale to any use case, from deep reasoning and math to autonomous coding.
Open-source AI model family built by Alibaba Cloud, ranging from sub-1B to 480B+ parameters, designed to scale to any use case, from deep reasoning and math to autonomous coding.
Model Name
Model Name
Context Length
Context Length
Input
Cached Input
Output
Actions
Model Name
Qwen3.8-2.4T-A95B
1049K
Context Length
Input (/M Tokens)
$
2.0
Cached Input (/M Tokens)
$
0.25
$
6.0
Output (/M Tokens)
Model Name
Qwen3.6-27B
262K
Context Length
Input (/M Tokens)
$
0.3
$
3.2
Output (/M Tokens)
Model Name
Qwen3.6-35B-A3B
262K
Context Length
Input (/M Tokens)
$
0.2
$
1.6
Output (/M Tokens)
Model Name
Qwen3.5-9B
262K
Context Length
Input (/M Tokens)
$
0.1
$
0.15
Output (/M Tokens)
Model Name
Qwen3.5-122B-A10B
262K
Context Length
Input (/M Tokens)
$
0.26
$
2.08
Output (/M Tokens)
Prices shown are per 1 million tokens.

Z.ai
Zhipu AI builds the ChatGLM family of LLMs, develops LLMs as Agents. The latest model, GLM-4.7, delivers frontier-level performance in coding, creative writing, and role-play scenarios.
Zhipu AI builds the ChatGLM family of LLMs, develops LLMs as Agents. The latest model, GLM-4.7, delivers frontier-level performance in coding, creative writing, and role-play scenarios.
Model Name
Model Name
Context Length
Context Length
Input
Cached Input
Output
Actions
Model Name
GLM-5.3
1049K
Context Length
Input (/M Tokens)
$
1.4
Cached Input (/M Tokens)
$
0.26
$
4.4
Output (/M Tokens)
Model Name
GLM-5.3-Flash
1049K
Context Length
Input (/M Tokens)
$
0.15
Cached Input (/M Tokens)
$
0.03
$
0.5
Output (/M Tokens)
Model Name
GLM-5.2
1049K
Context Length
Input (/M Tokens)
$
1.302
Cached Input (/M Tokens)
$
0.26
$
4.092
Output (/M Tokens)
Model Name
GLM-5.1
205K
Context Length
Input (/M Tokens)
$
1.19
Cached Input (/M Tokens)
$
0.6
$
3.74
Output (/M Tokens)
Model Name
GLM-5
205K
Context Length
Input (/M Tokens)
$
0.95
Cached Input (/M Tokens)
$
0.2
$
2.55
Output (/M Tokens)
Prices shown are per 1 million tokens.

Moonshot AI
Moonshot AI stands out for breakthroughs in long-context language models. Its flagship product, Kimi, is especially well suited for research, legal work, and complex information synthesis. The latest release, Kimi K2 Thinking, is a state-of-the-art thinking agent with deep reasoning and tool orchestration.
Moonshot AI stands out for breakthroughs in long-context language models. Its flagship product, Kimi, is especially well suited for research, legal work, and complex information synthesis. The latest release, Kimi K2 Thinking, is a state-of-the-art thinking agent with deep reasoning and tool orchestration.
Model Name
Model Name
Context Length
Context Length
Input
Input
Cached Input
Output
Output
Actions
Actions
Model Name
Kimi-K2.5
262K
Context Length
Input (/M Tokens)
$
0.45
Cached Input (/M Tokens)
$
0.07
$
2.25
Output (/M Tokens)
Model Name
Kimi-K2.6
262K
Context Length
Input (/M Tokens)
$
0.77
Cached Input (/M Tokens)
$
0.14
$
3.4
Output (/M Tokens)
Model Name
Kimi-K2.7-Code
262K
Context Length
Input (/M Tokens)
$
0.85916
Cached Input (/M Tokens)
$
0.17993
$
3.8
Output (/M Tokens)
Model Name
Kimi-K3
1049K
Context Length
Input (/M Tokens)
$
2.7
Cached Input (/M Tokens)
$
0.27
$
13.5
Output (/M Tokens)
Prices shown are per 1 million tokens.

MiniMaxAI
Specialized in multimodal capabilities, MiniMax develops advanced models that seamlessly integrate text, voice, and vision, with notable achievements in natural-sounding text-to-speech and voice cloning.
Specialized in multimodal capabilities, MiniMax develops advanced models that seamlessly integrate text, voice, and vision, with notable achievements in natural-sounding text-to-speech and voice cloning.
Model Name
Model Name
Context Length
Context Length
Input
Input
Cached Input
Output
Output
Actions
Actions
Model Name
MiniMax-M2.5
197K
Context Length
Input (/M Tokens)
$
0.3
Cached Input (/M Tokens)
$
0.03
$
1.2
Output (/M Tokens)
Prices shown are per 1 million tokens.
OpenAI
OpenAI is a pioneering AI research organization that helped spark today's generative AI revolution. Its GPT series brought LLMs into the mainstream and is currently led by GPT-5.2 and o3, which set industry benchmarks for natural language understanding, generation, and reasoning.
OpenAI is a pioneering AI research organization that helped spark today's generative AI revolution. Its GPT series brought LLMs into the mainstream and is currently led by GPT-5.2 and o3, which set industry benchmarks for natural language understanding, generation, and reasoning.
Model Name
Model Name
Context Length
Context Length
Input
Input
Cached Input
Output
Output
Actions
Actions
Model Name
gpt-oss-120b
131K
Context Length
Input (/M Tokens)
$
0.05
$
0.45
Output (/M Tokens)
Prices shown are per 1 million tokens.
Others
Model Name
Model Name
Context Length
Context Length
Input
Input
Cached Input
Output
Output
Actions
Actions
Model Name
gemma-4-12B-it
262K
Context Length
Input (/M Tokens)
$
0.1
$
0.3
Output (/M Tokens)
Model Name
gemma-4-26B-A4B-it
262K
Context Length
Input (/M Tokens)
$
0.12
$
0.4
Output (/M Tokens)
Model Name
gemma-4-31B-it
262K
Context Length
Input (/M Tokens)
$
0.13
$
0.4
Output (/M Tokens)
Model Name
Hunyuan-A13B-Instruct
131K
Context Length
Input (/M Tokens)
$
0.14
$
0.57
Output (/M Tokens)
Model Name
Hy3
262K
Context Length
Input (/M Tokens)
$
0.132
Cached Input (/M Tokens)
$
0.033
$
0.528
Output (/M Tokens)
Prices shown are per 1 million tokens.
Image Generation
Generate high-quality images from text prompts with our state-of-the-art image generation models.
Model Name
Price (/image)
Actions
Model Name
FLUX 1.1 [pro] Ultra
$
0.06
Price (
)
/ Image
Model Name
FLUX.1 Kontext [max]
$
0.08
Price (
)
/ Image
Model Name
FLUX.1 Kontext [pro]
$
0.04
Price (
)
/ Image
Prices shown are per image generated or edited.
Video Generation
Create dynamic videos from text descriptions with our cutting-edge video generation models.
Audio Models
Process and generate audio with our high-quality speech recognition and synthesis models.
Model Name
Output (/M UTF-8 bytes)
Actions
Model Name
Fish-Speech-1.5
$
15.0
Price (
)
/ M UTF-8 bytes
Model Name
FunAudioLLM/CosyVoice2-0.5B
$
7.15
Price (
)
/ M UTF-8 bytes
Prices for transcription and translation are per minute of audio. Text-to-Speech prices are per 1,000 characters.

Frequently Asked Questions
How does billing work?
You're billed based on your usage. For chat models, you're charged per token for both input and output. For image, video, and audio models, pricing varies based on the specific task and output quality.
Are there any minimum commitments?
No, there are no minimum commitments. You only pay for what you use, and you can start with $1 in free credits.
Can I set spending limits?
Yes, you can set monthly spending limits in your account dashboard to control costs and prevent unexpected charges.
Do you offer volume discounts?
Yes, we offer volume discounts for high-usage customers. If your usage is substantial, please contact our sales team who can create a custom pricing plan tailored to your needs.
How do I get started?
Sign up for an account, get your API key, and start using our models right away. We provide comprehensive documentation and code examples to help you integrate quickly.
Ready to accelerate your AI development?

Ready to accelerate your AI development?

Ready to accelerate your AI development?



