Models & pricing
The NexoAI catalog is priced through one prepaid USD wallet. Token prices are in USD per one million input or output tokens; image and moderation models with a request price are billed per request instead. The rates below are NexoAI's live customer-facing rates.
Use each model ID exactly as shown. IDs are case-sensitive in API requests and CLI configuration, and endpoint badges show which compatible API shapes accept that model.
Not every catalogued model is served yet
The catalog lists NexoAI's planned lineup with committed pricing. Only a subset is currently routable; the rest are marked Coming soon in the dashboard model list. Calling a model that is not yet served returns a model-not-found error even with a valid key.
Check https://nexoai.dev/models for what is live right now before you wire a model ID into production.
One key, the right model access
Use a named NexoAI key with access to the model family you call. Separate keys per app make budget caps, rate limits, usage, and revocation easier to manage.
Alibaba
9 models| Model ID | Input / 1M | Output / 1M | Endpoints |
|---|---|---|---|
qwen3-coder-next | $1.00 | $4.00 | openaiopenai-responseanthropic |
qwen3-max | $2.50 | $10.00 | openaiopenai-responseanthropic |
qwen3-vl-flash | $0.15 | $1.50 | openaiopenai-responseanthropic |
qwen3.5-flash | $0.20 | $2.00 | openaiopenai-responseanthropic |
qwen3.5-plus | $0.80 | $4.80 | openaiopenai-responseanthropic |
qwen3.6-max-preview | $9.00 | $54.00 | openaiopenai-responseanthropic |
qwen3.6-plus | $2.00 | $12.00 | openaiopenai-responseanthropic |
qwen3.7-max | $12.00 | $36.00 | openaiopenai-responseanthropic |
qwen3.7-plus | $2.00 | $8.00 | openaiopenai-responseanthropic |
Anthropic
10 models| Model ID | Input / 1M | Output / 1M | Endpoints |
|---|---|---|---|
claude-fable-5 | $10.00 | $50.00 | anthropicopenai |
claude-haiku-4-5-20251001 | $1.00 | $5.00 | anthropicopenai |
claude-opus-4-1-20250805 | $15.00 | $75.00 | anthropicopenai |
claude-opus-4-5-20251101 | $5.00 | $25.00 | anthropicopenai |
claude-opus-4-6 | $5.00 | $25.00 | anthropicopenai |
claude-opus-4-7 | $5.00 | $25.00 | anthropicopenai |
claude-opus-4-8 | $5.00 | $25.00 | anthropicopenai |
claude-sonnet-4-5-20250929 | $3.00 | $15.00 | anthropicopenai |
claude-sonnet-4-6 | $3.00 | $15.00 | anthropicopenai |
claude-sonnet-5 | $2.00 | $10.00 | anthropicopenai |
DeepSeek
2 models| Model ID | Input / 1M | Output / 1M | Endpoints |
|---|---|---|---|
deepseek-v4-flash | $1.00 | $2.00 | openaianthropic |
deepseek-v4-pro | $12.00 | $24.00 | openaianthropic |
| Model ID | Input / 1M | Output / 1M | Endpoints |
|---|---|---|---|
gemini-2.5-flash | $0.30 | $2.50 | geminiopenai |
gemini-2.5-flash-image | $0.07 per request | geminiopenai | |
gemini-2.5-pro | $1.25 | $10.00 | geminiopenai |
gemini-3-flash-preview | $0.50 | $3.00 | geminiopenai |
gemini-3-pro-image-preview | $0.134 per request | openai | |
gemini-3-pro-preview | $2.00 | $12.00 | geminiopenai |
gemini-3.1-flash-image-preview | $0.067 per request | geminiopenai | |
gemini-3.1-pro-preview | $2.00 | $12.00 | geminiopenai |
gemini-3.5-flash | $1.50 | $9.00 | geminiopenai |
Hunyuan
1 model| Model ID | Input / 1M | Output / 1M | Endpoints |
|---|---|---|---|
hy3 | $1.00 | $4.00 | openaianthropic |
MiniMax
3 models| Model ID | Input / 1M | Output / 1M | Endpoints |
|---|---|---|---|
MiniMax-M2.7 | $2.10 | $8.40 | openaiopenai-responseanthropic |
MiniMax-M3 | $4.20 | $16.80 | anthropicopenai |
minimax-m2.5 | $2.10 | $8.40 | openaiopenai-responseanthropic |
Moonshot
4 models| Model ID | Input / 1M | Output / 1M | Endpoints |
|---|---|---|---|
kimi-k2.5 | $4.00 | $21.00 | openaiopenai-responseanthropic |
kimi-k2.6 | $6.50 | $27.00 | anthropic |
kimi-k2.7-code | $6.50 | $27.00 | anthropic |
kimi-k3 | $20.00 | $100.00 | anthropic |
OpenAI
12 models| Model ID | Input / 1M | Output / 1M | Endpoints |
|---|---|---|---|
codex-auto-review | $5.00 | $30.00 | openai-responseopenai |
gpt-4.1 | $2.00 | $4.00 | openai |
gpt-5.3-codex | $1.75 | $14.00 | openai |
gpt-5.4 | $2.50 | $15.00 | openai-responseopenai |
gpt-5.4-mini | $0.75 | $4.50 | openai-responseopenai |
gpt-5.4-pro | $30.00 | $180.00 | openai |
gpt-5.5 | $5.00 | $30.00 | openai-responseopenai |
gpt-5.6-luna | $1.00 | $6.00 | openai-response |
gpt-5.6-sol | $5.00 | $30.00 | openai-responseopenai |
gpt-5.6-terra | $2.50 | $15.00 | openai-responseopenai |
gpt-image-2 | $0.08 per request | openai | |
omni-moderation-latest | $0.01 per request | openai | |
Xiaomi MiMo
5 models| Model ID | Input / 1M | Output / 1M | Endpoints |
|---|---|---|---|
mimo-v2-flash | $0.70 | $2.10 | openaiopenai-responseanthropic |
mimo-v2-omni | $2.80 | $14.00 | openaiopenai-responseanthropic |
mimo-v2-pro | $7.00 | $21.00 | openaiopenai-responseanthropic |
mimo-v2.5 | $1.00 | $2.00 | openaiopenai-responseanthropic |
mimo-v2.5-pro | $3.00 | $6.00 | openaiopenai-responseanthropic |
Zhipu AI
3 models| Model ID | Input / 1M | Output / 1M | Endpoints |
|---|---|---|---|
glm-4.7 | $4.00 | $16.00 | openaianthropic |
glm-5 | $4.00 | $18.00 | openaiopenai-responseanthropic |
glm-5.2 | $8.00 | $28.00 | openaianthropic |
xAI
1 model| Model ID | Input / 1M | Output / 1M | Endpoints |
|---|---|---|---|
grok-4.5 | $2.00 | $6.00 | openaiopenai-response |
How token billing works
Input and output are metered separately against the rates in the catalog:
text
input cost = input tokens / 1,000,000 x input rate
output cost = output tokens / 1,000,000 x output rate
request cost = input cost + output costPer-request entries do not use that formula; the listed USD request price is deducted once for each billed request. Your wallet balance and key budget caps provide the final usage boundary.
Rates and availability are live
Model availability and customer-facing rates can change. Check this catalog and the dashboard before fixing a production cost estimate, and keep a fallback for preview models.
USD end to end
Usage, model prices, and top-ups are all denominated in USD. You fund the same prepaid wallet by card; the dashboard shows the payment amount and resulting wallet credit before you confirm.
For workload planning:
- Estimate token or request volume against the USD rates above.
- Include retries, tool loops, and longer-than-expected outputs.
- Review the checkout quote before topping up the wallet.
- Put a hard USD budget cap on the named key used by the workload.
NexoAI has no subscription and no separate vendor balance. Review wallet and key controls ->
