Models
Catalog models, their categories and endpoints, and differences in behavior.
The list of models available to your key is returned by GET /v1/models. Take names from there: the catalog changes, and a key can be limited to a list of models. See prices in the console.
Text
These models accept any of the three formats: /v1/chat/completions, /v1/messages, /v1/responses. You do not need to write code for a specific model.
| Model | Notes |
|---|---|
claude-opus-5.5 | temperature is supported |
claude-fable-5.1 | explicit prompt cache through cache_control |
claude-sonnet-5 | explicit prompt cache through cache_control |
gpt-6-astra, gpt-6-sol, gpt-6-luna | support seed, implicit cache |
gemini-3.1-pro | temporarily unavailable |
gpt-5-mini, gpt-4o | old names for existing integrations |
The context of Claude models is 1 million tokens, of GPT-6 1.05 million. The maximum response length of flagship models is 128,000 tokens. When the client does not send max_tokens, the money reserve is calculated from the model's response ceiling.
Embeddings
Endpoint /v1/embeddings.
| Model | Dimensions |
|---|---|
text-embedding-3-small | 1536 |
text-embedding-3-large | 3072 |
bge-m3 | 1024, works well with Russian |
Speech
Endpoint /v1/audio/transcriptions: whisper-large-v3. The price is counted per second of audio.
Images
Endpoint /v1/images/generations, and hybrid models also have /v1/chat/completions with modalities.
| Model |
|---|
gpt-image-2 |
gpt-image-1-mini |
seedream-4.5 |
flux-2-pro |
flux-2-klein |
qwen-image-3 |
recraft-v4.1-flash |
Video
Endpoint /v1/videos. The allowed duration, resolution and aspect ratio depend on the model, and unsuitable values give 400 unsupported_video_parameter.
| Model | Resolutions |
|---|---|
grok-imagine-video | 480p, 720p |
wan-3.0 | 480p, 720p, 1080p |
seedance-2.0-mini | 480p, 720p |
kling-3.0 | 720p, with or without sound |
Differences in behavior
- A parameter the model does not support is ignored without an error.
temperaturehas no effect onclaude-fable-5.1,claude-sonnet-5and GPT-6 models, andseedexists only on GPT-6. - Reasoning counts toward output tokens in all formats. Models with reasoning, for example
gpt-5-miniandgpt-6-luna, can spend the whole limit on it with a smallmax_tokensand return empty text. Leave headroom. - The GPT-6 price rises when the input is over 272,000 tokens.
- All text models accept images in the input in all three formats.
You can also call any model by its full vendor/model identifier, for example anthropic/claude-sonnet-5. Feature compatibility: Compatibility.