Skip to navigationSkip to main contentSkip to footerScaleway Docs HomepageAsk our AI
Ask our AI

Generative APIs supported models

This page provides a quick overview of available models in Scaleway's catalog and their core attributes. To see usage examples and detailed capabilities, click the name of the model in the summary table.

All models comply with OpenAI and come with their own support guarantees and a clear lifecycle.

Models served through the Serverless offering are subject to rate limits.

Models technical summary

Do not see a model you want to use? Tell us or vote for what you would like to add.

Model nameAvailable in Serverless?ModalitiesMaximum context window (tokens)License*
glm-5.2YesText, Code256k**MIT
deepseek-v4-flash-0731YesText, Code256k**MIT
deepseek-r1-distill-llama-70bEOLText16k (Serverless) / 128k (Dedicated)MIT and Llama 3.3 Community
deepseek-r1-distill-llama-8bNoText128kMIT and Llama 3.1 Community
qwen3.6-35b-a3bYesText, Code, Vision256kApache 2.0
qwen3.5-397b-a17bYesText, Code, Vision250kApache 2.0
qwen3.5-35b-a3bNoText, Code, Vision262kApache 2.0
qwen3.5-122b-a10bNoText, Code, Vision262kApache 2.0
qwen3-235b-a22b-instruct-2507YesText250kApache 2.0
qwen3-235b-a22b-thinking-2507NoText262kApache 2.0
qwen3-embedding-8bYesEmbeddings32kApache 2.0
qwen3-coder-30b-a3b-instructYesCode128kApache 2.0
qwen2.5-coder-32b-instructEOLCode32kApache 2.0
gemma-4-31b-itNoText, Vision262kApache 2.0
gemma-4-26b-a4b-itYesText, Vision256kApache 2.0
gemma-3-27b-itYesText, Vision40kGemma
gpt-oss-120bYesText128kApache 2.0
gpt-oss-20bNoText131kApache 2.0
whisper-large-v3YesAudio transcription-Apache 2.0
llama-3.3-70b-instructYesText100k (Serverless)/ 128k (Dedicated)Llama 3.3 Community
llama-3.1-70b-instructEOLText128kLlama 3.1 Community
llama-3.1-8b-instructEOLText128kLlama 3.1 Community
llama-3-8b-instructNoText8kMeta Llama 3
llama-3-70b-instructNoText8kLlama 3 Community
llama-3.1-nemotron-70b-instructNoText128kLlama 3.1 Community
mistral-7b-instruct-v0.3NoText32kApache 2.0
mistral-large-3-675b-instruct-2512NoText, Vision250kApache 2.0
mistral-medium-3.5-128bYesText, Vision180k***Modified MIT License
mistral-small-3.2-24b-instruct-2506YesText, Vision128kApache 2.0
mistral-small-3.1-24b-instruct-2503EOLText, Vision128kApache 2.0
mistral-small-24b-instruct-2501NoText32kApache 2.0
voxtral-small-24b-2507YesText, Audio32kApache 2.0
mistral-nemo-instruct-2407EOLText128kApache 2.0
mixtral-8x7b-instruct-v0.1NoText32kApache 2.0
magistral-small-2506NoText32kApache 2.0
devstral-2-123b-instruct-2512YesText, Code200k (Serverless)/ 260k (Dedicated)Modified MIT
devstral-small-2505EOLText128kApache 2.0
pixtral-12b-2409YesText, Vision128kApache 2.0
molmo-72b-0924NoText, Vision50kApache 2.0 and Twonyi Qianwen license
holo2-30b-a3bYesText, Vision22kCC-BY-NC-4.0
bge-multilingual-gemma2YesEmbeddings8kGemma
sentence-t5-xxlEOLEmbeddings512kApache 2.0
minimax-m2.5NoCode197kMIT

*Licences which are not open-weight and may restrict commercial usage (such as CC-BY-NC-4.0), do not apply to usage through Scaleway Products due to existing partnerships between Scaleway and the corresponding providers. Original licenses are provided for transparency only.

**GLM 5.2 and DeepSeek-V4-Flash support 1 million context. During the preview stage, the context is limited to 256k to ensure consistent token generation speed.

***Mistral Medium 3.5 supports 1 million context. During the preview stage, the context is limited to 180k to ensure consistent token generation speed.

End of Life (EOL) models

Between the Deprecation date and the End Of Life date, models can still be accessed in Generative APIs - Serverless, but their End of Life (EOL) is planned according to our model lifecycle policy. Deprecated models should not be queried anymore. Scaleway recommends using newer models available in Generative APIs or to deploy these models in dedicated Generative APIs deployments.

After the End Of Life date, models are not accessible anymore from Generative APIs - Serverless. They can still however be deployed on dedicated Generative APIs deployments.

ProviderModel stringDeprecation dateEOL dateRequests routed to model
Mistralmistral-small-3.1-24b-instruct-2503August 14, 2025November 14, 2025mistral-small-3.2-24b-instruct-2506
Mistraldevstral-small-2505August 14, 2025November 14, 2025qwen3-coder-30b-a3b-instruct
Qwenqwen2.5-coder-32b-instructAugust 14, 2025November 14, 2025qwen3-coder-30b-a3b-instruct
Metallama-3.1-70b-instructFebruary 25, 2025May 25, 2025llama-3.3-70b-instruct
SBERTsentence-t5-xxlNovember 26, 2024February 26, 2025None
Deepseekdeepseek-r1-distill-llama-70bJanuary 16, 2026April 16, 2026llama-3.3-70b-instruct
Mistralmistral-nemo-instruct-2407January 16, 2026April 16, 2026mistral-small-3.2-24b-instruct-2506
Googlegemma-3-27b-itJuly 1, 2026August 1, 2026gemma-4-26b-a4b-it
Mistraldevstral-2-123b-instruct-2512July 1, 2026August 1, 2026qwen3.5-397b-a17b
Mistralvoxtral-small-24b-2507July 1, 2026August 1, 2026whisper-large-v3 (/audio/transcription only)
Mistralpixtral-12b-2409July 1, 2026October 1, 2026mistral-small-3.2-24b-instruct-2506
Qwenqwen3-coder-30b-a3b-instructJuly 1, 2026October 1, 2026qwen3.6-35b-a3b
H Companyholo2-30b-a3bJuly 9, 2026August 9, 2026qwen3.6-35b-a3b

Multimodal models (Text and Vision)

Note

Vision models can understand and analyze images, not generate them. You will use vision models through the /v1/chat/completions endpoint.

Qwen3.6-35b-a3b

Released in April 2026, Qwen3.6-35b-a3b is a state-of-the-art small-sized model optimized for agentic tasks and logical reasoning.

AttributeValue
ProviderQwen
Supports structured outputYes
Supports function callingYes
Supports parallel tool-callingYes
Supported reasoning effortsnone, low, medium, high
Supported image formatsPNG, JPEG, GIF
Maximum image resolution (pixels)4096x4096
Token dimension (pixels)32x32
Supported languagesEnglish, French, Portuguese, German, Romanian, Swedish, and 70 additional languages and dialects
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100, H100-2, H100-SXM-2
Maximum output (tokens) - Serverless32k
Hugging Face model cardqwen3.6-35b-a3b

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model name

qwen/qwen3.6-35b-a3b:bf16
qwen/qwen3.6-35b-a3b:fp8

Qwen3.5-397b-a17b

Qwen3.5-397b-a17b is a model developed by Qwen to perform text processing, agentic coding, image, and video analysis in several languages. This model was released as a frontier reasoning model on February 16, 2026.

AttributeValue
ProviderQwen
Supports structured outputYes
Supports function callingYes
Supports parallel tool-callingYes
Supported reasoning effortsnone, low, medium, high
Supported image formatsPNG, JPEG, WEBP, and non-animated GIFs
Supported video formatsMP4, MPEG, MOV, OGG and WEBM
Maximum image resolution (pixels)4096x4096
Token dimension (pixels)32x32
Supported languagesEnglish, French, German, Chinese, Japanese, Korean, and 113 additional languages and dialects
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100-SXM-8
Maximum output (tokens) - Serverless16k
Hugging Face model cardqwen3.5-397b-a17b

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model names

qwen/qwen3.5-397b-a17b:int4

Qwen3.5-35b-a3b

Released in March 2026, Qwen3.5-35b-a3b is a state-of-the-art small-sized model optimized for agentic and coding tasks, as well as logical reasoning.

AttributeValue
ProviderQwen
Supports parallel tool-callingYes
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100-2, H100-SXM-2, H100-SXM-4, H100-SXM-8
Hugging Face model cardQwen3.5-35B-A3B-GPTQ-Int4

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model name

qwen/qwen3.5-35b-a3b:int4

Qwen3.5-122b-a10b

Released in March 2026, Qwen3.5-122b-a10b is a state-of-the-art medium-sized model optimized for agentic and coding tasks, as well as logical reasoning.

AttributeValue
ProviderQwen
Supports parallel tool-callingYes
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100-2, H100-SXM-2, H100-SXM-4, H100-SXM-8
Hugging Face model cardQwen/Qwen3.5-122B-A10B-GPTQ-Int4

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model name

qwen/qwen3.5-122b-a10b:int4

Gemma-4-31b-it

Released in April 2026, Gemma-4-31b-it is a frontier small-sized model to perform agentic and reasoning tasks on many languages.

AttributeValue
ProviderGoogle
Supports structured outputYes
Supports function callingYes
Supports parallel tool callingYes
Supported reasoning efforts*none, low, medium, high
Supported image formatsPNG, JPEG, WEBP, and non-animated GIFs
Maximum image resolution (pixels)896x896
Token dimension (pixels)64x64
Supported languagesEnglish, Chinese, Japanese, Korean, and 136 additional languages
Compatible Instances (max context in tokens**) - Dedicated DeploymentH100 (66k), H100-2 (262k), H100-SXM-2 (262k), H100-SXM-4
Hugging Face model cardgoogle/gemma-4-31B-it

*Some providers recommend enabling reasoning by using additional arguments such as "chat_template_kwargs": {"enable_thinking": true}. These fields are not supported on Generative APIs, and we recommend using the reasoning_effort property, which is equivalent.

**Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model name

google/gemma-4-31b-it:bf16

Gemma-3-27b-it

Gemma-3-27b-it is a model developed by Google to perform text processing and image analysis on many languages. The model was not trained specifically to output function / tool call tokens. Therefore, function calling is currently supported, but reliability remains limited.

Pan & Scan is not yet supported for Gemma 3 images. This means that high-resolution images are currently resized to a resolution of 896x896, which may generate artifacts and lead to lower accuracy.

AttributeValue
ProviderGoogle
Supports structured outputYes
Supports function callingPartial
Supports parallel tool-callingNo
Supported image formatsPNG, JPEG, WEBP, and non-animated GIFs
Maximum image resolution (pixels)896x896
Token dimension (pixels)56x56
Supported languagesEnglish, Chinese, Japanese, Korean, and 31 additional languages
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100, H100-2
Maximum output (tokens) - Serverless8k
Hugging Face model cardgemma-3-27b-it

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model names

google/gemma-3-27b-it:bf16

Gemma-4-26b-a4b-it

Released in April 2026, Gemma-4-26b-a4b-it is a frontier small-sized model to perform agentic and reasoning tasks on many languages. This model has a Mixture-of-Expert (MoE) architecture, providing significant throughput and fitting on a single H100 GPU while supporting its maximum context size.

AttributeValue
ProviderGoogle
Supports structured outputYes
Supports function callingYes
Supports parallel tool-callingYes
Supported reasoning efforts*none, low, medium, high
Supported image formatsPNG, JPEG, WEBP, and non-animated GIFs
Maximum image resolution (pixels)896x896
Token dimension (pixels)64x64
Supported languagesEnglish, Chinese, Japanese, Korean, and 136 additional languages
Compatible Instances (max context in tokens**) - Dedicated DeploymentH100 (262k), H100-2 (262k), H100-SXM-2 (262k), H100-SXM-4
Maximum output (tokens) - Serverless32k
Hugging Face model cardgemma-4-26B-A4B-it

*Some providers recommend enabling reasoning by using additional arguments such as "chat_template_kwargs": {"enable_thinking": true}. These fields are not supported on Generative APIs, and Scaleway recommends using the reasoning_effort property, which is equivalent.

**Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model name

google/gemma-4-26b-a4b-it:bf16

Mistral-large-3-675b-instruct-2512

Mistral-large-3-675b-instruct-2512 is a frontier model, ideal for agentic workflows and image understanding.

AttributeValue
ProviderMistral
Supports structured outputYes
Supports function callingYes
Supports parallel tool-callingYes
Supported image formatsPNG, JPEG, WEBP, and non-animated GIFs
Maximum image resolution (pixels)1540x1540
Token dimension (pixels)28x28
Supported languagesEnglish, French, German, Spanish, Italian, Portuguese, Dutch, Chinese, Japanese, Korean, Arabic
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100-SXM-8 (180k)

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model names

mistral/mistral-large-3-675b-instruct-2512:fp4

Mistral-medium-3.5-128b

Mistral-medium-3.5 is a unified model from Mistral with strong performance for instruct, reasoning, and coding tasks.

AttributeValue
ProviderMistral
Supports structured outputYes
Supports function callingYes
Supports parallel tool-callingYes
Supported reasoning effortsnone, high
Supported image formatsPNG, JPEG, WEBP, GIF
Maximum image resolution (pixels)1540x1540
Token dimension (pixels)28x28
Supported languagesEnglish, French, German, Spanish, Portuguese, Italian, and 18 additional languages and dialects
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100-SXM-4 (256k), H100-SXM-8 (256k)
Maximum output (tokens) - Serverless16k
Hugging Face model cardmistralai/Mistral-Medium-3.5-128B

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model name

mistral/mistral-medium-3.5-128b:fp8

Mistral-small-3.2-24b-instruct-2506

Mistral-small-3.2-24b-instruct-2506 is an improved version of Mistral-small-3.1, which performs better on tool-calling. This model was optimized to have a dense knowledge and faster token throughput compared to its size.

AttributeValue
ProviderMistral
Supports structured outputYes
Supports function callingYes
Supports parallel tool-callingYes
Supported image formatsPNG, JPEG, WEBP, and non-animated GIFs
Maximum image resolution (pixels)1540x1540
Token dimension (pixels)28x28
Supported languagesEnglish, French, German, Greek, Hindi, Indonesian, Italian, Japanese, Korean, Malay, Nepali, Polish, Portuguese, Romanian, Russian, Serbian, Spanish, Swedish, Turkish, Ukrainian, Vietnamese, Arabic, Bengali, Chinese, Farsi
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100, H100-2
Maximum output (tokens) - Serverless32k
Hugging Face model cardmistral-small-3.2-24b-instruct-2506

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model names

mistral/mistral-small-3.2-24b-instruct-2506:fp8

Mistral-small-3.1-24b-instruct-2503

Mistral-small-3.1-24b-instruct-2503 is a model developed by Mistral to perform text processing and image analysis on many languages. This model was optimized to have a dense knowledge and faster token throughput compared to its size.

The model supports bitmap (or raster) image formats, that is, storing images as grids of individual pixels. Vector image formats (SVG, PSD), PDFs, and videos are not supported.

Image size is limited in the following ways:

  • Directly by the maximum context window. As an example, since tokens are squares of 28x28 pixels, the maximum context window taken by a single image is 3025 tokens (i.e., (1540*1540)/(28*28)).
  • Indirectly by model accuracy: resolution above 1540x1540 will not increase model output accuracy. Images above a width or height of 1540 pixels will be automatically downscaled to fit within the 1540x1540 dimension. Note that image ratio and overall aspect is preserved (images are not cropped, only additionally compressed).
AttributeValue
ProviderMistral
Supports structured outputYes
Supports function callingYes
Supports parallel tool-callingYes
Supported image formatsPNG, JPEG, WEBP, and non-animated GIFs
Maximum image resolution (pixels)1540x1540
Token dimension (pixels)28x28
Supported languagesEnglish, French, German, Greek, Hindi, Indonesian, Italian, Japanese, Korean, Malay, Nepali, Polish, Portuguese, Romanian, Russian, Serbian, Spanish, Swedish, Turkish, Ukrainian, Vietnamese, Arabic, Bengali, Chinese, Farsi
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100, H100-2

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model names

mistral/mistral-small-3.1-24b-instruct-2503:bf16
mistral/mistral-small-3.1-24b-instruct-2503:fp8

Pixtral-12b-2409

Pixtral is a vision language model introducing a novel architecture: a 12B-parameter multimodal decoder and a 400M-parameter vision encoder. It can analyze images and offer insights from visual content and text.

AttributeValue
ProviderMistral
Supports structured outputYes
Supports function callingYes
Supports parallel tool-callingYes
Supported image formatsPNG, JPEG, WEBP, and non-animated GIFs
Maximum image resolution (pixels)1024x1024
Token dimension (pixels)16x16
Maximum images per request12
Supported languagesEnglish
Compatible Instances (max context in tokens*) - Dedicated DeploymentL40S (50k), H100, H100-2
Maximum output (tokens) - Serverless4k
Hugging Face model cardpixtral-12b-2409

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model name

mistral/pixtral-12b-2409:bf16

Holo2-30b-a3b

Holo2 30B is a text and vision model optimized to analyze a Graphical User Interface, such as a web browser or software, and take actions.

AttributeValue
ProviderH
Supports structured outputYes
Supports function callingNo
Supports parallel tool-callingYes
Supported image formatsPNG, JPEG, WEBP, and non-animated GIFs
Token dimension (pixels)16x16
Supported languagesEnglish
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100-SXM-2
Maximum output (tokens) - Serverless32k
Hugging Face model cardholo2-30b-a3b

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model name

hcompany/holo2-30b-a3b:bf16

Molmo-72b-0924

Molmo 72B is the powerhouse of the Molmo family of multimodal models developed by the renowned research lab Allen Institute for AI. Vision-language models like Molmo can analyze an image and offer insights from visual content alongside text. This multimodal functionality creates new opportunities for applications that need both visual and textual comprehension.

AttributeValue
ProviderAllen Institute for AI
Supports structured outputYes
Supports function callingNo
Supported languagesEnglish
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100-2

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model name

allenai/molmo-72b-0924:fp8

Multimodal models (Text and Audio)

Voxtral-small-24b-2507

Voxtral-small-24b-2507 is a model developed by Mistral to perform text processing and audio analysis on many languages. This model was optimized to enable transcription in several languages while keeping conversational capabilities (translations, classification, etc.).

The model supports mono and stereo audio formats. For stereo formats, both left and right channels are merged before being processed.

The model processes audio files in 30-second chunks:

  • If an audio file of less than 30 seconds is sent, the rest of the chunk will be considered silent.
  • 80ms is equal to 1 input token.
AttributeValue
ProviderMistral
Supports structured outputYes
Supports function callingYes
Supports parallel tool-callingYes
Supported audio formatsflac, m4a, mpeg, mp2, mp3, mp4, ogg, wav, webm
Audio chunk duration30 seconds
Token duration (audio)80ms
Maximum transcription duration30 minutes
Maximum understanding duration40 minutes
Maximum file size - Serverless25 MB
Supported languagesEnglish, French, German, Dutch, Spanish, Italian, Portuguese, Hindi
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100, H100-2
Maximum output (tokens) - Serverless16k
Hugging Face model cardvoxtral-small-24b-2507

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model names

mistral/voxtral-small-24b-2507:bf16
mistral/voxtral-small-24b-2507:fp8

Audio transcription models

Whisper-large-v3

Whisper-large-v3 is a model developed by OpenAI to transcribe audio on many languages. This model is optimized for audio transcription tasks.

The model supports mono and stereo audio formats. For stereo formats, left and right channels are merged before being processed.

The model processes audio files in 30-second chunks. If an audio file of less than 30 seconds is sent, the rest of the chunk will be considered silent.

AttributeValue
ProviderOpenAI
Supports structured output-
Supports function calling-
Supports parallel tool-calling-
Supported audio formatsflac, m4a, mpeg, mp2, mp3, mp4, ogg, wav, webm
Audio chunk duration30 seconds
Maximum file size - Serverless25 MB
Supported languagesEnglish, French, German, Chinese, Japanese, Korean, and 81 additional languages
Compatible Instances (max context in tokens*) - Dedicated DeploymentL4, L40S, H100, H100-SXM-2
Hugging Face model cardwhisper-large-v3

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model names

openai/whisper-large-v3:bf16

Text models

Glm-5.2

Released in June 2026, GLM 5.2 is the best open-weight model for long-horizon and coding tasks at the time of release.

AttributeValue
ProviderZ.ai
Supports structured outputYes
Supports function callingYes
Supports parallel tool-callingYes
Supported reasoning effortsnone, high, max
Supported languagesEnglish, Chinese
Compatible Instances (max context in tokens*) - Dedicated DeploymentB300-SXM-8
Maximum output (tokens) - Serverless16k
Hugging Face model cardglm-5.2

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

DeepSeek-V4-Flash-0731

Released in July 2026, DeepSeek-V4-Flash is the best cost-efficient model for long-horizon and coding tasks at the time of release.

AttributeValue
ProviderDeepSeek
Supports structured outputYes
Supports function callingYes
Supports parallel tool-callingYes
Supported reasoning effortsnone, low ,high, max
Supported languagesEnglish, Chinese
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100-SXM-4, H100-SXM-8
Maximum output (tokens) - Serverless32k
Hugging Face model carddeepseek-v4-flash-0731

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

DeepSeek-R1-Distill-Llama-70B

Released in January 2025, DeepSeek’s R1 Distilled Llama 70B is a distilled version of the Llama model family based on DeepSeek R1. DeepSeek R1 Distill Llama 70B is designed to improve the performance of Llama models in reasoning use cases, such as mathematics and coding tasks.

AttributeValue
ProviderDeepSeek
Supports structured outputYes
Supports function callingYes
Supported languagesEnglish, Chinese
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100 (13k), H100-2
Maximum output (tokens) - Serverless4k
Hugging Face model carddeepseek-r1-distill-llama-70b

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model name

deepseek/deepseek-r1-distill-llama-70b:fp8
deepseek/deepseek-r1-distill-llama-70b:bf16

DeepSeek-R1-Distill-Llama-8B

Released in January 2025, DeepSeek’s R1 Distilled Llama 8B is a distilled version of the Llama model family based on DeepSeek R1. DeepSeek R1 Distill Llama 8B is designed to improve the performance of Llama models in reasoning use cases, such as mathematics and coding tasks.

AttributeValue
ProviderDeepSeek
Supports structured outputYes
Supports function callingYes
Supported languagesEnglish, Chinese
Compatible Instances (max context in tokens*) - Dedicated DeploymentL4 (90k), L40S, H100, H100-2

Model names

deepseek/deepseek-r1-distill-llama-8b:fp8
deepseek/deepseek-r1-distill-llama-8b:bf16

Qwen3-235b-a22b-instruct-2507

Released in July 2025, Qwen 3 235B A22B is an open-weight model, competitive in multiple benchmarks (such as LM Arena for text use cases) compared to Gemini 2.5 Pro and GPT4.5.

AttributeValue
ProviderQwen
Supports structured outputYes
Supports function callingYes
Supports parallel tool-callingYes
Supported languagesEnglish, French, German, Chinese, Japanese, Korean, and 113 additional languages and dialects
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100-SXM-2 (40k), H100-SXM-4
Maximum output (tokens) - Serverless16k
Hugging Face model cardqwen3-235b-a22b-instruct-2507

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model name

qwen/qwen3-235b-a22b-instruct-2507

Qwen3-235b-a22b-thinking-2507

Qwen3-235b-a22b-thinking-2507 is a highly capable and versatile language model, optimized for instruction following, logical reasoning through thinking capabilities, and long-context understanding, with enhanced performance across various domains and preferences.

AttributeValue
ProviderQwen
Supports parallel tool-callingYes
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100-2 (100k), H100-SXM-2 (100k), H100-SXM-4 (262k)
Hugging Face model cardQwen3-235B-A22B-Thinking-2507-AWQ

Model name

qwen/qwen3-235b-a22b-thinking-2507:awq

Gpt-oss-120b

Released in August 2025, GPT OSS 120B is an open-weight model providing significant throughput performance and reasoning capabilities. Currently, this model should be used through the Responses API, because the Chat Completions API does not yet support tool-calling for this model.

AttributeValue
ProviderOpenAI
Supports structured outputYes
Supports function callingYes
Supports parallel tool-callingYes
Supported reasoning effortslow, medium, high
Supported languagesEnglish
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100
Maximum output (tokens) - Serverless32k
Hugging Face model cardgpt-oss-120b

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model name

openai/gpt-oss-120b:fp4

Gpt-oss-20b

GPT OSS 20b is an OpenAI open-weight model designed for powerful reasoning, agentic tasks, and versatile developer use cases.

AttributeValue
ProviderOpenAI
Supports parallel tool-callingYes
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100, H100-2, H100-SXM-2, H100-SXM-4, H100-SXM-8
Hugging Face model cardgpt-oss-20b

Model name

openai/gpt-oss-20b:fp4

Llama-3.3-70b-instruct

Released in December 2024, Llama 3.3 70b by Meta is a fine-tune of the Llama 3.1 70b model. This model is text-only (text in/text out). However, Llama 3.3 was designed to approach the performance of Llama 3.1 405B on some applications.

AttributeValue
ProviderMeta
Supports structured outputYes
Supports function callingYes
Supports parallel tool-callingYes
Supported languagesEnglish, German, French, Italian, Portuguese, Hindi, Spanish, and Thai
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100 (15k), H100-2
Maximum output (tokens) - Serverless16k
Hugging Face model cardllama-3.3-70b-instruct

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model name

meta/llama-3.3-70b-instruct:fp8
meta/llama-3.3-70b-instruct:bf16

Llama-3.1-70b-instruct

Released in July 2024, Llama 3.1 by Meta is an iteration of the open-access Llama family. Llama 3.1 was designed to match the best proprietary models and outperform many of the available open-source common industry benchmarks.

AttributeValue
ProviderMeta
Supports structured outputYes
Supports function callingYes
Supports parallel tool-callingYes
Supported languagesEnglish, German, French, Italian, Portuguese, Hindi, Spanish, and Thai
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100 (15k), H100-2

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model names

meta/llama-3.1-70b-instruct:fp8
meta/llama-3.1-70b-instruct:bf16

Llama-3.1-8b-instruct

Released in July 2024, Llama 3.1 by Meta is an iteration of the open-access Llama family. Llama 3.1 was designed to match the best proprietary models and outperform many of the available open-source common industry benchmarks.

AttributeValue
ProviderMeta
Supports structured outputYes
Supports function callingYes
Supports parallel tool-callingYes
Supported languagesEnglish, German, French, Italian, Portuguese, Hindi, Spanish, and Thai
Compatible Instances (max context in tokens*) - Dedicated DeploymentL4 (90k), L40S, H100, H100-2
Hugging Face model cardllama-3.1-8b-instruct

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model names

meta/llama-3.1-8b-instruct:fp8
meta/llama-3.1-8b-instruct:bf16

Llama-3-8b-instruct

Llama-3-8b-instruct is the first generation of 8B-param models by Meta, fine-tuned for instruction and automation.

AttributeValue
ProviderMeta
Compatible Instances (max context in tokens*) - Dedicated DeploymentL4, L40S, H100, H100-2, H100-SXM-2, H100-SXM-4, H100-SXM-8
Hugging Face model cardMeta-Llama-3-8B-Instruct

Model name

meta/llama-3-8b-instruct:bf16

Llama-3-70b-instruct

Llama 3 by Meta is an iteration of the open-access Llama family. Llama 3 was designed to match the best proprietary models, enhanced by community feedback for greater utility and responsibly spearheading the deployment of LLMs. With a commitment to open-source principles, this release marks the beginning of a multilingual, multimodal future for Llama 3, pushing the boundaries in reasoning and coding capabilities.

AttributeValue
ProviderMeta
Supports structured outputYes
Supports function callingNo
Supported languagesEnglish
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100, H100-2

Model name

meta/llama-3-70b-instruct:fp8

Llama-3.1-Nemotron-70b-instruct

Introduced in October 2024, Nemotron 70B Instruct by NVIDIA is a specialized version of the Llama 3.1 model designed to follow complex instructions. NVIDIA employed Reinforcement Learning from Human Feedback (RLHF) to fine-tune the ability of the model to generate relevant and informative responses.

AttributeValue
ProviderNvidia
Supports structured outputYes
Supports function callingYes
Supported languagesEnglish
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100 (15k), H100-2

Model name

nvidia/llama-3.1-nemotron-70b-instruct:fp8

Mixtral-8x7b-instruct-v0.1

Mixtral-8x7b-instruct-v0.1, developed by Mistral, is tailored for instructional platforms and virtual assistants. Trained on vast instructional datasets, it provides clear and concise instructions across various domains, enhancing user learning experiences.

AttributeValue
ProviderDeepSeek
Supports structured outputYes
Supports function callingNo
Supported languagesEnglish, French, German, Italian, Spanish
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100-2

Model names

mistral/mixtral-8x7b-instruct-v0.1:fp8
mistral/mixtral-8x7b-instruct-v0.1:bf16

Mistral-7b-instruct-v0.3

The first dense model released by Mistral AI, perfect for experimentation, customization, and quick iteration. At the time of the release, it matched the capabilities of models up to 30B parameters. This model is open-weight and distributed under the Apache 2.0 license.

AttributeValue
ProviderMistral
Supports structured outputYes
Supports function callingYes
Supported languagesEnglish
Compatible Instances (max context in tokens*) - Dedicated DeploymentL4, L40S, H100, H100-2

Model name

mistral/mistral-7b-instruct-v0.3:bf16

Mistral-small-24b-instruct-2501

Mistral Small 24B Instruct is a state-of-the-art transformer model of 24B parameters, built by Mistral. This model is open-weight and distributed under the Apache 2.0 license.

AttributeValue
ProviderMistral
Supports structured outputYes
Supports function callingYes
Supported languagesEnglish, French, German, Dutch, Spanish, Italian, Polish, Portuguese, Chinese, Japanese, Korean
Compatible Instances (max context in tokens*) - Dedicated DeploymentL40S (20k), H100, H100-2

Model name

mistral/mistral-small-24b-instruct-2501:fp8
mistral/mistral-small-24b-instruct-2501:bf16

Mistral-nemo-instruct-2407

Mistral Nemo is a state-of-the-art transformer model of 12B parameters, built by Mistral in collaboration with NVIDIA. This model is open-weight and distributed under the Apache 2.0 license. It was trained on a large proportion of multilingual and code data.

AttributeValue
ProviderMistral
Supports structured outputYes
Supports function callingYes
Supports parallel tool-callingYes
Supported languagesEnglish, French, German, Spanish, Italian, Portuguese, Russian, Chinese, Japanese
Compatible Instances (max context in tokens*) - Dedicated DeploymentL40S, H100, H100-2
Maximum output (tokens) - Serverless8k
Hugging Face model cardmistral-nemo-instruct-2407

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model name

mistral/mistral-nemo-instruct-2407:fp8

Magistral-small-2506

Magistral Small is a reasoning model optimized to perform well on reasoning tasks, such as academic or scientific questions. It is well suited for complex tasks requiring multiple reasoning steps.

AttributeValue
ProviderMistral
Supports structured outputYes
Supports function callingYes
Supported languagesEnglish, French, German, Spanish, Portuguese, Italian, Japanese, Korean, Russian, Chinese, Arabic, Persian, Indonesian, Malay, Nepali, Polish, Romanian, Serbian, Swedish, Turkish, Ukrainian, Vietnamese, Hindi, Bengali
Compatible Instances (max context in tokens*) - Dedicated DeploymentL40S, H100, H100-2

Model name

mistral/magistral-small-2506:fp8
mistral/magistral-small-2506:bf16

Code models

Qwen3-coder-30b-a3b-instruct

Qwen3-coder is an improved version of Qwen2.5 with better accuracy and throughput. Thanks to its a3b architecture, only a subset of its weights is activated for a given generation, leading to much faster input and output token processing, ideal for code completion.

AttributeValue
ProviderQwen
Supports structured outputYes
Supports function callingYes
Supports parallel tool-callingYes
Supported languagesEnglish, French, German, Chinese, Japanese, Korean, and 113 additional languages and dialects
Compatible Instances (max context in tokens*) - Dedicated DeploymentL40S, H100, H100-2
Maximum output (tokens) - Serverless32k
Hugging Face model cardqwen3-coder-30b-a3b-instruct

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model name

qwen/qwen3-coder-30b-a3b-instruct:fp8

Qwen2.5-coder-32b-instruct

Qwen2.5-coder is your intelligent programming assistant familiar with more than 40 programming languages. With Qwen2.5-coder deployed at Scaleway, your company can benefit from code generation, AI-assisted code repair, and code reasoning.

AttributeValue
ProviderQwen
Supports structured outputYes
Supports function callingYes
Supports parallel tool-callingNo
Supported languagesEnglish, French, Spanish, Portuguese, German, Italian, Russian, Chinese, Japanese, Korean, Vietnamese, Thai, Arabic, and 16 additional languages
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100, H100-2

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model name

qwen/qwen2.5-coder-32b-instruct:int8

Devstral-2-123b-instruct-2512

Devstral 2 is a coding model released in December 2025, which excels at using tools to explore codebases, editing multiple files, and powering software engineering agents.

AttributeValue
ProviderMistral
Supports structured outputYes
Supports function callingYes
Supported languagesEnglish
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100-SXM-2 (75k), H100-SXM-4, H100-SXM-8
Maximum output (tokens) - Serverless16k
Hugging Face model carddevstral-2-123b-instruct-2512

Model name

mistral/devstral-2-123b-instruct-2512:fp8

Devstral-small-2505

Devstral Small is a fine-tune of Mistral Small 3.1, optimized to perform software engineering tasks. It is a good fit to be used as a coding agent, for instance in an IDE.

AttributeValue
ProviderMistral
Supports structured outputYes
Supports function callingYes
Supported languagesEnglish, French, German, Spanish, Portuguese, Italian, Japanese, Korean, Russian, Chinese, Arabic, Persian, Indonesian, Malay, Nepali, Polish, Romanian, Serbian, Swedish, Turkish, Ukrainian, Vietnamese, Hindi, Bengali
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100, H100-2

Model name

mistral/devstral-small-2505:fp8
mistral/devstral-small-2505:bf16

MiniMax-M2.5

A coding and agentic model excelling in tool use, search, and office tasks with exceptional efficiency and cost-effectiveness.

AttributeValue
ProviderMiniMaxAI
Supports parallel tool-callingYes
Compatible Instances (max context in tokens*) - Dedicated DeploymentH100-SXM-4, H100-SXM-8
Hugging Face model cardlukealonso/MiniMax-M2.5-NVFP4

*Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model name

minimaxai/minimax-m2.5:nvfp4

Embeddings models

Qwen3-embedding-8b

Qwen/Qwen3-Embedding-8B is an embedding model ranking third on the METB leaderboard as of November 2025, supporting custom dimensions between 32 and 4096.

AttributeValue
ProviderQwen
Supports structured outputNo
Supports function callingNo
Embedding dimensions (maximum)4096
Embedding dimensions (minimum)32
Matryoshka* embeddingYes
Supported languagesEnglish, French, German, Chinese, Japanese, Korean, and 113 additional languages and dialects
Compatible Instances (max context in tokens**) - Dedicated DeploymentL4, L40S, H100, H100-2
Hugging Face model cardqwen3-embedding-8b

*Matryoshka embeddings refers to embeddings trained on multiple dimension numbers. Consequently, resulting vector dimensions will be sorted by most meaningful first. For example, a 4096-dimension vector can be truncated to its 768 first dimensions and used directly.

**Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Bge-multilingual-gemma2

BGE-Multilingual-Gemma2 tops the MTEB leaderboard, scoring the number one spot in French and Polish, and number seven in English, as of Q4 2024. As its name suggests, the training data of the model spans a broad range of languages, including English, Chinese, Polish, French, and more.

AttributeValue
ProviderBAAI
Supports structured outputNo
Supports function callingNo
Embedding dimensions (maximum)3584
Embedding dimensions (minimum)3584
Matryoshka* embeddingNo
Supported languagesEnglish, French, Chinese, Japanese, Korean
Compatible Instances (max context in tokens**) - Dedicated DeploymentL4, L40S, H100, H100-2
Hugging Face model cardbge-multilingual-gemma2

*Matryoshka embeddings refers to embeddings trained on multiple dimension numbers. Consequently, resulting vector dimensions will be sorted by most meaningful first. For example, a 4096-dimension vector can be truncated to its 768 first dimensions and used directly.

**Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model name

baai/bge-multilingual-gemma2:fp32

Sentence-t5-xxl

The Sentence-T5-XXL model represents a significant evolution in sentence embeddings, building on the robust foundation of the Text-To-Text Transfer Transformer (T5) architecture. Designed for performance in various language processing tasks, Sentence-T5-XXL leverages the strengths of T5's encoder-decoder structure to generate high-dimensional vectors that encapsulate rich semantic information.

This model has been meticulously tuned for tasks such as text classification, semantic similarity, and clustering, making it a useful tool in the Retrieval-Augmented Generation (RAG) framework. It excels in sentence similarity tasks, but its performance in semantic search tasks is less optimal.

AttributeValue
ProviderSBERT
Supports structured outputNo
Supports function callingNo
Embedding dimensions768
Matryoshka* embeddingNo
Supported languagesEnglish
Compatible Instances (max context in tokens**) - Dedicated DeploymentL4

*Matryoshka embeddings refers to embeddings trained on multiple dimension numbers. Consequently, resulting vector dimensions will be sorted by most meaningful first. For example, a 4096-dimension vector can be truncated to its 768 first dimensions and used directly.

**Maximum context length is only mentioned when the VRAM size of an instance limits context length. Otherwise, maximum context length is the one defined by the model.

Model name

sentence-transformers/sentence-t5-xxl:fp32
Still need help?

Create a support ticket
No Results