---
title: How to query reasoning models
description: Learn how to interact with powerful reasoning models using Scaleway's Generative APIs service.
tags: generative-apis ai-data language-models chat-completions-api reasoning think
dates:
  validation: 2025-10-07
  posted: 2025-10-07
---
import Requirements from '@macros/iam/requirements.mdx'

Scaleway's Generative APIs service allows users to interact with language models benefiting from additional reasoning capabilities.

A reasoning model is a language model that is capable of carrying out multiple inference steps and systematically verifying intermediate results before producing answers. You can specify how much effort it should put into reasoning via dedicated parameters, and access reasoning content in its outputs. Even with default parameters, such models are designed to perform better on reasoning tasks like maths and logic problems than non-reasoning language models. 

Language models supporting the reasoning feature include `gpt-oss-120b`. See [Supported models](/generative-apis/reference-content/supported-models/) for a full list.

You can interact with reasoning models in the following ways:

- Use the [playground](/generative-apis/how-to/query-reasoning-models/#accessing-the-playground) in the Scaleway [console](https://console.scaleway.com) to test models, adapt parameters, and observe how your changes affect the output in real-time
- Use the [Chat Completions API](https://www.scaleway.com/en/developers/api/generative-apis/chat-completions) or the [Responses API](https://www.scaleway.com/en/developers/api/generative-apis/responses)
- Use your own [dedicated deployment](/generative-apis/how-to/create-deployment/) of a chosen model

<Requirements />

- A Scaleway account logged into the [console](https://console.scaleway.com)
- [Owner](/iam/concepts/#owner) status or [IAM permissions](/iam/concepts/#permission) allowing you to perform actions in the intended Organization
- A valid [API key](/iam/how-to/create-api-keys/) for API authentication
- Python 3.7+ installed on your system

## Querying reasoning language models via the playground

### Accessing the playground

Scaleway provides a web playground for instruct-based models hosted on Generative APIs.

1. Navigate to **Generative APIs** under the **AI** section of the [Scaleway console](https://console.scaleway.com/) side menu. The list of models you can query displays.
2. Click the name of the reasoning model you want to try. Alternatively, click **Try** next to the model's name. Ensure that you choose a model with [reasoning capabilities](/generative-apis/reference-content/supported-models/).

The web playground displays.

### Using the playground

1. Enter a prompt at the bottom of the page, or use one of the suggested prompts in the conversation area.
2. Edit the parameters listed in the right column, for example the default temperature for more or less randomness in the outputs. 
3. Switch models at the top of the page, to observe the capabilities of chat models offered via Generative APIs. 
4. Click **Deploy**, then select the **Serverless** option to get code snippets configured according to your settings in the playground. 
    
    You can also choose to deploy a model on your own dedicated Instance by selecting the **Dedicated** option. In this case, you can access the playground after completing the steps in the deployment wizard. Once in the playground of your deployment, click **View code** to get code snippets that match your settings in the playground. 

<Message type="note">
You cannot currently set values for parameters such as `reasoning_effort`, or access reasoning metadata in the model's output, via the console playground. Query the models programmatically as shown below in order to access the full reasoning feature set.
</Message>

## Querying reasoning language models via API

You can query models programmatically using your favorite tools or languages.
In the example that follows, we will use the OpenAI Python client.

### Chat Completions API or Responses API?

Both the [Chat Completions API](https://www.scaleway.com/en/developers/api/generative-apis/chat-completions) and the [Responses API](https://www.scaleway.com/en/developers/api/generative-apis/responses) allow you to access and control reasoning for supported models.

For more details on Chat Completions versus Responses API, see the information provided in the [querying language models](/generative-apis/how-to/query-language-models/#chat-completions-api-or-responses-api) documentation.

### Installing the OpenAI SDK

Install the OpenAI SDK using pip:

```bash
pip install openai
```

### Initializing the client

Initialize the OpenAI client with your base URL and API key:

<Message type="tip">
    In the case of a dedicated Generative APIs deployment, the `base_url` value is the **Public Endpoint URL** displayed on the Overview tab of the deployment's dashboard.
</Message>

```python
from openai import OpenAI

# Initialize the client with your base URL and API key
client = OpenAI(
    base_url="https://api.scaleway.ai/v1",  # Scaleway's Generative APIs service URL
    api_key="<SCW_SECRET_KEY>"  # Your unique API secret key from Scaleway
)
```

### Generating a chat completion with reasoning

You can now create a chat completion with reasoning, using either the Chat Completions or Responses API, as shown in the following examples:

<Tabs>

    <TabsTab label="Chat Completions API">

    ```python
    # Create a chat completion using the 'gpt-oss-120b' model
    response = client.chat.completions.create(
      model="gpt-oss-120b",
      messages=[{"role": "user", "content": "Describe a futuristic city with advanced technology and green energy solutions."}],
      temperature=0.2,  # Adjusts creativity
      max_completion_tokens=512,   # Limits the length of the output
      top_p=0.7,         # Controls diversity through nucleus sampling. You usually only need to use temperature.
      reasoning_effort="medium"
    )

    # Print the generated response
    print(f"Reasoning: {response.choices[0].message.reasoning_content}")
    print(f"Answer: {response.choices[0].message.content}")
    ```

    This code sends a message to the model, specifies the effort to make with reasoning, and returns an answer based on your input. You can access the model's reasoning metadata and its answer in the model's output. Here is an example output:

    ```python
    Reasoning: The user wants a description of a futuristic city with advanced tech and green energy solutions. Should be creative, vivid, detailed. No disallowed content. Provide description.

    Answer: **City of Luminara – A Blueprint for the Future** <rest of answer truncated> 
   ```

    </TabsTab>

    <TabsTab label="Responses API">

    ```python
      response = client.responses.create(
        model="gpt-oss-120b",
        input=[{"role": "user", "content": "Briefly describe a futuristic city with advanced technology and green energy solutions."}],
        temperature=0.2,  # Adjusts creativity
        max_output_tokens=512,   # Limits the length of the output
        top_p=0.7,         # Controls diversity through nucleus sampling. You usually only need to use temperature.
        reasoning={"effort":"medium"}
      )
      # Print the generated response. Here, the last output message will contain the final content.
      # Previous outputs will contain reasoning content.
      for output in response.output:
        if output.type == "reasoning":
            print(f"Reasoning: {output.content[0].text}") # output.content[0].text can only be used with openai >= 1.100.0
        if output.type == "message":
            print(f"Answer: {output.content[0].text}")
    ```
    This code sends a message to the model, specifies the effort to make with reasoning, and returns an answer based on your input. You can access the model's reasoning metadata and its answer in the model's output. Here is an example output:

    ```python
    Reasoning: The user asks: "Briefly describe a futuristic city with advanced technology and green energy solutions." They want a brief description. Should be concise but vivid. Provide details: architecture, transport, energy, AI, and sustainability. Probably a paragraph or a few sentences. Ensure it's brief. Let's produce a short description.
    
    Answer: **Solaris Arcadia** rises from a reclaimed river delta, its skyline a lattice of translucent, self‑healing bioglass towers that
    ```
    </TabsTab>
</Tabs>

### Configuring reasoning

All models with reasoning capabilities have reasoning enabled by default (i.e., if the field `reasoning_effort` is not provided). You can disable reasoning for most models (except for `gpt-oss-120b`) by using `reasoning_effort=none`.

For a quick test, issue the following simple API calls.

Reasoning is disabled:

```bash
curl https://api.scaleway.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $SCW_SECRET_KEY" \
  -d '{
    "model": "qwen3.5-397b-a17b",
    "messages": [
      {
        "role": "system",
        "content": "You are a helpful assistant."
      },
      {
        "role": "user",
        "content": "Hello!"
      }
    ],
   "reasoning_effort": "none"
  }'
```
Reasoning is set to medium effort:

```bash
curl https://api.scaleway.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $SCW_SECRET_KEY" \
  -d '{
    "model": "qwen3.5-397b-a17b",
    "messages": [
      {
        "role": "system",
        "content": "You are a helpful assistant."
      },
      {
        "role": "user",
        "content": "Hello!"
      }
    ],
   "reasoning_effort": "medium"
  }' 
```

The supported `reasoning_effort` value (values such as `low`, `medium`, `high`) differs by model.


## Exceptions and legacy models

Some legacy models, such as `deepseek-r1-distill-llama-70b`, do not output reasoning data as described above, but make it available in the `content` field of the response inside special tags, as shown in the example below:

```
response.content = "<think> The user asks for questions about mathematics (...) </think>  Answer is 42."
```

The reasoning content is inside the `<think>`...`</think>` tags, and you can parse the response accordingly to access such content. There is, however, a known bug that can lead the model to omit the opening `<think>` tag, so we suggest taking care when parsing such outputs.

Note that the `reasoning_effort` parameter is not available for this model.

## Distinguishing between reasoning data and answer content in streaming mode

In streaming mode, some models (for example, Qwen3.5-397b-a17b) output two different server-side events for reasoning data and answer content. You receive all the reasoning chunks one by one, followed by all the answer chunks one by one (each chunk is a server‑side event of one type or the other).

Take the following example:

```json
data: {
    "id": "chatcmpl-d6804e7b-8099-41f8-8486-3233eb11178d",
    "object": "chat.completion.chunk",
    "created": 1777583471,
    "model": "qwen3.5-397b-a17b",
    "choices": [
        {
            "index": 0,
            "delta": {
                "reasoning": ")*"
            },
            "logprobs": null,
            "finish_reason": null,
            "token_ids": null
        }
    ]
}

data: {
    "id": "chatcmpl-d6804e7b-8099-41f8-8486-3233eb11178d",
    "object": "chat.completion.chunk",
    "created": 1777583471,
    "model": "qwen3.5-397b-a17b",
    "choices": [
        {
            "index": 0,
            "delta": {
                "reasoning": "\n"
            },
            "logprobs": null,
            "finish_reason": null,
            "token_ids": null
        }
    ]
}

data: {
    "id": "chatcmpl-d6804e7b-8099-41f8-8486-3233eb11178d",
    "object": "chat.completion.chunk",
    "created": 1777583471,
    "model": "qwen3.5-397b-a17b",
    "choices": [
        {
            "index": 0,
            "delta": {
                "content": "\n\n"
            },
            "logprobs": null,
            "finish_reason": null,
            "token_ids": null
        }
    ]
}

data: {
    "id": "chatcmpl-d6804e7b-8099-41f8-8486-3233eb11178d",
    "object": "chat.completion.chunk",
    "created": 1777583471,
    "model": "qwen3.5-397b-a17b",
    "choices": [
        {
            "index": 0,
            "delta": {
                "content": "The"
            },
            "logprobs": null,
            "finish_reason": null,
            "token_ids": null
        }
    ]
}
```


## Impact on token generation

Reasoning models generate reasoning tokens, which are billable. Generally these are in the model's output as part of the reasoning content. To limit the generation of reasoning tokens, you can adjust settings for the `reasoning_effort` and `max_completion_tokens` / `max_output_tokens` parameters. Alternatively, use a non-reasoning model to avoid the generation of reasoning tokens and subsequent billing.
