Scaleway logo
Generative APIs - Dedicated Deployment API

Generative APIs - Dedicated Deployment API

Scaleway Generative APIs - Dedicated Deployment allows you to deploy and run machine learning models on Scaleway infrastructure. This service provides scalable and efficient endpoints for your model inference needs. The Scaleway Generative APIs - Dedicated Deployment API enables you to manage these endpoints and perform inference operations with any OpenAI API compatible software.

Tip

For details about the available models, refer to the Generative APIs supported modelsOpen in new context page.

Concepts

Refer to the Generative APIs - ConceptsOpen in new context page to find the definitions of all concepts and terminology related to Generative APIs - Dedicated Deployment.

Quickstart

Requirement

  1. Configure your environment variables. This is an optional step to simplify your usage of the Generative APIs - Dedicated Deployment API. You can find your Project ID in the Scaleway consoleOpen in new context.

    TerminalCode
    export SCW_SECRET_KEY="<YOUR_API_SECRET_KEY>" export SCW_DEFAULT_REGION="fr-par" export SCW_PROJECT_ID="<YOUR_SCALEWAY_PROJECT_ID>"
  2. List available models: Run the following command to get a list of all the models available for deployment, with their details.

    TerminalCode
    curl -X GET \ -H "Content-Type: application/json" \ -H "X-Auth-Token: $SCW_SECRET_KEY" \ "https://api.scaleway.com/inference/v1/regions/$SCW_DEFAULT_REGION/endpoints"
  3. Create a model deployment: Run the following command to create a deployment. Customize the details in the payload.

    TerminalCode
    curl -X POST https://api.scaleway.com/inference/v1/regions/$SCW_DEFAULT_REGION/deployments \ -H "Content-Type: application/json" \ -H "X-Auth-Token: $SCW_SECRET_KEY" \ -d '{ "project_id": "'"$SCW_PROJECT_ID"'", "name": "<YOUR_INFERENCE_DEPLOYMENT>", "model_id": "<YOUR_CHOSEN_MODEL_ID>", "node_type": "L4", "min_size": 1, "max_size": 1, "accept_eula": true, "endpoints": [ { "public": {} } ] }'
    ParameterDescriptionValid values
    project_idThe Project in which the deployment should be created (string)Any valid Scaleway Project ID, e.g., "b4bd99e0-b389-11ed-afa1-0242ac120002"
    nameA name of your choice for the deployment (string)Any string containing only alphanumeric characters, dots, spaces, and dashes, e.g., "my-inference-deployment"
    model_idThe model to deploy (string)Any valid model ID found in your model library (see models listing)
    node_typeThe type of node to use for the deployment (string)Example: "L4"
    min_sizeMinimum number of replicas for the deployment (integer)Any integer, e.g., 1
    max_sizeMaximum number of replicas for the deployment (integer)Any integer, e.g., 3
    accept_eulaIndicates acceptance of the End User License Agreement (boolean)true
    endpointsDefines the endpoints for the deployment (array)At least one endpoint, e.g., [ { "public": {} } ]
  4. Create a model endpoint: Run the following command to create an inference endpoint for the deployment. Customize the details in the payload.

    Example for creating a public endpoint

    TerminalCode
    curl -X POST https://api.scaleway.com/inference/v1/regions/$SCW_DEFAULT_REGION/endpoints \ -H "Content-Type: application/json" \ -H "X-Auth-Token: $SCW_SECRET_KEY" \ -d '{ "project_id": "'"$SCW_PROJECT_ID"'", "deployment_id": "your-deployment-id", "endpoint": { "disable_auth": false, "public": {} } }'

    Example for creating a private endpoint

    TerminalCode
    curl -X POST https://api.scaleway.com/inference/v1/regions/$SCW_DEFAULT_REGION/endpoints \ -H "Content-Type: application/json" \ -H "X-Auth-Token: $SCW_SECRET_KEY" \ -d '{ "project_id": "'"$SCW_PROJECT_ID"'", "deployment_id": "your-deployment-id", "endpoint": { "disable_auth": false, "private_network": { "private_network_id": "your-private-network-id" } } }'
    ParameterDescriptionValid values
    project_idThe Project in which the endpoint should be created (string)Any valid Scaleway Project ID, e.g., "b4bd99e0-b389-11ed-afa1-0242ac120002"
    deployment_idThe deployment ID to which the endpoint will be associated (string)Any valid deployment ID, e.g. "bcb0976d-98d6-49c1-b6b5-17804941c0b7"
    disable_authSpecifies whether to disable authentication (boolean)true or false
    publicPublic endpoint configuration (object){} for public endpoint
    private_networkPrivate endpoint configuration including the private network ID (object){ "private_network_id": "private-network-id" }
  5. List your deployments: Run the following command to get a list of all the deployments in your account, with their details.

    TerminalCode
    curl -X GET \ -H "Content-Type: application/json" \ -H "X-Auth-Token: $SCW_SECRET_KEY" \ "https://api.scaleway.com/inference/v1/regions/$SCW_DEFAULT_REGION/deployments"
  6. List your endpoints: Run the following command to get a list of all the inference endpoints in your account, with their details.

    TerminalCode
    curl -X GET \ -H "Content-Type: application/json" \ -H "X-Auth-Token: $SCW_SECRET_KEY" \ "https://api.scaleway.com/inference/v1/regions/$SCW_DEFAULT_REGION/endpoints"
  7. Delete an endpoint: Run the following command to delete an inference endpoint, specified by its endpoint ID.

    TerminalCode
    curl -X DELETE \ -H "X-Auth-Token: $SCW_SECRET_KEY" \ -H "Content-Type: application/json" \ "https://api.scaleway.com/inference/v1/regions/$SCW_DEFAULT_REGION/endpoints/<endpoint-ID>"

    The expected successful response is empty.

    Important

    Dedicated Generative APIs deployments must have at least one endpoint, either public or private.

Technical information

Region

Generative APIs - Dedicated Deployment endpoints are available in the following region:

NameAPI ID
Parisfr-par

Pagination

Most listing requests receive a paginated response. Requests against paginated endpoints accept two query arguments:

  • page, a positive integer to choose which page to return
  • per_page, a positive integer lower or equal to 100 to set the number of items returned per page. The default value is 50.

Paginated endpoints usually also accept filters to search and sort results. These filters are documented in each endpoint's documentation.

The X-Total-Count header contains the total number of items returned.

Creating a deployment: the model object

When creating a deployment, the model_id parameter is required to specify the model to deploy. Use the List ModelsOpen in new context endpoint to retrieve available model IDs.

Note

This information is designed to help you correctly configure the model_id parameter when using the Create a deploymentOpen in new context method.

Going further

For more information about Generative APIs - Dedicated Deployment, refer to the following resources: