Scaleway logo
Clusters for Apache Spark™ API

Each cluster is composed of one or more dedicated compute nodes running Apache Spark™. It can optionally include a dedicated JupyterLab Notebook. These resources are fully manageable through this API.


GET
https://api.scaleway.com
/datalab/v1beta1/regions/{region}/cluster-versions

List the Apache Spark™ versions the product is compatible with.

path Parameters

region
​string · enum · required

The region you want to target

Enum values:
fr-par
it-mil

query Parameters

page
​integer · int32

The page number.

page_size
​integer · uint32

The page size.

The order by field.

Enum values:
name_asc
name_desc
Default: name_asc

Responses

200

The list of cluster versions.

total_count
​integer · uint64

The total count of cluster versions.


GET
https://api.scaleway.com
/datalab/v1beta1/regions/{region}/datalabs

List information about cluster within a project or an organization.

path Parameters

region
​string · enum · required

The region you want to target

Enum values:
fr-par
it-mil

query Parameters

organization_id
​string

The unique identifier of the organization whose clusters you want to list.

project_id
​string

The unique identifier of the project whose clusters you want to list.

name
​string

The name of the cluster you want to list.

tags
​string[]

The tags associated with the cluster you want to list.

page
​integer · int32

The page number for pagination.

page_size
​integer · uint32

The page size for pagination.

The order by field, available options are name_asc, name_desc, created_at_asc, created_at_desc, updated_at_asc, updated_at_desc.

Enum values:
name_asc
name_desc
created_at_asc
created_at_desc
updated_at_asc
updated_at_desc
Default: name_asc

Responses

200

The list of clusters. This is a list composed of messages of type DataLab.

total_count
​integer · uint64

The total count of clusters.


POST
https://api.scaleway.com
/datalab/v1beta1/regions/{region}/datalabs

Create a new cluster. In this call, one can personalize the node counts, add a notebook, choose the private network, define the persistent volume storage capacity.

path Parameters

region
​string · enum · required

The region you want to target

Enum values:
fr-par
it-mil

Request Body

project_id
​string

The unique identifier of the project where the cluster will be created.

name
​string

The name of the cluster.

description
​string

The description of the cluster.

tags
​string[]

The tags of the cluster.

The cluster main node specification. It holds the parameters node_type which specifies the node type of the main node. See ListNodeTypes for available options. See ListNodeTypes for available options.

The cluster worker node specification. It holds the parameters node_type which specifies the node type of the worker node and node_count for specifying the amount of nodes.

has_notebook
​boolean

Select this option to include a notebook as part of the cluster.

spark_version
​string

The version of Apache Spark™ running inside the cluster, available options can be viewed at ListClusterVersions.

The maximum persistent volume storage that will be available during workload.

private_network_id
​string

The unique identifier of the private network the cluster will be attached to.

Responses

200
id
​string

The unique identifier of the cluster. (UUID format)

project_id
​string

The unique identifier of the project where the cluster has been created. (UUID format)

name
​string

The name of the cluster.

description
​string

The description of the cluster.

tags
​string[]

The tags of the cluster.

The Apache Spark™ Main node specification of cluster. It holds the parameters node_type, spark_ui_url (available to reach Apache Spark™ UI), spark_master_url (used to reach the cluster within a VPC), root_volume (size of the volume assigned to the cluster).

The cluster worker nodes specification. It holds the parameters node_type, node_count, root_volume (size of the volume assigned to the cluster).

The status of the cluster. For a working cluster the status is marked as ready.

Enum values:
unknown_status
creating
updating
ready
error
deleting
locked
deleted
Default: unknown_status
created_at
​string | null · date-time

The creation timestamp of the cluster. (RFC 3339 format)

updated_at
​string | null · date-time

The last update date of the cluster. (RFC 3339 format)

region
​string

The region of the cluster.

has_notebook
​boolean

Whether a JupyterLab notebook is associated with the cluster or not.

notebook_url
​string | null

The URL of the notebook if available.

spark_version
​string

The version of Apache Spark™ running inside the cluster.

The total persistent volume storage selected to run Apache Spark™.

private_network_id
​string

The unique identifier of the private network to which the cluster is attached to. (UUID format)

notebook_master_url
​string | null

The URL that is used to reach the cluster from the notebook when available. This URL cannot be used to reach the cluster from a server.


GET
https://api.scaleway.com
/datalab/v1beta1/regions/{region}/datalabs/{datalab_id}

Retrieve information about a given cluster, specified by the region and datalab_id parameters. Its full details, including name, status, node counts, are returned in the response object.

path Parameters

region
​string · enum · required

The region you want to target

Enum values:
fr-par
it-mil
datalab_id
​string · required

The unique identifier of the cluster.

Responses

200
id
​string

The unique identifier of the cluster. (UUID format)

project_id
​string

The unique identifier of the project where the cluster has been created. (UUID format)

name
​string

The name of the cluster.

description
​string

The description of the cluster.

tags
​string[]

The tags of the cluster.

The Apache Spark™ Main node specification of cluster. It holds the parameters node_type, spark_ui_url (available to reach Apache Spark™ UI), spark_master_url (used to reach the cluster within a VPC), root_volume (size of the volume assigned to the cluster).

The cluster worker nodes specification. It holds the parameters node_type, node_count, root_volume (size of the volume assigned to the cluster).

The status of the cluster. For a working cluster the status is marked as ready.

Enum values:
unknown_status
creating
updating
ready
error
deleting
locked
deleted
Default: unknown_status
created_at
​string | null · date-time

The creation timestamp of the cluster. (RFC 3339 format)

updated_at
​string | null · date-time

The last update date of the cluster. (RFC 3339 format)

region
​string

The region of the cluster.

has_notebook
​boolean

Whether a JupyterLab notebook is associated with the cluster or not.

notebook_url
​string | null

The URL of the notebook if available.

spark_version
​string

The version of Apache Spark™ running inside the cluster.

The total persistent volume storage selected to run Apache Spark™.

private_network_id
​string

The unique identifier of the private network to which the cluster is attached to. (UUID format)

notebook_master_url
​string | null

The URL that is used to reach the cluster from the notebook when available. This URL cannot be used to reach the cluster from a server.


DELETE
https://api.scaleway.com
/datalab/v1beta1/regions/{region}/datalabs/{datalab_id}

Delete a cluster based on its region and id.

path Parameters

region
​string · enum · required

The region you want to target

Enum values:
fr-par
it-mil
datalab_id
​string · required

The unique identifier of the cluster.

Responses

200
id
​string

The unique identifier of the cluster. (UUID format)

project_id
​string

The unique identifier of the project where the cluster has been created. (UUID format)

name
​string

The name of the cluster.

description
​string

The description of the cluster.

tags
​string[]

The tags of the cluster.

The Apache Spark™ Main node specification of cluster. It holds the parameters node_type, spark_ui_url (available to reach Apache Spark™ UI), spark_master_url (used to reach the cluster within a VPC), root_volume (size of the volume assigned to the cluster).

The cluster worker nodes specification. It holds the parameters node_type, node_count, root_volume (size of the volume assigned to the cluster).

The status of the cluster. For a working cluster the status is marked as ready.

Enum values:
unknown_status
creating
updating
ready
error
deleting
locked
deleted
Default: unknown_status
created_at
​string | null · date-time

The creation timestamp of the cluster. (RFC 3339 format)

updated_at
​string | null · date-time

The last update date of the cluster. (RFC 3339 format)

region
​string

The region of the cluster.

has_notebook
​boolean

Whether a JupyterLab notebook is associated with the cluster or not.

notebook_url
​string | null

The URL of the notebook if available.

spark_version
​string

The version of Apache Spark™ running inside the cluster.

The total persistent volume storage selected to run Apache Spark™.

private_network_id
​string

The unique identifier of the private network to which the cluster is attached to. (UUID format)

notebook_master_url
​string | null

The URL that is used to reach the cluster from the notebook when available. This URL cannot be used to reach the cluster from a server.


PATCH
https://api.scaleway.com
/datalab/v1beta1/regions/{region}/datalabs/{datalab_id}

Update a cluster node counts. Allows for up- and downscaling on demand, depending on the expected workload.

path Parameters

region
​string · enum · required

The region you want to target

Enum values:
fr-par
it-mil
datalab_id
​string · required

The unique identifier of the cluster.

Request Body

name
​string | null

The updated name of the cluster.

description
​string | null

The updated description of the cluster.

tags
​string[]

The updated tags of the cluster.

node_count
​integer | null · uint32

The updated node count of the cluster. Scale up or down the number of worker nodes.

Responses

200
id
​string

The unique identifier of the cluster. (UUID format)

project_id
​string

The unique identifier of the project where the cluster has been created. (UUID format)

name
​string

The name of the cluster.

description
​string

The description of the cluster.

tags
​string[]

The tags of the cluster.

The Apache Spark™ Main node specification of cluster. It holds the parameters node_type, spark_ui_url (available to reach Apache Spark™ UI), spark_master_url (used to reach the cluster within a VPC), root_volume (size of the volume assigned to the cluster).

The cluster worker nodes specification. It holds the parameters node_type, node_count, root_volume (size of the volume assigned to the cluster).

The status of the cluster. For a working cluster the status is marked as ready.

Enum values:
unknown_status
creating
updating
ready
error
deleting
locked
deleted
Default: unknown_status
created_at
​string | null · date-time

The creation timestamp of the cluster. (RFC 3339 format)

updated_at
​string | null · date-time

The last update date of the cluster. (RFC 3339 format)

region
​string

The region of the cluster.

has_notebook
​boolean

Whether a JupyterLab notebook is associated with the cluster or not.

notebook_url
​string | null

The URL of the notebook if available.

spark_version
​string

The version of Apache Spark™ running inside the cluster.

The total persistent volume storage selected to run Apache Spark™.

private_network_id
​string

The unique identifier of the private network to which the cluster is attached to. (UUID format)

notebook_master_url
​string | null

The URL that is used to reach the cluster from the notebook when available. This URL cannot be used to reach the cluster from a server.


GET
https://api.scaleway.com
/datalab/v1beta1/regions/{region}/node-types

List the available compute node types for creating a new cluster.

path Parameters

region
​string · enum · required

The region you want to target

Enum values:
fr-par
it-mil

query Parameters

page
​integer · int32

The page number.

page_size
​integer · uint32

The page size.

The order by field. Available fields are name_asc, name_desc, vcpus_asc, vcpus_desc, memory_gigabytes_asc, memory_gigabytes_desc, vram_bytes_asc, vram_bytes_desc, gpus_asc, gpus_desc.

Enum values:
name_asc
name_desc
vcpus_asc
vcpus_desc
memory_gigabytes_asc
memory_gigabytes_desc
vram_bytes_asc
vram_bytes_desc
Default: name_asc

Filter based on the target of the nodes. Allows to filter the nodes based on their purpose which can be main or worker node.

Enum values:
unknown_target
notebook
worker

Filter based on node type ( cpu/gpu/all ).

Enum values:
all
gpu
cpu
Default: all

Responses

200

The list of node types.

total_count
​integer · uint64

The total count of node types.


GET
https://api.scaleway.com
/datalab/v1beta1/regions/{region}/notebook-versions

Lists available notebook versions.

path Parameters

region
​string · enum · required

The region you want to target

Enum values:
fr-par
it-mil

query Parameters

page
​integer · int32

The page number.

page_size
​integer · uint32

The page size.

The order by field. Available options are name_asc and name_desc.

Enum values:
name_asc
name_desc
Default: name_asc

Responses

200

The list of notebook versions.

total_count
​integer · uint64

The total count of notebook versions.