Use Private Networks with your Apache Spark™ cluster
Private Networks allow your Clusters for Apache Spark™ cluster to communicate in an isolated and secure network without needing to be connected to the public internet.
At the moment, Apache Spark™ clusters can only be attached to a Private Network during their creation, and cannot be detached and reattached to another Private Network afterward.
For full information about Scaleway Private Networks and VPC, see our dedicated documentation and best practices guide.
Before you start
To complete the actions presented below, you must have:
- A Scaleway account logged into the console
- Owner status or IAM permissions allowing you to perform actions in the intended Organization
- Created a Private Network
- Created an Ubuntu Instance attached to a Private Network
Use a cluster through a Private Network
Set up your Instance with Python 3.13 and Spark tools
-
Run the following commands from the shell of your Instance to install the required dependencies:
sudo apt update sudo apt install -y \ build-essential zlib1g-dev libssl-dev libbz2-dev libsqlite3-dev \ libreadline-dev libncurses-dev liblzma-dev libffi-dev \ openjdk-21-jre-headless -
Run the following command to install
pyenv:curl https://pyenv.run | bash -
Run the following commands to add
pyenvto your Bash configuration:echo 'export PATH="$HOME/.pyenv/bin:$PATH"' >> ~/.bashrc echo 'eval "$(pyenv init -)"' >> ~/.bashrc echo 'eval "$(pyenv virtualenv-init -)"' >> ~/.bashrc -
Run the following command to reload your shell:
exec $SHELL -
Run the following command to install Python 3.13:
# Download and install the latest Python 3.13 pyenv install 3.13 # Sets the default python version to 3.13 # If you prefer not to change this setting globally, use a pyenv local config. pyenv global 3.13 -
Run the following command to install Apache Spark™:
cd ~ wget https://archive.apache.org/dist/spark/spark-4.0.0/spark-4.0.0-bin-hadoop3.tgz sudo mkdir -p /opt/spark sudo tar -xzf spark-4.0.0-bin-hadoop3.tgz -C /opt/spark --strip-components=1 -
Run the following commands to add Apache Spark™ to your Bash configuration:
echo 'export SPARK_HOME=/opt/spark' >> ~/.bashrc echo 'export PATH="$SPARK_HOME/bin:$PATH"' >> ~/.bashrc -
Run the following command to reload your shell:
exec $SHELL
Run a Python application using spark-submit
-
Verify that you are using Python 3.13.x:
python3 --version -
Run
spark-submitwith the following command to calculate pi for 1000 iterations. Do not forget to replace the placeholders with the appropriate values.export SPARK_LOCAL_IP=<INSTANCE_PN_IP> spark-submit \ --master spark://<SPARK_MASTER_ENDPOINT>:7077 \ --deploy-mode client \ $SPARK_HOME/examples/src/main/python/pi.py 1000 -
Access the Apache Spark™ UI of your cluster. The list of completed applications displays. From here, you can inspect the jobs previously started using
spark-submit.
Run a Jvm application using spark-submit (Java or Scala)
-
Run the SparkPi application in client mode with the following command:
export SPARK_LOCAL_IP=<INSTANCE_PN_IP> spark-submit \ --master spark://<SPARK_MASTER_ENDPOINT>:7077 \ --deploy-mode client \ --class org.apache.spark.examples.SparkPi \ $SPARK_HOME/examples/jars/spark-examples_2.13-4.0.0.jar 1000 -
Run the SparkPi application in cluster mode with the following command:
UNSET SPARK_LOCAL_IP spark-submit \ --master spark://<SPARK_MASTER_ENDPOINT>:7077 \ --deploy-mode cluster \ --conf spark.driver.port=37000 --conf spark.blockManager.port=38000 \ --class org.apache.spark.examples.SparkPi \ $SPARK_HOME/examples/jars/spark-examples_2.13-4.0.0.jar 1000