Skip to navigationSkip to main contentSkip to footerScaleway Docs HomepageAsk our AI
Ask our AI

Use Private Networks with your Apache Spark™ cluster

Private Networks allow your Clusters for Apache Spark™ cluster to communicate in an isolated and secure network without needing to be connected to the public internet.

At the moment, Apache Spark™ clusters can only be attached to a Private Network during their creation, and cannot be detached and reattached to another Private Network afterward.

For full information about Scaleway Private Networks and VPC, see our dedicated documentation and best practices guide.

Note

You can attach this product to a Private Network, but it currently does not support VPC routing. It is therefore not compatible with VPC Peering.

Before you start

To complete the actions presented below, you must have:

Use a cluster through a Private Network

Set up your Instance with Python 3.13 and Spark tools

  1. Connect to your Instance via SSH.

  2. Run the following commands from the shell of your Instance to install the required dependencies:

    sudo apt update
    sudo apt install -y \
      build-essential zlib1g-dev libssl-dev libbz2-dev libsqlite3-dev \
      libreadline-dev libncurses-dev liblzma-dev libffi-dev \
      openjdk-21-jre-headless
    Note
    • Pyenv requires some dev packages for building a python version.
    • Spark 4.0 runs on Java 17/21 (a too recent version of the Java runtime might not work).
  3. Run the following command to install pyenv:

    curl https://pyenv.run | bash
  4. Run the following commands to add pyenv to your Bash configuration:

    echo 'export PATH="$HOME/.pyenv/bin:$PATH"' >> ~/.bashrc
    echo 'eval "$(pyenv init -)"' >> ~/.bashrc
    echo 'eval "$(pyenv virtualenv-init -)"' >> ~/.bashrc
  5. Run the following command to reload your shell:

    exec $SHELL
  6. Run the following command to install Python 3.13:

    # Download and install the latest Python 3.13
    pyenv install 3.13
    # Sets the default python version to 3.13
    # If you prefer not to change this setting globally, use a pyenv local config.
    pyenv global 3.13
    Note

    Your Instance's Python version must be 3.13. If you encounter an error due to a mismatch between the worker and driver Python versions, run the following command to display minor versions, then reinstall using the exact one:

    pyenv install -l | grep 3.13
  7. Run the following command to install Apache Spark™:

    cd ~
    wget https://archive.apache.org/dist/spark/spark-4.0.0/spark-4.0.0-bin-hadoop3.tgz
    sudo mkdir -p /opt/spark
    sudo tar -xzf spark-4.0.0-bin-hadoop3.tgz -C /opt/spark --strip-components=1
  8. Run the following commands to add Apache Spark™ to your Bash configuration:

    echo 'export SPARK_HOME=/opt/spark' >> ~/.bashrc
    echo 'export PATH="$SPARK_HOME/bin:$PATH"' >> ~/.bashrc
  9. Run the following command to reload your shell:

    exec $SHELL

Run a Python application using spark-submit

  1. Connect to your Instance via SSH.

  2. Verify that you are using Python 3.13.x:

    python3 --version
  3. Run spark-submit with the following command to calculate pi for 1000 iterations. Do not forget to replace the placeholders with the appropriate values.

    export SPARK_LOCAL_IP=<INSTANCE_PN_IP>
    spark-submit \
    --master spark://<SPARK_MASTER_ENDPOINT>:7077 \
    --deploy-mode client \
    $SPARK_HOME/examples/src/main/python/pi.py 1000
    Note
    • <SPARK_MASTER_ENDPOINT> can be found in the Overview tab of your cluster, under Private endpoint in the Network section.
    • <INSTANCE_PN_IP> can be found in the Private Networks tab of your Instance. Make sure to only copy the IP, and not the /22 part.
  4. Access the Apache Spark™ UI of your cluster. The list of completed applications displays. From here, you can inspect the jobs previously started using spark-submit.

Run a Jvm application using spark-submit (Java or Scala)

  1. Connect to your Instance via SSH.

  2. Run the SparkPi application in client mode with the following command:

    export SPARK_LOCAL_IP=<INSTANCE_PN_IP>
    spark-submit \
    --master spark://<SPARK_MASTER_ENDPOINT>:7077 \
    --deploy-mode client \
    --class org.apache.spark.examples.SparkPi \
    $SPARK_HOME/examples/jars/spark-examples_2.13-4.0.0.jar 1000
  3. Run the SparkPi application in cluster mode with the following command:

    UNSET SPARK_LOCAL_IP
    spark-submit  \
    --master spark://<SPARK_MASTER_ENDPOINT>:7077 \
    --deploy-mode cluster \
    --conf spark.driver.port=37000 --conf spark.blockManager.port=38000 \
    --class org.apache.spark.examples.SparkPi \
    $SPARK_HOME/examples/jars/spark-examples_2.13-4.0.0.jar 1000
    Note
    • spark.driver.port and spark.blockManager.port need to be in the 37000-38999 port range since only these ports are allowed in Spark clusters.
    • The spark-submit command will exit as soon as the driver has been started on the cluster. Check out the Spark UI to get the Pi result.
Still need help?

Create a support ticket
No Results