Announcing Spark Connect on Amazon EMR on EKS: Interactive PySpark development, anywhere

Post Syndicated from Amit Maindola original https://aws.amazon.com/blogs/big-data/announcing-spark-connect-on-amazon-emr-on-eks/

Today, we’re announcing support for Spark Connect on Amazon EMR on EKS, starting from EMR release 7.14 (Apache Spark 3.5.8) and emr-spark-8.1 (Apache Spark 4.1.1). You can now build, test, and debug Spark applications from your preferred tools, such as VS Code, PyCharm, Jupyter notebooks, Amazon SageMaker Unified Studio. At the same time, your full-scale Spark operations run on Amazon Elastic Kubernetes Service (Amazon EKS).

Deploying Spark applications from a local development environment to a remote Amazon EKS cluster often means dealing with environment differences, dependency conflicts, and performance gaps at scale. Spark Connect removes this friction. It separates your application client from the Spark server, so you develop and debug locally while Spark Connect routes your operations to a scalable Spark cluster running on Amazon EKS.

This client-server architecture supports a range of use cases, including interactive development from notebooks and IDEs, embedded Spark in web services, and continuous integration and continuous delivery (CI/CD) data-quality tests. All of these run on your existing EKS infrastructure. Each Spark Connect session uses its own AWS Identity and Access Management (IAM) execution role, custom tags, and cost tracking. For more information, see the Amazon EMR on EKS documentation.

Here are two demonstrations of using Spark Connect in Amazon SageMaker Unified Studio Notebooks and in a VS Code local IDE:

Amazon SageMaker Unified Studio Notebooks demo:

Local IDE demo:

For a runnable end-to-end example in an IDE, try the Spark Connect sample notebook in the aws-emr-utilities repository. It includes a client wrapper solution, built by AWS architects, for simplified connectivity:

How Spark Connect works on Amazon EMR on EKS

Spark Connect uses a client-server architecture that separates application code from the Spark engine:

  1. Client – A lightweight PySpark library running in your environment (such as an IDE or notebook). It doesn’t need Spark installed, direct access to data, or resources sized for the workload.
  2. Connection (EMR managed endpoint) – The client sends Spark operations over a secure gRPC/TLS channel to the Spark Connect server.
  3. Server – Runs Spark pods in your Amazon EMR on EKS namespace, starting from a minimum of two executors (adjustable) with autoscaling. The server performs Spark operations using the EKS compute resources and accesses data stores, such as an Amazon Simple Storage Service (Amazon S3) bucket, through job execution roles.
  4. Results – The server streams query results back to the client through gRPC as Apache Arrow-encoded row batches.
Client-server data flow from a local PySpark client through a gRPC channel to Spark pods on Amazon EMR on EKS

Figure 1: Spark Connect’s client-server architecture

On endpoint creation, Amazon EMR on EKS launches the Spark Connect server as pods on EKS and returns an Elastic Load Balancing (ELB)-backed endpoint and a short-lived token. You don’t need to provision any server or networking manually. Because the Spark Connect server runs on the EKS cluster you already operate, it inherits the node types, container images, and Spark configurations. What you see while developing Spark applications on the client side is what runs in the EKS environment at scale.

To provide a secure, simplified experience, Amazon EMR on EKS provisions two additional components on first use of Spark Connect on the EKS cluster:

Shared Envoy authentication-proxy router and Secret Agent service on the EKS cluster

Figure 2: Shared Envoy router and Secret Agent service on the EKS cluster

  • Managed authentication-proxy router – a shared Envoy router with three replicas by default (adjustable), fronted by a Network Load Balancer (NLB). It routes client traffic to the correct server pods, terminates TLS, and validates the session token. One router serves Spark Connect endpoints on the EKS cluster.
  • Secret Agent service – a lightweight, long-running pod that manages the short-lived credentials for session authentication. One service per EMR security configuration.

These components are long-running and shared across endpoints. Amazon EMR on EKS creates them automatically with the first endpoint on the cluster. Because the router is cluster-scoped and Secret Agent is namespace-scoped, deleting a managed endpoint doesn’t remove them. They keep running so that new endpoints can start within a minute. The router’s replica count is tunable. Scale down for non-production environments to reduce cost or scale up for higher throughput.

To fully remove these components:

  • Terminate all active managed endpoints and their virtual cluster that reference the Secret Agent’s security configuration, then delete the security configuration.
  • Once the last session-enabled virtual cluster is deleted, the authentication-proxy router and its underly resources, including the NLB and VPC endpoint, are removed automatically.
  • Alternatively, delete the EKS cluster to remove all in-cluster components at once.

Why use Spark Connect on Amazon EMR on EKS

With Amazon EMR on EKS, teams can run Spark alongside other applications on shared Kubernetes clusters with existing infrastructure, operational tooling, and system expertise. Spark Connect extends that value to interactive, embedded, and self-service Spark workloads. Your client stays lightweight while Spark code runs in governed, scalable server pods on EKS.

Interactive development on shared Kubernetes clusters

Data engineers and scientists iterate on Spark code cell-by-cell in notebooks or local IDEs. The Spark engine runs remotely on EKS, so validation runs on the same engine as your batch workloads. After validation on the Spark Connect client, the same Spark code deploys as a batch StartJobRun with no changes.

Spark Connect sessions run as pods on your existing cluster. They reuse your EKS RBAC, network policies, node autoscaling, and observability stack (Prometheus, Grafana, Amazon CloudWatch Container Insights). There are no separate compute and monitoring layers to operate.

Embedded Spark in applications and services

The Spark Connect client is a compact PySpark library. Teams can embed Spark operations directly into Python applications such as web services, dashboards, automation scripts, or backend APIs. The heavy processing runs on EKS while the application stays lightweight.

Teams can also expose Spark Connect as a self-service capability on their internal application. Business users submit Spark SQL scripts from a web UI. The compute runs on Spark Connect server on EKS, so the team manages capacity, security, and upgrades centrally.

Multi-tenant data exploration with governance

Each Spark Connect session uses the data user’s IAM permissions that you configure, limiting their access to authorized AWS services, data lake tables, and S3 paths. Every session carries tags with user, project, endpoint and virtual cluster IDs, feeding directly into billing and compliance reports. Meanwhile, data producers maintain guardrails on source data without blocking self-service exploration.

To manage resource consumption across teams, Amazon EMR on EKS virtual clusters provide namespace-level isolation. Each tenant binds their Spark Connect endpoints to a virtual cluster (a namespace) with independent IAM roles. Using resource quotas and limit ranges on EKS, you can protect each virtual cluster by controlling the compute resources that Spark Connect sessions can consume. Importantly, activating EKS split-cost allocation tags helps with chargeback reporting in a multi-tenant environment.

Reusable container images and scalable deployment

Teams often maintain custom container images with proprietary libraries, including internal feature stores, compliance toolkits, UDFs, or machine learning (ML) frameworks. With Spark Connect on Amazon EMR on EKS, teams reuse those same images as the Spark runtime for interactive sessions. No separate dependency lists needed. The same image works for both batch jobs and Spark Connect sessions.

Beyond the image itself, you can control Spark pod scheduling in Amazon EMR on EKS through pod templates and managed endpoint APIs, scaling across your environment. For example, you can:

  • Pin server pods to specific node types through pod templates. For example, Spot for cost savings.
  • Apply Spark Dynamic Resource allocation (DRA) to right-size each interactive session.
  • Use GPU node pools for accelerated Spark RAPIDS or ML.
cat > /tmp/spark-connect-endpoint.json << EOF
{
  "name": "spark-connect-custom-config",
  "virtualClusterId": "$VC_ID",
  "type": "SPARK_CONNECT",
  "releaseLabel": "emr-7.14.0-latest",
  "executionRoleArn": "$ROLE_ARN",
  "configurationOverrides": {
    "applicationConfiguration": [{
      "classification": "spark-defaults",
      "properties": {
        "spark.kubernetes.container.image": "${CUSTOM_IMAGE_URI}",
        "spark.kubernetes.executor.podTemplateFile": "s3://$S3BUCKET/exec-pod-template.yaml",
        "spark.kubernetes.node.selector.karpenter.sh/nodepool": "gpu-pool",
        "spark.dynamicAllocation.enabled": "true",
        "spark.dynamicAllocation.minExecutors": "0"
      }
    }]
  }
}
EOF

aws emr-containers create-managed-endpoint \
--cli-input-json file:///tmp/spark-connect-endpoint.json

Multi-cluster, multi-Region, and hybrid architectures

Enterprises running EKS clusters across multiple AWS accounts, AWS Regions, or hybrid environments with on-premises Kubernetes can use Spark Connect to query data wherever it’s processed. The lightweight client only needs to reach the Spark Connect endpoint, not the underlying S3 buckets or AWS Glue data catalogs. This means no VPC peering or direct network paths to every data store.

The client-server split is the core architectural advantage of Spark Connect on Amazon EMR on EKS. A developer on a laptop behind a VPN, a CI/CD deployment pipeline in a centralized service account, or an Airflow DAG orchestrating across Regions can all connect to a remote Spark server on EKS. This works regardless of where the client itself runs. This decoupling simplifies cross-Region or cross-account analytics without duplicating data or requiring direct access to each data store.

Getting started

To create a Spark Connect endpoint on Amazon EMR on EKS, complete the following steps:

  1. Create EMR namespaces on EKS.
  2. Create an EMR security configuration.
  3. Create a virtual cluster with the security configuration.
  4. Create a Spark Connect managed endpoint.
  5. Obtain a session token.
  6. Connect from your application.

Prerequisites

To proceed with this post, make sure you have the following:

Step 1: Create EMR namespaces

# set environment variables
export EKS_CLUSTER_NAME=my-eks-cluster
export USER_NAMESPACE=spark-connect-1
export SYS_NAMESPACE=spark-connect-1-system
export AWS_REGION=us-west-2
# connect to your EKS cluster
aws eks update-kubeconfig --name $EKS_CLUSTER_NAME --region $AWS_REGION
kubectl create namespace $USER_NAMESPACE
kubectl create namespace $SYS_NAMESPACE

Step 2: Create a security configuration

cat > /tmp/sec-config.json << EOF
{
  "name": "spark-connect-1-sc",
  "securityConfigurationData": {
    "authenticationConfiguration": {
      "identityCenterConfiguration": { "enableIdentityCenter": false }
    }
  },
  "containerProvider": {
    "type": "EKS",
    "id": "$EKS_CLUSTER_NAME",
    "info": { "eksInfo": { "namespace": "$SYS_NAMESPACE" } }
  }
}
EOF

SEC_CONFIG_ID=$(aws emr-containers create-security-configuration \
--region $AWS_REGION \
--cli-input-json file:///tmp/sec-config.json \
--query id \
--output text)
echo "Security Configuration ID: $SEC_CONFIG_ID"

Step 3: Create a virtual cluster with the security configuration

cat > /tmp/vc.json << EOF
{
  "name": "spark-connect-demo",
  "containerProvider": {
    "id": "$EKS_CLUSTER_NAME",
    "type": "EKS",
    "info": {"eksInfo": {"namespace": "$USER_NAMESPACE"}}
  },
  "securityConfigurationId": "$SEC_CONFIG_ID",
  "sessionEnabled": true
}
EOF

VC_ID=$(aws emr-containers create-virtual-cluster \
--region $AWS_REGION \
--cli-input-json file:///tmp/vc.json \
--query 'id' \
--output text)
# validate the virtual cluster
echo "Virtual Cluster ID: $VC_ID"
aws emr-containers describe-virtual-cluster --region $AWS_REGION --id $VC_ID

Step 4: Create a Spark Connect managed endpoint

Start an interactive session on your virtual cluster. Provide a job execution role that grants the session access to your data sources.

# reuse an existing execution role
ROLE_ARN="arn:aws:iam::YOUR_ACCOUNT_ID:role/EMRonEKSExecutionRole"
cat > /tmp/spark-connect-endpoint.json << EOF
{
  "name": "spark-connect-demo",
  "virtualClusterId": "$VC_ID",
  "type": "SPARK_CONNECT",
  "releaseLabel": "emr-7.14.0-latest",
  "executionRoleArn": "$ROLE_ARN",
  "sessionIdleTimeoutInMinutes": 1440,
  "configurationOverrides": {
    "applicationConfiguration": [{
      "classification": "spark-defaults",
      "properties": {
        "spark.dynamicAllocation.enabled": "true",
        "spark.dynamicAllocation.minExecutors": "0",
        "spark.dynamicAllocation.maxExecutors": "2"
      }
    }]
  }
}
EOF

EP_ID=$(aws emr-containers create-managed-endpoint \
--region $AWS_REGION \
--cli-input-json file:///tmp/spark-connect-endpoint.json \
--query 'id' \
--output text)
# Validate
echo "Endpoint ID: $EP_ID"
export EP_URL=$(aws emr-containers describe-managed-endpoint \
--region $AWS_REGION \
--virtual-cluster-id $VC_ID \
--id $EP_ID \
--query 'endpoint.authProxyUrl' \
--output text)
echo "Endpoint URL: $EP_URL"
Managed endpoint creation output showing the endpoint ID and endpoint URL

Figure 3: Managed endpoint creation output with the endpoint ID and URL

You can optionally pass some custom configuration overrides and tags:

aws emr-containers create-managed-endpoint \
--type SPARK_CONNECT \
--virtual-cluster-id $VC_ID \
--name more-endpoint \
--execution-role-arn $ROLE_ARN \
--release-label emr-7.14.0-latest \
--configuration-overrides '{
"applicationConfiguration": [{
"classification": "spark-defaults",
"properties": {
"spark.executor.instances": "1",
"spark.executor.memory": "4g",
"spark.executor.cores": "1",
"spark.sql.extensions": "org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions"
}
}]
}' \
--tags '{
"team": "data-engineering",
"project": "customer-analytics"
}'

Step 5: Obtain a session token

Request a session token after the managed endpoint is active:

# get a session token with a 12-hour expiry (adjustable)
export TOKEN=$(aws emr-containers get-managed-endpoint-session-credentials \
--region $AWS_REGION \
--virtual-cluster-identifier $VC_ID \
--endpoint-identifier $EP_ID \
--execution-role-arn $ROLE_ARN \
--credential-type TOKEN \
--duration-in-seconds 43200 \
--query 'credentials.token' \
--output text)
echo "Session Token: $TOKEN"

Security note: Communication between your environment and the Spark Connect server is encrypted using TLS. The authentication token is time-limited (15 minutes by default). For long-running sessions, refresh the token periodically by calling get-managed-endpoint-session-credentials again. Consider using AWS Secrets Manager to store and retrieve tokens programmatically.

Step 6: Connect from your application

Use the returned endpoint URL and token to connect from a PySpark-compatible environment. The following Python code shows how to establish a Spark Connect session:

import os
from pyspark.sql import SparkSession
session_endpoint = os.environ["EP_URL"]
auth_token = os.environ["TOKEN"]
spark_conn_url = (f"{session_endpoint};use_ssl=true;x-aws-proxy-auth={auth_token}")
spark = SparkSession.builder
.remote(spark_conn_url)
.getOrCreate()
# verify the connection
print(f"Connected remotely! Spark version: {spark.version}")
# query data through the AWS Glue Data Catalog
df = spark.sql("SELECT * FROM my_catalog.my_database.my_table LIMIT 10")
df.show()

After you’re connected, you can:

  • Debug interactively – Set breakpoints, inspect DataFrames, and step through Spark code in your IDE or notebook while the operations run remotely on EKS.
  • Combine local and remote processing – Pull query results back to the client as a pandas or PyArrow DataFrame for local analysis, visualization, or ML (scikit-learn, notebook widgets), then push further Spark operations back to the server in the same session. Heavy processing stays on Amazon EMR on EKS. Only the results you request cross the wire.
  • Reconnect without losing state – A managed endpoint runs independently of single clients for a configurable idle timeout (default: 60 minutes). Your Spark session, cached data, and temporary views are preserved on the server between connections. When a session token expires (default: 15 minutes, configurable up to 12 hours), request a new token and reconnect to the same endpoint to resume where you left off.
  • Reuse across workload types – The same client connection pattern works everywhere Python runs: notebooks, IDEs, batch scripts, Airflow operators, or web services. One endpoint, one connection pattern, many workload types.

Validation

After you create the endpoint, verify that the Spark Connect server is running and reachable through Amazon EMR on EKS API and standard Kubernetes tooling:

# get endpoint status
aws emr-containers describe-managed-endpoint --virtual-cluster-id $VC_ID --id $EP_ID
# inspect the server pods (driver + executors) in your namespace
kubectl get pods -n $USER_NAMESPACE -l "emr-containers.amazonaws.com/managed-endpoint-id=$EP_ID"
# View driver logs
kubectl logs -n $USER_NAMESPACE <driver-pod-name> -c spark-kubernetes-driver
Terminal output showing endpoint status and the running driver and executor pods

Figure 4: Endpoint status and the running driver and executor pods

# to view the live Spark UI, port-forward your driver pod:
DRIVER_POD=$(kubectl get pods -n $USER_NAMESPACE \
-l "emr-containers.amazonaws.com/managed-endpoint-id=$EP_ID,emr-containers.amazonaws.com/component=driver" \
-o name)
kubectl port-forward -n $USER_NAMESPACE "$DRIVER_POD" 4040:4040
# Open http://localhost:4040 in your browser
Live Spark UI for the Spark Connect session viewed in a browser through port forwarding

Figure 5: Live Spark UI for the Spark Connect session

Spark Connect endpoints run as pods on your EKS cluster. The existing Kubernetes observability stack, such as CloudWatch Container Insights, Prometheus, and Grafana, captures Spark Connect endpoint metrics alongside other cluster workloads.

Clean up resources

Terminate your session when you’re done to avoid ongoing costs:

# (OPTIONAL) Endpoints are auto-deleted after the idle timeout (default: 60 minutes).
aws emr-containers delete-managed-endpoint \
--virtual-cluster-id $VC_ID \
--id $EP_ID
# Delete the virtual cluster only when no active endpoints remain
aws emr-containers delete-virtual-cluster --id $VC_ID
# Delete Security Configuration
aws emr-containers delete-security-configuration --id $SEC_CONFIG_ID
# remove the remaining EKS namespaces
kubectl delete namespace $USER_NAMESPACE $SYS_NAMESPACE spark-connect-router

Deleting or timing out a managed endpoint automatically removes its corresponding driver and executor pods. The Envoy router and Secret Agent service are shared across endpoints on the EKS cluster and remain running when individual endpoints are terminated. To fully remove these shared components, delete the virtual cluster to remove its corresponding Secret Agent service. Before doing so, ensure that no managed endpoints in the virtual cluster are active. Terminating the last session-enabled virtual cluster automatically removes the Envoy router from the EKS cluster.

Availability and pricing

Spark Connect on Amazon EMR on EKS is available with EMR release 7.14 (Apache Spark 3.5) and emr-spark-8.1 (Apache Spark 4.1), in all AWS Regions where Amazon EMR on EKS is available, except the AWS GovCloud (US) Regions and the China Regions. The Amazon SageMaker Unified Studio experience is available in supported Regions.

There is no additional charge for Spark Connect managed endpoints beyond the standard Amazon EMR on EKS pricing. You pay for underlying Amazon EKS resources such as EC2 and ELB. For timed-out or terminated managed endpoints, EMR automatically removes their Spark pods from the EKS cluster.

Recommendations for cost efficiency:

  • Use Karpenter (or Cluster Autoscaler) to right-size cluster capacity to session workload demand. This provisions nodes when endpoints need them and removes them when idle, which keeps cost aligned to actual usage.
  • Schedule interactive session pods on On-Demand instances for persistent compute.
  • Use AWS Graviton processors for better performance on Spark workloads.
  • Activate Amazon EMR on EKS Cost Allocation tags to track per-team and per-project spending at granular level.
  • Keep a single, shared Envoy router and NLB serving all Spark Connect endpoints (the default) on the cluster. Right-size the router replica count (three by default) for your availability requirements.

Considerations and limitations

Before you build on Spark Connect for Amazon EMR on EKS, review the Considerations and limitations in the Amazon EMR on EKS documentation.

Conclusion

In this post, we showed how, with Spark Connect on Amazon EMR on EKS, you can build, test, and debug Spark applications from the tools you already use: IDEs, notebooks, Amazon SageMaker Unified Studio or Airflow. Your workloads run at scale on your existing Kubernetes clusters, with no application code changes.

For teams already running Amazon EMR on EKS, Spark Connect extends your virtual clusters to interactive and embedded workloads. The same virtual cluster that runs your batch StartJobRun jobs now also serves Spark Connect sessions. Each session runs as pods on your EKS cluster, inheriting your node groups, container images, and Spark configurations. Each session also carries its own IAM execution role and cost tags. This extends the security, multi-tenancy, and observability of your Amazon EMR on EKS investment to a broader set of users and use cases.

To get started, visit the Spark Connect on Amazon EMR on EKS documentation, try the Amazon SageMaker Unified Studio Getting Started guide, and review the Amazon EMR on EKS release notes for EMR 7.14.


About the authors

Amit Maindola

Amit Maindola

Amit is a Senior Data & AI Architect with AWS ProServe team focused on data engineering, analytics, and AI/ML at Amazon Web Services. He helps customers in their digital transformation journey and enables them to build highly scalable, robust, and secure cloud-based analytical solutions on AWS to gain timely insights and make critical business decisions.

Melody Yang

Melody Yang

Melody Yang is a Principal Analytics Specialist Solution Architect at AWS with expertise in Big Data technologies. She is an experienced analytics leader working with AWS customers to provide best practice guidance and technical advice in order to assist their success in data transformation. Her areas of interests are open-source frameworks and automation, data engineering and DataOps.

Al MS

Al MS

Al is a product manager for Amazon EMR at AWS.