AI Advertising Company Hacked

Post Syndicated from Bruce Schneier original https://www.schneier.com/blog/archives/2025/12/ai-advertising-company-hacked.html

At least some of this is coming to light:

Doublespeed, a startup backed by Andreessen Horowitz (a16z) that uses a phone farm to manage at least hundreds of AI-generated social media accounts and promote products has been hacked. The hack reveals what products the AI-generated accounts are promoting, often without the required disclosure that these are advertisements, and allowed the hacker to take control of more than 1,000 smartphones that power the company.

The hacker, who asked for anonymity because he feared retaliation from the company, said he reported the vulnerability to Doublespeed on October 31. At the time of writing, the hacker said he still has access to the company’s backend, including the phone farm itself.

Slashdot thread.

Със или без утре – ние решаваме

Post Syndicated from Емилия Милчева original https://www.toest.bg/sys-ili-bez-utre-nie-reshavame/

Със или без утре – ние решаваме

За да бъде разкачена България от модела „Борисов – Пеевски“, политическата задача на антистатуквото е да разкаже какво предлага отвъд него, а гражданите да решат кой разказ за бъдещето им харесва. Защото сега опозицията няма с какво друго да се конкурира помежду си, освен кой да нарита Пеевски – буквално и преносно.

В интервю за LSE авторът на книгата In the Long Run: The Future as a Political Idea Джонатан Уайт обяснява, че за да функционира демокрацията, от решаващо значение е хората да вярват, че е възможно „отворено бъдеще“: че съществуват алтернативи на статуквото; че обществото може да се развива по различни начини и хората могат да избират измежду тях. 

В наши дни, казва Уайт, е по-трудно да се поддържа подобно търпение и вяра. Бъдещето се възприема като заплашително и задушаващо. Уайт го описва като „затварящо се“, а катастрофата – общественият разпад, конфликтът, екологичният срив – изглежда трудно предотвратима. 

В България усещането за „отворено бъдеще“ е силно ерозирало, което си е болестотворен симптом за състоянието на демокрацията. Формално има избори, партии и смяна на властта, но част от обществото е уморена и за нея резултатът изглежда предрешен: „Нищо няма да се промени, пак ще е същото.“ В протестите срещу задкулисието едни граждани участват, за да победят страха, че следващият шанс може да не дойде, а други – защото вярват, че ще пробият затворения кръг на формалната демокрация и ще наложат дълбока промяна отвъд поредната смяна на лица.

Българското общество живя твърде дълго в режим на перманентна криза – икономическа, финансова, политическа. Представата му за бъдещето се върти около връщане към „по-нормално вчера“, вместо да се насочи към „по-добро утре“. Последните седем парламентарни избори от 2021-ва насам са все в резултат на сблъсък, но не и етап от демократичен процес. Осмите, които са на път, ще са поредните от тази серия, предшествани от поредните промени в Изборния кодекс (ИК).

Избори и доверие

След като коалиционното правителство обяви оставка, следващият етап са изборите – и не просто самият вот, а механизмите за неговото провеждане и контрол, така че да бъде честен. Ще бъде трудно чрез протести да се наложи гласуването изцяло чрез машини, както предлага ПП–ДБ със законопроект. Тоест за тях решението е в технократизацията на демокрацията. За същото, но с добавка за възможност за електронно гласуване призовава и „Спаси София“, макар да няма парламентарно представителство.

Когато оставката вече е подадена и предстои президентът да насрочи дата за избори (вероятно на 5 април догодина), промените в ИК могат да се реализират единствено чрез парламентарно мнозинство. Към момента такова изглежда невъзможно да се намери, макар всички партии да декларират, че са за честни избори.

Алиансът за права и свободи (групата около Доган) настоява за машинно броене и отчитане на резултатите. МЕЧ е против този парламент да променя ИК. В „Има такъв народ“ също са пълни с идеи за промени – чрез нови машини за гласуване, сканиращи устройства, макар да са убедени в изключителната вредност на подобни законодателни дейности преди вот. 

Мандатоносителят на кабинета в оставка – ГЕРБ–СДС, засега не е излязъл със свои предложения за промени. Но партията на Бойко Борисов ще има решаващата дума за съдбата на машинното гласуване в зависимост от това дали ще подкрепи ПП–ДБ с гласовете си. ДПС – Ново начало, партията, която се владее от санкционирания за корупция от САЩ и Великобритания Делян Пеевски, оферира хартиен вот и отчитане с преброителни машини. 

За намаляване „в максимална степен на субективността на изборния процес“ се обяви президентът Румен Радев на консултациите с политическите сили.

Това се дава от машините, но не във формата на тяхното използване като принтери, както е към момента, а с електронно отчитане на резултата и контролно преброяване на разписките. Това е въпрос, който трябва да реши българският законодател. Призовавам парламента да премине към тази стъпка.

През декември 2022 г. депутатите от ГЕРБ, ДПС и БСП обезсмислиха машинното гласуване с промени в ИК, които превърнаха устройствата в принтери и задължиха секционните избирателни комисии да броят ръчно гласовете. А сезираният за промените Конституционен съд не видя никакъв проблем.

Каквито и „клъцни-срежи“ да се обмислят за ИК, остават за след Нова година. А демократичността на вота зависи и от контрола на МВР. Честността на изборите е стъпка към отвореното бъдеще, но не го гарантира. Гарантира единствено, че политическото представителство в парламента ще бъде основано в максимална степен на свободната воля на гражданите. 

Отвъд „Борисов – Пеевски“

Раздалечаването от Пеевски обаче започна, макар че той не е крайното зло, а фирмената му табела. Изглежда, и той усети ветровете на промяната, освободи бившия кабинет на Тодор Живков под предлог, че го декорира като бутафорен „Музей на сглобката“ (макар стаята да остава във владение на ДПС до разпускането на този парламент), и декларира уж отказ от охраната си от НСО. Не на шега обаче парламентарната вътрешна комисия предприе първи стъпки да го лиши от нея. С единодушие членовете ѝ одобриха на първо четене Националната служба за охрана да не охранява депутати, освен председателя на Народното събрание. Същите предложения на ПП–ДБ бяха отхвърлени по-рано тази година.

Свалянето на държавната охрана от Пеевски е знаково събитие, сигнал за избирателите на ДПС – Ново начало (турци, помаци, роми, по-малко етнически българи), че техният „султан“ вече не е всемогъщ. Доскоро това беше немислим ход. 

Следващата стъпка от обезсилването му е да се превърне в „електорално джудже“, по определението на Явор Божанков (ПП–ДБ). Вече му го предричат, а социологическите агенции го и отчитат. По данни на „Алфа Рисърч“ ДПС – Ново начало пада на четвърто място с 9,4%, ако изборите бяха днес, а Пеевски остава рекордьор по неодобрение от избирателите: 5,6% доверие срещу 80,3% недоверие. Очаквано ГЕРБ–СДС губи от протестите, но все още запазва първото място, а ПП–ДБ отбелязва ръст.

Способността на олигархията да се прегрупира, но не и да се разпада, е добре известна. В този контекст ключов въпрос е дали ПП–ДБ са готови да използват момента, за да предложат структурни реформи, които надхвърлят символичните ходове срещу отделни фигури. Досегашните действия – премахване на охраната на Пеевски, законодателните инициативи за машинно гласуване и засилването на прозрачността, показват, че обединението има амбиция за промяна. 

Но липсва пакетът от дълбоки реформи в съдебната система, здравеопазването, образованието, публичните финанси, държавната администрация, пенсионната система. Само така биха се легитимирали като алтернатива на сегашното статукво.

В противен случай алтернативата ще дойде не от парламента, а от Президентството. Но Румен Радев не би могъл да предложи „отворено бъдеще“ в плуралистичния смисъл на Уайт, а по-скоро обещание за поредната стабилност и поредния ред. То може да изглежда привлекателно за уморено общество, което търси изход от перманентната криза, но свежда избора до доверие в една-единствена фигура вместо в демократичния процес.

Не такъв разказ искат младите хора, изпълващи отново и отново площади и улици. Те търсят не държава, подредена от един властови център, а перспектива, в която имат място, глас и избор – бъдеще чрез конкуренция на идеи, политики и визии. Жокерът на политическия сезон е в техни ръце и промени играта. 

Остава да просветли и бъдещето.

Modernize Apache Spark workflows using Spark Connect on Amazon EMR on Amazon EC2

Post Syndicated from Philippe Wanner original https://aws.amazon.com/blogs/big-data/modernize-apache-spark-workflows-using-spark-connect-on-amazon-emr-on-amazon-ec2/

Apache Spark Connect, introduced in Spark 3.4, enhances the Spark ecosystem by offering a client-server architecture that separates the Spark runtime from the client application. Spark Connect enables more flexible and efficient interactions with Spark clusters, particularly in scenarios where direct access to cluster resources is limited or impractical.

A key use case for Spark Connect on Amazon EMR is to be able to connect directly from your local development environments to Amazon EMR clusters. By using this decoupled approach, you can write and test Spark code on your laptop while using Amazon EMR clusters for execution. This capability reduces development time and simplifies data processing with Spark on Amazon EMR.

In this post, we demonstrate how to implement Apache Spark Connect on Amazon EMR on Amazon Elastic Compute Cloud (Amazon EC2) to build decoupled data processing applications. We show how to set up and configure Spark Connect securely, so you can develop and test Spark applications locally while executing them on remote Amazon EMR clusters.

Solution architecture

The architecture centers on an Amazon EMR cluster with two node types. The primary node hosts both the Spark Connect API endpoint and Spark Core components, serving as the gateway for client connections. The core node provides additional compute capacity for distributed processing. Although this solution demonstrates the architecture with two nodes for simplicity, it scales to support multiple core and task nodes based on workload requirements.

In Apache Spark Connect version 4.x, TLS/SSL network encryption is not inherently supported. We show you how to implement secure communications by deploying an Amazon EMR cluster with Spark Connect on Amazon EC2 using an Application Load Balancer (ALB) with TLS termination as the secure interface. This approach enables encrypted data transmission between Spark Connect clients and Amazon Virtual Private Cloud (Amazon VPC) resources.

The operational flow is as follows:

  1. Bootstrap script – During Amazon EMR initialization, the primary node fetches and executes the start-spark-connect.sh file from Amazon Simple Storage Service (Amazon S3). This script starts the Spark Connect server.
  2. Server availability – When the bootstrap process is complete, the Spark Server enters a waiting state, ready to accept incoming connections. The Spark Connect API endpoint becomes available on the configured port (typically 15002), listening for gRPC connection from remote clients.
  3. Client interaction – Spark Connect clients can establish secure connections to an Application Load Balancer. These clients translate DataFrame operations into unresolved logical query plans, encode these plans using protocol buffers, and send them to the Spark Connect API using gRPC.
  4. Encryption in transit – The Application Load Balancer receives incoming gRPC or HTTPS traffic, performs TLS termination (decrypting the traffic), and forwards the requests to the primary node. The certificate is stored in AWS Certificate Manager (ACM).
  5. Request processing – The Spark Connect API receives the unresolved logical plans, translates them into Spark’s built-in logical plan operators, passes them to Spark Core for optimization and execution, and streams results back to the client as Apache Arrow-encoded row batches.
  6. (Optional) Operational access – Administrators can securely connect to both primary and core nodes through Session Manager, a capability of AWS Systems Manager, enabling troubleshooting and maintenance without exposing SSH ports or managing key pairs.

The following diagram depicts the architecture of this post’s demonstration for submitting Spark unresolved logical plans to EMR clusters using Spark Connect.

Apache Spark Connect on Amazon EMR solution architecture diagram

Apache Spark Connect on Amazon EMR solution architecture diagram

Prerequisites

To proceed with this post, ensure you have the following:

Implementation steps

In this recipe, through AWS CLI commands, you will:

  1. Prepare the bootstrap script, a bash script starting Spark Connect on Amazon EMR.
  2. Set up the permissions for Amazon EMR to provision resources and perform service-level actions with other AWS services.
  3. Create the Amazon EMR cluster with these associated roles and permissions and eventually attach the prepared script as a bootstrap action.
  4. Deploy the Application Load Balancer and certificate with ACM secure data in transit over the internet.
  5. Modify the primary node’s security group to allow Spark Connect clients to connect.
  6. Connect with a test application connecting the client to Spark Connect server.

Prepare the bootstrap script

To prepare the bootstrap script, follow these steps:

  1. Create an Amazon S3 bucket to host the bootstrap bash script:
    REGION=
    BUCKET_NAME=
    aws s3api create-bucket \
       --bucket $BUCKET_NAME \ 
       --region $REGION \
       --create-bucket-configuration LocationConstraint=$REGION

  2. Open your preferred text editor, add the following commands in a new file with a name such start-spark-connect.sh. If the script runs on the primary node, it starts Spark Connect server. If it runs on a task or core node, it does nothing:
    #!/bin/bash
    if grep isMaster /mnt/var/lib/info/instance.json | grep false;
    then
        echo "This is not master node, do nothing."
        exit 0
    fi
    echo "This is master, continuing to execute script"
    SPARK_HOME=/usr/lib/spark
    SPARK_VERSION=$(spark-submit --version 2>&1 | grep "version" | head -1 | awk '{print $NF}' | grep -oE '[0-9]+\.[0-9]+\.[0-9]+')
    SCALA_VERSION=$(spark-submit --version 2>&1 | grep -o "Scala version [0-9.]*" | awk '{print $3}' | grep -oE '[0-9]+\.[0-9]+')
    echo "Spark version ${SPARK_VERSION} is installed under ${SPARK_HOME} running with scala version ${SCALA_VERSION}"
    sudo "${SPARK_HOME}"/sbin/start-connect-server.sh --packages org.apache.spark:spark-connect_"${SCALA_VERSION}:${SPARK_VERSION}"

  3. Upload the script into the bucket created in step 1:
    aws s3 cp start-spark-connect.sh s3://$BUCKET_NAME
    

Set up the permissions

Before creating the cluster, you must create the service role, and instance profile. A service role is an IAM role that Amazon EMR assumes to provision resources and perform service-level actions with other AWS services. An EC2 instance profile for Amazon EMR assigns a role to every EC2 instance in a cluster. The instance profile must specify a role that can access the resources for your bootstrap action.

  1. Create the IAM role:
    aws iam create-role \
    --role-name AmazonEMR-ServiceRole-SparkConnectDemo \
    --assume-role-policy-document '{
    	"Version": "2012-10-17",
    	"Statement": [{
    		"Effect": "Allow",
    		"Principal": {"Service": "elasticmapreduce.amazonaws.com"},
    		"Action": "sts:AssumeRole"
    		}]
    }'
    

  2. Attach the necessary managed policies to the service role to allow Amazon EMR to manage the underlying services Amazon EC2 and Amazon S3 on your behalf and optionally grant an instance to interact with Systems Manager:
    aws iam attach-role-policy \
    --role-name AmazonEMR-ServiceRole-SparkConnectDemo \
    --policy-arn arn:aws:iam::aws:policy/service-role/AmazonEMRServicePolicy_v2
    
    aws iam attach-role-policy \
    --role-name AmazonEMR-ServiceRole-SparkConnectDemo \
    --policy-arn arn:aws:iam::aws:policy/AmazonSSMManagedInstanceCore
    
    aws iam attach-role-policy \
    --role-name AmazonEMR-ServiceRole-SparkConnectDemo \
    --policy-arn arn:aws:iam::aws:policy/service-role/AmazonElasticMapReduceRole
    

  3. Create an Amazon EMR instance role to grant permissions to EC2 instances to interact with Amazon S3 or other AWS services:
    aws iam create-role \
    --role-name EMR_EC2_SparkClusterNodesRole \
    --assume-role-policy-document '{
    "Version": "2012-10-17",
    "Statement": [{
       "Effect": "Allow",
       "Principal": {"Service": "ec2.amazonaws.com"},
       "Action": "sts:AssumeRole"
       }]
    }'
    

  4. To allow the primary instance to read from Amazon S3, attach the AmazonS3ReadOnlyAccess policy to the Amazon EMR instance role. For production environments, this access policy should be reviewed and replaced with a custom policy following the principle of least privilege, granting only the specific permissions needed for your use case:
    aws iam attach-role-policy \
    --role-name EMR_EC2_SparkClusterNodesRole \
    --policy-arn arn:aws:iam::aws:policy/AmazonS3ReadOnlyAccess
    

  5. Attaching AmazonSSMManagedInstanceCore policy enables the instances to use core Systems Manager features, such as Session Manager, and Amazon CloudWatch:
    aws iam attach-role-policy \
    --role-name EMR_EC2_SparkClusterNodesRole \
    --policy-arn arn:aws:iam::aws:policy/AmazonSSMManagedInstanceCore
    

  6. To pass the EMR_EC2_SparkClusterInstanceProfile IAM role information to the EC2 instances when they start, create the Amazon EMR EC2 instance profile:
    aws iam create-instance-profile \
    --instance-profile-name EMR_EC2_SparkClusterInstanceProfile
    

  7. Attach the role EMR_EC2_SparkClusterNodesRole created in step 3 to the newly instance profile:
    aws iam add-role-to-instance-profile \
    --instance-profile-name EMR_EC2_SparkClusterInstanceProfile \
    --role-name EMR_EC2_SparkClusterNodesRole
    

Create the Amazon EMR cluster

To create the Amazon EMR cluster, follow these steps:

  1. Set the environment variables, where your EMR cluster and load-balancer must be deployed:
    VPC_ID=<vpc-emr-and-alb>
    EMR_PRI_SB_ID_1=<emr-private-subnet-id-az1>
    ALB_PUB_SB_ID_1=<alb-public-subnet-id-az1>
    ALB_PUB_SB_ID_2=<alb-public-subnet-id-az2>
    

  2. Create the EMR cluster with the latest Amazon EMR release. Replace the placeholder value with your actual S3 bucket name where the bootstrap action script is stored:
    CLUSTER_ID=$(aws emr create-cluster \
    --name "Spark Connect cluster demo" \
    --applications Name=Spark \
    --release-label emr-7.9.0 \
    --service-role AmazonEMR-ServiceRole-SparkConnectDemo \
    --ec2-attributes InstanceProfile=EMR_EC2_SparkClusterInstanceProfile,SubnetId=$EMR_PRI_SB_ID_1 \
    --instance-groups InstanceGroupType=MASTER,InstanceCount=1,InstanceType=m5.xlarge InstanceGroupType=CORE,InstanceCount=1,InstanceType=m5.xlarge \
    --bootstrap-actions Path="s3://$BUCKET_NAME/start-spark-connect.sh" \
    --query 'ClusterId' --output text)
    echo CLUSTER_ID="$CLUSTER_ID"
    

    To modify primary node’s security group to allow Systems Manager to start a session.

  3. Get the primary node’s security group identifier. Record the identifier because you’ll need it for subsequent configuration steps in which primary-node-security-group-id is mentioned:
    PRIMARY_NODE_SG=$(aws emr describe-cluster \
    --cluster-id $CLUSTER_ID \
    --query 'Cluster.Ec2InstanceAttributes.EmrManagedMasterSecurityGroup' \
    --output text)
    echo PRIMARY_NODE_SG=$PRIMARY_NODE_SG
    

  4. Find the EC2 instance connect prefix list ID for your Region. You can use the EC2_INSTANCE_CONNECT filter with the describe-managed-prefix-lists command. Using a managed prefix list provides a dynamic security configuration to authorize Systems Manager EC2 instances to connect the primary and core nodes by SSH:
    IC_PREFIX_LIST=$(aws ec2 describe-managed-prefix-lists \
    --filters Name=prefix-list-name,Values=com.amazonaws.$REGION.ec2-instance-connect \
    --query 'PrefixLists[0].PrefixListId' \
    --output text)
    echo IC_PREFIX_LIST=$IC_PREFIX_LIST
    

  5. Modify the primary node security group inbound rules to allow SSH access (port 22) to the EMR cluster’s primary node from resources that are part of the specified Instance Connect service contained in the prefix list:
    aws ec2 authorize-security-group-ingress \
    --region $REGION \
    --group-id $PRIMARY_NODE_SG \
    --ip-permissions "[{\"IpProtocol\":\"tcp\",\"FromPort\":22,\"ToPort\":22,\"PrefixListIds\":[{\"PrefixListId\":\"$IC_PREFIX_LIST\"}]}]"
    

Optionally, you can repeat the preceding steps 1–3 for the core (and tasks) cluster’s nodes to allow Amazon EC2 Instance Connect to access the EC2 instance through SSH.

Deploy the Application Load Balancer and certificate

To deploy the Application Load Balancer and certificate, follow these steps:

  1. Create a load balancer’s security group:
    ALB_SG_ID=$(aws ec2 create-security-group \
    --group-name spark-connect-alb-sg \
    --description "Security group for Spark Connect ALB" \
    --region $REGION \
    --vpc-id $VPC_ID \
    --query 'GroupId' \
    --output text)
    

  2. Add rule to accept TCP traffic from a trusted IP on port 443. We recommend that you use the local development machine’s IP address. You can check your current public IP address here: https://checkip.amazonaws.com:
    aws ec2 authorize-security-group-ingress \
    --group-id $ALB_SG_ID \
    --protocol tcp \
    --port 443 \
    --cidr <replace-with-trusted-IP>/32
    

  3. Create a new target group with gRPC protocol, which targets the Spark Connect server instance and the port the server is listening to:
    ALB_TG_ARN=$(aws elbv2 create-target-group \
    --name spark-connect-tg \
    --protocol HTTP \
    --protocol-version GRPC \
    --port 15002 \
    --target-type instance \
    --health-check-enabled \
    --health-check-protocol HTTP \
    --health-check-path / \
    --vpc-id $VPC_ID \
    --query 'TargetGroups[0].TargetGroupArn' \
    --output text)
    echo "ALB TG created (ARN)=$ALB_TG_ARN"
    

  4. Create the Application Load Balancer:
    ALB_ARN=$(aws elbv2 create-load-balancer \
    --name spark-connect-alb \
    --type application \
    --scheme internet-facing \
    --subnets $ALB_PUB_SB_ID_1 $ALB_PUB_SB_ID_2 \
    --security-groups $ALB_SG_ID \
    --query 'LoadBalancers[0].LoadBalancerArn' \
    --output text)
    echo "ALB created (ARN)=$ALB_ARN"
    

  5. Get the load balancer DNS name:
    ALB_DNS=$(aws elbv2 describe-load-balancers \
    --load-balancer-arns $ALB_ARN \
    --query 'LoadBalancers[0].DNSName' \
    --output text)
    echo "ALB DNS=$ALB_DNS"
    

  6. Retrieve the Amazon EMR primary node ID:
    PRIMARY_NODE_ID=$(aws emr list-instances --cluster-id $CLUSTER_ID --instance-group-types MASTER --query 'Instances[0].Ec2InstanceId' --output text)
    echo PRIMARY_NODE_ID=$PRIMARY_NODE_ID
    

  7. (Optional) To encrypt and decrypt the traffic, the load balancer needs a certificate. You can skip this step if you already have a trusted certificate in ACM. Otherwise, create a self-signed certificate:
    PRIVATE_KEY_PATH=./sc-private-key.key
    CERTIFICATE_PATH=./sc-certificate.cert
    sudo openssl req -x509 -nodes -days 365 -newkey rsa:2048 -keyout $PRIVATE_KEY_PATH -out $CERTIFICATE_PATH -subj "/CN=$ALB_DNS"
    

  8. Upload to ACM:
    ACM_CERT_ARN=$(aws acm import-certificate \
    --certificate fileb://$CERTIFICATE_PATH \
    --private-key fileb://$PRIVATE_KEY_PATH \
    --region $REGION \
    --query CertificateArn \
    --output text)
    echo "Certificate created (ARN)=$ACM_CERT_ARN"
    

  9. Create the load balancer listener:
    ALB_LISTENER_ARN=$(aws elbv2 create-listener \
    --load-balancer-arn $ALB_ARN \
    --protocol HTTPS \
    --port 443 \
    --certificates CertificateArn=$ACM_CERT_ARN \
    --ssl-policy ELBSecurityPolicy-TLS13-1-2-2021-06 \
    --default-actions Type=forward,TargetGroupArn=$ALB_TG_ARN \
    --region $REGION \
    --query 'Listeners[0].ListenerArn' \
    --output text)
    echo "ALB listener created (ARN)=$ALB_LISTENER_ARN"
    

  10. After the listener has been provisioned, register the primary node to the target group:
    aws elbv2 register-targets \
    --target-group-arn $ALB_TG_ARN \
    --targets Id=$PRIMARY_NODE_ID
    

Modify the primary node’s security group to allow Spark Connect clients to connect

To connect to Spark Connect, amend only the primary security group. Add an inbound rule to the primary’s node security group to accept Spark Connect TCP connection on port 15002 from your chosen trusted IP address:

aws ec2 authorize-security-group-ingress \
--group-id $PRIMARY_NODE_SG \
--protocol tcp \
--port 15002 \
--source-group $ALB_SG_ID

Connect with a test application

This example demonstrates that a client running a newer Spark version (4.0.1) can successfully connect to an older Spark version on the Amazon EMR cluster (3.5.5), showcasing Spark Connect’s version compatibility feature. This version combination is for demonstration only. Running older versions might pose security risks in production environments.

To test the client-to-server connection, we provide the following test Python application. We recommend that you create and activate a Python virtual environment (venv) before installing the packages. This helps isolate the dependencies for this specific project and prevents conflicts with other Python projects. To install packages, run the following command:

pip install pyspark-client==4.0.1

In your integrated development environment (IDE), copy and paste the following code, replace the placeholder, and invoke it. The code creates a Spark DataFrame containing two rows and it shows its data:

from pyspark.sql import SparkSession
import os
os.environ['GRPC_DEFAULT_SSL_ROOTS_FILE_PATH'] = os.path.expanduser('sc-certificate.cert')
spark = SparkSession.builder \
    .remote("sc://:443/;use_ssl=true") \
    .config('spark.sql.execution.pandas.inferPandasDictAsMap', True) \
    .config('spark.sql.pyspark.legacy.inferMapTypeFromFirstPair.enabled', True) \
    .getOrCreate()
spark.createDataFrame([("sue", 32),("li", 3)],["first_name", "age"]).show()

The following shows the application output:

+----------+---+
|first_name|age|
+----------+---+
|       sue| 32|
|        li|  3|
+----------+---+

Clean up

When you no longer need the cluster, release the following resources to stop incurring charges:

  1. Delete the Application Load Balancer listener, target group, and the load balancer.
  2. Delete the ACM certificate.
  3. Delete the load balancer and Amazon EMR node security groups.
  4. Terminate the EMR cluster.
  5. Empty the Amazon S3 bucket and delete it.
  6. Remove AmazonEMR-ServiceRole-SparkConnectDemo and EMR_EC2_SparkClusterNodesRole roles and EMR_EC2_SparkClusterInstanceProfile instance profile.

Considerations

Security considerations with Spark Connect:

  • Private subnet deployment – Keep EMR clusters in private subnets with no direct internet access, using NAT gateways for outbound connectivity only.
  • Access logging and monitoring – Enable VPC Flow Logs, AWS CloudTrail, and bastion host access logs for audit trails and security monitoring.
  • Security group restrictions – Configure security groups to allow Spark Connect port (15002) access only from bastion host or specific IP ranges.

Conclusion

In this post, we showed how you can adopt modern development workflows and debug Spark applications from local IDEs or notebooks, so you can step through code execution. With Spark Connect’s client-server architecture, the Spark cluster can run on a different version than the client applications, so operations teams can perform infrastructure upgrades and patches independently.

As the cluster operators gain experience, they can customize the bootstrap actions and add steps to process data. Consider exploring Amazon Managed Workflows for Apache Airflow (MWAA) for orchestrating your data pipeline.


About the authors

Philippe Wanner

Philippe Wanner

Philippe is EMEA Tech Lead at AWS. His role is to accelerate the digital transformation for large organizations. His current focus is in a multidisciplinary area involving business transformation, technical strategy, and distributed systems.

Ege Oguzman

Ege Oguzman

Ege is a Software Development Engineer at AWS, and previously he was a Solutions Architect in the public sector. As a builder and cloud enthusiast, he specializes in distributed systems and dedicates his time to infrastructure development and helping organizations build solutions on AWS.

How Taxbit achieved cost savings and faster processing times using Amazon S3 Tables

Post Syndicated from Larry Christensen original https://aws.amazon.com/blogs/big-data/how-taxbit-achieved-cost-savings-and-faster-processing-times-using-amazon-s3-tables/

In this post, we discuss how Taxbit partnered with Amazon Web Services (AWS) to streamline their crypto tax analytics solution using Amazon S3 Tables, achieving 82% cost savings and five times faster processing times.

Taxbit is a leading tax compliance suite serving cryptocurrency exchanges, digital platforms, and government agencies, generating more than 100 million forms for users and reconciling more than 500 billion digital asset transactions. The suite powers a complex environment that handles real-time pricing data from 29 cryptocurrency exchanges covering over 10,000 digital assets.

Recently, Taxbit experienced challenges with their pricing data infrastructure. As data volumes continued to expand, infrastructure costs rose sharply, putting pressure on operational budgets. At the same time, the system struggled to efficiently ingest the growing number of pricing data points, creating persistent bottlenecks in their data pipeline. These technical limitations led to customers missing data and experiencing slow processing times, leading to dissatisfaction. In addition to these operational challenges, Taxbit has strict regulatory compliance requirements to be considered when designing solutions. This combination of issues led Taxbit to modernize their pricing data infrastructure with a focus on helping to meet regulatory standards.

“During peak workloads, our solutions process hundreds of millions of digital asset transactions across blockchain and cryptocurrency exchanges,”

– says Clark Roberts, CTO at Taxbit.

“Our legacy database architecture was becoming a bottleneck, leading to increased costs and slower response times for our enterprise and government customers.”

Solution overview

Taxbit’s modernized architecture uses Amazon S3 Tables with Apache Iceberg as the foundation, combined with purpose-built AWS services for data ingestion, processing, and analytics. The solution processes real-time pricing data from 29 cryptocurrency exchanges including over 10,000 digital assets. This architecture is shown in the following diagram.

This AWS cloud architecture diagram illustrates a comprehensive data pipeline for processing digital assest market data.

The data pipeline architecture uses AWS services to deliver a comprehensive solution. At its foundation, Amazon S3 Tables provides the scalable storage infrastructure necessary for managing large volumes of pricing data. For data processing and transformation, the solution combines Amazon EMR and AWS Glue, handling both extract, transform, and load (ETL) operations and asynchronous API requirements efficiently.

Real-time data handling is managed through Amazon Kinesis, enabling streaming of pricing updates. AWS Lambda functions perform multiple tasks, including periodic polling of vendor APIs, transformation of streaming data, and data enrichment. The orchestration of these components is managed by AWS Step Functions, helping to ensure coordination of data workflows. Completing the architecture, Amazon Athena provides query capabilities, supporting both synchronous APIs and one-time analytical queries. This approach creates a scalable system built to handle both real-time and batch processing workflows while maintaining high performance and reliability.

Data ingestion layer

The ingestion layer operates through two key components: API integration and stream processing. The API integration uses Lambda functions to systematically poll multiple external APIs. These polling operations are orchestrated by Amazon EventBridge, which manages the scheduled data collection tasks. Additionally, WebSocket listeners maintain continuous connections to capture real-time price updates as they occur.

On the stream processing side, Amazon Kinesis Data Streams serves as the backbone for handling real-time data ingestion at scale. As data flows in, Lambda functions perform transformations and enrichment operations to prepare the data for downstream use. Throughout this process, custom validation checks are applied to help ensure the quality and completeness of the data, helping to maintain the integrity of the pricing information pipeline.

Data storage layer

At the storage layer, Taxbit uses Amazon S3 Tables because of its optimized storage format designed for analytical queries. Amazon S3 Tables is designed to automatically handle table optimization and compaction, helping to streamline data management processes. The system also incorporates time-travel capabilities, allowing Taxbit to meet audit requirements and their need for historical data analysis.

The data organization strategy is designed to maximize efficiency and accessibility. Data is systematically partitioned by date and exchange, allowing for targeted data retrieval and improved query performance. The implementation of columnar storage further enhances query efficiency by minimizing unnecessary data scans. Additionally, version control mechanisms are in place to maintain clear data lineage, enabling precise tracking of data changes and transformations over time.

Analytics layer

At the analytics layer, the query engine forms the foundation, using Amazon Athena to facilitate flexible ad-hoc analysis of the pricing data. This is complemented by Presto-based queries that handle complex aggregations efficiently. The system includes carefully crafted execution plans optimized for common query patterns, designed to provide consistent and reliable performance.

To maximize efficiency, the analytics layer incorporates several key performance optimizations. The system uses an Athena reuse query result to minimize redundant processing and parallel query execution capabilities to handle multiple simultaneous requests effectively.

Security and compliance

The data protection strategy implements multiple layers of security, starting with AWS Key Management Service (AWS KMS) encryption for all data at rest. This is complemented by TLS encryption for data in transit, helping to secure data movement throughout the system. Access to data and resources is controlled through AWS Identity and Access Management (IAM), providing fine-grained permissions that enforce the principle of least privilege.

The audit trail component provides comprehensive monitoring and compliance capabilities. AWS CloudTrail logging captures detailed records of system activities, enabling thorough security analysis and incident investigation. Data lineage tracking maintains clear records of data movement and transformations throughout the pipeline. These features are augmented by robust compliance reporting capabilities, helping the system demonstrate adherence to regulatory requirements and internal governance policies. Together, these security controls create an environment that protects sensitive data, maintains transparency, and provides accountability.

Business impact

Most notably, Taxbit achieved an 82% reduction in storage infrastructure costs, while simultaneously delivering processing speeds five times faster than their previous architecture. Data completeness for calculations achieved approximately 99.99% accuracy and the workload can now successfully support over 10,000 digital assets.The benefits extended beyond these quantitative improvements. Customer experience has improved, with transaction pricing times shrinking from hours to minutes. Higher throughput capabilities increased operational efficiency, enabling faster data loading while reducing compute costs. The new architecture also established a scalable foundation that provides faster data access and the flexibility to expand into new markets. The modern infrastructure has also enabled Taxbit to pursue new product offerings by supporting advanced analytics and real-time insights that were previously unattainable. These capabilities created new business opportunities and revenue streams that weren’t possible under the constraints of the legacy system.

Conclusion

Taxbit’s implementation of Amazon S3 Tables has transformed their cryptocurrency tax compliance solutions, delivering 82% cost savings and five times faster processing speeds. The modernized architecture, combining Amazon EMR, AWS Glue, Amazon Kinesis, and Lambda, now processes transactions in minutes instead of hours. Additionally, the architecture has helped Taxbit maintain approximately 99.99% data accuracy across more than 10,000 digital assets. Beyond operational improvements, this transformation has enabled new product offerings and real-time analytics capabilities. By partnering with AWS, Taxbit addressed their scaling challenges and built a foundation for continued innovation in the digital asset space.

For more information, see Amazon S3 Tables.


About the authors

Larry Christensen

Larry Christensen

Larry is a Principal Engineer at Taxbit based in the Salt Lake City area. He’s spearheaded many architectural, big data, and AI transformations across Taxbit.

Washim Nawaz

Washim Nawaz

Washim is an Analytics Specialist Solutions Architect at AWS with extensive professional experience building and tuning data warehouse and data lake solutions. He is passionate about helping customers modernize their data platforms with efficient, performant, and scalable analytics solutions. Outside of work, he enjoys watching sports and traveling.

Derek Ziehl

Derek Ziehl

Derek is a Senior Technical Account Manager (TAM) at AWS. He has a background designing large-scale network systems and managing cloud migrations. As a TAM he enjoys enabling customers to run resilient, optimized workloads on AWS.

Pranjal Gururani

Pranjal Gururani

Pranjal is a Solutions Architect at AWS based out of Seattle. Pranjal works with various customers to architect cloud solutions that address their business challenges. He enjoys hiking, kayaking, skydiving, and spending time with family during his spare time.

Create and update Apache Iceberg tables with partitions in the AWS Glue Data Catalog using the AWS SDK and AWS CloudFormation

Post Syndicated from Aarthi Srinivasan original https://aws.amazon.com/blogs/big-data/create-and-update-apache-iceberg-tables-with-partitions-in-the-aws-glue-data-catalog-using-the-aws-sdk-and-aws-cloudformation/

In recent years, we’ve witnessed a significant shift in how enterprises manage and analyze their ever-growing data lakes. At the forefront of this transformation is Apache Iceberg, an open table format that’s rapidly gaining traction among large-scale data consumers.

However, as enterprises scale their data lake implementations, managing these Iceberg tables at scale becomes challenging. Data teams often need to manage table schema evolution, its partitioning, and snapshots versions. Automation streamlines these operations, provides consistency, reduces human error, and helps data teams focus on higher-value tasks.

The AWS Glue Data Catalog now supports Iceberg table management using the AWS Glue API, AWS SDKs, and AWS CloudFormation. Previously, users had to create Iceberg tables in the Data Catalog without partitions using CloudFormation or SDKs and later add partitions from Amazon Athena or other analytics engines. This prevents the table lineage from being tracked in one place and adds steps outside automation in the continuous integration and delivery (CI/CD) pipeline for table maintenance operations. With the launch, AWS Glue customers can now use their preferred automation or infrastructure as code (IaC) tools to automate Iceberg table creation with partitions and use the same tools to manage schema updates and sort order.

In this post, we show how to create and update Iceberg tables with partitions in the Data Catalog using the AWS SDK and CloudFormation.

Solution overview

In the following sections, we illustrate the AWS SDK for Python (Boto3) and AWS Command Line Interface (AWS CLI) usage of Data Catalog APIs—CreateTable() and UpdateTable()—for Amazon Simple Storage Service (Amazon S3) based Iceberg tables with partitions. We also provide the CloudFormation templates to create and update an Iceberg table with partitions.

Prerequisites

The Data Catalog API changes are made available in the following versions of the AWS CLI and SDK for Python:

  • AWS CLI version of 2.27.58 or above
  • SDK for Python version of 1.39.12 or above

AWS CLI usage

Let’s create an Iceberg table with one partition, using CreateTable() in the AWS CLI:

aws glue create-table --cli-input-json file://createicebergtable.json

The createicebergtable.json is as follows:

{
    "CatalogId": "123456789012",
    "DatabaseName": "bankdata_icebergdb",
    "Name": "transactiontable1",
    "OpenTableFormatInput": { 
      "IcebergInput": { 
         "MetadataOperation": "CREATE",
         "Version": "2",
         "CreateIcebergTableInput": { 
            "Location": "s3://sampledatabucket/bankdataiceberg/transactiontable1/",
            "Schema": {
                "SchemaId": 0,
                "Type": "struct",
                "Fields": [ 
                    { 
                        "Id": 1,
                        "Name": "transaction_id",
                        "Required": true,
                        "Type": "string"
                    },
                    { 
                        "Id": 2,
                        "Name": "transaction_date",
                        "Required": true,
                        "Type": "date"
                    },
                    { 
                        "Id": 3,
                        "Name": "monthly_balance",
                        "Required": true,
                        "Type": "float"
                    }
                ]
            },
            "PartitionSpec": { 
                "Fields": [ 
                    { 
                        "Name": "by_year",
                        "SourceId": 2,
                        "Transform": "year"
                    }
                ],
                "SpecId": 0
            },
            "WriteOrder": { 
                "Fields": [ 
                    { 
                        "Direction": "asc",
                        "NullOrder": "nulls-last",
                        "SourceId": 1,
                        "Transform": "none"
                    }
                ],
                "OrderId": 1
            }  
        }
      }
   }
}

The preceding AWS CLI command creates the metadata folder for the Iceberg table in Amazon S3, as shown in the following screenshot.

Amazon S3 bucket interface showing metadata folder containing single JSON file dated November 6, 2025

You can populate the table with values as follows and verify the table schema using the Athena console:

SELECT * FROM "bankdata_icebergdb"."transactiontable1" limit 10;
insert into bankdata_icebergdb.transactiontable1 values
    ('AFTERCREATE1234', DATE '2024-08-23', 6789.99),
    ('AFTERCREATE5678', DATE '2023-10-23', 1234.99);
SELECT * FROM "bankdata_icebergdb"."transactiontable1";

The following screenshot shows the results.

Amazon Athena query editor showing SQL queries and results for bankdata_icebergdb database with transaction data

After populating the table with data, you can inspect the S3 prefix of the table, which will now have the data folder.

Amazon S3 bucket interface displaying data folder with two subfolders organized by year: 2023 and 2024

The data folders partitioned according to our table definition and Parquet data files created from our INSERT command are available under each partitioned prefix.

Amazon S3 bucket interface showing by_year=2023 folder containing single Parquet file of 575 bytes

Next, we update the Iceberg table by adding a new partition, using UpdateTable():

aws glue update-table --cli-input-json file://updateicebergtable.json

The updateicebergtable.json is as follows.

{
  "CatalogId": "123456789012",
  "DatabaseName": "bankdata_icebergdb",
  "Name": "transactiontable1",
  "UpdateOpenTableFormatInput": {
    "UpdateIcebergInput": {
      "UpdateIcebergTableInput": {
        "Updates": [
          {
            "Location": "s3://sampledatabucket/bankdataiceberg/transactiontable1/",
            "Schema": {
              "SchemaId": 1,
              "Type": "struct",
              "Fields": [
                {
                  "Id": 1,
                  "Name": "transaction_id",
                  "Required": true,
                  "Type": "string"
                },
                {
                  "Id": 2,
                  "Name": "transaction_date",
                  "Required": true,
                  "Type": "date"
                },
                {
                  "Id": 3,
                  "Name": "monthly_balance",
                  "Required": true,
                  "Type": "float"
                }
              ]
            },
            "PartitionSpec": {
              "Fields": [
                {
                  "Name": "by_year",
                  "SourceId": 2,
                  "Transform": "year"
                },
                {
                  "Name": "by_transactionid",
                  "SourceId": 1,
                  "Transform": "identity"
                }
              ],
              "SpecId": 1
            },
            "SortOrder": {
              "Fields": [
                {
                  "Direction": "asc",
                  "NullOrder": "nulls-last",
                  "SourceId": 1,
                  "Transform": "none"
                }
              ],
              "OrderId": 2
            }
          }
        ]
      }
    }
  }
}

UpdateTable() modifies the table schema by adding a metadata JSON file to the underlying metadata folder of the table in Amazon S3.

Amazon S3 bucket interface showing 5 metadata objects including JSON and Avro files with timestamps

We insert values into the table using Athena as follows:

insert into bankdata_icebergdb.transactiontable1 values
    ('AFTERUPDATE1234', DATE '2025-08-23', 4536.00),
    ('AFTERUPDATE5678', DATE '2022-10-23', 23489.00);
SELECT * FROM "bankdata_icebergdb"."transactiontable1";

The following screenshot shows the results.

Amazon Athena query editor with SQL statements and results after iceberg partition update and insert data

Inspect the corresponding changes to the data folder in the Amazon S3 location of the table.

Amazon S3 prefix showing new partitions for the Iceberg table

This example has illustrated how to create and update Iceberg tables with partitions using AWS CLI commands.

SDK for Python usage

The following Python scripts illustrate using CreateTable() and UpdateTable() for an Iceberg table with partitions:

CloudFormation usage

Use the following CloudFormation templates for CreateTable() and UpdateTable(). After the CreateTable template is complete, update the same stack with the UpdateTable template by creating a new changeset for your stack and executing it.

Clean up

To avoid incurring costs on the Iceberg tables created using the AWS CLI, delete the tables from the Data Catalog.

Conclusion

In this post, we illustrated how to use the AWS CLI to create and update Iceberg tables with partitions in the Data Catalog. We also provided the SDK for Python and CloudFormation sample code and templates. We hope this helps you automate the creation and management of your Iceberg tables with partitions in your CI/CD pipelines and production environments. Try it out for your own use case and share your feedback in the comments section.


About the authors

Acknowledgements: A special thanks to everyone who contributed to the development and launch of this feature – Purvaja Narayanaswamy, Sachet Saurabh, Akhil Yendluri and Mohit Chandak.

Aarthi Srinivasan

Aarthi Srinivasan

Aarthi is a Senior Big Data Architect with AWS. She works with AWS customers and partners to architect data lake house solutions, enhance product features, and establish best practices for data governance.

Pratik Das

Pratik Das

Pratik is a Senior Product Manager with AWS. He is passionate about all things data and works with customers to understand their requirements and build delightful experiences. He has a background in building data-driven solutions and machine learning systems in production.

Touring the Center of the Internet and an AI Data Center at Equinix Silicon Valley

Post Syndicated from Patrick Kennedy original https://www.servethehome.com/touring-the-center-of-the-internet-and-an-ai-data-center-at-equinix-silicon-valley-sv1-sv11-nvidia-dgx-superpod/

We tour the “Center of the Internet,” Equinix SV1, and a new modern AI data center SV11, where an important NVIDIA DGX SuperPOD is based

The post Touring the Center of the Internet and an AI Data Center at Equinix Silicon Valley appeared first on ServeTheHome.

CVE-2025-37164: Critical unauthenticated RCE affecting Hewlett Packard Enterprise OneView

Post Syndicated from Rapid7 original https://www.rapid7.com/blog/post/etr-cve-2025-37164-critical-unauthenticated-rce-affecting-hewlett-packard-enterprise-oneview

Overview

On December 17, 2025, Hewlett Packard Enterprise (HPE) published an advisory for CVE-2025-37164, a CVSS 10.0 vulnerability in HPE OneView. The vulnerability, which was reported to HPE by security researcher Nguyen Quoc Khanh, facilitates unauthenticated remote code execution (RCE) on versions of HPE OneView before 11.0. Defenders are advised to prioritize upgrading to version 11.0 or applying the emergency hotfixes (HPE OneView virtual appliance hotfix, HPE Synergy hotfix) as soon as possible.

OneView sits at a privileged control plane for enterprise infrastructure, so successful exploitation isn’t just about establishing remote code execution, it’s about gaining centralized control over servers, firmware, and lifecycle management at scale. The real concern here is exposure and trust assumptions. Management platforms are often deployed deep inside the network with broad privileges and minimal monitoring because they’re ‘supposed’ to be trusted. When an unauthenticated RCE shows up in that layer, defenders need to treat it as an assumed-breach scenario, prioritize patching immediately, and review access paths and segmentation.

Hotfix analysis

Rapid7 Labs has begun an initial analysis of the vendor-supplied hotfix HPE_OneView_CVE_37164_Z7550-98077.bin. This hotfix applies a new HTTP rule to the appliance’s webserver to block access to a specific REST API endpoint. This endpoint is /rest/id-pools/executeCommand. Initial inspection of the appliance code indicates this endpoint is reachable without authentication. Rapid7 Labs assesses with a high degree of confidence that this is the access vector for triggering the vulnerability and achieving remote code execution.

Mitigation guidance

According to HPE, CVE-2025-37164 affects HPE OneView versions below 11.0, version 5.20 through version 10.20, unless a security hotfix (HPE OneView virtual appliance hotfix, HPE Synergy hotfix) has been applied.

For the latest mitigation guidance for HPE OneView, please refer to the vendor’s security advisory.

Rapid7 customers

Exposure Command, InsightVM, and Nexpose

Exposure Command, InsightVM, and Nexpose customers can assess exposure to CVE-2025-37164 with an unauthenticated vulnerability check expected to be available in today’s (December 18) content release.

Updates

  • December 18, 2025: Initial publication.

Bookblaze 2025: Backblaze Employee Recommended Reads

Post Syndicated from Stephanie Doyle original https://www.backblaze.com/blog/bookblaze-2025-backblaze-employee-recommended-reads/

A decorative image showing several books on a holiday background.

Sure, we may be a global tech company who spends our days on the front lines of helping our customers solve their toughest data storage challenges, but that doesn’t mean we don’t ever power down the devices and curl up with a good book. Welcome to the third annual Bookblaze, Backblaze’s much-anticipated book guide where our team shares the stories, insights, and adventures that shaped their reading year. 

From thought-provoking nonfiction to immersive fiction and unexpected gems, these recommendations are curated by the people who read, think, and create here at Backblaze—offering you a cozy companion for winter nights, inspiration for your 2026 reading list, and maybe even the perfect gift idea along the way. Whether you’re reconnecting with old favorites or discovering your next great read, we hope this year’s picks spark joy, curiosity, and conversation.

Chris McGranahan, Director, Information Security Architecture

An image of the cover of The Story of CO2 Is the Story of Everything by Peter Brannen.

The Story of CO2 Is the Story of Everything, by Peter Brannen

It’s an exhaustive but entertaining explanation of how our world came to be the way it is, why CO2 is so important to it and how the path we’re currently on is likely to create a world that hasn’t existed in millions of years and never supported humans. And, if you want a fun fiction read, check out any of the Murderbot Diaries series by Martha Wells.

Maddie Presland, Product Marketing Manager

An image of the cover of the book Clytemnestra, by Costanza Casati.

Clytemnestra, by Costanza Casati

I love deeply flawed female protagonists. I also love the fact that myth retellings have been so popular for the last few years, and this year, I decided to tackle my hyper-specific TBR I was neglecting. (Editor’s note: For those of you not afflicted with chronic book collecting, TBR = to be read.) 

Clytemnestra tells the story of one of the most reviled women in Greek mythology. She’s cunning, ruthless, and quite possibly the original champion of playing the long game to seek revenge, as a key player in the Trojan War you’ve probably never heard of. And yet, you can’t help but root for her. It’s shocking that this is a debut novel because, though it can be a slow burn in parts, the characterization and completely immersive writing provides a different perspective of how the Trojan War unfolded for those left at home.

Bala Krishna Gangisetty, Sr. Product Manager

An image of the cover of the book Positive Intelligence, by Shirzad Chamine.

Positive Intelligence, by Shirzad Chamine

I appreciate how Positive Intelligence translates mindset and emotional intelligence into practical exercises for building mental fitness. The Saboteur framework makes it easy to spot negative thinking and shift toward a more productive mindset. It’s a great balance of psychology, neuroscience, and real-world application that supports personal and professional growth.

AJ Sedlak, Director, GTM and Marketing Operations

An image of the cover the book Essentialism: The Disciplined Pursuit of Less, by Greg McKeown.

Essentialism: The Disciplined Pursuit of Less, by Greg McKeown

In work and in life, we often struggle with saying “no”, even to ourselves. We take on more than we can effectively manage and execute. As a result, we’re burned out—frustrated by our ever-growing to-do lists and disappointed in the quality of what we do get done. This book helped me understand why doing less actually results in accomplishing more and better things. In addition, it gave me a sense of how to make this case to others—whether it’s about prioritization of work projects or helping a loved one who’s feeling overwhelmed.

Amy Kunde, Sr. Executive Assistant

An image of the cover of the book The Extraordinary Life of Sam Hell, by Robert Dugoni.

The Extraordinary Life of Sam Hell, by Robert Dugoni

Although there were some dark moments in the book, in general it was a feel good book of a boy coming of age through adulthood, paying his dues and then paying it forward.  The setting is right in the Backblaze neighborhood so it’s always interesting to picture the local intersections, and schools referenced in the novel.

Elisa Miller, Sr. Organizational Development Partner

An image of the cover of the book Parable of the Sower, by Octavia Butler.

Parable of the Sower, by Octavia Butler

Octavia Butler is a masterful sci fi writer who has woven a tale in the 1990s set to present time about a dystopian reality oddly similar to the one we are living/heading towards currently of lawlessness, greed, and the quest for survival. Focusing on the power of community, togetherness, and nature, this book was an epic (and scary) adventure into what happens when people gather together to fight the status quo while lifting one another up. It’s not for the faint of heart, but it really was a game changer for me to read, offering solutions beyond capitalism and towards empowerment of humanity, spirituality, and purpose. 

Molly Clancy, Sr. Manager, Content & Creative

An image of the cover of the book Project Hail Mary, by Andy Weir.

Project Hail Mary, by Andy Weir

It’s a buddy comedy, but also the fate of the human race is at stake. There’s lots of science to nerd out on if that’s your thing. And Andy Weir is a former software engineer, so you could kinda sorta say it’s “for work.”

Beth Grey, Sr. Risk & Regulatory Compliance Specialist

An image of the cover of the book The Happy Sleeper: The Science-Backed Guide to Helping Your Baby Get a Good Night's Sleep-Newborn to School Age, by Heather Turgeon, MFT, and Julie Wright, MFT.

The Happy Sleeper: The Science-Backed Guide to Helping Your Baby Get a Good Night’s Sleep-Newborn to School Age, by Heather Turgeon, MFT, and Julie Wright, MFT

Sleep training my child without having to resort to too much crying seemed daunting, but this book helped inform our process with evidence based guidance. I am happy to report that I have a great little sleeper because of it. This book will help any parent gain the skills and confidence to effectively sleep train their child.

Yev Pusin, Head of Communications and Community

An image of the cover of the book Dungeon Crawler Carl, by Matt Dinniman.

Dungeon Crawler Carl, by Matt Dinniman

This book series is absolute insanity, and if you are able to get the audiobook, you will not regret it. I even got my sister’s mother-in-law to listen to it on audio and she loved it. The premise is that it’s the end of the world, and Carl is sucked into a dungeon to fight for the entertainment of the universe at large—plus there’s a talking cat! What’s not to like? It’s a genre known as LitRPG (editor’s note: Literary role-playing game) which follows Carl and his friends’ progression as they work through the dungeon and try to topple the powers that be.

Nicole Gale, Sr. Marketing Operations Manager

An image of the cover of the book The Ballad of Songbirds and Snakes, by Suzanne Collins.

The Ballad of Songbirds and Snakes, by Suzanne Collins

The new “Hunger Games” book is the first prequel I’ve read in years that genuinely adds something meaningful to its original series. It pulled me right back into Panem, had me rewatching all the movies, and had me loving characters that I didn’t expect to get attached to. A fantastic return to a world I thought I already knew.

Stephanie Doyle, Writer and Content Operations Strategist

An image of the cover of the book Children of Time, by Adrian Tchaikovsky.

Children of Time, by Adrian Tchaikovsky

This series reminds me of old-school science fiction in all the best ways. Without giving too much away, a terraforming project is sabotaged, leading to unexpected outcomes for the targeted planet. Meanwhile, back on Earth, the world ends, and a race begins for the remainder of humanity to find a new home. Tchaikovsky’s brilliance thrives in the details of understanding systems, people, biology, engineering, and science, and each new revelation about what’s happening—in this new world with a new sentient species, and with the humans on their ever-devolving arc ship—stems from each of those details showing up in ways that feel both expected and unexpected at the same time.

The post Bookblaze 2025: Backblaze Employee Recommended Reads appeared first on Backblaze Blog | Cloud Storage & Cloud Backup

Someone Boarded a Plane at Heathrow Without a Ticket or Passport

Post Syndicated from Bruce Schneier original https://www.schneier.com/blog/archives/2025/12/someone-boarded-a-plane-at-heathrow-without-a-ticket-or-passport.html

I’m sure there’s a story here:

Sources say the man had tailgated his way through to security screening and passed security, meaning he was not detected carrying any banned items.

The man deceived the BA check-in agent by posing as a family member who had their passports and boarding passes inspected in the usual way.

A change of maintainership for linux-next

Post Syndicated from corbet original https://lwn.net/Articles/1051179/

Stephen Rothwell, who has maintained the kernel’s linux-next integration
tree from its inception, has announced his
retirement from that role:

I will be stepping down as Linux-Next maintainer on Jan 16, 2026.
Mark Brown has generously volunteered to take up the challenge. He
has helped in the past filling in when I have been unavailable, so
hopefully knows what he is getting in to. I hope you will all
treat him with the same (or better) level of respect that I have
received.

It has been a long but mostly interesting task and I hope it has
been helpful to others. It seems a long time since I read Andrew
Morton’s “I have a dream” email and decided that I could help out
there – little did I know what I was heading for.

Over the last two decades or so, the kernel’s development process has evolved
from an unorganized mess with irregular releases to a smooth machine with a
new release every nine or ten weeks. That would not have happened without
linux-next; thanks are due to Stephen for helping to make the current
process possible.

[$] Episode 29 of the Dirk and Linus show

Post Syndicated from corbet original https://lwn.net/Articles/1050317/

Linus Torvalds is famously averse to presenting prepared talks, but the
wider community is always interested in what he has to say about the
condition of the Linux kernel. So, for some time now, his appearances have
been in the form of an informal conversation with Dirk Hohndel. At the
2025 Open Source Summit Japan, the pair followed that tradition for the
29th time. Topics covered include the state of the development process,
what Torvalds actually does, and how machine-learning tools might fit into
the kernel project.

How Kaltura Accelerates CI/CD Using AWS CodeBuild-hosted Runners

Post Syndicated from Michael Shapira original https://aws.amazon.com/blogs/devops/how-kaltura-accelerates-ci-cd-using-aws-codebuild-hosted-runners/

Kaltura, a leading AI video expirience cloud and corporate communications technology provider, transformed CI/CD infrastructure by migrating to AWS CodeBuild-hosted runners for GitHub Actions. This migration reduced DevOps operational overhead by 90%, accelerated build queue times by 66%, and cut infrastructure costs by 60%. Most importantly, the migration achieved these results while supporting Kaltura’s scale: over 1,000 repositories, 100+ distinct workflow types, and 1,300 daily builds across multiple development teams.

As organizations scale their engineering operations, maintaining efficient CI/CD infrastructure becomes increasingly critical. While tools like GitHub Actions simplify pipeline creation, managing the underlying infrastructure can become a significant burden for engineering teams, particularly when dealing with security requirements and private network access needs. For Kaltura, this challenge became acute as the company rapidly grew its engineering teams and onboarded new microservices weekly.

In this post, you’ll see how Kaltura modernized CI/CD infrastructure by migrating from self-managed Amazon Elastic Kubernetes Service (Amazon EKS) runners to CodeBuild-hosted runners, implementing enhanced security features while dramatically improving performance and reducing operational overhead.

Overview of Challenge and Solution

Understanding Self-Hosted Runners

GitHub-hosted runners offer zero operational overhead, automatic scaling, and a clean slate for each job, making them an excellent choice for many development teams. However, for enterprises like Kaltura with specific security and operational requirements, self-hosted runners provided a better fit. GitHub-hosted runners operate in a shared environment that, while secure, doesn’t offer the same level of granular control that enterprises may need for sensitive workloads. By moving to self-hosted runners on AWS, Kaltura gained access to robust security controls like Amazon Virtual Private Cloud (Amazon VPC) isolation, AWS Identity and Access Management (IAM) policies, and fine-grained access management. Additionally, self-hosted runners allowed Kaltura to customize hardware configurations for their specialized needs, optimize costs for their specific usage patterns, and maintain direct access to private network resources essential for their operations.

Self-hosted runners, which were initially implemented, offered the control Kaltura needed. By deploying runners within Amazon VPC, Kaltura gained crucial capabilities for enterprise-scale operations. The implementation enabled direct access to internal resources while implementing granular permissions through IAM roles. Using Amazon endpoints allowed Kaltura to avoid public API requests, ensuring all traffic remained within the organization’s secure private network.

The initial solution based on Amazon EKS

Kaltura’s initial solution deployed self-hosted GitHub Actions runners on Amazon EKS, using Karpenter for node auto-scaling. Kaltura implemented a custom controller that would poll the GitHub API for queued workflows and spin up necessary runners. While this solution provided the security and control Kaltura needed, it introduced substantial operational challenges.

The heart of the problems stemmed from Kaltura’s polling mechanism. As the solution’s scale grew, Kaltura frequently hit GitHub’s API rate limits, forcing a reduction of polling frequency to two-minute intervals. These circumstances created a cascading effect of operational issues. The DevOps teams spent considerable time maintaining runner images, infrastructure, and scaling mechanisms. Each new repository required manual configuration updates, creating bottlenecks in the development process. To meet performance SLAs, Kaltura maintained warm runner pools, significantly increasing infrastructure costs.

Architecture diagram showing Kaltura's initial CI/CD solution with GitHub repositories triggering workflows that are polled by a custom controller, which provisions GitHub Actions runners on Amazon EKS with Karpenter for auto-scaling, all operating within an Amazon VPC for secure access to internal resources

Figure 1: The initial solution was based on Amazon EKS and Karpenter spinning up GitHub Runners.

The impact on development teams was substantial. Every workflow execution faced a minimum two-minute delay between queuing and execution. These delays accumulated throughout the day, severely impacting developer productivity. The DevOps team found themselves constantly pulled away from other initiatives to handle infrastructure maintenance tasks. The situation became increasingly untenable as Kaltura continued to scale.

Kaltura’s Solution – AWS CodeBuild-hosted Runners

After evaluating several options, Kaltura chose CodeBuild-hosted runners to resolve infrastructure challenges while maintaining the security and control benefits of self-hosted solution. This new architecture fundamentally changed how the CI/CD solution operated, moving from a poll-based to a webhook-based system.

Architecture diagram showing Kaltura's modernized CI/CD solution using AWS CodeBuild-hosted runners, where GitHub repositories send webhook notifications through AWS CodeConnections to trigger CodeBuild, which provisions runners within an Amazon VPC with IAM role-based access to AWS services for executing GitHub Actions workflows.

Figure 2: The solution based on AWS CodeBuild is fully managed and is based on Webhooks.

The new architecture operates through a straightforward but powerful flow. When developers push code to GitHub, GitHub sends an immediate webhook notification to AWS CodeConnections. This triggers CodeBuild, which provisions a runner within Kaltura’s Amazon VPC. The GitHub Actions workflow then executes on this CodeBuild runner, leveraging fine-grained IAM roles that follow the principle of least privilege to access AWS services.

Key Architectural Components

The webhook-based architecture eliminates previous polling challenges entirely. Instead of waiting for a periodic check, workflows begin executing immediately when triggered. CodeBuild and CodeConnections use a GitHub App with webhooks, configurable at the repository, organization, or enterprise level. This integration allows true CI/CD auto-discovery, a significant advancement from previous manual configuration requirements.

Security remains one of the major components of the new architecture. Each runner operates within Amazon VPC, maintaining strict network security requirements. Kaltura implemented fine-grained access control through IAM roles, ensuring runners access only the specific AWS services they need, such as AWS Systems Manager Parameter Store, Amazon CloudWatch, and AWS Secrets Manager. This maintains security posture while simplifying access management.

Infrastructure Management

CodeBuild’s serverless nature transformed the infrastructure management approach. Rather than maintaining a complex Amazon EKS cluster with custom controllers and scaling logic, Kaltura now leverages AWS’s managed service. This shift eliminated the need to patch runner images, maintain infrastructure, or optimize scaling mechanisms.

The system’s flexibility proved particularly valuable for diverse workflow requirements. CodeBuild supports various compute configurations, from standard instances to multi-architecture builds and specialized ARM and GPU runners. Kaltura can easily match compute resources to workflow needs through simple label configurations, without managing different runner pools or maintaining separate infrastructure stacks.

Docker Workflow Improvements

One unexpected benefit emerged in Docker build processes. Previous Amazon EKS-based solutions required complex Docker-in-Docker (DinD) configurations or alternative tools like Kaniko for container builds. CodeBuild’s native Docker support eliminated these complications. The service provides isolated build environments where Docker can run directly, with built-in layer caching capabilities that significantly improve build performance.

Auto-Discovery and Self-Service

A key benefit of the new architecture is its self-service capability. When development teams create new repositories or modify existing ones, no manual DevOps intervention is required. The system automatically provisions appropriate runners based on predefined configurations and the workflow’s runs-on label. This self-service approach has dramatically reduced Kaltura’s operational overhead while improving developer productivity.

Here’s a typical workflow configuration demonstrating new approach:

name: Hello World

on: [push]

jobs:

  Hello-World-Job:

    runs-on:

      - codebuild-myProject-${{ github.run_id }}-${{ github.run_attempt }}

      - image:${{ matrix.os }}

      - instance-size:${{ matrix.size }}

      - fleet:myFleet

      - buildspec-override:true

    strategy:

      matrix:

        include:

          - os: arm-3.0

            size: small

          - os: linux-5.0

            size: large

    steps:

      - run: echo "Hello World!"

This configuration shows how Kaltura leverages CodeBuild’s flexibility while maintaining simple, declarative workflow definitions. Teams can specify their compute needs through labels, and the system handles all the underlying provisioning and management.

Migration Approach

The migration to CodeBuild runners involved a seamless transition with minimal workflow changes. The key to its successful migration was its simplicity – most workflows required only a single change to the runs-on label:

runs-on: codebuild-myProject-${{ github.run_id }}-${{ github.run_attempt }} 

Because of the 1-by-1 compatibility, it meant existing workflows continued to function without further modification.

Results

The new architecture successfully handles over 1,300 daily builds across more than 1,000 repositories and 100 workflow types while serving multiple development teams with varying security requirements. The results of the migration to CodeBuild-hosted runners delivered significant improvements across all key metrics:

Operational impact:

  • 90% reduction in DevOps operational overhead
  • 66% decrease in build queue times
  • 60% reduction in infrastructure costs
  • 30 minutes of daily time savings per developer

Most importantly, developer satisfaction has improved due to faster builds, reduced friction, and consistent performance. The self-service nature of the system has eliminated onboarding bottlenecks and accelerated the development lifecycle.

Conclusion

The transformation of Kaltura’s CI/CD infrastructure through CodeBuild-hosted runners demonstrates how modern cloud services solves complex enterprise-scale development challenges. The journey from managing self-hosted runners on Amazon EKS to leveraging AWS managed services delivered a 90% reduction in operational overhead, 66% faster build queues, and 60% cost savings while maintaining enterprise-grade security requirements.

For organizations considering a similar path, we recommend starting with a pilot program using non-critical repositories. Focus on understanding your workflow requirements, security needs, and performance bottlenecks to shape an effective migration strategy. Implement cost allocation tags and monitoring early to ensure visibility into the migration’s impact and demonstrate ROI to stakeholders.

Additional Resources:

About the Authors

Adi Ziv is a Senior Platform Engineer at Kaltura with over a decade of experience designing and building scalable, resilient, and optimized cloud-native applications and infrastructure. He specializes in serverless, containerized, and event-driven architectures.
MIchael Shapira photo Michael Shapira is a Senior Solution Architect at AWS specializing in Machine Learning and Generative AI solutions. With 19 years of software development experience, he is passionate about leveraging cutting-edge AI technologies to help customers transform their businesses and accelerate their cloud adoption journey. Michael is also an active member of the AWS Machine Learning community, where he contributes to innovation and knowledge sharing while helping customers scale their AI and cloud infrastructure at enterprise level. When he’s not architecting cloud solutions, Michael enjoys capturing the world through his camera lens as an avid photographer.
Maya Morav Freiman is a Technical Account Manager at AWS helping customers maximize value from AWS services and achieve their operational and business objectives. She is part of the AWS Serverless community and has 10 years experience as a DevOps engineer.

The content and opinions in this post are those of the third-party author and AWS is not responsible for the content or accuracy of this post.

Systemd v259 released

Post Syndicated from jzb original https://lwn.net/Articles/1051163/

Systemd
v259
has been released. Notable changes include a new
“--empower” option for run0 that provides elevated
privileges to a user without switching to root, ability to propagate a
user’s home directory into a VM with systemd-vmspawn, and
more. Support for System V service scripts has been deprecated, and
will be removed in v260. See the release notes for other changes,
feature removals, and deprecated features.

The collective thoughts of the interwebz