What shaped computing education in 2025 — and what comes next

Post Syndicated from Liz Eaton original https://www.raspberrypi.org/blog/what-shaped-computing-education-in-2025-and-what-comes-next/

To mark the start of 2026, we’re releasing a special episode of our Hello World podcast, which reflects on the key developments in computing education during 2025 and considers the trends likely to shape the year ahead.

Hosted by James Robinson, the episode brings together a conversation between three Foundation team members — Rehana Al-Soltane, Dr Bobby Whyte, and Laura James — and perspectives from colleagues and partners in Kenya, South Africa, and Greece.

The Hello World Podcast team

The podcast is framed around three major themes that defined 2025: data science, AI literacy, and digital literacy, all of which continue to play an increasingly important role in education systems worldwide.

Looking back at 2025

In the podcast, Rehana reflects on a year characterised by research, collaboration, and community, highlighting the importance of global partnerships in developing and localising AI literacy resources for diverse educational contexts.

From a research perspective, Bobby explains that 2025 was about pulling together what we already know and making sense of it, to better understand what good data science education should look like, including curriculum design, pedagogy, and appropriate tools.

Laura focuses on resilience and creativity in computing education, as well as the growing presence of more personalised forms of artificial intelligence, which present both significant opportunities and complex ethical challenges.

The new set!

A key concern raised throughout the episode is the risk of cognitive offloading, whereby learners rely on AI tools to bypass critical thinking processes. The speakers emphasise the need for learning experiences and assessments that value process, reasoning, and reflection rather than solely final outputs.

The episode also examines barriers to the adoption of computing and AI education, including teacher confidence, limited access to devices, restrictive school IT policies, and the need for translated and localised resources.

Contributions from our colleagues around the world highlight stark contrasts in educational contexts, with challenges such as funding constraints, connectivity issues, and teacher training needs, alongside examples of innovation where educators are adequately supported.

What’s ahead

Looking ahead to 2026, Rehana outlines the potential of interdisciplinary approaches to AI literacy, integrating AI concepts into subjects such as geography, history, languages, and the arts to increase relevance and engagement (look out for our upcoming research seminar series on the topic).

The cast on set

Bobby anticipates a gradual shift towards more data-informed approaches to computing education, with greater emphasis on classroom-based trials and research that directly informs practice.

Laura offers a strong call to renew focus on cybersecurity education, arguing that security and safety must remain central as digital systems and AI technologies continue to evolve.

In a series of concise predictions, the speakers point to increased attention on explainable AI, wider integration of AI literacy across the curriculum, and renewed concern for digital safety and security.

More from Hello World

You can subscribe to Hello World and listen to the full podcast episodes from wherever you get your podcasts. Or you can find this and previous Hello World podcasts on our podcast page.

Also check out Hello World magazine, our free digital and print magazine from computing educators for computing educators.

The post What shaped computing education in 2025 — and what comes next appeared first on Raspberry Pi Foundation.

AMD Teases Ryzen AI Halo, a ROCm Ecosystem AI Development Mini-PC

Post Syndicated from Ryan Smith original https://www.servethehome.com/amd-teases-ryzen-ai-halo-a-rocm-ecosystem-ai-development-mini-pc/

Among a spate of AMD announcements during the company’s CES 2026 opening keynote, CEO Dr. Lisa Su briefly teased a forthcoming AMD-branded AI development box, dubbed the Ryzen AI Halo. Seemingly taking a page from similar AI development boxes that have popped up over the last year – both those built using AMD’s hardware and […]

The post AMD Teases Ryzen AI Halo, a ROCm Ecosystem AI Development Mini-PC appeared first on ServeTheHome.

Amazon EMR Serverless eliminates local storage provisioning, reducing data processing costs by up to 20%

Post Syndicated from Karthik Prabhakar original https://aws.amazon.com/blogs/big-data/amazon-emr-serverless-eliminates-local-storage-provisioning-reducing-data-processing-costs-by-up-to-20/

At AWS re:Invent 2025, Amazon Web Services (AWS) announced serverless storage for Amazon EMR Serverless, a new capability that eliminates the need configure local disks for Apache Spark workloads. This reduces data processing costs by up to 20% while eliminating job failures from disk capacity constraints.

With serverless storage, Amazon EMR Serverless automatically handles intermediate data operations, such as shuffle, on your behalf. You pay only for compute and memory—no storage charges. By decoupling storage from compute, Spark can release idle workers immediately, reducing costs throughout the job lifecycle. The following image shows the serverless storage for EMR Serverless announcement from the AWS re:Invent 2025 keynote:

The challenge: Sizing local disk storage

Running Apache Spark workloads requires sizing local disk storage for shuffle operations—where Spark redistributes data across executors during joins, aggregations, and sorts. This requires analyzing job histories to estimate disk requirements, leading to two common problems: overprovisioning wastes money on unused capacity, and under provisioning causes job failures when disk space runs out. Most customers overprovision local storage to ensure jobs complete successfully in production.

Data skew compounds this further. When one executor handles a disproportionately large partition, that executor takes significantly longer to complete while other workers sit idle. If you didn’t provision enough disk for that skewed executor, the job fails entirely—making data skew one of the top causes of Spark job failures. However, the problem extends beyond capacity planning. Because shuffle data couples tightly to local disks, Spark executors pin to worker nodes even when compute requirements drop between job stages. This prevents Spark from releasing workers and scaling down, inflating compute costs throughout the job lifecycle. When a worker node fails, Spark must recompute the shuffle data stored on that node, causing delays and inefficient resource usage.

How it works

Serverless storage for Amazon EMR Serverless addresses these challenges by offloading shuffle operations from individual compute workers onto a separate, elastic storage layer. Instead of storing critical data on local disks attached to Spark executors, serverless storage automatically provisions and scales high-performance remote storage as your job runs.

The architecture provides several key benefits. First, compute and storage scale independently—Spark can acquire and release workers as needed across job stages without worrying about preserving locally stored data. Second, shuffle data is evenly distributed across the serverless storage layer, eliminating data skew bottlenecks that occur when some executors handle disproportionately large shuffle partitions. Third, if a worker node fails, your job continues processing without delays or reruns because data is reliably stored outside individual compute workers.

Serverless storage is provided at no additional charge, and it eliminates the cost associated with local storage. Instead of paying for fixed disk capacity sized for maximum potential I/O load—capacity that often sits idle during lighter workloads—you can use serverless storage without incurring storage costs. You can focus your budget on compute resources that directly process your data, not on managing and overprovisioning disk storage.

Technical innovation brings three breakthroughs

Serverless storage introduces three fundamental innovations that solve Spark’s shuffle bottlenecks: multi-tier aggregation architecture, purpose-built networking, and true storage-compute decoupling. Apache Spark’s shuffle mechanism has a core constraint: each mapper independently writes output as small files, and each reducer must fetch data from potentially thousands of workers. In a large-scale job with 10,000 mappers and 1,000 reducers, this creates 10 million individual data exchanges. Serverless storage aggregates early and intelligently—mappers stream data to an aggregation layer that consolidates shuffle data in memory before committing to storage. Whereas individual shuffle write and fetch operations might show slightly higher latency due to network round-trips compared to local disk I/O, the overall job performance improves by transforming millions of tiny I/O operations into a smaller number of large, sequential operations.

Traditional Spark shuffle creates a mesh network where each worker maintains connections to potentially hundreds of other workers, spending significant CPU on connection management rather than data processing. We built a custom networking stack where each mapper opens a single persistent remote procedure call (RPC) connection to our aggregator layer, eliminating the mesh complexity. Although individual shuffle operations might show slightly higher latency due to network round trips compared to local disk I/O, overall job performance improves through better resource utilization and elastic scaling. Workers no longer run a shuffle service—they focus entirely on processing your data.

Traditional Amazon EMR Serverless jobs store shuffle data on local disks, coupling data lifecycle to worker lifecycle—idle workers can’t terminate without losing shuffle data. Serverless storage decouples these entirely by storing shuffle data in AWS managed storage with opaque handles tracked by the driver. Workers can terminate immediately after completing tasks without data loss, enabling elastic scaling. In funnel-shaped queries where early stages require massive parallelism that narrows as data aggregates, we’re seeing up to 80% compute cost reduction in benchmarks by releasing idle workers instantly. The following diagram illustrates instant worker release in funnel-shaped queries.

Our aggregator layer integrates directly with AWS Identity and Access Management (IAM), AWS Lake Formation, and fine-grained access control systems, providing job-level data isolation with access controls that match source data permissions.

Getting started

Serverless storage is available in multiple AWS Regions. For the current list of supported Regions, refer to the Amazon EMR User Guide.

New applications

Serverless storage can be enabled for new applications starting with Amazon EMR release 7.12. Follow these steps:

  1. Create an Amazon EMR Serverless application with Amazon EMR 7.12 or later:
aws emr-serverless create-application \
  --type "SPARK" \
  --name my-application \
  --release-label emr-7.12.0 \
  --runtime-configuration '[{
      "classification": "spark-defaults",
        "properties": {
          "spark.aws.serverlessStorage.enabled": "true"
        }
    }]' \
  --region us-east-1
  1. Submit your Spark job:
aws emr-serverless start-job-run \
  --application-id <application-id> \
  --execution-role-arn <execution-role-arn> \
  --job-driver '{
    "sparkSubmit": {
      "entryPoint": "s3://<bucket>/<your_script.py>",
      "sparkSubmitParameters": "--conf spark.executor.cores=4 --conf spark.executor.memory=20g --conf spark.driver.cores=4 --conf spark.driver.memory=8g --conf spark.executor.instances=10"
    }
  }'

Existing applications

You can enable serverless storage for existing applications on Amazon EMR 7.12 or later by updating your application settings.

To enable serverless storage using AWS Command Line Interface (AWS CLI), enter the following command:

aws emr-serverless update-application \
  --application-id <application-id> \
  --runtime-configuration '[{
      "classification": "spark-defaults",
        "properties": {
          "spark.aws.serverlessStorage.enabled": "true"
        }
    }]'

To enable serverless storage using Amazon EMR Studio UI, navigate to your application in Amazon EMR Studio, go to Configuration, and add the Spark property spark.aws.serverlessStorage.enabled=true in the spark-defaults classification.

Job-level configuration

You can also enable serverless storage for specific jobs, even when it’s not enabled at the application level:

aws emr-serverless start-job-run \
  --application-id <application-id> \
  --execution-role-arn <execution-role-arn> \
  --job-driver '{
    "sparkSubmit": {
      "entryPoint": "s3://<bucket>/<your_script.py>",
      "sparkSubmitParameters": "--conf spark.executor.cores=4 --conf spark.executor.memory=20g --conf spark.aws.serverlessStorage.enabled=true"
    }
  }'

(Optional) Disabling serverless storage

If you prefer to continue using local disks, you can disable serverless storage by omitting the spark.aws.serverlessStorage.enabled configuration or setting it to false at either the application or job level:

spark.aws.serverlessStorage.enabled=falseTo use traditional local disk provisioning, configure the appropriate disk type and size for your application workers.

Monitoring and cost tracking

You can monitor elastic shuffle usage through standard Spark UI metrics and track costs at the application level in AWS Cost Explorer and AWS Cost and Usage Reports. The service automatically handles performance optimization and scaling, so you don’t need to tune configuration parameters.

When to use serverless storage

Serverless storage delivers the most value for workloads with substantial shuffle operations—typically jobs that shuffle more than 10 GB of data (and less than 200 G per job, the limitation as of this writing). These include:

  • Large-scale data processing with heavy aggregations and joins
  • Sort-heavy analytics workloads
  • Iterative algorithms that repeatedly access the same datasets

Jobs with unpredictable shuffle sizes benefit particularly well because serverless storage automatically scales capacity up and down based on real-time demand. For workloads with minimal shuffle activity or very short duration (under 2–3 minutes), the benefits might be limited. In these cases, the overhead of remote storage access might outweigh the advantages of elastic scaling.

Security and data lifecycle

Your data is stored in serverless storage only while your job is running and is automatically deleted when your job is completed. Because Amazon EMR Serverless batch jobs can run for up to 24 hours, your data will be stored for no longer than this maximum duration. Serverless storage encrypts your data both in transit between your Amazon EMR Serverless application and the serverless storage layer and at rest while temporarily stored, using AWS managed encryption keys. The service uses an IAM based security model with job-level data isolation, which means that one job can’t access the shuffle data of another job. Serverless storage maintains the same security standards as Amazon EMR Serverless, with enterprise-grade security controls throughout the processing lifecycle.

Conclusion

Serverless storage represents a fundamental shift in how we approach data processing infrastructure, eliminating manual configuration, aligning costs to actual usage, and improving reliability for I/O intensive workloads. By offloading shuffle operations to a managed service, data engineers can focus on building analytics rather than managing storage infrastructure.

To learn more about serverless storage and get started, visit the Amazon EMR Serverless documentation.


About the authors

Karthik Prabhakar

Karthik Prabhakar

Karthik is a Data Processing Engines Architect for Amazon EMR at AWS. He specializes in distributed systems architecture and query optimization, working with customers to solve complex performance challenges in large-scale data processing workloads. His focus spans engine internals, cost optimization strategies, and architectural patterns that enable customers to run petabyte-scale analytics efficiently.

Ravi Kumar

Ravi Kumar

Ravi is a Senior Product Manager Technical at Amazon Web Services, specializing in exabyte-scale data infrastructure and analytics platforms. He helps customers unlock insights from structured and unstructured data using open-source technologies and cloud computing. Outside of work, Ravi enjoys exploring emerging trends in data science and machine learning.

Matt Tolton

Matt Tolton

Matt is a Senior Principal Engineer at Amazon Web Services.

author name

Neil Mukerje

Neil is a Principal Product Manager at Amazon Web Services.

Building scalable AWS Lake Formation governed data lakes with dbt and Amazon Managed Workflows for Apache Airflow

Post Syndicated from Abhilasha Agarwal original https://aws.amazon.com/blogs/big-data/building-scalable-aws-lake-formation-governed-data-lakes-with-dbt-and-amazon-managed-workflows-for-apache-airflow/

Organizations often struggle with building scalable and maintainable data lakes—especially when handling complex data transformations, enforcing data quality, and monitoring compliance with established governance. Traditional approaches typically involve custom scripts and disparate tools, which can increase operational overhead and complicate access control. A scalable, integrated approach is needed to simplify these processes, improve data reliability, and support enterprise-grade governance.

Apache Airflow has emerged as a powerful solution for orchestrating complex data pipelines in the cloud. Amazon Managed Workflows for Apache Airflow (MWAA) extends this capability by providing a fully managed service that eliminates infrastructure management overhead. This service enables teams to focus on building and scaling their data workflows while AWS handles the underlying infrastructure, security, and maintenance requirements.

dbt enhances data transformation workflows by bringing software engineering best practices to analytics. It enables analytics engineers to transform warehouse data using familiar SQL select statements while providing essential features like version control, testing, and documentation. As part of the ELT (Extract, Load, Transform) process, dbt handles the transformation phase, working directly within a data warehouse to enable efficient and reliable data processing. This approach allows teams to maintain a single source of truth for metrics and business definitions while enabling data quality through built-in testing capabilities.

In this post, we show how to build a governed data lake that uses modern data tools and AWS services.

Solution overview

We explore a comprehensive solution that includes:

  • A metadata-driven framework in MWAA that dynamically generates directed acyclic graphs (DAGs), significantly improving pipeline scalability and reducing maintenance overhead.
  • dbt with Amazon Athena adapter to implement modular, SQL-based data transformations directly on a data lake, enabling well-structured, and thoroughly tested transformations.
  • An automated framework that proactively identifies and segregates problematic records, maintaining the integrity of data assets.
  • AWS Lake Formation to implement fine-grained access controls for Athena tables, ensuring proper data governance and security throughout a data lake environment.

Together, these components create a robust, maintainable, and secure data management solution suitable for enterprise-scale deployments.

The following architecture illustrates the components of the solution.

The workflow contains the following steps:

  1. Multiple data sources (PostgreSQL, MySQL, SFTP) push data to an Amazon S3 raw bucket
  2. S3 event triggers AWS Lambda Function
  3. Lambda function triggers the MWAA DAG to convert file formats to parquet
  4. Data is stored in Amazon S3 formatted bucket under formatted_stg prefix
  5. Crawler crawls the data in formatted_stg prefix in the formatted bucket and creates catalog tables
  6. dbt using Athena adapter processes the data and puts the processed data after data quality checks under formatted prefix in Formatted bucket
  7. dbt using Athena adapter can perform further transformations on the formatted data and put the transformed data in Published bucket

Prerequisites

To implement this solution, the following prerequisites need to be met.

Deploy the solution

For this solution, we provide an AWS CloudFormation (CFN) template that sets up the services included in the architecture, to enable repeatable deployments.

Note:

  • US-EAST-1 Region is required for the deployment.
  • Deploying this solution will involve costs associated with AWS services.

To deploy the solution, complete the following steps:

  1. Before deploying the stack, open the AWS Lake Formation console. Add your console role as a Data Lake Administrator and choose Confirm to save the changes.
  2. Download the CloudFormation template.
    After the file is downloaded to the local machine, follow the steps below to deploy the stack using this template:

    1. Open the AWS CloudFormation Console.
    2. Choose Create stack and choose With new resources (standard).
    3. Under Specify template, select Upload a template file.
    4. Select Choose file and upload the CFN template that was downloaded earlier.
    5. Choose Next to proceed.

  3. Enter a stack name (for example, bdb4834-data-lake-blog-stack) and configure the parameters (bdb4834-MWAAClusterName can be left as the default value and update SNSEmailEndpoints with your email address), then choose Next.
  4. Select “I acknowledge that AWS CloudFormation might create IAM resources with custom names” and choose Next

  5. Review all the configuration details on the next page, then choose Submit.
  6. Wait for the stack creation to complete in the AWS CloudFormation console. The process typically takes approximately 35 to 40 minutes to provision all required resources.

    The following table shows resources available in the AWS Account after CloudFormation template deployment is successfully completed:

    Resource Type Description Example Resource Name
    S3 Buckets For storing raw, processed data and assets bdb4834-mwaa-bucket-<AWS_ACCOUNT>-<AWS_REGION>,bdb4834-raw-bucket-<AWS_ACCOUNT>-<AWS_REGION>,bdb4834-formatted-bucket-<AWS_ACCOUNT>-<AWS_REGION>,bdb4834-published-bucket-<AWS_ACCOUNT>-<AWS_REGION>
    IAM Role Role assumed by MWAA for permissions bdb4834-mwaa-role
    MWAA Environment Managed Airflow environment for orchestration bdb4834-MyMWAACluster
    VPC Network setup required by MWAA bdb4834-MyVPC
    Glue Catalog Databases Logical grouping of metadata for tables bdb4834_formatted_stg,bdb4834_formatted_exception, bdb4834_formatted, bdb4834_published
    Glue Crawlers Automatically catalog metadata from S3 bdb4834-formatted-stg-crawler
    Lambda Lambda to Trigger MWAA DAG on file arrival and to setup Lake Formation Permissions bdb4834_mwaa_trigger_process_s3_files,bdb4834-lf-tags-automation
    Lake Formation Setup Centralized governance and permissions LF-Setup for the above Resources
    Airflow DAGs Airflow DAGs are stored in the S3 bucket named mwaa-bucket-<AWS_ACCOUNT>-<AWS_REGION> under the dags/ prefix. These DAGs are responsible for triggering data pipelines based on either file arrival events or scheduled intervals. The exact functionality of each DAG is explained in the following sections. blog-test-data-processingcrawler-daily-runcreate-audit-tableprocess_raw_to_formatted_stage
  7. When the stack is complete perform the below steps:
    1. Open the Amazon Managed Workflows for Apache Airflow (MWAA) console, choose on Open Airflow UI
    2. In the DAGs console, locate the following DAGs and unpause them by unchecking the toggle switch (radio button) next to each DAG.

Add sample data to raw S3 bucket and create catalog tables

In this section, we upload sample data to raw S3 bucket (bucket name starting with bdb4834-raw-bucket) and convert the file formats to parquet and run AWS Glue crawler to create catalog tables that are used by dbt in the ELT Process. Glue Crawler automatically scans the data in S3 and creates or updates tables in the Glue Data Catalog, making the data queryable and accessible for transformation.

  1. Download the sample data.
  2. Zip folder contains two sample data files, cards.json and customers.json
    Schema for cards.json

    Field Data Type Description
    cust_id String Unique customer identifier
    cc_number String Credit card number
    cc_expiry_date String Credit card expiry date

    Schema for customers.json

    Field Data Type Description
    cust_id String Unique customer identifier
    fname String First name
    lname String Last name
    gender String Gender
    address String Full address
    dob String Date of birth (YYYY/MM/DD)
    phone String Phone number
    email String Email address
  3. Open S3 console, choose General purpose buckets in the navigation pane.
  4. Locate the S3 bucket with a name starting with bdb4834-raw-bucket. This bucket is created by the CloudFormation stack and can also be found under the stack’s Resources tab in the CloudFormation console.
  5. Choose the bucket name to open it, and follow these steps to create the required prefix:
    1. Choose Create folder.
    2. Enter the folder name as mwaa/blog/partition_dt=YYYY-MM-DD/, replacing YYYY-MM-DD with the actual date to be used for the partition.
    3. Choose Create folder to confirm.
  6. Upload the sample data files from the location to the s3 raw bucket prefix.
  7. As soon as the files are uploaded, the on_put object event on the raw bucket invokes thebdb4834_mwaa_trigger_process_s3_files lambda which triggers the process_raw_to_formatted_stg MWAA DAG.
    1. In the Airflow UI, choose the process_raw_to_formatted_stg DAG to view execution status. This DAG converts the file formats to parquet and typically completes within a few seconds.
    2. (Optional) To check the Lambda execution details:
      1. On the AWS Lambda Console, choose Functions in the navigation pane.
      2. Select the function named bdb4834_mwaa_trigger_process_s3_files.
  8. Validate the parquet files are created in formatted bucket (bucket name starting with bdb4834-formatted) under the respective data object prefix.
  9. Before proceeding further, re-upload the Lake Formation metadata file in MWAA bucket.
    1. Open the S3 console, choose General purpose buckets in the navigation pane.
    2. Search for the bucket starting with bdb4834-mwaa-bucket
    3. Choose the bucket name and go to the lakeformation prefix. Download the file named lf_tags_metadata.json. Now, re-upload the same file to the same location.
      Note: This re-upload is necessary because the Lambda function is configured to trigger on file arrival. When the resources were initially created by the CloudFormation stack, the files were simply moved to S3 and did not trigger the Lambda. Re-uploading the file ensures the Lambda function is executed as intended.
    4. As soon as the file is uploaded, the on_put object event on the MWAA bucket invokes the lf_tags_automation lambda, which creates the Lake Formation (LF) tags as defined in the metadata file and grants access to the specified AWS Identity and Access Management (IAM) roles for read/write.
    5. Validate that the LF-Tags have been created by visiting the Lake Formation Console. In the left navigation pane, choose Permissions, and then select LF-Tags and permissions.
  10. Now, run the crawler DAG to create/update the catalog tables: crawler-daily-run
    1. In the Airflow UI select the crawler-daily-run DAG and choose Trigger DAG to execute it.
    2. This DAG is configured to trigger Glue Crawler which crawls the formatted_stg prefix under the bdb4834-formatted s3 bucket to create catalog tables as per the prefixes available under the formatted_stg prefix.
      bdb4834-formatted-bucket-<aws-account-id>-<region>/formatted_stg/
      

    3. Monitor the execution of the crawler-daily-run DAG until it completes, which typically takes 2 to 3 minutes. The crawler run status can be verified in the AWS Glue Console by following these steps:
      1. Open the AWS Glue Console.
      2. In the left navigation pane, choose Crawlers.
      3. Search for the crawler named bdb4834-formatted-stg-crawler.
      4. Check the Last run status column to confirm the crawler executed successfully.
      5. Choose the crawler name to view additional run details and logs if needed.

    4. Once the crawler has completed successfully, in the left-hand panel, choose Databases and select the bdb4834_formatted_stg database to view the created tables, which should appear as showing in the following image. Optionally, select the table’s name to view its schema, and then select Table data to open Athena for data analysis. (An error may appear when querying data using Athena due to Lake Formation permissions. Review the Governance using Lake Formation section in this post to resolve the issue.)

Note: If this is the first time Athena is being used, a query result location must be configured by specifying an S3 bucket. Follow the instructions in the AWS Athena documentation to set up the S3 staging bucket for storing query results.

Run model through DAG in MWAA

In this section, we cover how dbt models run in MWAA using Athena adapter to create Glue-catalogued tables and how auditing is done for each run.

  1. After creating the tables in the Glue database using the AWS Glue Crawler in the previous steps, we can now proceed to run the dbt models in MWAA. These models are stored in S3 in the form of SQL files, located at the S3 prefix: bdb4834-mwaa-bucket-<account_id>-us-east-1/dags/dbt/models/
    The following are the dbt models and their functionality:

    • mwaa_blog_cards_exception.sql This model reads data from the mwaa_blog_cards table in the bdb4834_formatted_stg database and writes records with data quality issues to the mwaa_blog_cards_exception table in the bdb4834_formatted_exception database.
    • mwaa_blog_customers_exception.sql This model reads data from the mwaa_blog_customers table in the bdb4834_formatted_stg database and writes records with data quality issues to the mwaa_blog_customers_exception table in the bdb4834_formatted_exception database.
    • mwaa_blog_cards.sql This model reads data from the mwaa_blog_cards table in the bdb4834_formatted_stg database and loads it into the mwaa_blog_cards table in the bdb4834_formatted database. If the target table does not exist, dbt automatically creates it.
    • mwaa_blog_customers.sql This model reads data from the mwaa_blog_customers table in the bdb4834_formatted_stg database and loads it into the mwaa_blog_customers table in the bdb4834_formatted database. If the target table does not exist, dbt automatically creates it.
  2. The mwaa_blog_cards.sql model processes credit card data and depends on the mwaa_blog_customers.sql model to complete successfully before it runs. This dependency is necessary because certain data quality checks—such as referential integrity validations between customer and card records—must be performed beforehand.
    • These relationships and checks are defined in the schema.yml file located in the same S3 path: bdb4834-mwaa-bucket-<account_id>-us-east-1/dags/dbt/models/. The schema.yml file provides metadata for dbt models, including model dependencies, column definitions, and data quality tests. It utilizes macros like get_dq_macro.sql and dq_referentialcheck.sql (found under the macros/ directory) to enforce these validations.

    As a result, dbt automatically generates a lineage graph based on the declared dependencies. This visual graph helps orchestrate model execution order—ensuring models like mwaa_blog_customers.sql run before dependent models such as mwaa_blog_cards.sql, and identifies which models can execute in parallel to optimize the pipeline.

  3. As a pre-step before running models, choose the trigger DAG button for create-audit-table to create audit table for storing run details for each model.
  4. Trigger the blog-test-data-processing DAG in the Airflow UI to start the Model run.
  5. Choose blog-test-data-processing to see the execution status. This DAG runs the models in order and creates Glue catalogued iceberg tables. The flow diagram of a DAG from Airflow UI can be found by choosing Graph after choosing DAG.

    1. The exception models puts the failed records under exception prefix in S3:
      bdb4834-formatted-bucket-<aws-account-id>-<region>/formatted_exception/

      Records that failed are found in an added column, tests_failed, where all the data quality checks that failed for that particular row are added, separated by a pipe (‘|’). (For the mwaa_blog_customers_exception two exception records are found in the table.)

    2. The passed records are put under formatted prefix in S3.
      bdb4834-formatted-bucket-<aws-account-id>-<region>/formatted/

    3. For each run, a run audit is captured in the audit table with execution details like model_nm, process_nm, execution_start_date, execution_end_date, execution_status, execution_failure_reason, rows_affected.
      Find the data in S3 under the prefix bdb4834-formatted-bucket-<aws-account-id>-<region>/audit_control/
    4. Monitor the execution until the DAG completes, which can take up to 2-3 mins. The execution status of the DAG can be seen in the left panel after opening the DAG.
    5. Once the DAG has completed successfully, open the AWS Glue console and select Databases. Select the bdb4834_formatted database, which should create three tables, as shown in the following image.
      Optionally, choose Table data to access Athena for data analysis.
    6. Choose bdb4834_formatted_exception database from under Databases in AWS Glue console, which should create two tables as shown in the following image.
    7. Each model is assigned LF tags through the config block of model itself. Therefore, when the iceberg tables are created through dbt, LF tags are attached to the tables after the run completes.

      Validate the LF tags attached to the tables by visiting the AWS Lake Formation console. In the left navigation pane, choose Tables and look for mwaa_blog_customers or mwaa_blog_cards table under bdb4834_formatted database. Select any table among the two and under Actions, choose Edit LF tags and the tags are attached, as shown in the following screen shot.

    8. Similarly, for the bdb4834_formatted_exception database, select any one of the exception tables under the bdb4834_formatted_exception database and the LF tags are attached.
    9. Run SQL queries on the tables created by opening the Athena console and running Analytical queries on the tables created above.Sample SQL queries:
      SELECT * FROM bdb4834_formatted.mwaa_blog_cards;
      Output: Total 30 rows

      SELECT * FROM bdb4834_formatted_exception.mwaa_blog_customers_exception;
      Output: Total 2 records

Governance using Lake Formation

In this section, we show how assigning Lake Formation permissions and creating LF tags is automated using the metadata file.Below is a metadata file structure, which is needed for reference when uploading the metadata file for Lake Formation in Airflow S3 bucket, inside the Lake Formation prefix.

Metadata file structure-
{
    "role_arn": "<<IAM_ROLE_ARN>>",
    "access_type": "GRANT",
    "lf_tags": [
      {
        "TagKey": "<<LF_tag_key>>",
        "TagValues": ["<<LF_tag_values>>"]
      }
    ],
	  "named_data_catalog": [
      {
        "Database": "<<Database_Name>>",
        "Table": ""<<Table_Name>>"
      }
    ],
    "table_permissions": ["SELECT", "DESCRIBE"]
  }

Components of the metadata file

  • role_arn: The IAM role that the Lambda function assumes to perform operations.
  • access_type: Specifies whether the action is to grant or revoke permissions (GRANT, REVOKE).
  • lf_tags: Tags used for tag-based access control (TBAC) in Lake Formation.
  • named_data_catalog: A list of databases and tables on which Lake Formation permissions or tags are applied to.
  • table_permissions: Lake Formation-specific permissions (e.g., SELECT, DESCRIBE, ALTER, etc.).

Lambda function bdb4834-lf-tags-automation parses this JSON and grants the required LF tags to the role with given table permissions.

  1. To update the metadata file, download it from the MWAA bucket (lakeformation prefix)
    bdb4834-mwaa-bucket-<<ACCOUNT_NO>>-<<REGION>>/lakeformation/lf_tags_metadata.json

  2. Add a JSON object with the metadata structure defined above, mentioning the IAM role ARN and the tags and tables to which access needs to be granted.
    Example:Let’s assume below is how the metadata file initially looks like:

    
    	[
    	{
        "role_arn": "arn:aws:iam::XXX:role/aws-reserved/sso.amazonaws.com/XX ",
        "access_type": "GRANT",
        "lf_tags": [
          {
            "TagKey": " blog",
            "TagValues": ["bdb-4834"]
          }
        ],
        "named_data_catalog": [],
        "table_permissions": ["SELECT", "DESCRIBE"]
      }
    ]

    Below is the json object that has to be added in the above metadata file:

    
    {
              "role_arn": "arn:aws:iam::XXX:role/aws-reserved/sso.amazonaws.com/XX ",
              "access_type": "GRANT",
              "lf_tags": [],
              "named_data_catalog": [
              {
                "Database": " bdb4834_formatted",
                "Table": "audit_control"
              },
              {
                "Database": " bdb4834_formatted_stg",
                "Table": "*"
              }
             ],
             "table_permissions": ["SELECT", "DESCRIBE"]}
    
    
    

    So now, the final metadata file should look like:

    
    [
      {
        "role_arn": "arn:aws:iam::XXX:role/aws-reserved/sso.amazonaws.com/XX ",
        "access_type": "GRANT",
        "lf_tags": [
          {
            "TagKey": "blog",
            "TagValues": ["bdb-4834"]
          }
        ],
        "named_data_catalog": [],
        "table_permissions": ["SELECT", "DESCRIBE"]
      },
      {
        "role_arn": "arn:aws:iam::XXX:role/aws-reserved/sso.amazonaws.com/XX ",
        "access_type": "GRANT",
        "lf_tags": [],
        "named_data_catalog": [
          {
            "Database": " bdb4834_formatted",
            "Table": "audit_control"
          },
          {
            "Database": " bdb4834_formatted_stg",
            "Table": "*"
          }
        ],
        "table_permissions": ["SELECT", "DESCRIBE"]
      }
    ]

  3. Upon uploading this file at the same location (bdb4834-mwaa-bucket-<<ACCOUNT_NO>>-<<REGION>>/lakeformation/) in S3, the lf_tags_automation lambda is triggered to create LF tags if they don’t exist and then it assigns those tags to the IAM role ARN and also grants permission to the IAM role ARN using named_data_catalog as defined.

    To verify the permissions, go to the Lake Formation console and choose Tables under Data Catalog and search for the table name.

To check LF-Tags, choose the table name and under the LF tags section, all the tags are found attached to this table.

This metadata file used as a structured input to an AWS Lambda function automates the following to perform automated, consistent, and scalable data access governance across the AWS Lake Formation environments:

  • Granting AWS Lake Formation (LF) permissions on Glue Data Catalog resources (like databases and tables).
  • Creating Lake Formation Tags and Applying Lake Formation tags (LF-Tags) for tag-based access control (TBAC).

Explore more on dbt

Now that the deployment includes a bdb4834-published S3 bucket and a published Catalog database, robust dbt models can be built for data transformation and curation.

Here’s how to implement a complete dbt workflow:

  • Start by developing models that follow this pattern:
    • Read from the formatted tables in the staging area
    • Apply business logic, joins, and aggregations
    • Write clean, analysis-ready data to the published schema
  • Tagging for automation: Use consistent dbt tags to enable automatic DAG generation. These tags trigger MWAA orchestration to automatically include new models in the execution pipeline.
  • Adding new models: When working with new datasets, refer to existing models for guidance. Apply appropriate LF tags for data access control. The new LF tags can also now be used for permissions.
  • Enable DAG execution: For new datasets, update the MWAA metadata file to include a new JSON entry. This step is necessary to generate a DAG that executes the new dbt models.

This approach ensures the dbt implementation scales systematically while maintaining automated orchestration and proper data governance.

Clean up

1. Open the S3 console and delete all objects from below buckets:

  • bdb4834-raw-bucket-<aws-account-id>-<region>
  • bdb4834-formatted -bucket-<aws-account-id>-<region>
  • bdb4834-mwaa-bucket-<aws-account-id>-<region>
  • bdb4834-published-bucket-<aws-account-id>-<region>

To delete all objects, choose the bucket name, select all objects and choose Delete.

After that, type ‘permanently delete’ in the text box and choose Delete Objects.

Do this for all three buckets mentioned above.

2. Go to the AWS Cloudformation console, choose you’re the stack name and select Delete. It may take approximately 40 mins for the deletion to complete.

Recommendations

When using dbt with MWAA, some typical challenges include worker resource exhaustion, dependency management issues, and in some rare cases, issues like DAGs disappearing and re-appearing when there are a large number of dynamic DAGs being created from a single python script.

To mitigate these issues, follow these best practices:

1. Scale the MWAA environment appropriately by upgrading the environment class as required.

2. Use custom requirements.txt and proper dbt adapter configuration to ensure consistent environments.

3. Set airflow configuration parameters to tune the performance of MWAA.

Conclusion

In this post, we explored the end-to-end setup of a governed data lake using MWAA and dbt which improved data quality, security, and compliance, leading to better decision-making and increased operational efficiency. We also covered how to build custom dbt frameworks for auditing and data quality, automate Lake Formation access control, and dynamically generate MWAA DAGs based on dbt tags. These capabilities enable a scalable, secure, and automated data lake architecture, streamlining data governance and orchestration.

For further exploring, refer to From data lakes to insights: dbt adapter for Amazon Athena now supported in dbt Cloud


About the authors

Muralidhar Reddy

Muralidhar Reddy

Muralidhar is a Delivery Consultant at Amazon Web Services (AWS), helping customers build and implement data analytics solution. When he’s not working, Murali is an avid bike rider and loves exploring new places.

Abhilasha Agarwal

Abhilasha Agarwal

Abhilasha is an Associate Delivery Consultant at Amazon Web Services (AWS), support customers in building robust data analytics solutions. Apart from work, she loves cooking and trying out fun outdoor experiences.

Тротоари (на цени) като в Германия

Post Syndicated from Боян Юруков original https://yurukov.net/blog/2026/trotoari/

Попаднах на мой стар пост от преди 8 години показващ как правят тротоар в тогавашния ми квартал във Франкфурт, Германия. Тогава го сравних с оплакванията на ремонтите на Дондуков и Цариградско. В продължение на темата от вчера реших да проверя, колко всъщност е струвал да се направи така. Намерих аналогичен тротоар в съседно село, който миналата година е бил направен за около 120 евро на кв.м. Това включва основа, бордьори, подобни очертания на зелени площи без осветление.

За сравнение, цената на тротоарите, които виждаме да се правят в София, при последните поръчки излиза около 100 евро на кв. м. Имайки предвид, че разходите за труд са по-ниски, цените са почти идентични Разликите в качеството на изпълнение, плочите и елементите се виждат с просто око. Бях във Франкфурт пет години след като завършиха този тротоар и не беше мръднал. Някои от тези в София вече ги виждам как пропадат.

По социалки и групи се възмущаваме на общината и липсата на контрол когато нещо е нескопосано. „Ремонт на ремонта“ стана нарицателно по време на Фандъкова за лошо свършена работа и криво-ляво закърпена след това. Доколкото си мислим, че това става само когато държавата и общината плаща, показах как частни компании правят път и се налага поне четири пъти да го поправят защото пропада. В случая това беше по поръчка и под зоркото око на Артекс, но важи с пълна сила при изпълнение на всякакви сгради и инфраструктура. При тротоарите обаче, особено ремонтите, наистина голяма отговорност носи общината в целия процес.

Има обаче поне седем фактора, които влияят дори повече на тези ремонти. Някои от тях оскъпяват работата значително, други блокират налагането на контрол, а трети в най-добрия случай удължават времето за изпълнение.

Подземна инфраструктура

В София и практически всеки български град който не е коп’нал, той не е направил шахта или прекарал нещо под земята. Понякога кабели, тръби и оптика са на сантиметри под краката на хората. Понякога са дори опасни. Спомняте си смъртта на детето, което го хвана ток. Столичния общински съвет обеща да се захване с решаване на проблема и създаване на структура, която да отговаря. Това не се случи.

Преди две години пуснах карта на шахтите на София. Това са поне тези, за които общината знае и са в ГИС системата на НАГ. Има още много кабели и трасета между тях и още доста, за които не се знае. Всички следва да са достатъчно дълбоко, за да не пречат на тротоари и улици. Проблемът е, че понякога са буквално сантиметри под плочки, корени и бордюри. Това се установява едва когато започне ремонта, т.е. когато работата е планирана, бюджетирана и поръчката е спечелена.

Дори когато са незаконни не могат просто да се отрежат. Когато са законни обаче, но не на правилната дълбочина, проблемът възниква когато някой чиновник някога е подписал, че приема обекта както си е. Дали е било въпрос на корупция или нехайство, факт е, че сега разходът за преместването им на правилна дълбочина е на общината години по-късно. Това оскъпява значително нещата.

Лошо или липса на планиране на улици и тротоари

ПУП-овете на парче, позволяването улиците да са с намалена широчина оставяйки никакво място за тротоарите, приватизирането на места за паркиране или повече ленти на булевардите за сметка на тротоари и велоалеи, странните архитектурни чупки по сградите усложнявайки достъпа до гаражи, входове и съседни сгради, липсата на отчуждаване на достатъчно от частните имоти преди урегулирането им години назад във времето за нормална инфраструктура и редица други проблеми около хаоса в презастрояването в София и редица други градове води до там, че тротоарите не са просто 2.5 метра широки алеи подходящи за хора с проблеми с придвижването, незрящи или родители с колички, а криволичеща какафония от неравности и абсурди.

Дори при липса на подземни „мини“, най-съвестно изпълнение на строителя и контрол от общината, каквото и да бъде направено на определени места ще е абсурдно. В Изгрев видях тротоар широк една педя с бордюр още 10 см. Опитах и сам не мога да стъпя на него. Но имаше знак „мини на отсрещния тротоар“, т.е. на тия 30 см., защото се строеше поредният голям комплекс наблизо.

Дори когато се ремонтират улиците, се прави единствено смяна на горния слой на асфалта. Не се преосмисля концепцията на улицата, новата натовареност, дали следва да има издадени части на тротоара за по-лесно преминаване на учениците по пътеките, дали трябва да се обособят паркоместа и стесни на места платното, за да се намали скоростта и шума. Именно това очаквах при поръчката за ремонт на улицата в най-лошо състояние в район Изгрев – Тинтява. Районният кмет се хвалеше години наред, че работи над промяната ѝ, че мисли и готови всичко и подава предложения как да е в общината. Накрая като излезе поръчката се оказа, че нищо не е подал като предложения, а се използва остарели и вече неверни скици и планове от преди 10 години. Това е пример как не се прави и накрая плащаме много по-вече от данъците си за нещо по-лошо.

Строителен надзор

При такива проекти се взима строителен надзор, който би следвало да е независим. Проблемът е, че често става въпрос за свързани фирми. Независимият надзор трябва да следи за изпълнението на параметрите и изискванията за качество. Както добре виждаме обаче, това почти никога не се случва. Общините често нямат експертен капацитет да преценят дали нещо следва да е така или не и чисто естетическата им оценка и това, че някои неща са очевидни не издържат в съда. Становището на „независимия“ надзор натежава.

Непостоянни съдебни практики с дъх на корупция

Всяка обществена поръчка, всяко искане за поправка, всяка наложена глоба за несвършена работа или лошо изпълнение, всичко може и често се обжалва пред административните съдилища. Те имат крайно противоречива практика. Отчасти това е заради неспазен процес от страна на общината въпреки наличието на нарушения. По-често е защото става дума за сериозни пари и съдиите или не им пука, защото „ма то така си беше“, или са заставени да не им пука. Аналогични проблеми има при отказите за разрешения за строеж или ПУП-ове.

В тази среда административните съдилища в България носят със себе си един от най-негативните ефекти на градската среда. Несигурността, че законът и публичният интерес ще бъдат запазени води до там, че общините масово не се възползват от възможността да рестарвират или поне укрепят разпадащи се паметници на културата и да заставят собствениците да платят, ако ще със самия имот. Страх ги е, че никога няма да си върнат парите, защото е по-евтино да подариш апартамент във Витоша на сина на съдийка в Административния съд, отколкото да върнеш парите, които сме дали от данъците си.

При поръчките за тротоарите ситуацията е подобна. Затова винаги се гледа да се разберат с добро, ако ще да е на база компромиси и ниско качество на резултатът. Най-лошото е, че следващия път най-вероятно същите хора ще спечелят поръчката отново и ще трябва да се работи с тях отново

Липса на конкуренция и некачествена работа

Доколкото има страшно много строителни фирми в България покрай манията за крипто-бетона, всъщност много малко фирми кандидатстват за подобни поръчки. Отчасти това е заради изискванията за опит и мащаб, отчасти защото много от фирмите работят на черно и не могат да докажат доходи от подобна работа, отчасти, защото хора като Вълка и Таки извиват оркестрират кой да кандидатства къде. Така разпределят територия и сфери и дори общините да не са „в кюпа“ и да се стараят да получат качество на разумна цена, често са изправени пред свършен факт. Пример за това беше боклука в София, където изведнъж всички камиони за смет в страната потънаха в дън земя, когато община София искаше да наеме такива.

Липсва конкуренция заради такива мафиотски тактики и малкия брой фирми по принцип. Отделно качеството на работата е изключително лошо. Въпреки възторгът на инфлуенсъри и брокери, ако се наблюдавате строежите преди да сложат фасадата виждате с просто око невъзможни неща. Залепят ли фасадата и покрият бетона на гаражите с боядисан в зелено тънък слой трева вече изглежда готино като за снимки.

При тротоарите това го забелязваме по-бързо, защото е пред нас. При първият лек дъжд се виждат и проблемите в основата като се разкривят и започнат да плюят плочките. Бордюрите започват да се чупят от качващи се коли. Плочите и бордюрите не са изрязани правилно, а се пълнят с малко бетон ронещ се на втората седмица. Липсват хора, но дори тези служители не са обучени, нямат достъп до правилната техника и не им се плаща достатъчно. Твърде малък пазар сме, че фирми да дойдат за няколко тротоара от други държави, а и мафиотските схеми по-горе ги гонят. Материалите не са достатъчно добри, дори да има сложени изисквания.

Специфични изисквания, схеми и планове

За да може да се аргументира използването на 5 см. тежки плочи вместо 1.5 см. или специални бордюри с изрязани плочи около тях или многослойна основа от точно определен чакъл, а не каквито строителни отпадъци фирмата не е успяла да изхвърли вероятно нелегално някъде, трябва да е разписано ясно и недвусмислено в стандарт и схеми за тротоарите. Предложения за такива е имало доста през времето, но или не са били технически издържани отвъд шарените картини, или просто не са били придвижвани в Столичния общински съвет.

Именно тяхна е отговорността за приемане на такива изисквания, схеми и планове. Аналогично могат да изискат разработването и да приемат разписана единна визия на фасадите в централната част или да засилят вместо да зачеркват, както явно се готви икономическото мнозинство, изискванията за озеленяване на нови сгради. За да може да се изисква качество като описаното горе, трябва да се стъпи на такава схема. Именно това имат във Франкфурт в различни варианти и всеки район си избира специфика и дори цветове.

Паркиране върху тротоарите

Никой тротоар не се прави с идеята да може да се паркира върху него. Това би оскъпило значително процеса, особено предвид увеличения брой тежки джипове из града. У нас всеки паркира където си иска и се кара като му направят забележка „ма къде да паркирам другаде“. Новите тротоари в София масово се използват за паркоместа. Това се вижда дори и особено когато тротоарите са пред училища и детски градини точно когато десетки деца трябва да минат точно от там. Доколкото изпълнението е лошо като цяло, именно паркирането е основна причина за скапването им.

В Германия за паркиране като този горе се плаща 100 евро глоба на момента. Няма нужда от катаджия на място – просто снимка от някой, където се вижда номера и мястото. Дори да стъпиш с гума на бордюра, което неизменно го чупи с времето, значи глоба. У нас имаме изключително странна нарочно бюрократизирана система с предразполагаща към корупция и отказ от отговорност, в която практически не се глобяват повечето нарушения, а събираемостта е плачевна. Столична община няма право да налага глоби, за това, че им чупи някой тротоарите. Само ако са паркирали в зона. СДВР не приема да използва камерите на общината, не следи за минаване на червено или престрояване на кръстовище и масово не налага глоби. Никога не са взимали подписаните с електронен подпис сигнали от Гражданите или собственият им мейл за сигнали, например, за разлика от останалата част от страната.

И не, колчетата не решават нищо. Тротоарите и без това са крайно тесни с всички лампи, стълбове, дървета, кофи за боклук, стълби на нечий магазин или щъркели на някой бар или направо маси на заведение наблъскани така че и сам човек да не може да мине, а какво остава за детска или инвалидна количка. Тук се налагат законодателни промени даващи повече лостове на общината и възможности за гражданско участие, както и автоматизация на процеса както на установяване, така и на връчване и налагане на глобите.

Какви други причини смятате, че са замесени тук? Ще ми е интересно да го обсъдим в коментарите.

[$] Questions for the Technical Advisory Board

Post Syndicated from daroc original https://lwn.net/Articles/1051768/

The nature and role of the Linux Foundation’s Technical Advisory Board (TAB) is
not well-understood, though
a recent LWN article shed some light on its
role and
history. At the 2025

Linux Plumbers Conference
(LPC), the TAB held a question and
answer session to address whatever it was the community wanted to know
(video).
Those questions ended up covering the role of large language models in kernel
development, what it is like to be on the TAB, how the TAB can help grease the
wheels of corporate bureaucracy, and more.

[$] The difficulty of safe path traversal

Post Syndicated from daroc original https://lwn.net/Articles/1050887/

Aleksa Sarai, as the maintainer of the
runc container runtime, faces a
constant battle against security problems. Recently, runc has seen

another
instance
of a security vulnerability that can be traced back to the difficulty
of handling file paths on Linux. Sarai spoke at the 2025
Linux Plumbers Conference
(slides;
video)
about
some of the problems runc has had with path-traversal vulnerabilities, and to
ask people to please use

libpathrs
, the library that he has been developing for
safe path traversal.

A Cyberattack Was Part of the US Assault on Venezuela

Post Syndicated from Bruce Schneier original https://www.schneier.com/blog/archives/2026/01/a-cyberattack-was-part-of-the-us-assault-on-venezuela.html

We don’t have many details:

President Donald Trump suggested Saturday that the U.S. used cyberattacks or other technical capabilities to cut power off in Caracas during strikes on the Venezuelan capital that led to the capture of Venezuelan President Nicolás Maduro.

If true, it would mark one of the most public uses of U.S. cyber power against another nation in recent memory. These operations are typically highly classified, and the U.S. is considered one of the most advanced nations in cyberspace operations globally.

Security updates for Tuesday

Post Syndicated from jzb original https://lwn.net/Articles/1052955/

Security updates have been issued by AlmaLinux (kernel, ruby, and thunderbird), Debian (libsodium and ruby-rmagick), Fedora (gnupg2 and proxychains-ng), Oracle (gcc-toolset-14-binutils, rsync, tar, and thunderbird), Red Hat (buildah, mariadb, mariadb10.11, podman, and tar), SUSE (alloy, apache2, buildah, erlang26, glib2, ImageMagick, kernel, libsoup, pgadmin4, python-tornado6, python3, python312, python313, qemu, webkit2gtk3, and xen), and Ubuntu (webkit2gtk).

A research-led framework for teaching about models in AI and data science

Post Syndicated from Manni Cheung original https://www.raspberrypi.org/blog/a-research-led-framework-for-teaching-about-models-in-ai-and-data-science/

Research indicates that teaching learners to use and create with data-driven technologies such as AI and machine learning (ML) requires an entirely different approach for solving problems compared to traditional programming activities.

Learner in a computing classroom.

In this blog, we share the new data paradigms framework that we have developed through research and used to help improve our understanding about how to teach and learn about AI and data science. We also invite you to register your interest in participating in our next collaborative study on the topic.

Knowledge-based approaches to systems design

Let’s start by highlighting an important distinction between different approaches to designing systems. In a knowledge-based approach to system design, a set of rules (e.g., if-then statements) are written for the system to execute. Every rule is explicitly defined. This approach is called ‘rule-based’, ‘symbolic’, or ‘logic-based’. For example, a developer could create a program that simulates dialogue by writing specific lines of code to handle a greeting, such as “IF user says “Hello” THEN output “Hi!”. If the user types “Greetings!” instead, the program fails because it has no rule for that specific word. 

An educator helps students with a coding task.

Knowledge-based models are often said to be explainable by design. This means the logic is accessible and interpretable and developers can trace the exact steps taken to produce an output. For example, if developers manually classify restaurant reviews as positive or negative using a pre-defined set of criteria, the rules their restaurant classifying system follows are entirely explicit, and the path from input to output is clear and explainable.

Data-driven approaches to systems design

By contrast, in a data-driven approach to system design developers do not write specific rules. Instead, they collect lots of data and train a model. In the dialogue simulator example, they would collect hundreds of examples of greetings and train a model to the pattern of a greeting. If the user types “Greetings!”, the system generates a response based on the patterns in its training data.

Photo focused on a young person working on a computer in a classroom.

Data-driven models are often opaque. In other words, the internal workings of these ML models are hidden. While we can see our input and the system’s output, the internal mathematical process is so complex — often involving layers of calculations and abstractions — that we cannot simply “explain” why a specific output was produced. For example, developers can create a classification model by training a neural network using thousands of images. Due to the large quantity of data used to train the model, and complex internal parameters and hidden layers, developers and users of the system cannot understand or explain the logic or features that lead to a specific output. These kinds of models are often referred to as a “black box” (as opposed to a “glass” or “clear” box).

Comparing knowledge-based and data-driven approaches

Researchers have argued that the move from knowledge-based (or rule-based) programming to data-driven system design represents a paradigm shift and creates unique challenges for educators. The challenge is helping students shift from the expectation that a system produces a single ‘right’ answer — characteristic of traditional rule-based programming — toward an understanding that systems trained on large quantities of data produce outcomes that aren’t always fixed or explainable. If the current instruction in the classroom still relies heavily on traditional rule-based programming approaches, we might be setting students up for misconceptions.

Data paradigms: A framework for analysing data science education approaches

In our research work on AI and data science at the Raspberry Pi Computing Education Research Centre, we analysed 84 research studies about the teaching and learning of data science. We categorised learning activities used in the studies to understand whether they were (i) knowledge-based or data-driven, and (ii) the extent to which the underlying models used were transparent or opaque. This led us to define four distinct data paradigms:

The data paradigms framework
The data paradigms framework
  1. Knowledge-based and transparent (KB + T): Activities in this paradigm are ones where students write rules for systems, or work with systems that use rules, where the logic is fully explainable by design. For example, if students manually classify data (e.g. creating simple ‘if-then’ statements to predict an outcome), the path from input to output is clear.
  2. Data-driven + Transparent (DD + T): In this paradigm, activities involve students working with models trained on data, but the trained model’s logic remains explainable and interpretable. For example these could be models using k-nearest neighbors (KNN) algorithm to group data points based on proximity, or using linear regression to predict a trend. Even though the model produces an output, the student can look at the inner workings of the model and see how the decision is made.
  3. Data-driven + Opaque (DD + O): This paradigm’s activities require students to work with data-driven ML models where the models’ internal logic is hidden, for example an image classification model using a type of neural network (e.g. CNN). The model produces an output (e.g. classifying an image as ‘This is a dog’), but the student cannot inspect the system to find a rule or clear path explaining why that specific output was produced. To understand these systems, it’s necessary to use additional testing and evaluation tools.
  4. Knowledge-based + Opaque (KB + O): Activities in this paradigm would involve systems with human-written rules that are not explainable. In our review of K–12 activities, we found no examples of activities within this paradigm.

The data paradigms framework helps us to distinguish between different kinds of modeling activities students take part in and how instructional approaches could be classified across one or more paradigms. For instance, we found that most data-driven activities were also opaque (DD + O), usually meaning that students collected and used data to train a model, but how the system worked was opaque. This pattern, where the data is visible but the model is not explainable, risks students forming misconceptions about the capabilities and limitations of data-driven systems. Without understanding how outputs are generated, students may expect data-driven ML systems to operate like fully explainable (or transparent) ones.

Learners at a Code Club.

We think that lessons are needed in the data-driven opaque (DD + O) quadrant to explicitly teach students about how data-driven systems work and the role they play in everyday contexts. However, when teaching data-driven opaque (DD + O) activities, learners’ attention needs to be directed to concepts such as model confidence, data quality, and model evaluation. Since an ML model is not inherently explainable, we need to teach students to use post-hoc explanation methods, such as testing different inputs to see how a system’s output changes. To prepare students for this learning experience, we think that first introducing activities about rule-based systems (knowledge-based + transparent; KB + T) or simple data exploration, such as linear regression or data visualisation (data-driven + transparent; DD + T) may serve as a ‘bridge’ to understanding data-driven modeling by helping students to distinguish between systems built from specific logical rules and systems trained on data.

We believe the idea of data paradigms can serve as a way of framing teaching activities about data science and help educators and students to consider the transition between different paradigms when engaging with the systems we interact with every day.

Teachers in England, participate in our new study

We’re launching a new study to explore how to teach learners aged 9 to 11 about data-driven computing. The study will take place in collaboration with upper key stage 2 teachers in England and look at:

  • What key ideas pupils need to understand
  • How teachers currently approach topics related to data-driven computing
  • How pupils make sense of data and probability

Our goal is to find practical ways to help teachers build children’s confidence in working with data in computing lessons. The study will be collaborative, with two workshops held throughout 2026, and we’re inviting upper KS2 teachers in England to take part.

You can express your interest in participating by filling in this form:

The post A research-led framework for teaching about models in AI and data science appeared first on Raspberry Pi Foundation.

Европис за отличници

Post Syndicated from original https://www.toest.bg/evropis-za-otlichnitsi/

Европис за отличници

Нова година – нова валута, казва народът. Поговорката, както знаем, е за деня и късмета, но защо да не погледнем малко по-ведро на промяната и ей така, за разнообразие да загърбим страховете, песимизма, инерцията, отгръщайки ненаписаната още страница на 2026 година.

Еврото вече е в портфейлите ни, в банковите сметки, монетите подрънкват в джоба ни – и да, все по-често ще присъства и в официалната, и в разговорната реч. Затова да видим на какви затруднения може да се натъкнем при употребата на тази, а и на други думи от финансовата сфера.

Има ли евро форма за множествено число?

Дори и от филолози ще чуете, че няма такава форма. Това според мен е некоректен отговор и бих призовала колегите поне малко да се замислят, преди да кажат категорично „не“. Всеки, който владее добре български, веднага може да образува тази форма – евра. Тя обаче не е приета в книжовната реч и

когато говорим или пишем публично, трябва да употребяваме формата за ед.ч. във всички случаи: едно евро, шест евро, няколко евро, много евро.

Добре е да обясним и защо мнозина са склонни да грешат в тези случаи. Причината е в системността на езика. Повечето съществителни нарицателни имена имат количествено измерение и съответно форми за единствено и множествено число¹. Това се отнася и за названията на повечето валути: лев – левове; долар – долари; паунд – паунди; крона – крони. Ето защо е логично да имаме и евро – евра².

В книжовната реч обаче сме ограничени и според мен е възможно за това решение да е повлиял и английският модел: (one) euro – (many) euro, макар логиката и в този език да предполага разлика във формите: (one) euro – (many) euros. Интересно е, че отделните европейски книжовни езици се справят по различен начин с приспособяването на думата към граматичната им система. Немският, италианският и полският например се държат като английския и българския – обща форма за двете числа, но френският, испанският и португалският в по-голяма степен са интегрирали съществителното и формите за ед. и мн.ч. се различават.

Докато сме все още на тази граматична тема, е добре да обърнем внимание на една грешка, за която мнозина дори не подозират, защото на практика е кажи-речи нормализирана. Изрази като ресто в лева, плащам в лева, сметка в лева са неправилни. Формата лева е бройна и най-общо се употребява след числително бройно име (по-просто казано, след число), както и след местоименията колко, няколко и николко: 20 лева, колко лева, няколко лева, николко лева. В останалите случаи се използва обикновената форма за мн.ч.: ресто в левове, плащам в левове, сметка в левове. Тези употреби – съответно и грешките – ще намалеят по естествен път, но в преходния месец януари със сигурност ще ги срещаме често.

Символът €

Дори и онези българи, които не са пътували много-много из Европа, вече би трябвало да са свикнали с него и да го разпознават – от половин година е на етикетите с цените на стоките в магазините. Допреди няколко дни не бях проучвала символиката на знака € и го асоциирах с първата буква на евро, а допълнителната хоризонтална черта – с пресечната черта в символите на други валути: $, £, ¥ (юан, йена). Оказва се, че в основата на € е гръцката буква ε (епсилон), а двете успоредни линии символизират стабилност.

При употребата на знака трябва да внимаваме за две неща.

Първо, поставяме € след, а не пред цифрите, с които означаваме сумата. Второ, оставяме интервал между цифрите и €.

Ето примери за правилна употреба: 25,73 €, 400 000 €. Вариантите €25,73, €400 000, 25,73€, 400 000€ са грешни.

С прилагането на тези две правила е лесно да се справим, но известни затруднения ще срещнем при самото изписване на знака. На мобилните си телефони би трябвало да имате € на клавиша за $, ако сте избрали като регион държава от еврозоната (вкл. България).

На компютрите е малко по-трудно. В интернет ще намерите съвети с различни начини за изписване на €, които не са универсални, за съжаление. Ако сте под Windows, пробвайте по следния начин: 1) уверете се, че сте на латиница; 2) активирайте нумлоковата част на клавиатурата (групата клавиши с цифри най-вдясно); 3) натиснете левия Alt и задръжте клавиша; 4) наберете 0128 на нумлоковата част; 5) пуснете левия Alt. Вече би трябвало на екрана да се е появил символът €.

Работещите под MacOS може да използват клавишната комбинация Option + 1 (ако са на кирилица) или Option + Shift + 2 (ако са на латиница).

А как да съкращаваме евро и евроцентове?

Съкращенията лв. и ст. ни вършеха много добра работа, защото бяха общоприети, но по отношение на новата валута все още се чудим дали да пишем евро, €, евроцент, цент, или да съкращаваме думите.

Ето какви начини за съкращаване посочва Институтът за български език на БАН в отговор на запитване на Министерството на образованието и науката:

евро – е.
цент – ц.
евроцент – е.ц.
стотинка – ст.

Не смея да прогнозирам как и доколко тези съкращения ще се установят в езиковата практика. Все пак ще отбележа, че това е. ми е доста странно – може би защото глаголната форма е се среща често в писмени текстове и поне в началото (ако изобщо съкращението е. започне да се употребява) би било озадачаващо да виждаме например Блузата струва 34 е. Всякак бих предпочела € пред е., ако опрем до краткостта.

Ще наблюдавам установяването на съкращенията с интерес, също както и борбата между центовете и стотинките – кой кого? От една страна, възможно е да оставим подразделенията на лева в миналото поради неразривната им връзка с предишната валута, но от друга, кой знае, може я носталгията, я пуризмът да заговори у нас, а защо не и да се вгледаме в нашите евромонети и да си кажем: ама тук си пише стотинки! Препоръката на ИБЕ е за назоваване на монетите с номинал, по-малък от 1 евро, да се използва тъкмо тази дума, а не (евро)центове.

Правописни неволи с други „финансови“ думи

С евроцентовете правим плавен преход към слятото и разделното писане на евронещата. По-голямата част от тези думи се пишат само слято – еврофонд, европрограма, еврокомисар, тъй като първата част е съкращение на прилагателното европейски. При евроцент, евровалута, еврозона ситуацията е друга – първата част представлява съществително име, което се употребява самостоятелно и пояснява втората част, също съществително име. В тези случаи е позволено думите да се пишат и слято, и разделно: евроцент/евро цент. В езиковата практика обаче по-често се среща слятото писане, вероятно под влияние на първия тип думи, които са по-многобройни и се пишат само по този начин. Ето защо ви препоръчвам да не се занимавате с тънки разграничения, а да пишете евродумите слято³.

Покрай евровалутата се сещам и за криптовалутите, които имат немалко почитатели, а самото съществително име се пише единствено слято. Използването на средства от еврофондовете пък често върви със съфинансиране. То може и да означава същото като co-financing, но английският правопис не бива да ни въвежда в заблуждение, защото в нашата дума съ- е представка и затова не се отделя с дефис (малко тире).

Рискувам да обидя четящите тази статия с бележката, че съществителното финанси се употребява само в тази форма и тя е за мн.ч. Да, редица думи, като лекции, технологии завършват на -ии в мн.ч., но срещу тях стоят лекция, технология в ед.ч., а финансия категорично липсва в българския език, така че финансиите няма откъде да се вземат. Факт е също така, че в граматично отношение съществителното е „дефектно“ – липсва му форма за ед.ч., подобно на еврото в книжовния език, което си няма специфична форма за мн.ч. Изключенията предразполагат към грешки и с това далеч не оправдавам грешащите, а се опитвам да разбера причините.

Останаха не една и две „финансови“ думи с колебаещ се правопис, но е време да си пожелаем безпроблемно обращение (не обръщение) на еврото в България, винаги да има какво да превеждаме (не да привеждаме) от своите сметки в други, тоест да сме платежоспособни (не платежноспособни), а най-малката купюра (не купюр) в портмонето ни да е от 100 евро!

1 Изключение правят абстрактни съществителни (доброта, егоизъм), названия на газове, метали, вещества (кислород, калций) и др., които имат само форма за ед.ч. Други съществителни се употребяват само с формата си за мн.ч., например названия на предмети от две части (очила, гащи), на вещества и други същности, които се мислят само като множества (въглища, джибри).

2 Факт е, че в момента в обращение е и друга парична единица, чието название е от среден род – песо, и книжовната формата за мн.ч. също е песо. От друга страна обаче, тази дума се среща рядко в нашата реч и съответно е трудно да се направи заключение доколко стабилна и безпроблемна е употребата на песо в мн.ч., още повече че в академичния тълковен речник са отбелязани формите песоси и песи.

3 Правописът на прилагателното евро-атлантически е изключение.

Езикът може да е вкусен и извън блюдото – онзи, българският език, на който говорим от малки и на който около 24 май се кълнем в обич. А той в същността си е средство за общуване и за да ни служи добре, непрекъснато се променя. Да го погледнем в неговата динамика и да се опитаме да разберем какво става и защо, кои са движещите механизми и как те са свързани с обществените процеси. И тъй като задачата не е лека, ще го правим постепенно – на порции.

Building an “Academy of Uptime” with Kristine Lamberte

Post Syndicated from Michael Kammer original https://blog.zabbix.com/building-an-academy-of-uptime-with-kristine-lamberte/31773/

If you’ve been working with Zabbix (or are planning to), you’re in luck – we’ve recently launched Zabbix Academy, a new learning platform designed to empower IT professionals and monitoring enthusiasts with self-paced, expert-led training.

Zabbix Academy is the brainchild of Kristine Lamberte, Head of Training at Zabbix. Kristine was gracious enough to participate in a short interview where she shares the vision behind it, goes into detail about who it’s targeted at (spoiler alert – everyone!), and ruminates about the future of learning and development at Zabbix.

In the beginning: The vision behind Zabbix Academy

Was there anything in particular that inspired the creation of Zabbix Academy, and how does it fit into Zabbix’s long-term vision for community and professional development?

Zabbix itself is an extremely flexible tool, so we want to offer the same level of flexibility in our professional services. At that moment, we had enough variety in the training offer in terms of different courses for different experience levels, and we no doubt had (and still have) a great level of quality as well as theory and hands-on balance, so this was a natural next step in how we can offer even more for our users.

As for the long-term vision, Zabbix Academy supports the growth of the Zabbix ecosystem, it strengthens our training portfolio without replacing live courses, and it also demonstrates our ongoing investment in the community.

Do you see a primary audience for Zabbix Academy (beginners, professionals, enterprise clients, etc.) or is it meant to be universal?

We are ready to meet you at any stage of your Zabbix journey. The Academy has free quick-start guides in the form of free courses and webinars for those who are just starting, as well as a variety of courses on different levels – introduction, fundamental, intermediate, and advanced. And we will add new paid and free material on a regular basis for all levels.

The Zabbix Academy learning experience

How does Zabbix Academy go about keeping learning hands-on and practical for complicated monitoring scenarios?

All the people involved in the creation of training materials are Zabbix Certified Trainers and Zabbix Certified Experts with multiple years of experience. We have a deep understanding that no training material is complete without good-quality real-life use cases and practical tasks. So, it is natural that in Zabbix Academy, for all the paid courses, you will get not only high-quality theory, but also an option to do labs. Learners can experiment in a safe sandbox setting — so they’re not just reading about Zabbix, they’re actually using it.

Instructors and expertise

How do you ensure consistency and quality across so many different topics and courses?

Practice makes perfect, doesn’t it? It all comes down to the people who are behind the course creation. As I already mentioned, they are experienced Zabbix trainers and experts, but most importantly, they have hands-on experience with Zabbix.

But it is not only about our training content creators; we have close collaboration with other teams, for example, internally, support, developers, integrators, etc., and we also have extremely knowledgeable training partners who are very much involved in the review of new courses and suggestions for what’s coming.

Additionally, learner feedback plays a key role — we continuously refine and update materials based on real-world experience and community input.

Career impact

Can you give a hypothetical scenario of how Zabbix Academy could help an IT professional advance their career or an organization strengthen their monitoring practices?

Personally, I put more emphasis on what knowledge brings to the company. Knowing how to work smarter, more efficiently, use effective automations, troubleshoot faster, and come up with new ways of what and how to monitor is something that everyone should want for their business, and these things come with knowledge. What we offer is structured knowledge, packed and passed down to Zabbix users in the most effective way.

The future

What are your goals for the first year of Zabbix Academy, and how will you measure its success?

In the first year, it is crucial to continuously grow and shape Zabbix Academy. The measure of success? I mean, we created this platform for our users, so the measure of success is based on their satisfaction, engagement, and the impact this platform will have on their day-to-day tasks.

What future expansions or features can learners expect?

We aim to continuously expand the course catalogue, both free and paid content, and establish Zabbix Academy as a trusted source of knowledge for both new and existing users. In short, our goal is for Zabbix Academy to evolve into a dynamic, living resource that grows alongside Zabbix itself.

A final thought

From your perspective as Head of Training, what has been the most rewarding or challenging part of launching Zabbix Academy?

There were no challenges worth mentioning, but when it comes to the most rewarding thing, I can name a few.
First thing is that we established right from the beginning that Zabbix Academy will have all kinds of content, including free content. This once again supports our effort in strengthening our community.

Secondly, we did not compromise on the course quality; we took our classroom-quality courses and transformed them for the self-paced training.

But the most rewarding part has been seeing how excited our community is about Zabbix Academy. The feedback from early users and our partners has been overwhelmingly positive. That shows me that we are on the right track, and now we just need to keep on delivering things we are good at – great quality, hands-on content that allows Zabbix users to reach and exceed their monitoring goals.

Continue reading Building an “Academy of Uptime” with Kristine Lamberte →

A closer look at a BGP anomaly in Venezuela

Post Syndicated from Bryton Herdes original https://blog.cloudflare.com/bgp-route-leak-venezuela/

As news unfolds surrounding the U.S. capture and arrest of Venezuelan leader Nicolás Maduro, a cybersecurity newsletter examined Cloudflare Radar data and took note of a routing leak in Venezuela on January 2.

We dug into the data. Since the beginning of December there have been eleven route leak events, impacting multiple prefixes, where AS8048 is the leaker. Although it is impossible to determine definitively what happened on the day of the event, this pattern of route leaks suggests that the CANTV (AS8048) network, a popular Internet Service Provider (ISP) in Venezuela, has insufficient routing export and import policies. In other words, the BGP anomalies observed by the researcher could be tied to poor technical practices by the ISP rather than malfeasance.

In this post, we’ll briefly discuss Border Gateway Protocol (BGP) and BGP route leaks, and then dig into the anomaly observed and what may have happened to cause it. 

Background: BGP route leaks

First, let’s revisit what a BGP route leak is. BGP route leaks cause behavior similar to taking the wrong exit off of a highway. While you may still make it to your destination, the path may be slower and come with delays you wouldn’t otherwise have traveling on a more direct route.

Route leaks were given a formal definition in RFC7908 as “the propagation of routing announcement(s) beyond their intended scope.” Intended scope is defined using pairwise business relationships between networks. The relationships between networks, which in BGP we represent using Autonomous Systems (ASes), can be one of the following: 

  • customer-provider: A customer pays a provider network to connect them and their own downstream customers to the rest of the Internet

  • peer-peer: Two networks decide to exchange traffic between one another, to each others’ customers, settlement-free (without payment)

In a customer-provider relationship, the provider will announce all routes to the customer. The customer, on the other hand, will advertise only the routes from their own customers and originating from their network directly.

In a peer-peer relationship, each peer will advertise to one another only their own routes and the routes of their downstream customers. 


These advertisements help direct traffic in expected ways: from customers upstream to provider networks, potentially across a single peering link, and then potentially back down to customers on the far end of the path from their providers. 

A valid path would look like the following that abides by the valley-free routing rule: 


A route leak is a violation of valley-free routing where an AS takes routes from a provider or peer and redistributes them to another provider or peer. For example, a BGP path should never go through a “valley” where traffic goes up to a provider, and back down to a customer, and then up to a provider again. There are different types of route leaks defined in RFC7908, but a simple one is the Type 1: Hairpin route leak between two provider networks by a customer. 


In the figure above, AS64505 takes routes from one of its providers and redistributes them to their other provider. This is unexpected, since we know providers should not use their customer as an intermediate IP transit network. AS64505 would become overwhelmed with traffic, as a smaller network with a smaller set of backbone and network links than its providers. This can become very impactful quickly. 

Route leak by AS8048 (CANTV)

Now that we have reminded ourselves what a route leak is in BGP, let’s examine what was hypothesized  in the newsletter post. The post called attention to a few route leak anomalies on Cloudflare Radar involving AS8048. On the Radar page for this leak, we see this information:


We see the leaker AS, which is AS8048 — CANTV, Venezuela’s state-run telephone and Internet Service Provider. We observe that routes were taken from one of their providers AS6762 (Sparkle, an Italian telecom company) and then redistributed to AS52320 (V.tal GlobeNet, a Colombian network service provider). This is definitely a route leak. 

The newsletter suggests “BGP shenanigans” and posits that such a leak could be exploited to collect intelligence useful to government entities. 

While we can’t say with certainty what caused this route leak, our data suggests that its likely cause was more mundane. That’s in part because BGP route leaks happen all of the time, and they have always been part of the Internet — most often for reasons that aren’t malicious.

To understand more, let’s look closer at the impacted prefixes and networks. The prefixes involved in the leak were all originated by AS21980 (Dayco Telecom, a Venezuelan company):


The prefixes are also all members of the same 200.74.224.0/20 subnet, as noted by the newsletter author. Much more intriguing than this, though, is the relationship between the originating network AS21980 and the leaking network AS8048: AS8048 is a provider of AS21980. 

The customer-provider relationship between AS8048 and AS21980 is visible in both Cloudflare Radar and bgp.tools AS relationship interference data. We can also get a confidence score of the AS relationship using the monocle tool from BGPKIT, as you see here: 

➜  ~ monocle as2rel 8048 21980
Explanation:
- connected: % of 1813 peers that see this AS relationship
- peer: % where the relationship is peer-to-peer
- as1_upstream: % where ASN1 is the upstream (provider)
- as2_upstream: % where ASN2 is the upstream (provider)

Data source: https://data.bgpkit.com/as2rel/as2rel-latest.json.bz2

╭──────┬───────┬───────────┬──────┬──────────────┬──────────────╮
│ asn1 │ asn2  │ connected │ peer │ as1_upstream │ as2_upstream │
├──────┼───────┼───────────┼──────┼──────────────┼──────────────┤
│ 8048 │ 21980 │ 9.9%   │ 0.6% │ 9.4%    │ 0.0%         │
╰──────┴───────┴───────────┴──────┴──────────────┴──────────────╯

While only 9.9% of route collectors see these two ASes as adjacent, almost all of the paths containing them reflect AS8048 as an upstream provider for AS21980, meaning confidence is high in the provider-customer relationship between the two.

Many of the leaked routes were also heavily prepended with AS8048, meaning it would have been potentially less attractive for routing when received by other networks. Prepending is the padding of an AS more than one time in an outbound advertisement by a customer or peer, to attempt to switch traffic away from a particular circuit to another. For example, many of the paths during the leak by AS8048 looked like this: “52320,8048,8048,8048,8048,8048,8048,8048,8048,8048,23520,1299,269832,21980”. 

You can see that AS8048 has sent their AS multiple times in an advertisement to AS52320, because by means of BGP loop prevention the path would never actually travel in and out of AS8048 multiple times in a row. A non-prepended path would look like this: “52320,8048,23520,1299,269832,21980”. 

If AS8048 was intentionally trying to become a man-in-the-middle (MITM) for traffic, why would they make the BGP advertisement less attractive instead of more attractive? Also, why leak prefixes to try and MITM traffic when you’re already a provider for the downstream AS anyway? That wouldn’t make much sense. 

The leaks from AS8048 also surfaced in multiple separate announcements, each around an hour apart on January 2, 2026 between 15:30 and 17:45 UTC, suggesting they may have been having network issues that surfaced in a routing policy issue or a convergence-based mishap. 


It is also noteworthy that these leak events begin over twelve hours prior to the U.S. military strikes in Venezuela. Leaks that impact South American networks are common, and we have no reason to believe, based on timing or the other factors I have discussed, that the leak is related to the capture of Maduro several hours later.

In fact, looking back the past two months, we can see plenty of leaks by AS8048 that are just like this one, meaning this is not a new BGP anomaly:


You can see above in the history of Cloudflare Radar’s route leak alerting pipeline that AS8048 is no stranger to Type 1 hairpin route leaks. Since the beginning of December alone there have been eleven route leak events where AS8048 is the leaker.

From this we can draw a more innocent possible explanation about the route leak: AS8048 may have configured too loose of export policies facing at least one of their providers, AS52320. And because of that, redistributed routes belong to their customer even when the direct customer BGP routes were missing. If their export policy toward AS52320 only matched on IRR-generated prefix list and not a customer BGP community tag, for example, it would make sense why an indirect path toward AS6762 was leaked back upstream by AS8048. 

These types of policy errors are something RFC9234 and the Only-to-Customer (OTC) attribute would help with considerably, by coupling BGP more tightly to customer-provider and peer-peer roles, when supported by all routing vendors. I will save the more technical details on RFC9234 for a follow-up blog post.

The difference between origin and path validation

The newsletter also calls out as “notable” that Sparkle (AS6762) does not implement RPKI (Resource Public Key Infrastructure) Route Origin Validation (ROV). While it is true that AS6762 appears to have an incomplete deployment of ROV and is flagged as “unsafe” on isbgpsafeyet.com because of it, origin validation would not have prevented this BGP anomaly in Venezuela. 

It is important to separate BGP anomalies into two categories: route misoriginations, and path-based anomalies. Knowing the difference between the two helps to understand the solution for each. Route misoriginations, often called BGP hijacks, are meant to be fixed by RPKI Route Origin Validation (ROV) by making sure the originator of a prefix is who rightfully owns it. In the case of the BGP anomaly described in this post, the origin AS was correct as AS21980 and only the path was anomalous. This means ROV wouldn’t help here.

Knowing that, we need path-based validation. This is what Autonomous System Provider Authorization (ASPA), an upcoming draft standard in the IETF, is going to provide. The idea is similar to RPKI Route Origin Authorizations (ROAs) and ROV: create an ASPA object that defines a list of authorized providers (upstreams) for our AS, and everyone will use this to invalidate route leaks on the Internet at various vantage points. Using a concrete example, AS6762 is a Tier-1 transit-free network, and they would use the special reserved “AS0” member in their ASPA signed object to communicate to the world that they have no upstream providers, only lateral peers and customers. Then, AS52320, the other provider of AS8048, would see routes from their customer with “6762” in the path and reject them by performing an ASPA verification process.

ASPA is based on RPKI and is exactly what would help prevent route leaks similar to the one we observed in Venezuela.

A safer BGP, built together 

We felt it was important to offer an alternative explanation for the BGP route leak by AS8048 in Venezuela that was observed on Cloudflare Radar. It is helpful to understand that route leaks are an expected side effect of BGP historically being based entirely on trust and carefully-executed business relationship-driven intent. 

While route leaks could be done with malicious intent, the data suggests this event may have been an accident caused by a lack of routing export and import policies that would prevent it. This is why to have a safer BGP and Internet we need to work together and drive adoption of RPKI-based ASPA, for which RIPE recently released object creation, on the wide Internet. It will be a collaborative effort, just like RPKI has been for origin validation, but it will be worth it and prevent BGP incidents such as the one in Venezuela. 

In addition to ASPA, we can all implement simpler mechanisms such as Peerlock and Peerlock-lite as operators, which sanity-checks received paths for obvious leaks. One especially promising initiative is the adoption of RFC9234, which should be used in addition to ASPA for preventing route leaks with the establishing of BGP roles and a new Only-To-Customer (OTC) attribute. If you haven’t already asked your routing vendors for an implementation of RFC9234 to be on their roadmap: please do. You can help make a big difference.

AMD Announces Zen 5-based Ryzen AI Embedded P100 Family of Chips

Post Syndicated from Ryan Smith original https://www.servethehome.com/amd-announces-zen-5-based-ryzen-ai-embedded-p100-family-of-chips/

Alongside AMD’s slew of consumer-related product announcements with desktop and mobile Ryzen processors, the company also took a moment of its time during its CES 2026 presentation to address the embedded market. The tangentially-related cousin to the consumer market, AMD’s embedded lineup of processors are aimed at the automotive and industrial markets, as well as […]

The post AMD Announces Zen 5-based Ryzen AI Embedded P100 Family of Chips appeared first on ServeTheHome.

AMD Reveals New Ryzen AI 400 Series, Ryzen AI Max+, and Ryzen 7 9850X3D Chips At CES 2026

Post Syndicated from Ryan Smith original https://www.servethehome.com/amd-reveals-new-ryzen-ai-400-series-ryzen-ai-max-and-ryzen-7-9850x3d-chips-at-ces-2026/

Kicking off AMD’s slate of consumer announcements for this year’s CES trade show, the company brought to the shows several new chip SKUs for the consumer market. Altogether the company is launching two new Ryzen AI Max+ processors for AI developers, a new desktop Ryzen 9000X3D chip for gamers, and for the mobile market a […]

The post AMD Reveals New Ryzen AI 400 Series, Ryzen AI Max+, and Ryzen 7 9850X3D Chips At CES 2026 appeared first on ServeTheHome.

The collective thoughts of the interwebz