The Attack Cycle is Accelerating: Announcing the Rapid7 2026 Global Threat Landscape Report

Post Syndicated from Rapid7 Labs original https://www.rapid7.com/blog/post/tr-accelerating-attack-cycle-2026-global-threat-landscape-report

The predictive window has collapsed.

In 2025, high-impact vulnerabilities weren’t quietly accumulating risk. They were operationalized, and often within days.

Today, Rapid7 Labs released the 2026 Global Threat Landscape Report, an in-depth analysis of how attacker behavior is evolving across vulnerability exploitation, ransomware operations, identity abuse, and AI-driven tradecraft. The data shows a clear pattern: exposure is being identified and weaponized faster than most organizations are set up to defend.

From disclosure to exploitation in days, not weeks

In 2025, confirmed exploitation of newly disclosed CVSS 7–10 vulnerabilities increased 105% year over year, rising from 71 to 146. The median time from publication to inclusion in CISA’s Known Exploited Vulnerabilities list fell from 8.5 days to 5.0 days.

At the same time, the number of high-probability vulnerabilities that remained unexploited dropped sharply. The buffer that once allowed teams to triage and schedule remediation is shrinking to the point where some severe flaws were seen to have been exploited almost immediately.

The broader trend is unmistakable: vulnerability management programs built around reactive remediation cycles are struggling to keep pace with adversaries operating at machine speed.

Cybercrime as a structured market

Cybercrime in 2025 no longer resembles chaotic hacking. It resembles platform capitalism.

The report highlights how the underground economy now mirrors legitimate SaaS ecosystems. Initial Access Brokers obtain and validate network footholds. Ransomware operators focus on encryption and extortion. Infostealer operators sell subscription-style access to fresh credential logs.

This specialization lowers barriers to entry and increases scale creating a supply chain in which access is acquired, packaged, priced, and sold to anyone who wants it. 

Ransomware is a good example of this business maturity. It was present in 42% of Rapid7 MDR investigations in 2025 with leak posts increasing 46.4% year over year, and the number of active groups growing from 102 to 140. That kind of growth is anything but random or coincidental: it is an indication of systemic changes to the ransomware ecosystem indicating growing sophistication, specialization, and, ultimately, risk. 

Logging in, not breaking in

Authentication-based attacks remain incredibly common as the lack of consistency across organizations can lead to easy exploitation. Valid accounts without multi-factor authentication (MFA) were responsible for 43.9% of incidents over that year. Rather than forcing their way past defenses, attackers increasingly authenticate with stolen credentials, hijacked sessions, or abused tokens. This is where the increase in AI-driven attacks is particularly acute with the benefits generative AI can play in improving the maturity and sophistication of social engineering attacks. 

As enterprises extend trust across cloud platforms, SaaS ecosystems, APIs, and remote work environments, authentication systems have become the backbone of operational control. This represents a structural shift with the control layer of cyber risk moving away from network perimeters toward authentication flows.

Attacks are using reliable vectors, just at alarming speeds

One hallmark of the attack landscape in 2025 was the use of tried and true attack vectors rather than novel exploits and zero-day vulnerabilities. CVE disclosures continued to climb last year, but confirmed exploitation clustered around dependable weakness types like deserialization, authentication bypass, and memory corruption vulnerabilities.

Attackers are targeting flaws that enable pre-authentication access, repeatable execution, and rapid data theft. They are not, necessarily, chasing every vulnerability. Just the ones they deem reliable. This pattern reinforces a key theme of the report: exploitability and context matter more than raw volume.

AI as an accelerant

AI is serving as a force multiplier and an expanding attack surface at the same time. 

Generative AI is accelerating established attack methods by reducing the time, skill, and coordination previously required to execute them at scale. Rather than introducing entirely new categories of exploitation, threat actors are integrating AI into existing workflows to industrialize phishing, automate reconnaissance, and refine malicious scripts with greater speed and precision. 

AI-assisted phishing campaigns were more polished and tailored to specific industries or executive roles, reflecting a measurable improvement in personalization and believability. They accelerated open-source intelligence collection to create details from fragmented data. AI was used to troubleshoot malware development in near real time, effectively compressing the cycle between initial research and malware deployment. The result is not radical technical innovation, but efficiency, speed, and fewer missed opportunities. 

Meanwhile, AI platforms themselves are emerging as targets with model servers, orchestration frameworks, and token-based integrations, inheriting familiar weaknesses such as unsafe deserialization and weak authentication. As organizations operationalize AI quickly, governance gaps create new high-impact pathways to risk.

The geography of attacks

When it comes to targeted regions, no area of the globe represents a better convergence of exposure and financial opportunity than North America. Organizations on this continent accounted for 82.04% of observed incidents, with the United States representing roughly 70% of leak posts on ransomware leak sites. Manufacturing, business services, and retail were among the most targeted industries as these sectors often combine operational dependence, sensitive data, and financial leverage making them fat targets for attackers looking for reliability not only in their attack vectors, but in gains available from their chosen targets. 

Across criminal and state-aligned activity, attackers are converging on identity systems, edge infrastructure, collaboration platforms, and cloud control planes where trust, scale, and business continuity intersect.

What this means for security leaders

There is a sobering reality in this year’s data: the underlying weaknesses remain familiar. Weak credentials. Social engineering. Exposed services. Unpatched edge infrastructure.

What has changed is the speed.

Security programs can no longer rely on moving slightly faster than attackers. The model must shift toward reducing exposure before it is operationalized.

That means:

  • Continuous exposure visibility with contextual prioritization

  • Strong MFA enforcement and hardened identity controls

  • Protected and monitored edge infrastructure

  • Governance around AI systems and integrations

  • AI-enabled security workflows capable of matching attacker velocity

The organizations that maintain clear, continuous insight into their exposure – and reduce it before it is monetized – will be best positioned to manage risk in this accelerated cycle.

The question is no longer whether exposure exists.
It is whether you can reduce it before attackers capitalize on it.

Read the full Rapid7 2026 Threat Landscape Report to explore the data and strategic implications in detail.

Introducing Custom Regions for precision data control

Post Syndicated from Andrew Berglund original https://blog.cloudflare.com/custom-regions/

A key part of our mission to help build a better Internet is giving our customers the tools they need to operate securely and efficiently, no matter their compliance requirements. Our Regional Services product helps customers do just that, allowing them to meet data sovereignty legal obligations using the power of Cloudflare’s global network.

Today, we’re taking two major steps forward: First, we’re expanding the pre-defined regions for Regional Services to include Turkey, the United Arab Emirates (UAE), IRAP (Australian compliance) and ISMAP (Japanese compliance). Second, we’re introducing the next evolution of our platform: Custom Regions.

Global security, local compliance: the Regional Services advantage

Before we dive into what’s new, let’s revisit how Regional Services provides the best of both worlds: local compliance and global-scale security. Our approach is fundamentally different from many sovereign cloud providers. Instead of isolating your traffic to a single geography (and a smaller capacity for attack mitigation), we leverage the full scale of our global network for protection and only inspect your data where you tell us to.

Here’s an overview of how it works:

  1. Global ingestion & L3/L4 DDoS defense: Traffic is ingested at the closest Cloudflare data center, wherever in the world that may be. At this initial entry point, we apply our massive-scale DDoS mitigation to block volumetric attacks at the network and transport layers. This happens outside your designated region, ensuring only clean traffic is forwarded.

  2. Intelligent in-region routing: Before any decryption occurs, we inspect the request’s metadata. If it has arrived at a data center outside your specified region, we route it across our secure, private backbone to a data center within your boundaries, using the most performant pathway.

  3. In-region TLS termination & L7 processing: Only once the traffic is confirmed to be within your chosen region do we decrypt the request. It is only then that we apply our application-layer security services, like our Web Application Firewall (WAF) or Bot Management, and execute any Cloudflare Workers logic.

  4. Secure transit to origin: Once processed, the request is re-encrypted and securely sent to your origin server.

This unique architecture means you can localize data inspection as needed to meet your legal obligations without sacrificing the robust DDoS protection that only a massive global network can provide.

New options available within Cloudflare Managed Regions

When we launched Regional Services in 2020, we started with just three regions: EU, UK, and U.S. Over time we have added regions that are shared across all accounts — we refer to these as Cloudflare Managed Regions.

A few more are newly available: Turkey, the United Arab Emirates (UAE), and IRAP (Australian compliance), bringing our total to 35 regions.

In addition, we are now giving our customers the ability to request a custom region that meets their account needs. These are Custom Regions, launching today.

Beyond pre-defined boundaries: introducing Custom Regions

While our 35 pre-defined regions serve many of our customers’ needs, the digital world isn’t one-size-fits-all. We’ve heard you loud and clear: you’ve asked for a specific country, unique combinations of countries, and the ability to exclude a set of countries from a region.

That’s why we’re excited to announce the next evolution of Regional Services: Custom Regions.

Simply put, Custom Regions give you the power to define your own geographical boundaries for traffic processing. Instead of choosing from a list of regions defined by us, you tell us precisely which locations constitute your region.

This flexibility unlocks a new level of control. Our early-access customers have already used Custom Regions to:

  • Regionalize AI inference: Keep LLM prompts and responses within a specific set of countries to optimize for performance and data localization legal obligations.

  • Launch hyper-targeted promotions: Serve marketing campaigns and content that are optimized for a unique combination of countries.

  • Scale government operations: Build regions that align with contractual commitments with government entities.

  • Mirror your corporate structure: Build regions that match your internal business units, like EMEA, MENA, or APAC, for perfectly aligned governance.

The core mechanism is the same; the only thing that changes is the boundary. Instead of Cloudflare defining the region, you do.

The possibilities are endless. For example, your region could be:

  • North America: Canada, United States, Mexico

  • Everywhere except North America: Not Canada, not United States, not Mexico

  • Countries that use Fahrenheit: USA, Bahamas, Cayman Islands, Marshall Islands, Liberia

How Regional Services works

At the core of Regional Services is enforcement of a simple rule: TLS termination and Layer 7 processing only happen inside your chosen region. Custom Regions expands this capability by allowing you to choose your own region definitions.

Cloudflare Managed Regions and Custom Regions rely on three building blocks: defining region membership, selecting an in-region destination, and enforcing the boundary at the edge.

Defining region membership

A region is ultimately a set of Cloudflare data centers.

  • Cloudflare managed regions use a pre-defined membership set.

  • Custom Regions define membership with an expression. The most common field is country_code: the ISO code where each data center is located:

Use case

Expression

Definition

Single country

country_code == "TR"

Turkey

Multiple countries

country_code in ["DE", "FR", "NL"]

Germany, France, and the Netherlands

Exclude countries

!(country_code in ["US", "CA", "MX"])

Everything except the U.S., Canada, and Mexico

That expression is evaluated against data centers’ metadata. Matches become your region’s membership set and are distributed globally, so every data center can quickly answer: “Am I in this region?”

As Cloudflare’s infrastructure evolves, membership updates, so new matching data centers can join automatically. You do not need to worry about when data centers are added or removed from the definition; Cloudflare takes care of that for you. 

Calculating optimal in-region routing

If a request enters Cloudflare outside your region, the next step is choosing the best in-region destination for that ingress location.

Cloudflare’s selection is a two-step process:

  1. Allowed destinations: the region’s membership set (which data centers are in-region)

  2. Best destination for this ingress: a performance-ranked list tailored to the data center where the request entered our network

These per-ingress rankings are computed centrally and distributed to the edge via Quicksilver. They are built from measured path quality across our network (not just physical distance), using signals like:

  • Network performance: Latency and reliability indicators (for example, loss and timeouts)

  • Capacity and load: Available resources and current utilization

  • Operational status: Health and availability

At routing time, we intersect the ranked list with the region membership set and choose from the top candidates. The final choice is validated against live availability: destinations that are disabled or otherwise unreachable are skipped, so traffic can fail over to the next best in-region option.

Enforcing the boundary

This is the process when a request arrives at Cloudflare:

  1. Ingress. The request lands at the nearest data center. Layer 3/4 DDoS mitigation is applied immediately.

  2. Configuration lookup. Is a region configured for this zone?

  3. Membership check. Is this data center in the configured region?

  4. Routing decision.

    • In region: Process locally. TLS termination and all Layer 7 services run here.

    • Out of region: An in-region data center is selected, and the request is forwarded over Cloudflare’s private backbone.

  5. In-region processing. TLS is terminated for the first time. Layer 7 services run here.

  6. Origin connection. The processed request is sent to your origin.

As noted above, Cloudflare does not decrypt the request outside your defined region. Instead, we forward it to the closest data center inside your region, where decryption and Layer 7 services occur. 

How we handle errors

Resilience is built in at multiple layers:

  • Multiple candidates: Routing considers multiple in-region options and selects an available destination in real time.

  • Health-aware routing: Unhealthy or disabled data centers are excluded.

  • Data quality gates: Fresh routing inputs are only published when sufficient monitoring data is available. 

  • Fail-close design: If no valid in-region destination exists, the connection fails rather than processing outside your region.


How to get started

The new Cloudflare managed regions are available now for customers using Regional Services. If you would like to use these, just follow the standard process to enable it via the Cloudflare Dashboard or via the Cloudflare API. Custom Regions are new and follow a different process.

To ensure a perfect fit for your needs, the initial setup for Custom Regions is a collaborative process. To get started, simply reach out to your account team. They will work with you to define your region and get it deployed. While the service is not yet self-serve, we are continuously developing the technology and will revisit this as the feature matures. Please note that some technical limitations may apply, and your solutions engineer is the perfect person to discuss the details with.

Interested in taking control of your data?

If you are interested in learning more about Regional Services, please contact your account team. If you’re not yet a Cloudflare customer, we would love to have you. Fill out this form, and we’ll be in touch with you soon.

Port Monitoring with the Zabbix Widget Switch

Post Syndicated from Patrik Uytterhoeven original https://blog.zabbix.com/port-monitoring-with-the-zabbix-widget-switch/32603/

If you’ve ever monitored network switches in Zabbix, you know perfectly well that the data is there. Interfaces are polled via SNMP, triggers fire when a port goes down, and events are logged. Technically, everything works.

But when you open a 24- or 48-port switch in Zabbix, you’re usually looking at a list of interface items or triggers. It’s accurate, but not visual. You still need to read through interface names to understand what’s happening. During an incident, that costs time.

Of course, you can create network maps for each switch, draw every port on it etc, but that takes a lot of time and is tedious if you have many devices.

That’s why we created the Zabbix Widget Switch.

Instead of presenting interface states as text, the widget renders a visual representation of the switch directly inside a Zabbix dashboard. Each port is displayed as a graphical element, color-coded according to its operational state.

Green means up.

Red means down.

Grey means disabled or unused.

You immediately see the health of the entire switch — without scrolling, filtering, or interpreting tables.

Designed for real environments

This widget is not just a static visual block. It was designed with real operational use in mind.

You can:

  • Reuse shareable switch profiles across devices
  • Save profile presets directly from the edit form
  • Configure rows, ports per row, SFP count, and port index start
  • Support mixed RJ45 + SFP layouts with realistic placement
  • Add port labels (uplinks, APs, user links, etc.)
  • Show a utilization heatmap overlay with configurable thresholds and colors
  • Display a live panel with IN/OUT sparklines, current utilization, 24h online state bar, and 24h errors/discards bars and trend summaries
  • Configure traffic/error/discard/speed item patterns
  • Use item-key suggestions in the edit UI for faster setup
  • Choose traffic unit display (B/s or bps)
  • Show switch summary context (CPU, uptime, serial, software, VLANs, monitoring state, maintenance badge)

This makes it practical for real-world deployments,  not just lab environments.

If you manage dozens of similar switches, profiles save time and keep dashboards consistent.

If you run a NOC screen, the legend ensures clarity for everyone.

If you monitor mixed copper and fiber ports, SFP support makes the layout realistic.

It adapts to how your network is built.

Why it matters

Monitoring is about reducing reaction time.

When a user says, “My connection dropped,” you don’t want to search through interface lists. You want to open the dashboard and immediately see that port 17 is red.

In NOC environments, a visual layout is even more powerful. A quick glance at the screen tells you whether everything is healthy or if some uplink needs attention.

And during maintenance windows, a before and after check becomes an instant visual validation.

How it looks

A visual overview like this makes the state of a switch immediately obvious, even from across the room.

Open source

The Zabbix Widget Switch is open source and available for Zabbix 7.0 on GitHub:

https://github.com/OpensourceICTSolutions/zabbix-widget-switch

Feel free to test it, adapt it, or contribute.

I hope this widget will be helpful! If you have any questions or need help configuring anything on your Zabbix setup feel free to contact us  at Opensource ICT Solutions.

Patrik Uytterhoeven

https://oicts.com

Disclaimer: Parts of this software were generated using Codex. We do not guarantee the total accuracy, security, or stability of the generated code.

 

The post Port Monitoring with the Zabbix Widget Switch appeared first on Zabbix Blog.

Building a scalable, transactional data lake using dbt, Amazon EMR, and Apache Iceberg

Post Syndicated from Umesh Pathak original https://aws.amazon.com/blogs/big-data/building-a-scalable-transactional-data-lake-using-dbt-amazon-emr-and-apache-iceberg/

Growing data volume, variety, and velocity has made it crucial for businesses to implement architectures that efficiently manage and analyze data, while maintaining data integrity and consistency. In this post, we show you a solution that combines Apache Iceberg, Data Build Tool (dbt), and Amazon EMR to create a scalable, ACID-compliant transactional data lake. You can use this data lake to process transactions and analyze data simultaneously while maintaining data accuracy and real-time insights for better decision-making.

Challenges, business imperatives, and technical advantages

Traditional data lakes have long struggled with fundamental limitations. For example, the lack of ACID compliance, data inconsistencies from concurrent writes, complex schema evolution, and the absence of time travel, rollback, and versioning capabilities. These shortcomings directly conflict with growing business demands for concurrent read/write support, robust data versioning and auditing, schema flexibility, and transactional capability within data lake environments. To address these gaps, modern solutions use ACID transactions at scale, optimized storage formats through Apache Iceberg, version control for data on Amazon Simple Storage Service (Amazon S3), and cost-effective, streamlined maintenance—delivering a reliable, enterprise-grade data lake architecture that meets both operational and analytical needs.

Solution overview

The solution is built around four tightly integrated layers that work together to deliver a scalable, transactional data lake.

Raw data is ingested and stored in Amazon S3, which serves as the foundational storage layer. This layer supports multiple data formats and enables efficient data partitioning through Apache Iceberg’s table format. This ensures that data is organized and accessible from the moment it lands. Then, Amazon EMR takes over as the distributed computing engine, using Apache Spark to process large-scale datasets in parallel, handling the heavy lifting of reading, transforming, and writing data across the lake.

Sitting within the processing layer, dbt drives the transformation logic. It applies SQL-based, version-controlled transformations that convert raw, unstructured data in the S3 raw layer into clean, curated datasets stored back in S3. This maintains ACID compliance and schema consistency throughout.

Finally, the curated data is available for consumption through Amazon Athena, which provides a serverless, one-time querying capability directly on S3. With this, analysts and business users can run interactive SQL queries without managing any infrastructure. Together, these components form a continuous pipeline: data flows from ingestion through distributed processing and structured transformation, ultimately surfacing as reliable, query-ready insights.

Amazon EMR is a cloud-based big data service that streamlines the deployment and management of open source frameworks like Apache Spark, Hive, and Trino. It provides a managed Apache Hadoop environment that organizations can use to process and analyze vast amounts of data efficiently.

Data Build Tool is an open source tool that data teams can use to transform and model data using SQL. It promotes best practices for data modeling, testing, and documentation, streamlining maintenance and collaboration on data pipelines.

Apache Iceberg is an open table format designed for large-scale analytics on data lakes. It supports features like transactions, time travel, and data partitioning, which are essential for building reliable and performant data lakes. By using Iceberg, organizations can maintain data integrity and enable efficient querying and processing of data.

When combined, these three technologies provide a powerful solution for building transactional data lakes. Amazon EMR provides the scalable and managed infrastructure for running big data workloads, dbt enables efficient data modeling and transformation, and Apache Iceberg provides data consistency and reliability within the data lake.

Prerequisites

Before proceeding with the solution walkthrough, make sure that the following are in place:

  • AWS Account – An active AWS account with sufficient permissions to create and manage EMR clusters, S3 buckets, Athena workgroups, and AWS Glue Data Catalog resources
  • IAM Roles – The following IAM roles must exist and have appropriate permissions:
    • EMR_DefaultRole – Service role for Amazon EMR
    • EMR_EC2_DefaultRole – Amazon Elastic Compute Cloud (Amazon EC2) instance profile for EMR nodes
  • AWS Command Line Interface (AWS CLI) – Installed and configured with credentials for your target AWS account and AWS Region (refer to Step 1.1 for setup instructions)
  • Python 3.8+ – Installed on your local machine or workspace for setting up the dbt virtual environment
  • Pip – Python package manager available for installing dbt and its dependencies
  • Git – Installed on the EMR primary node or local environment for version control and dbt package management
  • Amazon Athena – Athena query editor access with a configured S3 output location for query results
  • AWS Glue Data Catalog – Enabled as the metastore for Amazon EMR and Athena (no additional setup required if using the default AWS Glue integration)
  • S3 Bucket Naming – Prepare a unique identifier to suffix S3 bucket names, ensuring global uniqueness across all three buckets created in Step 1.3
  • Network Access – Make sure that your local machine can reach the Amazon EMR primary node’s DNS over port 10001 (Thrift/HiveServer2) for dbt connectivity; configure security groups accordingly

Solution walkthrough

Step 1: Environment setup

  1. Install the AWS CLI on your workspace by following the instructions in Installing or updating the latest version of the AWS CLI. To configure AWS CLI interaction with AWS, refer to Quick setup.
  2. Create EMR cluster.

    Create the following JSON file with the following contents emr-config.json:

    [
      {
        "Classification": "iceberg-defaults",
        "Properties": {
          "iceberg.enabled": "true"
        }
      },
      {
        "Classification": "spark-hive-site",
        "Properties": {
          "hive.metastore.client.factory.class": "com.amazonaws.glue.catalog.metastore.AWSGlueDataCatalogHiveClientFactory"
        }
      }
    ]

    Run the following command on your AWS CLI, updating the preferred AWS Region:

    aws emr create-cluster \
    --name "Iceberg-DBT-Cluster" \
    --release-label emr-7.7.0 \
    --applications Name=Spark Name=Hive Name=Livy \
    --ec2-attributes InstanceProfile=EMR_EC2_DefaultRole \
    --instance-type c3.4xlarge \
    --instance-count 1 \
    --service-role EMR_DefaultRole \
    --configurations file://emr-config.json \
    --region [region]

  3. Set up S3 buckets.
    Create the following S3 bucket using the AWS CLI after updating the bucket name.

    aws s3 mb s3://amzn-s3-demo-transactional-datalake-raw-[unique-identifier]
    aws s3 mb s3://amzn-s3-demo-transactional-datalake-curated-[unique-identifier]
    aws s3 mb s3://amzn-s3-demo-transactional-datalake-analytics-[unique-identifier]
    

Step 2. Raw layer implementation

The raw layer serves as the foundation of our data lake, ingesting and storing data in its original form. This layer is important for maintaining data lineage and enabling reprocessing if needed. We use Apache Iceberg tables to store our raw data, which provides benefits such as ACID transactions, schema evolution, and time travel capabilities.

In this step, we create a dedicated database for our raw data and set up tables for customers, products, and sales using Amazon Athena. These tables are configured to use the Iceberg table format and are compressed using the ZSTD algorithm to optimize storage. The LOCATION property specifies where the data will be stored in Amazon S3 so that data is organized and accessible.

After creating the tables, we insert sample data to simulate real-world scenarios. We use this data throughout the rest of the implementation to demonstrate the capabilities of our data lake architecture.

Update the respective bucket name in each create table bucket name from the previous step:

  1. Create database and tables
    -- Create Raw Database
    CREATE SCHEMA raw_sales_analytics_data_layer;
    
    -- Create Customers Table
    CREATE TABLE raw_sales_analytics_data_layer.customers (
        CustomerID string,
        CustomerName string,
        Region string,
        inserted_timestamp timestamp
    )
    LOCATION 's3://[bucket_name]/raw_sales_analytics_data_layer/customers'
    TBLPROPERTIES (
        'table_type'='iceberg', 
        'write_compression'='zstd'
    );
    
    -- Create Products Table
    CREATE TABLE raw_sales_analytics_data_layer.products (
        productid string,
        productname string,
        category string,
        supplier string,
        inserted_timestamp timestamp
    )
    LOCATION 's3://[bucket_name]/raw_sales_analytics_data_layer/products'
    TBLPROPERTIES (
        'table_type'='iceberg', 
        'write_compression'='zstd'
    );

  2. Insert sample data
    -- Insert Customers
    INSERT INTO raw_sales_analytics_data_layer.customers
    VALUES 
        ('201', Jane Doe', 'Central', current_timestamp),
        ('202', Arnav Desai, 'North', current_timestamp),
        ('203', Kwaku Mensah, 'West', current_timestamp);
    
    -- Insert Products
    INSERT INTO raw_sales_analytics_data_layer.products
    VALUES
        ('1', 'Laptop', 'Electronics', 'AnyAuthority', current_timestamp),
        ('2', 'Smartphone', 'Electronics', 'AnyCompany', current_timestamp);
    
    -- Insert Sales
    INSERT INTO raw_sales_analytics_data_layer.sales
    VALUES
        ('ORD001', '1', '201', '2025-04-01', 1299.99, current_timestamp),
        ('ORD002', '2', '202', '2025-04-02', 899.99, current_timestamp);

Step 3: dbt setup and configuration

Setting up dbt involves installing the necessary packages, configuring the connection to the data warehouse (in this case, Amazon EMR), and setting up the project structure.

We start by creating a Python virtual environment to isolate our dbt installation. Then, we install dbt-core and the Spark adapter, which allows dbt to connect to the EMR cluster. The profiles.yml file is configured to connect to the EMR cluster using the Thrift protocol, while the dbt_project.yml file defines the overall structure of the dbt project, including model materialization strategies and file formats.

  1. Install prerequisites
    # Create Python virtual environment
    python -m venv dbt-env
    source dbt-env/bin/activate
    
    # Install required packages
    pip install dbt-core dbt-spark[PyHive]
    
    # Install git
    yum install git

  2. Configure dbt profiles
    # ~/.dbt/profiles.yml
    sales_analytics:
      target: dev
      outputs:
        dev:
          type: spark
          method: thrift
          host: your-emr-master-dns
          port: 10001
          schema: curated_sales_analytics_data_layer
          threads: 4

  3. Project configuration
    # dbt_project.yml
    name: 'sales_analytics'
    version: '1.0.0'
    config-version: 2
    
    profile: 'sales_analytics'
    
    model-paths: ["models"]
    analysis-paths: ["analyses"]
    test-paths: ["tests"]
    seed-paths: ["seeds"]
    macro-paths: ["macros"]
    
    target-path: "target"
    clean-targets:
        - "target"
        - "dbt_packages"
    
    models:
      sales_analytics:
        dim:
          +materialized: table
          +file_format: iceberg
        ads:
          +materialized: table
          +file_format: iceberg

Step 4: dbt models implementation

In this step, we implement dbt models, which define the transformations that we will apply to raw data. We start by configuring data sources in the sources.yml file, which allows dbt to reference raw tables easily.

We then create dimension models for customers and products, and a fact model for sales.

These models use incremental materialization strategies to efficiently update data over time. The incremental strategy processes only new or updated records, significantly reducing the time and resources required for each run.

  1. Source configuration
    # models/sources.yml
    version: 2
    sources:
      - name: raw_sales
        database: raw_sales_analytics_data_layer
        schema: raw_sales_analytics_data_layer
        tables:
          - name: customers
            columns:
              - name: CustomerID
                tests:
                  - unique
                  - not_null
          - name: products
          - name: sales

  2. Dimension models
    -- models/dim/dim_customers.sql
    {{ config(
        materialized='incremental',
        unique_key='customerid',
        incremental_strategy='merge'
    ) }}
    
    WITH source_data AS (
        SELECT 
            customerid,
            customername,
            region,
            inserted_timestamp,
            ROW_NUMBER() OVER (
                PARTITION BY customerid 
                ORDER BY inserted_timestamp DESC
            ) as row_number
        FROM {{ source('raw_sales_analytics_data_layer', 'customers') }}
        {% if is_incremental() %}
        WHERE inserted_timestamp > (SELECT MAX(inserted_timestamp) FROM {{ this }})
        {% endif %}
    )
    
    SELECT 
        customerid,
        customername,
        region,
        inserted_timestamp
    FROM source_data
    WHERE row_number = 1

  3. Product models
    -- models/dim/dim_products.sql
    {{ config(
        materialized='incremental',
        unique_key='productid',
        incremental_strategy='merge'
    ) }}
    
    WITH source_data AS (
        SELECT
            productid,
            productname,
            category,
            supplier,
            inserted_timestamp,
            ROW_NUMBER() OVER (
                PARTITION BY productid
                ORDER BY inserted_timestamp DESC
            ) as row_number
        FROM {{ source('raw_sales_analytics_data_layer', 'products') }}
        {% if is_incremental() %}
        WHERE inserted_timestamp > (SELECT MAX(inserted_timestamp) FROM {{ this }})
        {% endif %}
    )
    
    SELECT
        s.productid,
        s.productname,
        s.category,
        s.supplier,
        s.inserted_timestamp
    FROM source_data s
    WHERE s.row_number = 1
    {% if is_incremental() %}
        AND NOT EXISTS (
            SELECT 1
            FROM {{ this }} t
            WHERE t.productid = s.productid
            AND t.inserted_timestamp >= s.inserted_timestamp
        )
    {% endif %}

  4. Fact models
    -- models/dim/fact_sales.sql
    {{ config(
        materialized='incremental',
        unique_key='orderid',
        incremental_strategy='merge'
    ) }}
    
    WITH source_data AS (
        SELECT
            orderid,
            productid,
            customerid,
            date,
            salesamount,
            inserted_timestamp,
            ROW_NUMBER() OVER (
                PARTITION BY orderid
                ORDER BY inserted_timestamp DESC
            ) as row_number
        FROM {{ source('raw_sales_analytics_data_layer', 'sales') }}
        {% if is_incremental() %}
        WHERE orderid NOT IN (SELECT orderid FROM {{ this }})  -- Changed condition
        {% endif %}
    )
    
    SELECT
        s.orderid,
        s.productid,
        s.customerid,
        s.date,
        s.salesamount,
        s.inserted_timestamp
    FROM source_data s
    WHERE s.row_number = 1

Step 5: Analytics layer

The analytics layer builds upon dimension and fact models to create more complex analyzes. In this step, we create a daily sales analysis model that combines data from fact_sales, dim_customers, and dim_products models.

We also implement a customer insights model that analyzes purchase patterns across different Regions and product categories.

These analytics models demonstrate how we can use our transformed data to generate valuable business insights. By materializing these models as Iceberg tables, we make sure that they benefit from the same ACID transactions and time travel capabilities as our raw and transformed data.

  1. Daily sales analysis

    The analytics layer introduces a fact_sales_analysis model that consolidates transactional sales data with customer and product dimensions to enable business-ready reporting. Built as an incremental model with a merge strategy, it efficiently processes data by deduplicating records using the latest inserted timestamp per order, enabling reliable downstream consumption without full table refreshes.

    -- models/ads/fact_sales_analysis.sql
    {{ config(
        materialized='incremental',
        unique_key='orderid',
        incremental_strategy='merge'
    ) }}
    
    WITH source_data AS (
        SELECT
            s.orderid,
            s.date,
            s.salesamount,
            c.customername,
            c.region,
            p.productname,
            p.category,
            p.supplier,
            s.inserted_timestamp,
            ROW_NUMBER() OVER (
                PARTITION BY s.orderid
                ORDER BY s.inserted_timestamp DESC
            ) as row_number
        FROM {{ ref('fact_sales') }} s
        JOIN {{ ref('dim_customers') }} c ON s.customerid = c.customerid
        JOIN {{ ref('dim_products') }} p ON s.productid = p.productid
        {% if is_incremental() %}
        WHERE s.orderid NOT IN (SELECT orderid FROM {{ this }})
        {% endif %}
    )
    
    SELECT
        s.orderid,
            s.date,
            s.salesamount,
            s.customername,
            s.region,
            s.productname,
            s.category,
            s.supplier,
            s.inserted_timestamp
    FROM source_data s
    WHERE s.row_number = 1

  2. Customer insights

    The customer_purchase_patterns model aggregates sales activity across customer Regions and product categories to surface revenue trends and buying behavior. Materialized as an Iceberg table in the analytics schema, it provides a performant and scalable foundation for customer segmentation, Regional performance analysis, and category-level revenue attribution.

    -- models/analytics/customer_purchase_patterns.sql
    {{
        config(
            materialized='table',
            file_format='iceberg',
            schema='analytics'
        )
    }}
    
    SELECT
        dc.Region,
        dp.category,
        COUNT(DISTINCT fs.orderid) as total_orders,
        COUNT(DISTINCT dc.customerid) as unique_customers,
        SUM(fs.salesamount) as total_revenue,
        SUM(fs.salesamount) / COUNT(DISTINCT dc.customerid) as revenue_per_customer
    FROM {{ ref('fact_sales') }} fs
    JOIN {{ ref('dim_customers') }} dc ON fs.customerid = dc.customerid
    JOIN {{ ref('dim_products') }} dp ON fs.productid = dp.productid
    GROUP BY dc.Region, dp.category

Step 6: Transactional operations and time travel with Apache Iceberg

This section demonstrates how to use Apache Iceberg’s time travel capabilities and transactional operations using actual snapshot data from our dim_customers table. We walk through querying data at different points in time and comparing changes between snapshots.

  1. Transactional capabilities

    Let’s first look at current data:

    Now, modify the raw layer data for customerid 201 and change the Region to East

    Run the dbt model for dim_customers to sync the changes

    Validate the data in curated layer for dim_customers dimension table

  2. Time-travel capabilities

    First, let’s fetch snapshots for customers dimension table in curated layer

    Now, find the data state before and after the modification.

Step 7: Data quality tests

Data quality is a critical pillar of any reliable data pipeline. In this step, we define and enforce quality checks directly within the dbt project using schema-level test configurations. Rather than relying on one-time validation scripts, with dbt’s built-in testing framework, we can declaratively specify expectations on our models, ensuring that key fields remain unique, non-null, and consistent across the data layer before they reach downstream consumers.

  1. Generic tests configuration

    The schema.yml file serves as the central contract for model integrity. Here, we apply generic tests on the fact_sales and dim_customers models to catch data anomalies early in the pipeline.

    # models/schema.yml
    version: 2
    
    models:
      - name: fact_sales
        columns:
          - name: orderid
            tests:
              - unique
              - not_null
          - name: salesamount
            tests:
              - not_null
    
      - name: dim_customers
        columns:
          - name: customerid
            tests:
              - unique
              - not_null

Step 8: Maintenance procedures

A well-functioning data pipeline requires ongoing maintenance to remain performant and auditable over time. This step covers two essential practices, table optimization to keep data storage efficient, and snapshot management to track historical changes in source data. Together, these procedures keep the pipeline reliable, cost-effective, and capable of supporting time-based analysis.

  1. Table optimization

    As data accumulates in Delta or Iceberg tables, small files and fragmented storage can degrade query performance. The optimize_table macro provides a reusable utility to run Databricks’ OPTIMIZE command on any target table, consolidating small files and improving read efficiency without manual intervention.

    -- macros/optimize_table.sql
    {% macro optimize_table(table_name) %}
        {% set query %}
            OPTIMIZE {{ table_name }}
        {% endset %}
        {% do run_query(query) %}
    {% endmacro %}

  2. Snapshot management

    To maintain a historical record of customer data changes, we use dbt snapshots with a timestamp-based strategy. The customers_snapshot model captures row-level changes from the raw source layer and persists them in a dedicated snapshots schema, enabling point-in-time analysis and audit trails.

    -- snapshots/customer_snapshot.sql
    {% snapshot customers_snapshot %}
    {{
        config(
          target_schema='snapshots',
          unique_key='CustomerID',
          strategy='timestamp',
          updated_at='inserted_timestamp'
        )
    }}
    
    SELECT * FROM {{ source('raw_sales_analytics_data_layer', 'customers') }}
    
    {% endsnapshot %}

Step 9: Monitoring and logging

Observability is an essential aspect of any production-grade data pipeline. This step establishes logging and monitoring practices within the dbt project to track pipeline runs, capture errors, and support debugging. With structured logging enabled, teams gain visibility into model execution, test results, and runtime behavior, streamlining issue diagnosis and maintaining operational confidence.

  1. dbt logging configuration

    The dbt_project.yml logging configuration directs dbt to write logs to a dedicated path and outputs them in JSON format. JSON-structured logs are particularly useful for integration with log aggregation tools and monitoring dashboards, enabling automated alerting and audit trail management.

    # dbt_project.yml
    logs:
      path: logs
      enable_json: true

Step 10. Deployment and running

With the pipeline fully built, tested, and maintained, the final step covers how to deploy and execute dbt models across different scenarios. Whether running a complete refresh, processing incremental updates, or validating data quality, these commands form the operational backbone of day-to-day pipeline management.

  1. Full refresh

    A full refresh rebuilds all models from scratch, reprocessing the entire dataset. This is typically used after significant schema changes, backfills, or when incremental state needs to be reset.

    dbt run --full-refresh

  2. Incremental update

    For routine pipeline runs, incremental updates process only new or changed data, significantly reducing compute time and cost. The following command targets specific models (dim_customers and fact_sales) allowing selective execution without triggering the full DAG.

    dbt run --select dim_customers fact_sales

  3. Testing

    After models are run, data quality tests defined in the schema configuration are executed to validate integrity across all models. This validates that constraints such as uniqueness and non-null checks are met before data reaches downstream consumers.

    dbt test

Step 11. Cleanup

  1. Infrastructure cleanup
    # Delete EMR cluster
    aws emr terminate-clusters --cluster-id <cluster-id>
    
    # Remove S3 buckets
    aws s3 rb s3://amzn-s3-demo-transactional-datalake-raw-bucket-[unique-identifier] --force
    aws s3 rb s3://amzn-s3-demo-transactional-datalake-curated-bucket-[unique-identifier] --force
    aws s3 rb s3://amzn-s3-demo-transactional-datalake-analytics-bucket-[unique-identifier] --force

  2. Database cleanup
    DROP SCHEMA raw_sales_analytics_data_layer CASCADE;
    DROP SCHEMA curated_sales_analytics_data_layer CASCADE;

Conclusion

In this post, you learned how to build a transactional data lake on Amazon EMR using dbt and Apache Iceberg, from environment setup and modeling raw data, to quality enforcing, snapshot management, and incremental pipeline deployment. The architecture brings together the scalability of Amazon EMR, dbt’s transformation capabilities, and Iceberg’s ACID-compliant table format to deliver a reliable, maintainable, and cost-efficient data platform.

To get started, see the Amazon EMR documentation to deploy this architecture in your own environment. Whether you’re modernizing a legacy data platform or building a new analytics foundation, this stack gives you the flexibility to scale with confidence.


About the authors

Umesh Pathak

Umesh Pathak

Umesh is a Data Analytics Lead Consultant at AWS ProServe, based in India. When not solving complex data challenges, Umesh is out on the trails — an avid runner and hiker who brings the same discipline and drive to fitness as he does to his work.

Amol Guldagad

Amol Guldagad

Amol is a Data Analytics Lead Consultant based in India. He helps customers to accelerate their journey to the cloud and innovate using AWS analytics services.

AWS completes the second GDV community audit with participant insurers in Germany

Post Syndicated from Flamur Abdyli original https://aws.amazon.com/blogs/security/aws-completes-the-second-gdv-community-audit-with-participant-insurers-in-germany/

We’re excited to announce that Amazon Web Services (AWS) has completed its second GDV (German Insurance Association) community audit with 36 members from the Germany insurance industry participating, corresponding to over 63% coverage of the German market in terms of insurance premiums. Community audits are an efficient method to provide additional assurance to a group of customers on security of the cloud as described in the AWS Shared Responsibility Model in addition to AWS Compliance Programs (for example, Cloud Computing Compliance Criteria Catalogue (C5)) and resources that are provided to customers through AWS Artifact.

At AWS, security is the highest priority. As customers embrace the scalability and flexibility of AWS, we’re helping them evolve security and compliance into key business enablers. We’re obsessed with earning and maintaining customer trust and providing our financial services customers and their regulatory bodies with assurance that AWS has the necessary controls in place to help protect their most sensitive material and regulated workloads.

With the increasing digitalization of the financial industry and the importance of cloud computing as a key enabling technology for digitalization, the financial services industry is experiencing greater regulatory scrutiny. Our engagement with GDV members is an example of how AWS supports customers’ risk management and regulatory efforts. For the second time, this pooled audit meticulously assessed the AWS controls that we use to help protect customers’ data and material workloads, while satisfying strict regulatory obligations.

GDV is the association of private insurers in Germany, representing around 470 members in the industry and a key player within German and European financial services industries. GDV’s members participating in this community audit have reached out to AWS to exercise their audit rights according to the Digital Operational Resilience Act (DORA), BaFin requirements, and EIOPA’s Guidelines on Outsourcing to Cloud Service Providers. For this cycle, the audit was performed by a single external audit service provider on behalf of 36 participant members within the German insurance industry.

Audit preparations

The scope of the audit has been defined with reference to the BSI’s (Federal Office for Information Security) C5 framework, including key domains and control areas, in addition to AWS services (such as Amazon Elastic Compute Cloud (Amazon EC2) and the AWS Region relevant to participant members—Europe (Frankfurt) Region (eu-central-1).

Audit fieldwork

This phase started after an initial discussion in Berlin, Germany, and used a remote approach, using videoconferencing and a secure audit portal for the inspection of evidence. Auditors assessed AWS policies, procedures, controls using evidence, deep-dive subject matter expert (SME) sessions, and follow-up questions to clarify provided evidence.

Audit results

The audit has been executed and completed according to the mutually agreed engagement set up between AWS, participant members, and external auditors during which participating members exercised their audit rights in line with contractual conditions. After AWS reviews to confirm factual accuracy of the contents, auditors finalized the audit report. The results of the GDV community audit are only available to the participaing members and their regulators. The audit provides GDV members with assurance regarding the AWS controls environment, enabling members to work to remove compliance blockers, accelerate their adoption of AWS services, and obtain confidence and trust in the security controls of AWS.

Voice of the GDV community

From the perspective of the participating insurance companies, the second joint audit at AWS was seen as efficient and beneficial, because it reduced individual audit burdens while delivering reliable assurance results. At the same time, extensive planning and coordination required a substantial effort. Coordination with GDV and engaging with the DCSO Deutsche Cybersicherheitsorganisation GmbH (DCSO) as a professional external audit service provider helped streamline communication with AWS and ensured a consistent approach across all participants. The cooperation between the GDV insurers, the DCSO auditors, and AWS was professional and constructive throughout the process. For the first time, two representatives from insurance companies were present at the interviews, thereby gaining an even better impression of the quality of the audit.

To learn more about our compliance and security programs, see AWS Compliance Programs. As always, we value your feedback and questions; reach out to the AWS Compliance team through the Contact Us page.

If you have feedback about this post, submit comments in the Comments section below.

Flamur Abdyli

Flamur Abdyli

Flamur is a Principal in Security Assurance at AWS, based in Berlin, Germany. He leads complex customer audits and regulatory assurance engagements across EMEA, with a strong focus on financial services, regulated industries, and large enterprise customers. With more than 18 years of experience, Flamur has built and led teams across multiple industries and sectors.

Andreas Terwellen

Andreas Terwellen

Andreas is a Senior Manager in Security Assurance at AWS, based in Frankfurt, Germany. His team is responsible for third-party and customer audits, attestations, certifications, and assessments across EMEA. Previously, he was a CISO in a DAX-listed telecommunications company in Germany. He also worked for different consulting companies managing large teams and programs across multiple industries and sectors.

 

Свобода без удавници

Post Syndicated from original https://www.toest.bg/svoboda-bez-udavnitsi/

Свобода без удавници

Един от любимите ми художествени персонажи е плувецът Харука Нанасе. Причината – когато го попитаха какъв иска да стане, той отговори просто: „Свободен.“ Така се казва и самата аниме поредица (Free!), в която той е главен герой. 

Често се улавям, че мисля за свободата – не просто като за абстрактна концепция, а като за динамично равновесие между личното желание и външните обстоятелства. В най-чистия си вид тя изглежда като пълна автономия. Реалността обаче е по-сложна. Свободата има много дефиниции: от правото на индивида да бъде оставен на мира до претенцията му за равен достъп до ресурсите, които правят избора му изобщо възможен.

Преди обаче да се изгубим в езиковите спорове, трябва да признаем, че

всяко говорене за свобода е „осъдено“ да бъде недостатъчно прецизно.

Защото лесно можем да объркаме свободата на волята с условията за нейното упражняване. Човекът е захвърлен в света, ако следваме Сартр, и е принуден да избира, но неговият избор винаги е ситуиран – той се прави в определен социален и материален контекст, който или отваря хоризонти, или ги затваря. Именно тези граници на „възможното“ превръщат свободата от философски блян в бойно поле на идеите.

Макар думата „свобода“ днес да се използва с почти религиозно благоговение, тя далеч не е универсална. Малцина биха се заявили като нейни противници, но това е така само защото съдържанието, което влагаме в понятието, се променя драстично според приоритетите на дадено общество.

Често това, което от една гледна точка изглежда като окови, от друга се припознава като необходим ред, сигурност или свещен дълг. Конфликтът тук не е между „свободни“ и „потиснати“, а между различни ценностни системи: едната поставя индивидуалната автономия в центъра, другата вижда смисъла в принадлежността към колективната йерархия. Парадоксът е, че дори най-репресивните идеологии рядко обещават тирания; те почти винаги поставят като цел „истинско освобождение“ чрез дисциплина или духовно пречистване. Именно затова дефинирането на понятието не е просто езикова задача, а избор на цивилизационен модел.

Гледната точка на другия

Трябва обаче да имаме предвид, че е изключително контрапродуктивно да нападаме чуждата позиция фронтално. Това само втвърдява убежденията и настройва човека срещу нас, задействайки рефлекса „бий се или бягай“ – фундаментален инстинкт още на първобитните хора.

Навремето в лекциите си по „Конфликтни ситуации“ доцент Светла Страшимирова ни каза съвсем сериозно: вместо да спорите с другия, по-добре приемете, че главата му „не го слуша“. Звучи смешно, скандално и дори обидно, но в действителност следва желязна логика – всеки е детерминиран от фактори извън своя контрол. Убежденията често са плод на идентичност, а не толкова на съзнателен избор. Висш пилотаж е да осъзнаеш, че през погледа на другия вероятно ти си този, който е – грубо казано – луд.

И все пак няма как да превърнем цялото общество в психиатрия, въпреки огромните усилия, които парламентарната група на „Има такъв народ“ полага в тази посока (извинявам се, не можах да се сдържа). По-конструктивно е да опитаме да разберем откъде извира „лудостта“ на другия; какво го води до изводите да отрича правото на ЛГБТ хората да се женят за човека, когото обичат, или да стигматизира цели народи, както правят например привържениците на антисемитизма и ислямофобите.

Затова е добре да изведем теоретични рамки: що е то свобода, какво съдържание се влага в това понятие и как различните хора обозначават с една и съща дума съвсем различни неща.

Теориите на Бърлин, Макалъм и Милър

В това отношение може да ни помогне британският философ от латвийски произход сър Исая Бърлин с прочутото си есе „Две концепции за свободата“. Първият тип свобода е отрицателната свобода. Тя се състои в липсата на намеса: човек прави каквото реши, а държавата не му се бърка. Тук влизат правото на свободно изразяване, свободното пътуване и липсата на принуда при избора на религия.

Вторият тип е положителната свобода. Това е способността реално да контролираш живота си, за да осъществиш своите мечти и да постигнеш целите си. Тук не става дума само за липса на пречки, но и за наличие на условия. Класическият пример е да притежаваш средства, за да упражниш правата си, или достъп до образование, за да избереш професията си.

Разликата понякога е трудно доловима, затова често се използва следната илюстрация: ако пред теб има отворена врата и можеш да преминеш през нея, това е отрицателна свобода (никой не те спира). Но ако кракът ти е счупен, ти липсва положителната свобода – вратата може и да е отворена, но това не ти помага, тъй като физически не си способен да се възползваш.

Накратко: отрицателната свобода се стреми пред теб да няма пречки, а положителната – да ти създаде условия.

Трябва да се отбележи, че не всички учени приемат диференциацията на Бърлин. Американският философ Джералд Макалъм-младши например настоява, че свободата е единен принцип. Според него, когато се поставя под въпрос нечия свобода, винаги става дума за триада: някой е свободен от ограничение, за да извърши нещо. Той признава нюансите, но ги разглежда не като изключващи се концепции, а като общо придържане към идеята, че хората са свободни тогава, когато могат да действат без пречки.

Английският професор Дейвид Милър пък предлага още по-сложна система, разделяйки свободата на три: либерална (идентична с отрицателната на Бърлин), републиканска (свързана с участието в обществения живот и правата на гражданите) и идеалистична – онази, при която човек преодолява вътрешните си „вериги“, като например пристрастяването към вредни вещества.

Рисковете на положителната свобода

Рамката на Бърлин е ценна и с това, че разделя големите мислители на две традиции: защитници на отрицателната свобода, като Джон Лок и Джон Стюарт Мил, и теоретици на положителната, като Русо и Маркс. Самият Бърлин предупреждава, че с концепцията за положителна свобода лесно се злоупотребява. Тя позволява на държавата да наложи тирания под претекст, че „освобождава“ гражданите си, както гласи надписът на един съветски концлагер: „С железен юмрук ще подгоним хората към щастие.“

В този контекст споровете около политическата коректност и културата на отмяната (т.нар. cancel culture) не са нищо друго освен съвременен прочит на сблъсъка между двата вида свобода. Докато защитниците на статуквото бранят своята отрицателна свобода (правото да не бъдат възпирани в словото си от външни регулации), техните критици настояват за положителна свобода – за създаването на такава среда, в която исторически потиснатите групи най-после да имат реалната възможност и капацитет да участват в публичния диалог като равни. В учебника Politics: An Introduction Браунинг дава пример, като признава, че шегите за ирландци в Англия може да са забавни, но системното им описване като глупави и склонни към насилие колективно ги лишава от правото да бъдат възприемани сериозно.

Политическата коректност възникна в американските университети с цел да коригира езика, считан за уязвяващ малцинствата. В резултат заглавието на най-известния роман на Агата Кристи бе променено, а се стигна и до редакция на „Пипи Дългото чорапче“ от Астрид Линдгрен. Лично аз смятам прекаленото избягване на определени думи за излишно, но преди да съдя, се опитвам да си представя аналогична ситуация. Вероятно бих махнал с ръка, ако някой напише книга, в която героят му се среща с „гяурски цар“ в България. Струва ми се обаче, че точно сред най-яростните противници на политкоректността това не би се приело безболезнено.

Сблъсъкът на права

В същото време има казуси, които чудесно илюстрират сложността на проблема. Професор Гари Браунинг дава пример със скандала около романа „Сатанински строфи“ на Салман Рушди. След издаването му аятолах Хомейни издава фетва, с която осъжда Рушди на смърт, и това сериозно ограничава свободата му, меко казано. Самият автор споделя, че е написал книгата с презумпцията, че е свободен. Но негови критици смятат, че тя ограничава колективните права на мюсюлманите, като оскърбява религията им в културна среда, в която те са малцинство.

Подобен е случаят и с Дж. К. Роулинг. Нейните резерви спрямо настоящата посока на активизма за правата на транс хората и влиянието му върху пространствата, включващи разделение по пол, я изправи дори срещу актьорите от филмите за „Хари Потър“. Защитниците ѝ виждат в опитите за „канселиране“ форма на цензура, докато критиците ѝ смятат, че тя засяга правото на транс хората да бъдат признати според собствената си идентичност. За съжаление, нюансираните мнения, с които се изразява несъгласие по отношение на определени аспекти на позицията ѝ, но се осъждат заплахите към личността ѝ, са рядкост.

Описването на тези казуси не цели да даде окончателна морална присъда на една или друга позиция, а да илюстрира неизбежността на сблъсъка, когато две фундаментални представи за свободата се срещнат на едно и също поле. 

Дисквалификация на човека в дигиталната ера

Според мен един от основните проблеми днес е липсата на желание дори да се опитаме да влезем в положението на другия. Споровете се водят предимно в социалните мрежи – платформи, проектирани за емоционално общуване, което държи потребителите ангажирани (и гледащи реклами) чрез постоянно поддържане на напрежението. Вторият фактор е гневът: отговорите се дават моментално и без мисъл. Реакцията е инстинктивна и едва след това мозъкът я облича в „рационални“ аргументи.

На трето място, несъгласието с дадена позиция често води до състояние, близко до описаното от японския писател Дадзай Осаму в романа му „Провалът на човека“, издаден на български и под заглавието „Дисквалифициран като човек“. Виждаме го както в крайностите на съвременната култура на отмяната вляво, така и в етикети като „либераст“ (libtard) и подобни, с които отдясно заливат опонентите си. Така се ражда двойна морална система: на „нашите“ прегрешенията се прощават, а на „чуждите“ се отричат дори достойнствата. Връщаме се към племенния морал, който е описан от Клод Леви-Строс и може да се обобщи така:

Добро е, когато ние нападнем съседното племе и избием воините му; лошо е, когато те постъпят така с нас.

Тонът прави музиката

Разбира се, понякога каузата ни се струва твърде справедлива, за да мълчим. Но дори тогава е добре да не забравяме, че от другата страна стоят живи хора (поне докато ботовете не превземат фийдовете ни окончателно). Баба ми казваше: „Тонът прави музиката“ – дори най-твърдата позиция може да бъде заявена възпитано. Едва ли това само по себе си ще излекува днешния екстремен климат, но поне задава модел, който има шанс – макар и крехък – да смекчи повсеместната ярост.

Защото напоследък, както виждаме от събитията в САЩ, все по-често се прокрадва идеята, че спорът се печели с куршум. А смъртта, макар и в някакъв метафизичен смисъл да е върховната свобода, извеждаща ни отвъд тленното, е твърде окончателно решение за политически дебат.

В крайна сметка, ако свободата е просто възможността да излезеш през отворената врата, може би е време да спрем да се блъскаме в стъклото като мушици и да си спомним, че от другата страна на вратата въздухът е общ. И че ако искаме да бъдем „свободни“ като Харука Нанасе, трябва първо да се научим да плуваме в дълбокото, без да се опитваме да удавим тези на съседната пътека. Защото, докато „адът – това са другите“, единственият ни шанс да излезем от него е да започнем да ги виждаме като хора.

The Sashiko patch-review system

Post Syndicated from corbet original https://lwn.net/Articles/1063292/

Roman Gushchin has announced the
existence of an LLM-driven patch-review system named Sashiko. It automatically creates reviews
for all patches sent to the linux-kernel mailing list (and some others).

In my measurement, Sashiko was able to find 53% of bugs based on a
completely unfiltered set of 1,000 recent upstream issues using
“Fixes:” tags (using Gemini 3.1 Pro). Some might say that 53% is
not that impressive, but 100% of these issues were missed by human
reviewers.

Sashiko is built on Chris Mason’s review prompts (covered here in October 2025), but the
implementation has evolved considerably.

Decoding the Future of Inference At NVIDIA: Groq LPUs Join Vera Rubin Platform For Low-Latency Inference

Post Syndicated from Ryan Smith original https://www.servethehome.com/decoding-the-future-of-inference-at-nvidia-groq-lpus-join-vera-rubin-platform-for-low-latency-inference/

With its upcoming Vera Rubin rackscale architecture, NVIDIA is going to be integrating LPUs from acquihire Groq, marking a major expansion beyond using GPUs alone for AI inference

The post Decoding the Future of Inference At NVIDIA: Groq LPUs Join Vera Rubin Platform For Low-Latency Inference appeared first on ServeTheHome.

Backblaze Pricing and Product Updates

Post Syndicated from Backblaze original https://www.backblaze.com/blog/backblaze-pricing-and-product-updates/

A decorative image showing the Backblaze logo on a cloud. A title reads Product Updates and Upgrades

As our customers scale larger AI workloads, build increasingly data-intensive applications, power more complex media workflows, and demand faster recovery times, we continue to invest in performance, infrastructure, and features to support them.

Today, we’re announcing updates to B2 Cloud Storage pricing and APIs that reflect those investments and reinforce our long-standing commitment to straightforward pricing and open cloud freedom.

Price updates

  • Free API calls: Effective May 1, we’re making API calls free for all B2 Cloud Storage customers.* This removes transaction costs and makes it easier to build, scale, and run high-volume workloads subject to our standard platform usage rules and Terms of Service.
  • Storage price: Also effective May 1, we are updating pricing from $6/TB to $6.95/TB.

Things that aren’t changing

Pricing on existing committed contracts will not change until renewal, and pricing for B2 Overdrive and B2 Reserve will remain the same. Also not changing: 3x free egress for all B2 customers and unlimited free egress between Backblaze B2 and many leading neoclouds, content delivery network (CDN) and compute partners. Our commitment to transparent, predictable pricing remains unchanged.

Why the changes for B2 Cloud Storage?

1. Continuing to deliver high-performance, cost-effective cloud storage

Backblaze B2 has long been the simplest and most cost-effective alternative to legacy cloud providers. That value proposition hasn’t changed.

Behind the scenes, however, we continuously invest in infrastructure capacity, durability, performance improvements, security enhancements, and product innovation to support increasingly demanding workloads, especially AI training, large-scale media processing, and application data at scale.

This pricing update allows us to continue making those investments while maintaining the simplicity and cost advantage our customers rely on.

2. Doubling down on openness and predictability

One of the biggest reasons customers choose B2 Cloud Storage is freedom: 

  • Freedom from surprise bills
  • Freedom from punitive egress fees
  • Freedom from vendor lock-in

By making API calls free for all customers*, we’re removing another layer of billing complexity and friction. Customers can build, iterate, and scale without worrying about transaction costs adding up in the background.

Cloud storage should enable innovation, not penalize it at the moment of success.

3. Why eliminate API charges now?

As workloads like AI training, media pipelines, and high-frequency application access grow, API activity increases. Removing API charges simplifies billing and ensures customers can innovate without worrying about transaction-level costs.

Thank you

We know you have a lot of choices when it comes to cloud storage. The trust you place in Backblaze to safeguard and power your data is something we never take lightly.

Our goal is to be a long-term partner in helping you build, protect, and scale your business. These updates ensure we can continue investing in performance, reliability, and simplicity for years to come.

Thank you for choosing Backblaze and for being part of our community. We’re honored to support the work you’re building.

*Fees apply for the Event Notifications feature as part of that specific offering.

FAQ

Am I affected by this B2 Cloud Storage pricing update?

The storage price increase applies to pay-as-you-go B2 Cloud Storage customers beginning on May 1. Existing committed contracts stay the same until renewal time. New committed contracts will start at $6.95/TB. Pricing for B2 Reserve and B2 Overdrive will remain the same.

When will I, as an existing B2 Cloud Storage pay-as-you-go customer, see this update in my monthly bill?

The updated rate will apply to storage usage beginning on May 1 and will be reflected in your next billing cycle after that date.

Are API calls now completely free?

Beginning May 1, all standard API calls for B2 Cloud Storage are free for all customers except transactions related to a product or features with a different price point. As of this publishing, the only transactions in this category are related to our Event Notifications feature.

Will Backblaze continue to offer unlimited free egress to neocloud, CDN and compute partners?

Yes. Unlimited free egress to approved neocloud, CDN, and compute partners remains unchanged.

Is Backblaze still much more affordable than other cloud providers like AWS?

Yes. B2 Cloud Storage remains substantially more affordable than traditional cloud providers, particularly when factoring in API fees, egress policies, and predictable pricing—often one-fifth the price.

Are you introducing minimum storage duration fees or other hidden charges?

No. Unlike some competitors, we are not introducing minimum storage duration fees or new complexity into our pricing model. Our commitment to straightforward pricing remains the same.

Will this impact performance, durability, or SLAs?

No changes are being made to durability or availability commitments. We continue to offer an SLA of 99.9% uptime and design for 11 9’s durability.

What sort of improvements do you plan alongside the increase in pricing?

We continue investing in performance optimizations, partner integrations, security enhancements, and features that support AI, media, backup, and application workloads. Stay tuned for upcoming announcements.

The post Backblaze Pricing and Product Updates appeared first on Backblaze Blog | Cloud Storage & Cloud Backup

FSFE reports trouble with payment provider

Post Syndicated from jzb original https://lwn.net/Articles/1063287/

The Free Software Foundation Europe (FSFE) is reporting
that payment provider Nexi has terminated its contract without prior
notice, which means that a number of FSFE supporters’ recurring
payments have been halted:

Over the past few months, our former payment provider Nexi
S.p.A. (“Nexi”) requested access to private data, which we understood
to be specifically the usernames and passwords of our supporters. We
have refused this request. All our attempts to clarify Nexi’s request,
or to understand how their need for such information was necessary and
legal, were met with what we consider to be vague and unsatisfactory
explanations relating to a general need for risk analysis.

[…] The decisions that Nexi has made are incomprehensible to
us. Over the last months, as part of a security audit that Nexi
claimed to be conducting, we have provided them with large amounts of
the FSFE’s financial documentation, which even included private
information of our executive staff. We have answered all of their
questions. But we have to draw a line when private companies like Nexi
demand access to the sensitive and private data of our supporters.

According to the blog post, more than 450 supporters have been
affected by this. The FSFE’s donation pages have been updated with its
new payment provider.

[$] Fedora ponders a “sandbox” technology lifecycle

Post Syndicated from jzb original https://lwn.net/Articles/1062579/

Fedora Project Leader (FPL) Jef Spaleta has issued
a “modest proposal” for a technology-innovation-lifecycle process
that would provide more formal structure for adopting technologies in
Fedora. The idea is to spur innovation in the project without having an adverse
impact on stability or the release process. Spaleta’s proposal is
somewhat light on details, particularly as far as specific examples of
which projects would benefit; however, the reception so far is mostly
positive and some think that it could make Fedora more “competitive” by being the
place where open-source projects come to grow.

Security updates for Tuesday

Post Syndicated from jzb original https://lwn.net/Articles/1063248/

Security updates have been issued by Fedora (mingw-openexr, vim, and yarnpkg), Oracle (freerdp), Red Hat (389-ds-base, container-tools:rhel8, libpng, libpng15, nginx, nginx:1.24, nginx:1.26, opencryptoki, python3, python3.11, python3.12, and python3.9), SUSE (ruby4.0-rubygem-activestorage, ruby4.0-rubygem-activesupport, ruby4.0-rubygem-glogalid, ruby4.0-rubygem-grpc, ruby4.0-rubygem-jquery-rails, ruby4.0-rubygem-loofah, and rubygem4.0-rubygem-fluentd), and Ubuntu (curl, linux, linux-aws, linux-aws-6.17, linux-gcp, linux-hwe-6.17, linux-oracle,
linux-oracle-6.17, linux, linux-aws, linux-gcp, linux-gcp-6.8, linux-gke, linux-gkeop,
linux-hwe-6.8, linux-ibm, linux-ibm-6.8, linux-lowlatency,
linux-lowlatency-hwe-6.8, linux-oracle, linux-oracle-6.8, linux, linux-aws, linux-gcp, linux-gkeop, linux-ibm, linux-ibm-5.15,
linux-intel-iotg, linux-kvm, linux-lowlatency, linux-nvidia,
linux-nvidia-tegra, linux-nvidia-tegra-5.15, linux-oracle,
linux-xilinx-zynqmp, linux-fips, linux-aws-fips, linux-gcp-fips, linux-gcp, linux-nvidia, linux-nvidia-6.8, linux-nvidia-lowlatency, python-cryptography, and roundcube).

The collective thoughts of the interwebz