Отново ще бъдем силна опозиция

Post Syndicated from Bozho original https://blog.bozho.net/blog/4589

В дебата за избор на правителството казах следното:

Българските граждани дадоха абсолютно мнозинство на една партия.

И това върви с недвусмислен мандат за демонтиране на модела на завладяната държава. С мандат не за общи приказки, а за отдавна закъсняла промяна, за смели реформи в затлачени системи и за визия за България.

За съжаление, на база на представения персонален състав, на заявките (и на липсата на такива) и на действията на мнозинството дотук, имаме всички основания да сме скептични, дали този мандат от българските граждани ще бъде изпълнен.

Защото едва ли ще видим смели реформи от опитни бюрократи, вероятно съучаствали в затлачването на редица публични системи.

Едва ли ще видим пълно демонтиране на модела, когато иззад някои министри надничат стари кръгове от икономически интереси.

Едва ли ще видим промяна от “квотата” на Има такъв народ, които се оказаха най-голямото недоразумение в българската политика – при доста силна конкуренция за този приз – и бяха изхвърлени от българските граждани.

Но въпреки тези основателни опасения сме длъжни да пожелаем успех на правителството, защото ако то се справи, това ще е успех за България. А всички ние сме тук, за да успее България, без значение от скептицизма си към опонента.

В дебата за гласуване на предходния кабинет, свален след по-малко от година с масови протести, казах следното:

“Ще бъдем силна опозиция. Не деструктивна, не креслива, а такава, която показва, че може по-добре.

Няма да злорадстваме за всеки неуспех, да профанизираме трудните компромиси и да атакуваме малките неволни грешки. Но управляващите добре знаят, че различаваме неволните грешки от съзнателното потъпкване на обществения интерес.

Ще се противопоставяме остро на всички опити за дозавладяване на държавата и за укрепване на нелегитимните влияния.

Но и ще подкрепяме всички правилни политики, свързани с борбата с корупцията, модернизацията на страната, намаляването на административната тежест, оптимизацията на публичните разходи. Ще предлагаме решения, които са в духа на постигнатите широки, принципни съгласия.”

Казвам това и днес със същата убеденост. Защото сме последователни. През миналата година подходихме именно така – с конструктивни, разумни предложения. И с остро противопоставяне тогава, когато интересите на задкулисието изместваха обществения интерес.

Ще участваме в избора на Висш съдебен съвет, защото това е конституционното ни задължение и защото демонтирането на модела започва от там. Но изобщо не свършва там.

Ще участваме с конкретни предложения за премахване на инструментите за нелегитимно упражняване на власт и корупционните кранчета, овладяни до съвършенство от модела Пеевски-Борисов, но чакащи своя нов господар и брокер.

Ще подкрепяме всяка воля за ограничаване на публичните разходи и всяка друга трудна и закъсняла реформа.

Но и ще се противопоставяме на всеки опит за прегрупиране на олигархии и за договорки с предишните властелини на модела под фанфарите на внушителната изборна победа.

Ще се противопоставяме на всяко отклоняване от мандата за реформи и за елиминиране на завладяната държава.

Ще се противопоставяме на всеки опит българските граждани да бъдат разделяни и насъсквсни едни срещу друти от някой умел вицепремиер по пропагандата.

Ще се противопоставяме на всеки опит за завой на изток или за превръщане на България в троянски кон на външни сили, прикрит зад клишета за българския глас в Европа или зад имагинерни ползи от подмазване на агресори и авторитарни режими.

Дано не ни се наложи. Но с гласа си не можем да носим отговорност за бъдещите действия на чуждо правителство, когато избирателите са ни отредили ролята на опозиция. Затова няма да подкрепим предложеното правителство.

Все пак му пожелавам успех в трудната задача да премахне системните предпоставки за политическата криза, която го доведе на власт. Демократична България ще изпълним нашия ангажимент да сме силна опозиция, за да може и управлението да бъде по-силно.

Материалът Отново ще бъдем силна опозиция е публикуван за пръв път на БЛОГодаря.

Dirty Frag: a zero-day universal Linux LPE

Post Syndicated from jzb original https://lwn.net/Articles/1071719/

Hyunwoo Kim has announced
the Dirty
Frag
security flaw, a
local-privilege-escalation (LPE) vulnerability similar to the
recently disclosed Copy Fail
flaw:

Because the embargo has now been broken, no patches or CVEs exist for
these vulnerabilities. After consultation with the [email protected]
maintainers, and at the maintainers’ request, I am publicly releasing this
Dirty Frag document.

As with the previous Copy Fail vulnerability, Dirty Frag likewise allows
immediate root privilege escalation on all major distributions.

Kim, who discovered the flaw and had attempted a coordinated
disclosure set for May 12, has released the code for an exploit, as well as a example
script to remove the vulnerable modules. A full
write-up
, with the disclosure timeline, is also available. It’s
unknown at this time whether this is an example of parallel discovery
or how the third party was able to disclose it prior to the end of the
embargo. We will be following up as more information comes to light.

Building for the future

Post Syndicated from Matthew Prince original https://blog.cloudflare.com/building-for-the-future/

This afternoon, we sent the following email to our global team. One of our core values at Cloudflare is transparency, and we believe it’s important that you hear this directly from us because it’s a major moment at Cloudflare. 

Team:

We are writing to let you know directly that we’ve made the decision to reduce Cloudflare’s workforce by more than 1,100 employees globally. 

The way we work at Cloudflare has fundamentally changed. We don’t just build and sell AI tools and platforms. We are our own most demanding customer. Cloudflare’s usage of AI has increased by more than 600% in the last three months alone. Employees across the company from engineering to HR to finance to marketing run thousands of AI agent sessions each day to get their work done. That means we have to be intentional in how we architect our company for the agentic AI era in order to supercharge the value we deliver to our customers and to honor our mission to help build a better Internet for everyone, everywhere. 

Today is a hard day. This decision unfortunately means saying goodbye to teammates who have contributed meaningfully to our mission and to building Cloudflare into one of the world’s most successful companies. We want to be clear that this decision is not a reflection of the individual work or talent of those leaving us. Instead, we are reimagining every internal process, team, and role across the company. Today’s actions are not a cost-cutting exercise or an assessment of individuals’ performance; they are about Cloudflare defining how a world-class, high-growth company operates and creates value in the agentic AI era. 

This is a moment we need to own as founders and leaders of the company. Matthew has personally sent out every offer letter we’ve extended. It is a practice he has always looked forward to because it represented our growth and the incredible talent joining our mission. It didn’t feel right for this message to come from anyone other than the two of us. Rather than trickling out notices through managers, we will be sending emails to every employee. 

Within the next hour, every member of our global team will receive an email from both of us clarifying how this change affects them. For those departing today, we will send this update to both their personal and Cloudflare addresses to ensure they receive the information immediately.

It’s important to us that we treat departing team members right and in a way that exceeds what we’ve seen from other companies. We believe acting with empathy isn’t about avoiding hard decisions but rather about how you treat people when those decisions are made. If we are asking our team to be world-class, we have a reciprocal obligation to be world-class in how we treat them. We are pairing the directness of these measures with severance packages that lead the industry. The packages for departing employees will include the equivalent of their full base pay through the end of 2026. Healthcare coverage is different across the globe, and if you’re in the United States, we’ll continue to provide support through the end of the year. We are also vesting equity for departing team members through August 15th, so they receive stock beyond their departure date. And, if departing team members haven’t hit their one-year cliffs, we are going to waive those and vest their pro-rated equity through August as well. 

We’ve asked the team to do this only once, as hard as that may be today. We don’t want to do it again for the foreseeable future. By taking decisive action now, we provide immediate clarity to those departing and protect the stability of the team that remains. We are making these changes now because making smaller, repeated cuts or dragging a reorganization out over multiple quarters creates prolonged emotional uncertainty for employees and stalls our ability to build. It’s the right thing to do; it’s the honest thing to do; and it reflects the values of the company we are continuing to build.

Cloudflare started as a digitally native company built in the cloud. That allowed us to catch up to and pass companies that had a head start of years or decades but were slowed down by outdated systems and processes. As we’ve now become the leader, we cannot rest on the workflows and organizational structures that worked yesterday. We’re confident that our reshaped organization will be even faster and more innovative as we continue building the future.

To those departing us: you’ve helped build the strong foundation Cloudflare stands on today. We have the utmost respect for your work and gratitude for the impact you have made. We’re confident you will land at other great places and build many future great companies, bringing with you a unique set of skills learned while building Cloudflare.

Transparency is a core principle at Cloudflare, and it was important that you hear this from us first. We will be heading to our earnings conference call at 2 PM PT, when we’ll share more. We also plan to address today’s announcements live with the team at our all-hands meeting. 

It’s not an easy day, but it’s the right decision. Our mission to help build a better Internet is more important now than ever, and there’s a lot of work left to be done.

Rapid7 and OpenAI: Helping Defenders Move at Machine Speed

Post Syndicated from Wade Woolwine original https://www.rapid7.com/blog/post/ai-rapid7-openai-helping-defenders-move-at-machine-speed

Wade Woolwine is Senior Director, Product Security at Rapid7.

Announcing OpenAI’s Trusted Access for Cyber program

CIOs and CISOs are telling us the same thing in different ways: Advances in frontier AI are accelerating the threat environment and putting pressure on security operating models built for a different pace. Vulnerabilities can be discovered faster, exploitation windows are shrinking, and attackers are increasingly using automation to move with greater speed and scale. For defenders, this changes the value equation. The premium is no longer only on detecting threats faster after they emerge, but on moving earlier: Reducing exposure, validating risk, strengthening detection, and remediating at scale before attackers can take advantage.

This is why Rapid7 is excited to be included in OpenAI’s Trusted Access for Cyber program and their announcement today. OpenAI’s approach recognizes that advanced AI can help verified security teams move faster on legitimate defensive work, from triage and detection to validation, patching, malware analysis, and detection engineering. It also recognizes that some specialized cyber workflows require stronger verification, monitoring, and feedback loops.

As Corey Thomas, CEO of Rapid7, shared:

“Security leaders are under pressure from every direction: More vulnerabilities, faster exploitation, and increasing business pressure. Through OpenAI’s Trusted Access for Cyber program, Rapid7 is exploring more ways to accelerate the shift from reactive to preemptive security. To stay ahead of attackers, defenders must proactively reduce exploitability and detect with machine-scale speed and precision. We’re working with OpenAI to equip security teams with advanced capabilities that will meaningfully improve their cyber resilience.”

AI in security: Not just faster discovery

For Rapid7, this moment is about more than faster vulnerability discovery. AI is creating new pressure across the entire security lifecycle, from vulnerability validation, prioritization, disclosure, and remediation to threat and exploitation detection. Security infrastructure built for human-speed discovery now needs to operate in a machine-speed world, with enough context, governance, and accountability to help defenders act with confidence.

Finding risk is only the beginning. Security teams need to understand which vulnerabilities and misconfigurations are truly exploitable, which systems and business services are affected, what compensating controls are in place, how remediation should be prioritized, and where detection coverage is needed. CISOs also need confidence that advanced AI is being applied responsibly, with clear guardrails, measurable outcomes, and accountability.

Our work with OpenAI will help us explore how frontier AI can strengthen three critical areas. First, it can support the identification of vulnerabilities in our own products and code earlier in the development lifecycle. By accelerating secure code review, surfacing risky patterns, supporting root cause analysis, reviewing patches, and giving engineering teams faster feedback, AI can help reduce risk before issues reach production.

Second, it can advance vulnerability research and exploitation analysis. Rapid7 has long-standing expertise in vulnerability intelligence, exploitability research, and offensive security with Rapid7 Labs. Frontier AI can help researchers reason across unfamiliar code, map affected surfaces, build safe reproduction harnesses, validate severity, and turn findings into practical remediation guidance.

Third, it can expand AI-driven red-teaming. As AI becomes more embedded in enterprise systems and security operations, it must also be tested adversarially. We see an opportunity to use AI to strengthen red-team workflows, explore attack paths, validate controls, and help defenders understand where exposure could become real-world risk.

Artificial intelligence in use at Rapid7

We are already seeing this potential inside our own security operations work. In support of our Agentic SOC initiatives, Rapid7 has designed and implemented a system that uses machine learning to surface threat- and risk-relevant events from raw log and telemetry data. By using frontier AI models, including OpenAI’s GPT-5.5, to support initial triage and escalate only relevant events to SOC analysts, we have seen a 25% reduction in time spent chasing false-positive events in the queue.

This is not about replacing human expertise. It is about giving defenders better leverage in a world where attackers, businesses, and technology are all moving faster. The shift from reactive to preemptive security, and from human-scale processes to machine-scale defense, is not a marketing reframe. It is becoming the only viable path for teams that need to anticipate where attackers will move next, prioritize the exposures that actually matter, and respond at the speed of modern attacks.

AI may accelerate discovery, but cyber resilience depends on what happens after discovery. Customers need to unify their data, apply AI with the right context, drive remediation at scale, and translate security activity into measurable outcomes. That is where Rapid7 is focused. Across the Command Platform, Rapid7’s AI capabilities are built to help security teams detect threats and anomalies at scale, reduce noise, optimize SOC workflows, and make faster, more confident decisions.

By unifying Exposure Management and Detection and Response on the Command Platform, and combining AI-driven operations with the depth of expertise we have built over 25 years, Rapid7 is giving customers a more coherent way to reduce risk, disrupt attackers, and build durable cyber resilience. Learn more about Rapid7’s AI capabilities.

Defense by Design: Building Infrastructure That Assumes Adversaries

Post Syndicated from Kari Rivas original https://www.backblaze.com/blog/defense-by-design-building-infrastructure-that-assumes-adversaries/

A decorative image showing a computer plus several icons that reference security.

Ransomware and other disruptive attacks rarely succeed because of a single catastrophic failure. More often, they succeed because a system was designed for availability and scale, but not for persistent, adaptive adversaries testing for weak points from the outside.

For infrastructure and architecture leaders, that creates a practical challenge: how do you build systems that remain performant, cost-efficient, and operable while also standing up to attackers who are probing your environment for opportunities through traffic abuse, credential attacks, vulnerability exploitation, and social engineering?

The answer is not a single tool or framework. It is an architectural mindset: assume adversaries exist, assume controls will be tested, and design systems that continue operating safely under pressure.

Security starts with architecture, not alerts

One of the most common mistakes organizations make is treating security as something layered onto infrastructure after it is built. In practice, resilience comes from decisions made much earlier:

  • How traffic is handled under stress.
  • How systems and services are segmented.
  • How identity and access are enforced.
  • How quickly vulnerabilities are surfaced and validated.
  • How failure is contained when something goes wrong.

This is what separates reactive security from resilient architecture. The strongest environments are not the ones with the most dashboards; they are the ones built so that no single weakness can easily cascade into a broader incident.

Designing the perimeter to buy time, not perfection

Even in a world shaped by zero trust, the perimeter still matters, especially for availability.

Large-scale traffic floods, automated scanning, and API abuse are often the opening move. These events may not be the full attack, but they can create noise, consume resources, and open the door for more targeted follow-on activity. Infrastructure teams need defenses that can:

  • Absorb unexpected traffic without cascading failures
  • Distinguish abusive patterns from legitimate use
  • Prevent noisy attacks from turning into operational incidents

These defenses can never provide perfect prevention; cybercriminals can attack with too much sophistication and velocity. Rather, the goal is resilience. Good perimeter design buys time, preserves service availability, and prevents external pressure from becoming internal disruption.

Continuous vulnerability discovery beats periodic assurance

Modern environments change too quickly for occasional reviews to be enough.

Attackers do not work on quarterly schedules, and neither should defensive programs. A stronger model is continuous vulnerability discovery: using multiple signals to understand what is exposed, what is exploitable, and what actually matters.

That can include a mix of:

  • Threat intelligence on active exploitation trends
  • External research programs such as bug bounties
  • Regular penetration testing
  • Internal testing and automated vulnerability scanning

Each of these presents different types of risk. Together, they reduce blind spots and help teams prioritize fixes based on real-world likelihood and impact, not just severity scores on paper.

Limiting blast radius is an architectural responsibility

A useful security question is not only “How do we stop every attack?” but also “What happens if one control fails?”

That shift changes how teams think about system design. It places greater emphasis on:

  • Hardening critical systems
  • Enforcing strict access controls
  • Separating environments and services
  • Reducing unnecessary trust relationships
  • Containing failure before it spreads

This is where architecture has an outsized role. Detection matters, but containment matters just as much. Systems built with clear boundaries are easier to defend and easier to recover operationally when incidents happen.

Identity is part of infrastructure

Many attacks do not begin with sophisticated exploits. They begin with compromised credentials, reused passwords, phishing, or other attempts to gain access through people rather than code.

That is why identity should be treated as a core infrastructure layer, not a separate administrative concern.

Strong identity practices often include:

  • Long, high-entropy passwords
  • Multi-factor authentication
  • Checks for compromised credentials
  • Clear access policies tied to real job needs

These controls reflect a simple truth: humans are part of the system. Security controls need to be strong enough to resist abuse and usable enough to work at scale.

Security as a system, not a checklist

No single control creates resilience on its own.

What matters is how controls reinforce one another: how traffic protections support availability, how vulnerability discovery informs remediation, how segmentation reduces impact, and how identity controls protect critical paths.

For infrastructure and architecture leaders, the takeaway is straightforward: the most resilient systems are not built on assumptions of safety. They are built on the expectation that adversaries will look for openings and that defenses need to hold up under real pressure.

That is why security works best as an architectural decision, not just an operational one.

A practical Backblaze perspective

At Backblaze, this is the lens we use when thinking about protection against bad actors: not as a single feature or isolated control, but as a layered systems problem that spans network protections, vulnerability discovery, access controls, and operational resilience. The important point is not any one safeguard in isolation. It is the way those safeguards work together so that a single weakness is less likely to become a customer-impacting event. 

Download the ebook on building an affordable, resilient disaster recovery strategy that matters when ransomware strikes.

The post Defense by Design: Building Infrastructure That Assumes Adversaries appeared first on Backblaze Blog | Cloud Storage & Cloud Backup

ICYMI: April 2026 @AWS Security

Post Syndicated from Rodolfo Brenes original https://aws.amazon.com/blogs/security/icymi-april-2026-aws-security/

Read all about the latest AWS security features, compliance updates, and hands-on resources in our new, monthly digest posts. You’ll find expert blog posts, new service capabilities, code samples, and workshops.

AWS Security Blog posts

This month’s AWS Security Blog posts covered AI security, identity and access management, threat intelligence, data protection, and multicloud operations. Whether you’re securing agentic AI systems, upgrading to post-quantum cryptography, or streamlining forensic collection, these posts offer practical guidance across the security landscape.

Identity

    Access control with IAM Identity Center session tags
    Author: Rashmi Iyer | Published: April 28, 2026
    Learn to combine AWS IAM Identity Center permission sets with session tags from Microsoft Entra ID to implement fine-grained attribute-based access control (ABAC) across multiple AWS accounts.

    Can I do that with policy? Understanding the AWS Service Authorization Reference
    Authors: Anshu Bathla, Prafful Gupta | Published: April 27, 2026
    Learn to use the AWS Service Authorization Reference to determine what’s achievable with IAM policies, recognize scenarios needing alternative solutions, and build more effective security controls.

    AI Security

    Secure AI agent access patterns to AWS resources using Model Context Protocol
    Author: Riggs Goodman III | Published: April 14, 2026
    Learn to secure AI agent access to AWS resources via MCP using three principles: least privilege, organizational role governance, and differentiating AI-driven from human-initiated actions.

    Four security principles for agentic AI systems
    Authors: Mark Ryland, Riggs Goodman III, Todd MacDermid | Published: April 2, 2026
    Learn four security principles from AWS’s NIST response for securing agentic AI: secure development lifecycle, traditional controls, deterministic external enforcement, and earned autonomy through evaluation.

    Designing trust and safety into Amazon Bedrock powered applications
    Author: Victor Lungu | Published: April 29, 2026
    Learn to integrate responsible AI concepts into Amazon Bedrock applications, including abuse detection, Amazon CloudWatch monitoring, Bedrock Guardrails configuration, and the abuse response process.

    Building AI defenses at scale: before the threats emerge
    Author: Amy Herzog | Published: April 7, 2026
    AWS CISO announces Project Glasswing with Anthropic, introducing Claude Mythos Preview for vulnerability research, plus the general availability of AWS Security Agent for autonomous penetration testing.

    Governance and compliance

      Shift-Left Tag Compliance using AWS Organizations and Terraform
      Authors: Welly Siauw, Sourav Kundu, Manu Chandrasekhar | Published: April 27, 2026
      Learn to validate tag compliance during development using AWS Organizations tag policies, a reusable Terraform tagging module, and a test-driven approach that dynamically validates against live organizational policies.

      Detection and incident response

      What the March 2026 Threat Technique Catalog update means for your AWS environment
      Authors: Shannon Brazil, Cydney Stude | Published: April 28, 2026
      The AWS CIRT’s latest Threat Technique Catalog update covers Amazon Cognito refresh token abuse, AMI image deletion targeting recovery, and trust policy modifications for persistence and privilege escalation.

      A framework for securely collecting forensic artifacts into S3 buckets
      Authors: Jason Garman, Vaishnav Murthy | Published: April 8, 2026
      Learn to securely collect forensic artifacts into Amazon S3 using time-limited, least-privilege credentials with AWS STS session policies and automated AWS Step Functions workflows.

      Transform security logs into OCSF format using a configuration-driven ETL solution
      Authors: Vivek Gautam, Arpit Gupta, Ryan Gomes | Published: April 17, 2026
      Learn to transform custom security logs into OCSF format using an AWS ProServe configuration-driven ETL solution with AWS Step Functions, AWS Glue or Amazon EMR Serverless, and Amazon Security Lake integration.

      A technical walkthrough of multicloud full-stack security using AWS Security Hub Extended
      Authors: Matt Meck, Michael Fuller | Published: April 22, 2026
      Learn how AWS Security Hub Extended simplifies multicloud security procurement and operations through curated partner solutions, unified billing, and OCSF-based findings consolidation.

      Data protection

        Protecting your secrets from tomorrow’s quantum risks
        Authors: Stéphanie Mbappe, Tobias Nickl | Published: April 24, 2026
        Learn to upgrade AWS Secrets Manager clients to use hybrid post-quantum TLS with ML-KEM, protecting secrets against harvest-now-decrypt-later attacks, and verify connections via AWS CloudTrail.

        How AWS KMS and AWS Encryption SDK overcome symmetric encryption bounds
        Authors: Panos Kampanakis, Matthew Campagna, Patrick Palmer | Published: April 3, 2026
        Learn how AWS Key Management Service and the AWS Encryption SDK use derived key methods to automatically handle AES-GCM encryption limits, eliminating the need to manually track bounds or rotate keys.

        How to clone an AWS CloudHSM cluster across Regions
        Authors: Desiree Brunner, Rickard Löfström | Published: April 20, 2026
        Learn to clone an AWS CloudHSM cluster to another Region using CopyBackupToRegion, then synchronize keys—including non-exportable keys—across cloned clusters for disaster recovery.

        April Security Bulletins

        Investigations of reported security vulnerabilities affecting Amazon and AWS services, software, and products.

        AWS Samples

        This month brings 16 new AWS samples spanning identity, governance, compliance, detection and incident response, AI Security, data protection, and infrastructure security. From beginner-friendly AI agent development on Amazon Bedrock to automated Control Tower re-registration at scale, these ready-to-deploy repositories help you implement security best practices across your AWS environment.

        Identity

          Amazon Cognito OAuth2 Token Proxy with Caching
          Learn to deploy an Amazon API Gateway proxy for Cognito’s OAuth2 token endpoint with intelligent caching and AWS WAF protection, reducing M2M authentication costs by over 90%.

          Cognito API Gateway Authorization Demo
          Learn to implement user-specific data protection using Amazon Cognito, API Gateway, and an AWS Lambda authorizer that enforces JWT sub claim matching to prevent cross-user data access.

          Securely Connecting On-Premises Data Systems to Amazon Redshift with IAM Roles Anywhere
          Learn to deploy a fully private environment connecting on-premises workloads to Amazon Redshift using X.509 certificate authentication via IAM Roles Anywhere for short-lived credentials.

          AWS IAM Access Key Lifecycle Management with Human Approval
          Learn to automate organization-wide detection, disabling, and deletion of unused IAM access keys using Step Functions, IAM Access Analyzer, and a secure human-in-the-loop approval workflow.

          Secrets Manager Audit
          Learn to resolve and report who can access your AWS Secrets Manager secrets—across accounts, through Identity Center, and down to the human behind the IAM role—in a single command.

          Governance

          Control Tower Organization Re-Registration Automation
          Learn to automate AWS Control Tower OU re-registration and account updates at scale using lifecycle events, Amazon EventBridge, and AWS Lambda to resolve mixed governance after landing zone changes.

          Sample Agent Skills for Builders
          A curated collection of installable agent skills that extend AI coding agents (Claude Code, Cursor, Copilot) with production-ready AWS, CDK, security scanning, and engineering workflows.

          How to Stop AI Agent Hallucinations: 5 Techniques + Production on Amazon Bedrock AgentCore
          Learn to detect, prevent, and self-correct AI agent hallucinations using Graph-RAG, semantic tool selection, multi-agent validation, neurosymbolic guardrails, and agent steering with Strands Agents.

          Compliance

          Compliance Lens
          Learn to deploy a serverless solution that analyzes AWS Config snapshots across an AWS Organization, compares them against conformance pack rule sets, and visualizes compliance posture via Amazon QuickSight dashboards.

          AWS Security Agent Terraform Configuration
          Learn to provision AWS Security Agent resources using the AWSCC Terraform provider, automating agent space creation, IAM roles, target domain registration, and penetration test setup.

          Detection and incident response

          AWS Security Agent Demo Suite
          Learn to use AWS Security Agent across three scenarios: automated design reviews, AI-generated infrastructure code review via GitHub, and penetration testing against intentionally vulnerable applications.

          Agentic SOC Workshop — CDK Infrastructure
          Learn to build an AI-powered Security Operations Center agent that investigates Amazon GuardDuty findings, queries CloudTrail logs, and takes automated containment actions using Amazon Bedrock AgentCore.

          Data Protection

          Implementing Kerberos Authentication for Apache Spark Jobs on Amazon EMR on EKS to Access a Kerberos-Enabled Hive Metastore
          Learn to configure Kerberos authentication for Spark jobs on Amazon EMR on Amazon Elastic Kubernetes Service, connecting to a Kerberos-enabled Hive Metastore using Microsoft Active Directory as the KDC.

          AWS Nitro Enclaves with Kubernetes – Hello World Example
          Learn to deploy a Hello World application inside an AWS Nitro Enclave on Amazon EKS, covering cluster creation, device plugin setup, and enclave image building.

          Infrastructure security

            Multi-Tenant OpenClaw on Firecracker
            Learn to deploy isolated, multi-tenant OpenClaw AI agents on AWS using Firecracker microVMs with per-tenant kernel/network isolation, auto-scaling, backup/restore, and a web management console.

            AI Security

            Amazon Bedrock for Beginners – From First Prompt to AI Agent
            Learn to build AI applications on Amazon Bedrock, from basic API calls to a full agent with RAG, guardrails, tool use, and the Strands Agents SDK.

            Conclusion

            April 2026 reinforces that securing AI workloads now requires the same rigor applied to traditional infrastructure. The posts and samples in this edition provide concrete patterns for enforcing least privilege on agentic systems, automating governance at organizational scale, and preparing cryptographic implementations for post-quantum requirements. The security bulletins address vulnerabilities across compute, networking, and developer tooling, reinforcing the need to apply patches consistently. Each resource includes deployment steps or runnable code so you can validate the approach in your own environment before adopting it. Subscribe to the AWS Security Blog RSS feed to receive updates as they publish, and revisit this digest monthly for a consolidated view of what changed and what to act on.


            If you have feedback about this post, submit comments in the Comments section below. If you have questions about this post, contact AWS Support.

            Rodolfo Brenes

            Rodolfo Brenes

            Rodolfo is a Principal Solutions Architect focused on Cloud Governance and Compliance. With over 18 years of experience, he currently leads a technical field community in AWS helping customers scale and improve their security and governance frameworks. Besides work, Rodolfo enjoys video games, playing with his four cats, and won’t say no to a good outdoor adventure.

            Anna Brinkmann

            Anna Brinkmann

            Anna is a project manager and editor with more than 18 years of experience with content management in the technology space. For the past 6 years, she has run the AWS Security Blog. In her free time, Anna gardens, spends time with family and friends, and learns new slang words from her kids.

            AWS achieves SNI 27017, SNI 27018, and SNI 9001 certifications for the AWS Asia Pacific (Jakarta) Region

            Post Syndicated from Ignatius Lee original https://aws.amazon.com/blogs/security/aws-achieves-sni-27017-sni-27018-and-sni-9001-certifications-for-the-aws-asia-pacific-jakarta-region/

            Amazon Web Services (AWS) achieved three Standar Nasional Indonesia (SNI) certifications for the AWS Asia Pacific (Jakarta) Region: SNI ISO/IEC 27017:2015, SNI ISO/IEC 27018:2019, and SNI ISO 9001:2015. SNI represents Indonesia’s national standards framework, comprising standards that are broadly applicable across industries within the country. These certifications further demonstrate that AWS services meet nationally recognized requirements.

            The certifications were assessed by an independent third-party auditor accredited by the Komite Akreditasi Nasional (KAN), Indonesia’s National Accreditation Committee, in accordance with applicable local regulatory requirements, helping customers rely on trusted, locally recognized validation for their compliance needs.

            All three certifications are based on international ISO standards adapted for Indonesia:

            • SNI 27017 adds cloud-specific security controls that complement ISO/IEC 27001, helping you run workloads securely while reducing security assessment overhead.
            • SNI 27018 focuses on protecting personally identifiable information (PII) in public clouds. This certification confirms that AWS handles your data according to international privacy standards.
            • SNI 9001 establishes quality management systems that ensure consistent service delivery and continuous improvement across AWS operations.

            Together with the existing SNI 27001 certification achieved in 2023, AWS is now the first cloud service provider (CSP) to hold all four SNI certifications—SNI 27001, SNI 27017, SNI 27018, and SNI 9001—demonstrating comprehensive alignment with Indonesia’s national standards for information security, cloud security, privacy, and quality management, and helping customers address a broad range of regulatory and risk management requirements.

            Customers can access the corresponding certificates through AWS Artifact, a self-service portal that provides on-demand access to AWS compliance documentation. For a full list of AWS services covered under the SNI certification, see the Services in Scope compliance page

            AWS continues to expand the scope of its compliance programs to help customers meet their architectural, business, and regulatory requirements. For more information regarding these certifications, contact your AWS Accounts team.

            Ignatius Lee

            Ignatius Lee

            Ignatius is a Security Assurance professional based in Singapore, responsible for third-party audits in Indonesia. He joined Security Assurance in early 2025 and has delivered and contributed to key audit programs across Hong Kong, Singapore, and Australia.

            [$] A new era for memory-management maintainership

            Post Syndicated from corbet original https://lwn.net/Articles/1070994/

            On April 21, Andrew Morton let
            it be known
            that he intends to begin stepping away from the
            maintainership of kernel’s memory-management subsystem — a responsibility
            he has carried since before memory management was even seen as its own
            subsystem. At the 2026 Linux Storage, Filesystem, Memory Management, and
            BPF Summit, one of the first sessions in the memory-management track was
            devoted to how the maintainership would be managed going forward. There
            are a lot of questions still to be answered.

            An update on KDE’s Union style engine

            Post Syndicated from jzb original https://lwn.net/Articles/1071703/

            Arjen Hiemstra has published
            an article on the status of the Union project: a
            single system to support all of KDE’s technologies used for styling
            applications.

            The work on Union’s Breeze implementation has progressed to the
            point where it is very hard to distinguish whether or not you are
            running the Union version. We have also tested with a bunch of
            applications and made sure that any differences were fixed. So we are
            at a stage where we need to get Union into the hands of more people,
            both to get extra people testing whether there are any major issues,
            but also to have interested people creating new styles.

            This means that with the upcoming Plasma 6.7 release, we plan to
            include Union. Discussion is currently ongoing whether we will enable
            it by default, but even if not there will be a way to try it out.

            See Hiemstra’s introductory
            article on Union
            , published in February 2025, for more about the
            project and its creation. KDE 6.7 is expected to be released in mid-June.

            Security updates for Thursday

            Post Syndicated from jzb original https://lwn.net/Articles/1071700/

            Security updates have been issued by AlmaLinux (dovecot, fence-agents, freeipmi, git-lfs, image-builder, kernel, libsoup, osbuild-composer, and python-tornado), Debian (apache2, libdatetime-timezone-perl, lrzip, tzdata, and wireshark), Fedora (dovecot, forgejo-runner, gh, gnutls, krb5, nano, pdns, pyOpenSSL, squid, vim, and xorg-x11-server-Xwayland), Mageia (graphicsmagick, kernel-linus, krb5-appl, libexif, libtiff, nano, nginx, ntfs-3g, opam, perl-Net-CIDR-Lite, perl-Starlet, perl-Starman, tcpflow, and virtualbox), Oracle (dovecot, fence-agents, freeipmi, image-builder, kernel, libcap, LibRaw, libsoup, openssh, osbuild-composer, python, python-tornado, python3, systemd, thunderbird, and tigervnc), SUSE (containerd, curl, erlang, flatpak, java-11-openjdk, java-21-openjdk, java-25-openjdk, liblxc-devel, libpng12, libthrift-0_23_0, openCryptoki, openexr, openssl-3, python3, python311-social-auth-core, rclone, skim, and thunderbird), and Ubuntu (apache2, coin3, editorconfig-core, insighttoolkit, linux, linux-aws, linux-aws-6.17, linux-gcp, linux-gcp-6.17, linux-hwe-6.17, linux-oracle, linux-realtime, linux-realtime-6.17, linux-azure, linux-azure-6.17, linux-oem-6.17, linux-azure-5.15, linux-gcp-6.8, nghttp2, python-dynaconf, slurm-wlm, swish-e, and webkit2gtk).

            AMD Intros Instinct MI350P Accelerator: CDNA 4 Comes to PCIe Cards

            Post Syndicated from Ryan Smith original https://www.servethehome.com/amd-intros-instinct-mi350p-accelerator-cdna-4-comes-to-pcie-cards/

            AMD has released a PCIe version of its flagship MI350 accelerators, the MI350P. Half of a MI350X, the card is aimed at customers who need to fit an modern AI accelerator into a traditional PCIe server

            The post AMD Intros Instinct MI350P Accelerator: CDNA 4 Comes to PCIe Cards appeared first on ServeTheHome.

            Why Security in 2026 Requires Continuous Threat and Exposure Management (CTEM) at Scale

            Post Syndicated from James Davis original https://www.rapid7.com/blog/post/em-2026-cybersecurity-requires-ctem-at-scale

            Let’s be honest, the patching window just shrank to something no practitioner or organization can keep up with. Organizations now need to operate in an environment that must assume breach, which means fundamentals like attack surface management, micro-segmentation, identity management, and attack path validation – aka a few core pillars of CTEM – just became the most important initiatives within the cybersecurity department. Rapid7 is the only vendor that provides a truly unified platform to master Continuous Threat Exposure Management (CTEM).

            How Rapid7 satisfies all 5 steps of the CTEM Framework

            Steps 1 and 2: Scoping and Discovery

            Achieving full visibility

            Rapid7 eliminates “unknown unknowns” by providing line-of-sight into 100% of your hybrid attack surface.

            • Surface Command (CAASM): We establish a single source of truth by unifying asset and identity inventory from over 200 third-party vendors and native sources.

            • Vulnerability Management: Our full-stack active scanning discovers shadow IT hidden within your enterprise network.

            • External Attack Surface Management (EASM): We scan the entire IPv4 space of the internet to automatically track changes to registered domains and public networks so you can map your external kingdom.

            • Unified CNAPP (Cloud Security): Our platform provides real-time, agentless visibility into every resource running across your multi-cloud environment (AWS, Azure, GCP, and Kubernetes). Through Event-Driven Harvesting (EDH), we identify infrastructure changes in under 60 seconds. This allows us to map not just the assets, but the complex identities and permissions that define your cloud risk.

            Step 3: Prioritization

            Moving beyond static scores

            We replace generic risk scores with Active Risk and Threat-Aware Context. Our platform automatically prioritizes vulnerabilities based on real-world exploitability data from Rapid7 Labs and the Exploit Prediction Scoring System (EPSS). We are also able to incorporate your own organization’s tagging infrastructure to properly contextualize your enterprise so you focus on what matters most. 

            Step 4: Validation

            Continuous human-led red teaming 

            This is where Rapid7 truly stands apart from automated-only vendors or point-in-time pen tests. Vector Command provides the expert human logic needed to bypass compensating controls like WAFs that stop automated tools cold. This gives Rapid7 the ability to answer the question: “How would an attacker get in?” We fully map the attack chain from the external to the internal so you have insight into where your controls are weakest.
            Ed Montgomery at Rapid7 has written extensively about the power of Vector Command – you can find his blogs here.
            Here’s a sampling of a couple of those stories: 

            • The Telerik UI Example: While a scanner flags an old version of Telerik, our operators discovered they could bypass a WAF by splitting a malicious payload into 118 individual, “harmless” fragments. We bypassed the WAF and this achieved full remote code execution that a time-boxed, two-week pentest would never have uncovered. An automated scan might have flagged the outdated telerik as something notable but it was really the configuration of the WAF that allowed us to bypass. Something an automated scan would never have found. 

            • SaaS Phishing: Our team used a misconfigured public Jira instance that allowed self-registration to hijack an Office 365 session and move laterally through internal trust. This validated that the true risk was a SaaS misconfiguration, not a patchable CVE.

            Step 5: Mobilization

            Instant response and remediation 

            We don’t just find problems; we close the loop with integrated action.

            • Cloud Runtime Security (CADR): Powered by our partnership with ARMO, our eBPF-based sensor can shut down an attack in seconds by killing malicious processes or pausing containers at the moment of detection.

            • Automation (SOAR): InsightConnect and our “Bot Factory” in CNAPP trigger automated remediation workflows to lock down S3 buckets or disable compromised users instantly.

            • Remediation Hub: We provide a centralized, vendor agnostic action-driven list of prioritized fixes to coordinate seamlessly with IT teams.

            CTEM-rapid7-framework.png

            The new standard: From weeks to minutes

            If your CTEM strategy relies on static tools and annual checkboxes, you are not just behind the curve. You are operating in a completely different era. By unifying the full visibility of Surface Command with the critical thinking of Vector Command and the instant response of our Cloud Runtime capabilities, Rapid7 empowers you to take command of your attack surface.

            Do not wait for a 118 single bit request bypass to prove your defenses are porous. Move from a posture of passive observation to one of preemptive security.

            How Cloudflare responded to the “Copy Fail” Linux vulnerability

            Post Syndicated from Chris J Arges original https://blog.cloudflare.com/copy-fail-linux-vulnerability-mitigation/

            On April 29, 2026, a Linux kernel local privilege escalation vulnerability was publicly disclosed under the name “Copy Fail” (CVE-2026-31431). Cloudflare’s Security and Engineering teams began assessing the vulnerability as soon as it was disclosed. We reviewed the exploit technique, evaluated exposure across our infrastructure, and validated that our existing behavioral detections could identify the exploit pattern within minutes. 

            There was no impact to the Cloudflare environment, no customer data was at risk, and no services were disrupted at any point. Read on to learn how our preparedness paid off. 

            Background

            Our Linux kernel release process

            Cloudflare operates a global Linux server infrastructure at an immense scale, with datacenters located across 330 cities. We maintain a custom Linux kernel build based on the community’s Long-Term Support (LTS) versions to manage updates effectively at this volume. At any given time, we may utilize multiple LTS versions from various series, such as 6.12 or 6.18, which benefit from extended update periods.

            The community regularly merges and releases security and stability updates which trigger an automated job to generate a new internal kernel build approximately every week. These builds undergo testing in our staging data centers to ensure stability before a global rollout. Following a successful release, the Edge Reboot Release (ERR) pipeline manages a systematic update and reboot of the edge infrastructure on a four-week cycle. Our control plane infrastructure typically adopts the most recent kernel, with reboots scheduled according to specific workload requirements.

            By the time a CVE becomes public knowledge, the necessary fix has typically been integrated into stable Linux LTS releases for several weeks. Our established procedures ensure that we have already deployed these patches.

            At the time of the “Copy Fail” disclosure, the majority of our infrastructure was running the 6.12 LTS version, while a subset of machines had begun transitioning to the newer 6.18 LTS release.

            About the Copy Fail vulnerability

            It helps to understand the vulnerability before getting to the response story. A comprehensive write-up can be found in the original Xint Code disclosure post.

            AF_ALG and the kernel crypto API

            The Linux kernel’s internal crypto API manages functions like kTLS and IPsec. Userspace programs access this via the AF_ALG socket family, allowing unprivileged processes to request encryption or decryption. The algif_aead module facilitates this for Authenticated Encryption with Associated Data (AEAD) ciphers.

            An unprivileged program follows these steps:

            1. Opens an AF_ALG socket and binds to an AEAD template.

            2. Sets a key and accepts a request socket.

            3. Submits input via sendmsg() or splice().

            4. Executes the operation using recvmsg().

            The splice() system call is critical here, as it moves data by passing page cache references.

            Memory mechanics: page cache and in-place crypto

            The page cache is a shared system cache for file contents. Modifying a page belonging to a setuid binary effectively edits that program for all users until the page is evicted.

            The crypto API utilizes scatterlists, which are structures linking various memory pages. In 2017, algif_aead was optimized for in-place operations, chaining destination and reference pages together. This design lacked enforcement to prevent algorithms from writing past intended boundaries.

            The vulnerability: out-of-bounds write

            When the user executes recvmsg(), the authencesn wrapper in the kernel performs a 4-byte write past the legitimate output region:

            scatterwalk_map_and_copy(tmp + 1, dst, assoclen + cryptlen, 4, 1);
            

            By using splice(), an attacker can chain a target file’s page cache pages to the scatterlist. The out-of-bounds write then taints the cached file, allowing an attacker to control which file is modified, the offset, and the specific 4 bytes written. This means the attacker can manipulate the following with this exploit:

            • File: Any readable file.

            • Offset: Tunable via assoclen and splice parameters.

            • Value: Controlled via AAD bytes 4-7 in sendmsg()

            The exploit, step by step


            The default exploit targets /usr/bin/su, a setuid-root binary present on essentially every distribution.

            1. Cache Reference: Open /usr/bin/su as O_RDONLY and read() to populate the page cache. Use splice() on the file descriptor to pass these page cache references into the crypto scatterlist.

            2. Setup: Create an AF_ALG socket, bind() to authencesn(hmac(sha256),cbc(aes)), set a key, and accept a request socket without needing privileges.

            3. Write Construction: For each 4-byte shellcode chunk:

              • sendmsg() with AAD bytes 4–7 containing the shellcode.

              • splice() the binary into a pipe then the AF_ALG socket so assoclen + cryptlen targets the desired .text offset.

            4. Trigger: recvmsg() initiates decryption. authencesn writes its scratch data to the target offset of /usr/bin/su in the page cache. Although the function returns -EBADMSG, the 4-byte write is now in the global page cache.

            5. Execution: Running execve("/usr/bin/su") loads the tainted page cache. Since the binary is setuid-root, the injected shellcode executes with root privileges.

            The upstream fix (commit a664bf3d603d) reverts the 2017 in-place optimization, removing the exploit.

            How we responded 

            When the vulnerability was disclosed, many workstreams started in parallel:

            • Mapping the blast radius: Our security team worked with kernel engineers to determine which kernel versions were vulnerable and assess the potential exposure.

            • Validating coverage: Security reviewed the exploit technique and confirmed that our existing behavioral detections could identify the exploit pattern during authorized internal validation.

            • Proactive threat hunting: Security began searching for signs that the vulnerability had been exploited before it was publicly known, going back 48 hours in our fleet-wide logs.

            • Engineering a mitigation: Kernel engineers began building a runtime mitigation that would protect the fleet without breaking production services.

            • Continuing software updates: Our engineering teams worked on delivering an updated Linux kernel, which required carefully rebooting and rolling it out across our servers.

            There was no customer impact at any point during this response.

            Validating detection coverage

            One of the first things our security team did was confirm that our existing endpoint detection would catch this exploit. Our servers run behavioral detection that continuously monitors process execution patterns. It doesn’t rely on knowing about specific vulnerabilities; it watches for anomalous behavior across the fleet.

            When our engineers validated the vulnerability internally as part of the response, the detection platform flagged it within minutes. The system linked the entire execution chain—starting at the script interpreter, moving through the kernel’s cryptographic subsystem, and ending at the privilege escalation binary—flagging it as malicious based on fleet-wide behavioral patterns.

            This happened without a signature update, without a rule change, and without human intervention. Our behavioral detection coverage existed before we wrote any custom logic for this particular Copy File exploit. 

            The confirmation was important because it meant we had coverage before writing a vulnerability-specific rule.

            Hunting for exploitation

            While our engineering team moved to a more targeted mitigation, our security investigation had been running since disclosure. This is our standard procedure for any critical vulnerability.

            Our security team operates on a simple principle for critical vulnerabilities: assume compromise until you can prove otherwise. The investigation started from the assumption that exploitation could have occurred before the vulnerability was public, and we worked systematically to either confirm or rule it out.

            The exploit leaves a distinctive trace in kernel logs when it runs. We searched for that trace across our centralized logging infrastructure, covering 48 hours before the vulnerability was publicly disclosed. If someone had exploited this before the world knew about it, we would have seen it.

            We pulled access logs for affected systems and reconstructed who connected, when, and what commands they ran. This gave us a complete forensic picture of interactive activity on potentially affected infrastructure.

            We checked that system binaries had not been tampered with, validated cryptographic hashes against known-good package manifests, looked for persistence mechanisms, and audited network connections for anything unusual. Everything was clean.

            Incident timeline and impact

            Time (UTC) Event
            2026-04-29 16:00 Copy Fail publicly disclosed.
            2026-04-29 ~21:00 Security and Engineering teams began assessing fleet exposure and mitigation options before full declaration of the Incident Response process
            2026-04-29 22:52 Security confirmed existing behavioral detection covered the Copy Fail exploit pattern. During authorized internal validation, detection flagged the activity within minutes.
            2026-04-29 23:01 Existing behavioral detection generated a high-severity alert for exploit-like activity, confirming detection coverage for the technique.
            2026-04-29 (evening) First mitigation attempt pushed to our staging datacenter. The deployment process surfaced a dependency conflict; the mitigation was rolled back. No production systems were affected.
            2026-04-29 (overnight) Engineering drafted bpf-lsm mitigation program.
            2026-04-30 03:14 Security incident declared to drive cross-functional collaboration and urgency. Security performed fleetwide threat hunting of historical data to confirm that no malicious activity was present on Cloudflare systems.
            2026-04-30 (morning) Engineering tested the bpf-lsm mitigation program and made it production-ready.
            2026-04-30 14:25 Engineering incident declared to coordinate mitigation program and Linux patch rollout.
            2026-04-30 ~17:00 Decision made: ship a patched build of the previous LTS line through reboot automation; do not accelerate the new LTS; lean on bpf-lsm in the meantime.
            2026-04-30 (afternoon) Visibility pipeline (eBPF tracing of AF_ALG socket usage) deployed fleet-wide. Gives a complete picture of all legitimate AF_ALG users.
            2026-04-30 (evening) bpf-lsm mitigation program rolled out behind a separate gate to fully mitigate the fleet. End-to-end verification on a previously-vulnerable test node confirms the exploit no longer works.
            2026-05-04 (morning) Reboot automation resumed at normal pace with the patched kernel.
            2026-05-04 onward Servers that had already passed through reboot automation earlier in the week manually rebooted to pick up the patched kernel. Unpatched servers update per our normal reboot automation.

            This graph shows the progress of our mitigation program as it progressed through our infrastructure.

            How did we mitigate it?

            Because of the long timeframe involved in deploying a patched Linux kernel, we also pursued mitigating this exploit without a reboot.

            Removing the module

            The bug was in the algif_aead kernel module. Therefore, the simple fix was to just remove this module and disallow it from being reloaded.

            This mitigation was therefore exactly what the Copy Fail write-up from the security researchers who identified it recommends.

            echo "install algif_aead /bin/false" > /etc/modprobe.d/disable-algif.conf
            rmmod algif_aead 2>/dev/null || true

            Unfortunately removing the module would have impacted software that leverages the kernel crypto API.  This meant that we had to figure out a more surgical mitigation.

            Bpf-lsm


            We’ve already developed and deployed such a tool for this exact scenario: bpf-lsm. Instead of removing the module, this tool leaves it loaded for legitimate users and uses a BPF Linux Security Module program to deny the socket_bind LSM hook for everyone else. This completely blocks the front door for any exploits.

            A draft of the eBPF program was put together overnight. Team members picked it up the following morning, ran validations, and made it production-ready. The program is fairly straightforward. On every socket_bind call:

            1. If the socket family is not AF_ALG, allow the call through unchanged.

            2. If the family is AF_ALG, check the calling binary’s path against an allow-list of the binaries we know to be legitimate users.

            3. If the binary is on the allow-list, allow the bind. Otherwise, deny it.

            To verify the mitigation on a given machine without exploiting it, the Copy Fail write-up gives a one-liner:

            python3 -c 'import socket; s = socket.socket(socket.AF_ALG, socket.SOCK_SEQPACKET, 0); s.bind(("aead","authencesn(hmac(sha256),cbc(aes))"));'

            On a mitigated machine you get PermissionError: [Errno 1] Operation not permitted (or FileNotFoundError, depending on which mitigation is active) instead of a successful bind.

            Rolling it out

            Before enabling enforcement, we verified that our known internal service was the sole legitimate AF_ALG user to avoid accidental outages. We used prometheus-ebpf-exporter to hook the socket() syscall and track AF_ALG usage per binary across the fleet. This required no kernel changes and provided aggregate data from hundreds of thousands of servers within hours. Results confirmed the identified service was indeed the only legitimate user.

            So the bpf-lsm rollout was deliberately staged in two steps:

            1. Get visibility first. Push the ebpf-exporter config gated by salt. Confirm at the metric layer that the known service is effectively the only thing creating AF_ALG sockets.

            2. Then enforce. Push the bpf-lsm program behind a separate enforcement gate.

            In parallel, the upstream backport for our majority LTS line finally became available, and our internal automation built a patched kernel against it.

            We started to test the patched kernel in our staging datacenters as soon as possible, then we resumed the longer reboot process in order to fully patch our fleet.

            Remediation and follow-up steps

            While we were prepared for this scenario, at Cloudflare we’re always learning and improving. Key areas we identified for improvement:

            • Better visibility into kernel-API dependencies. We will review kernel-subsystem usage across production services, so we can continue to quickly mitigate exploits without service disruption.

            • Better runtime mitigation. bpf-lsm is a valuable tool for mitigations, but we want to make this tool even better. This will include looking into faster deployments, better playbooks, and better logging and visibility of the tool. 

            • Reduce attack surface of Linux Kernel. Review and audit our kernel configuration. Proactively identify unused modules or features so that we can remove them from our build entirely.

            Conclusion

            The “Copy Fail” vulnerability presented a unique challenge for us. Despite our practice of deploying Linux patch updates every two weeks, we remained vulnerable because a month-old mainline fix had yet to be backported to our primary kernel line. Despite that, we were still able to roll out patched kernels within hours of the backport’s release. In the interim, bpf-lsm provided a surgical, no-reboot mitigation that secured our fleet. While our initial attempt to disable the problematic module failed, it did so safely within our internal staging environment rather than production, allowing us to identify this dependency.

            By the end of the rollout, every machine in our fleet was protected by either a patched kernel or a bpf-lsm program denying the vulnerable code path to non-allow-listed binaries. There was no customer impact at any point during this incident, and we have committed to the follow-up work above to make our response faster and our visibility better the next time something like this lands. Responsible disclosure works, in-kernel visibility tooling pays off in moments exactly like this one, and bpf-lsm continues to be one of the most useful primitives we have for runtime kernel mitigation.

            At Cloudflare, critical vulnerability response is a coordinated effort across Security, Engineering, Product, and many other teams. Special thanks to Ali Adnan, Ivan Babrou, Frederik Baetens, Curtis Bray, Piers Cornwell, Everton Didone Foscarini, Rob Dinh, Elle Dougherty, Kevin Flansburg, Matt Fleming, Kimberley Hall, Brandon Harris, Jerry Ho, Oxana Kharitonova, Marek Kroemeke, Fred Lawler, James Munson, Nafeez Nazer, Walead Parviz, Miguel Pato, Evan Pratten, Josh Seba, June Slater, Ryan Timken, Michael Wolf, Jianxin Zeng and everyone else who contributed to the investigation, mitigation, and remediation of Copy Fail. We’d also like to thank the Linux upstream maintainers and Copy Fail researchers whose work helped make a rapid response possible.

            The collective thoughts of the interwebz