Quality-of-service (QoS) mechanisms attempt to prioritize some processes (or
network traffic, disk I/O, etc.) over others in order to meet a system’s
performance goals. This is a difficult topic to handle in the world of Linux,
where workloads, hardware, and user expectations vary wildly. Qais Yousef spoke
at the 2025 Linux Plumbers Conference, alongside his collaborators John Stultz,
Steven Rostedt, and Vincent Guittot, about their plans for introducing a
high-level QoS API for Linux in a way that leaves end users in control of its
configuration. The talk focused specifically on a QoS mechanism for the
scheduler, to prioritize access to CPU resources differently for different kinds
of processes.
(slides; video)
Amazon Elastic Kubernetes Service (Amazon EKS) on AWS Outposts brings the power of managed Kubernetes to your on-premises infrastructure. Use Amazon EKS on Outposts rack to create hybrid cloud deployments that maintain consistent AWS experiences across environments. As organizations increasingly adopt edge computing and hybrid architectures, storage optimization and performance tuning become critical for successful workload deployment.
Outposts extend AWS infrastructure, services, APIs, and tools to virtually any datacenter, co-location space, or on-premises facility. In this blog post you will learn about your storage options and their performance characteristics which is essential for building resilient, high-performing applications using Amazon EKS on Outposts.
Amazon EKS extended clusters on Outposts provide a powerful solution for organizations seeking to use the benefits of Kubernetes while maintaining certain workloads on-premises, as shown in the following figure. This hybrid architecture allows businesses to extend their EKS clusters from the AWS Cloud to their own data centers or edge locations using Outposts. The Kubernetes control plane remains in the AWS Region, providing centralized management and benefiting from the AWS infrastructure in the cloud and on the Outpost.
Outposts is designed to be a connected service, and needs reliablenetwork connectivity to the AWS Region using the Outposts service link.
Figure 1 : Extended cluster
Amazon EKS local cluster architecture
Amazon EKS local clusters deploy the Kubernetes control plane on your Outpost, as shown in the following figure. This provides greater network resilience against outages as cluster operations run entirely on the Outposts and reduces the dependency on network connectivity to the AWS Region. Having the Kubernetes control plane hosted on your Outpost also reduces latency for cluster operations.
Figure 2: Local cluster
Storage options for Amazon EKS extended clusters on Outposts
Persistent Volumes (PV) and Persistent Volume Claims (PVC) serve as a critical abstraction layer in Kubernetes, separating the storage consumption details from storage provisioning, and allowing administrators to manage storage resources independently from how applications consume them. PVs and PVCs make sure of data persistence across pod restarts and rescheduling events, making them essential for applications that need to maintain state, such as databases, file storage systems, and other data-intensive workloads. The abstraction provided by PV and PVC enables platform-agnostic storage management, where applications can request storage through PVCs without needing to know the underlying storage implementation details. PVs and PVCs support dynamic provisioning through Storage Classes, allowing for automated storage allocation based on application demands, while also providing features such as access modes, capacity management, and reclaim policies to effectively manage the storage lifecycle in a Kubernetes cluster.
Integrating Amazon EBS with Amazon EKS
Amazon Elastic Block Store (Amazon EBS) provides high-performance block storage that’s ideal for low-latency applications providing consistent performance. When deployed on Outposts racks, EBS volumes are stored on the Outposts hardware, providing significant performance advantages over network-attached storage solutions, as shown in the following figure.
Figure 3 : Integrating Amazon EBS with Amazon EKS on Outposts
Benefits and use cases
Storage: EBS volumes on Outposts racks provide data access without dependency on external connectivity.
Performance: Local storage delivers consistent latency and high IOPS/throughput.
Cost: On-premises storage eliminates data transfer costs and reduces bandwidth needs, lowering the total cost of ownership.
Implementation considerations
Consider the following when using EBS on Outposts rack:
EBS volumes on Outposts are tied to a single rack and the availability zone the Outpost is homed to, needing applications to address single-point-of-failure risks.
Protect data using EBS snapshots in the parent Region and schedule regular backups.
Capacity on Outposts is finite, monitor Outposts storage usage and plan expansions proactively to avoid insufficient capacity errors.
Amazon Elastic File System (Amazon EFS) provides scalable, shared file storage that can be accessed across multiple AWS Availability Zones (AZs) and on-premises environments. Although Amazon EFS with Amazon EKS on Outposts maintains the same setup procedures as standard cloud deployments, there is a critical dependency on the service link connection between your Outposts and the AWS Region. Amazon EFS is not a locally supported service on Outposts, so connectivity to the AWS Region is required to use this service with your Outpost.
Amazon EFS allows multiple pods to concurrently access shared file systems. It is well-suited for applications that need collaborative data access, content management, and distributed processing workloads.
Amazon EFS as a persistent storage solution for Amazon EKS extended cluster instances
Amazon EFS as a PV for your Amazon EKS extended cluster operates through a hybrid architecture where the Amazon EFS file system resides in the Region, but mount points can be created on the worker nodes running on Outposts subnets through the service link as shown in the following figure.
Figure 4 : Amazon EFS as a persistent storage solution for extended clusters
Benefits and use cases
Shared storage capabilities: multiple pods can access a centralized file system, enabling shared data, code, and assets across instances.
Scalability: storage capacity and performance automatically scale with usage, eliminating manual provisioning and upfront planning.
Compliance: Amazon EFS provides full file system features and compatibility for traditional applications, such as locking, permissions, and directory structure.
Challenges and limitations
Consider the following when using Amazon EFS with Outposts:
Network latency: file access involves network traversal to Amazon EFS in the Region, adding more latency and making small or metadata operations potentially slow for latency-sensitive applications.
Throughput: aggregate throughput is restricted by the available bandwidth on the service link between the Outposts and AWS Region. This impacts concurrent access and large file transfers during peak usage.
Dependency on AWS Region connectivity: Amazon EFS needs continuous connectivity to the parent Region. Disruptions may affect file system availability, operations, and disaster recovery processes.
Data Transfer charges: Since EFS is in AWS Parent region and EKS worker nodes and pods are in Outpost additional charges are applicable.
Deploying pods on extended clusters using Amazon EFS as PV
Refer to Use Elastic File System Storage with Amazon EFS for deployment guidance. Note, Create Amazon EFS mount targets in subnets that are in the same Availability Zone (AZ) as the Outposts subnets.
Amazon S3 with Amazon EKS extended cluster
Amazon Simple Storage Service (Amazon S3) on Outposts delivers local object storage on your Outposts, allowing applications to use Amazon S3 APIs for storing and retrieving data while keeping it onsite. It is ideal for workloads that need Amazon S3 compatibility, low latency access to object data, and local data residency.
You should use Amazon S3 access point Amazon Resource Names (ARNs) and not bucket ARNs for proper integration with Amazon EKS workloads.
Figure 5 : Amazon S3 with Amazon EKS extended cluster on Outposts
Benefits and use cases
Data archiving and compliance: Enables cost-effective, locally retained storage for logs, audit trails, regulatory compliance, backups, and sensitive healthcare data with strict residency requirements.
Content distribution and media: Provides ultra-low latency local storage for serving static content, media streaming, digital asset management, and gaming asset delivery.
Data lake and analytics: Supports local data processing for analytics, ETL, machine learning (ML), real-time Internet of Things (IoT) data handling, and business intelligence with reduced latency and transfer costs.
Application integration: Seamlessly integrates with Amazon S3 compatible apps for backup, synchronization, microservices storage, API-driven workflows, and container image management on-premises.
Optimizing performance starts with selecting the right storage type for your workload: Amazon EBS for low-latency, high-throughput block storage; Amazon EFS for shared POSIX-compliant file systems; and Amazon S3 for scalable object storage with API compatibility. Ensure proper volume sizing, monitor usage proactively, and configure CPU and memory requests accurately to balance performance and efficiency—auto scaling and QoS classes can further optimize resource management. Improve data locality by using local storage, apply caching with intelligent eviction, and design for efficient, asynchronous, and compressed data access patterns.
Monitoring and observability
Monitoring key performance metrics is essential to maintain storage efficiency and application reliability. For Amazon EBS, track IOPS, throughput, latency, burst balance, queue depth, and snapshot performance to avoid degradation—see the Amazon CloudWatch metrics for Amazon EBS for the full list. For Amazon EFS, monitor total I/O, throughput, client connections, metadata operations, burst credits, and Regional data transfers to support effective capacity planning—refer to CloudWatch metrics for Amazon EFS. For Amazon S3, observe request and error rates, data transfer, storage usage, latency, multipart upload efficiency, and access patterns to optimize performance and cost—see Metrics and dimensions.
Security considerations
Strong security practices are critical for Amazon EKS on Outposts. Use AWS Key Management Service (AWS KMS) for Amazon EBS encryption, encrypt Amazon EFS data at rest and in transit, and enable server- or client-side encryption for Amazon S3. Enforce TLS for all data transfers and apply key rotation with compliance controls. Implement least privilege IAM policies, scoped roles, and Kubernetes Role-Based Access Control (RBAC) for granular pod access. Secure traffic with security groups and NACLs, and maintain audit logs for all storage operations.
Cost optimization strategies
Manage storage costs by right-sizing volumes, automating lifecycle policies, selecting appropriate storage classes, monitoring data transfer, and using de-duplication and compression where applicable. Lower operational expenses through automated backups, infrastructure as code (IaC), monitoring automation, leveraging managed services, applying cost allocation tags, and conducting regular usage reviews.
Conclusion
Amazon EKS on Outposts empowers organizations to build hybrid applications with storage options that align to performance, compliance, and data residency needs. By selecting the right storage solution for each workload and leveraging Outposts’ local infrastructure, you can reduce latency, minimize network dependencies, and maintain consistency across environments. As Outposts capabilities continue to evolve, they offer a strong foundation for modern, resilient, and cost-efficient hybrid cloud architectures.
Reach out to your AWS account team, or fill out this form to learn more about running containarized applications on Outposts.
A new version of AWS Security Hub, is now generally available, introducing new ways for organizations to manage and respond to security findings. The enhanced Security Hub helps you improve your organization’s security posture and simplify cloud security operations by centralizing security management across your Amazon Web Services (AWS) environment. The new Security Hub transforms how organizations handle security findings through advanced automation capabilities with real-time risk analytics, automated correlation, and enriched context that you can use to prioritize critical issues and reduce response times. Automation also helps ensure consistent response procedures and helps you meet compliance requirements.
AWS Security Hub CSPM (cloud security posture management) is now an integral part of the detection engines for Security Hub. Security Hub provides centralized visibility across multiple AWS security services to give you a unified view of your cloud environment, including risk-based prioritization views, attack path visualization, and trend analytics that help you understand security patterns over time.
This is the third post in our series on the new Security Hub capabilities. In our first post, we discussed how Security Hub unifies findings across AWS services to streamline risk management. In the second post, we shared the steps to conduct a successful Security Hub proof of concept (PoC).
We walk through the setup and configuration of automation rules, share best practices for creating effective response workflows, and provide real-world examples of how these tools can be used to automate remediation, escalate high-severity findings, and support compliance requirements.
Security Hub automation enables automatic response to security findings to help ensure critical findings reach the right teams quickly, so that they can reduce manual effort and response time for common security incidents while maintaining consistent remediation processes.
Note: Automation rules evaluate new and updated findings that Security Hub generates or ingests after you create them, not historical findings. These automation capabilities help ensure critical findings reach the right teams quickly.
Why automation matters in cloud security
Organizations often operate across hundreds of AWS accounts, multiple AWS Regions, and diverse services—each producing findings that must be triaged, investigated, and acted upon. Without automation, security teams face high volumes of alerts, duplication of effort, and the risk of delayed responses to critical issues.
Manual processes can’t keep pace with cloud operations; automation helps solve this by changing your security operations in three ways. Automation filters and prioritizes findings based on your criteria, showing your team only relevant alerts. When issues are detected, automated responses trigger immediately—no manual intervention needed.
If you’re managing multiple AWS accounts, automation applies consistent policies and workflows across your environment through centralized management, shifting your security team from chasing alerts to proactively managing risk before issues escalate.
Designing routing strategies for security findings
With Security Hub configured, you’re ready to design a routing strategy for your findings and notifications. When designing your routing strategy, ask whether your existing Security Hub configuration meets your security requirements. Consider whether Security Hub automations can help you meet security framework requirements like NIST 800-53 and identify KPIs and metrics to measure whether your routing strategy works.
Security Hub automation rules and automated responses can help you meet the preceding requirements, however it’s important to understand how your compliance teams, incident responders, security operations personnel, and other security stakeholders operate on a day-to-day basis. For example, do teams use the AWS Management Console for AWS Security Hub regularly? Or do you need to send most findings downstream to an IT systems management (ITSM) tool (such as Jira or ServiceNow) or third-party security orchestration, automation, and response (SOAR) platforms for incident tracking, workflow management, and remediation?
Next, create and maintain an inventory of critical applications. This helps you adjust finding severity based on business context and your incident response playbooks.
Consider the scenario where Security Hub identifies a medium-severity vulnerability on an Elastic Compute Cloud instance. In isolation, this might not trigger immediate action. When you add business context—such as strategic objectives or business criticality—you might discover that this instance hosts a critical payment processing application, revealing the true risk. By implementing Security Hub automation rules with enriched context, this finding can be upgraded to critical severity and automatically routed to ServiceNow for immediate tracking. In addition, by using Security Hub automation with Amazon EventBridge, you can trigger an AWS Systems Manager Automation document to isolate the EC2 instance for security forensics work to then be carried out.
Because Security Hub offers OCSF format and schema, you can use the extensive schema elements that OCSF offers you to target findings for automation and help your organization meet security strategy requirements.
Example use cases
Security Hub automation supports many use cases. Talk with your teams to understand which fit your needs and security objectives. The following are some examples of how you can use security hub automation:
Automated finding remediation
Use automated finding remediation to automatically fix security issues as they’re detected.
Supporting patterns:
Direct remediation: Trigger AWS Lambda functions to fix misconfigurations
Resource tagging: Add tags to non-compliant resources for tracking
Configuration correction: Update resource configurations to match security policies
Step 2: Create automation rules to update finding details and third-party integration
After Security Hub collects findings you can create automation rules to update and route the findings to the appropriate teams. The steps to create automation rules that update finding details or to a set up a third-party integration—such as Jira or ServiceNow—based on criteria you define can be found in Creating automation rules in Security Hub.
With automation rules, Security Hub evaluates findings against the defined rule and then makes the appropriate finding update or calls the APIs to send findings to Jira or ServiceNow. Security Hub sends a copy of every finding to Amazon EventBridge so that you can also implement your own automated response (if needed) for use cases outside of using Security Hub automation rules.
In addition to sending a copy of every finding to EventBridge, Security Hub classifies and enriches security findings according to business context, then delivers them to the appropriate downstream services (such as ITSM tools) for fast response.
Best practices
AWS Security Hub automation rules offer capabilities for automatically updating findings and integrating with other tools. When implementing automation rules, follow these best practices:
Centralized management: Only the Security Hub administrator account can create, edit, delete, and view automation rules. Ensure proper access control and management of this account.
Regional deployment: Automation rules can be created in one AWS Region and then applied across configured Regions. When using Region aggregation, you can only create rules in the home Region. If you create an automation rule in an aggregation Region, it will be applied in all included Regions. If you create an automation rule in a non-linked Region, it will be applied only in that Region. For more information, see Creating automation rules in Security Hub.
Define specific criteria: Clearly define the criteria that findings must match for the automation rule to apply. This can include finding attributes, severity levels, resource types, or member account IDs.
Understand rule order: Rule order matters when multiple rules apply to the same finding or finding field. Security Hub applies rules with a lower numerical value first. If multiple findings have the same RuleOrder, Security Hub applies a rule with an earlier value for the UpdatedAt field first (that is, the rule which was most recently edited applies last). For more information, see Updating the rule order in Security Hub.
Provide clear descriptions: Include a detailed rule description to provide context for responders and resource owners, explaining the rule’s purpose and expected actions.
Use automation for efficiency: Use automation rules to automatically update finding fields (such as severity and workflow status), suppress low-priority findings, or create tickets in third-party tools such as Jira or ServiceNow for findings matching specific attributes.
Consider EventBridge for external actions: While automation rules handle internal Security Hub finding updates, use EventBridge rules to trigger actions outside of Security Hub, such as invoking Lambda functions or sending notifications to Amazon Simple Notification Service (Amazon SNS) topics based on specific findings. Automation rules take effect before EventBridge rules are applied. For more information, see Automation rules in EventBridge.
Manage rule limits: This is a maximum limit of 100 automation rules per administrator account. Plan your rule creation strategically to stay within this limit.
Regularly review and refine: Periodically review automation rules, especially suppression rules, to ensure they remain relevant and effective, adjusting them as your security posture evolves.
Conclusion
You can use Security Hub automation to triage, route, and respond to findings faster through a unified cloud security solution with centralized management. In this post, you learned how to create automation rules that route findings to ticketing systems integrations and upgrade critical findings for immediate response. Through the intuitive and flexible approach to automation that Security Hub provides, your security teams can make confident, data-driven decisions about Security Hub findings that align with your organization’s overall security strategy.
With Security Hub automation features, you can centrally manage security across hundreds of accounts while your teams focus on critical issues that matter most to your business. By implementing the automation capabilities described in this post, you can streamline response times at scale, reduce manual effort, and improve your overall security posture through consistent, automated workflows.
The rise of the neocloud and open cloud ecosystem are dethroning the major cloud providers. Companies like Vultr, Akamai, and CoreWeave are proving that developers don’t need a walled garden to build world-class applications. Instead, teams can build their own best-of-breed stacks and choose specialized providers that do just one thing exceptionally well, such as high-performance compute, AI inference, or databases.
But this new stack creates a critical decision point: What about storage?
For the neocloud model to work, data needs to be open and flexible. Consider the economics of the split-stack. If a neocloud lacks a robust storage layer, the customer often defaults to keeping their data with a major cloud company like AWS. This creates a financial trap where every time the neocloud application needs to process that data, it must retrieve it from the big three cloud providers’ walled gardens. The resulting egress fees can negate the cost savings of moving to a neocloud provider in the first place.
For a technical-first neocloud CTO, the default answer is almost reflexive. “We’ll build it. It’s just storage. How hard can it be?”
It’s a fair question, and it usually kicks off an internal engineering debate that sounds something like this:
Should we use Ceph on commodity hardware, or design our own purpose-built architecture?
Do we currently have the data center capacity to stand up multi-petabyte-scale storage clusters?
Can we repurpose our existing older generation hardware, or do we need to order all-new equipment?
Is commodity object storage really the best use of our precious rack space, or would offering this service require an expensive data center expansion?
These are the right questions. But the answers often lead to a trap.
The build trap: When storage becomes the wrong problem
Whether you choose the software route (Ceph) or the hardware route (purpose-built), you aren’t just adding a new service feature. Both paths risk distracting your best engineers from your core business by forcing them to master a new (and expensive) specialty: storage infrastructure.
This guide reveals the true costs of both “build” paths, and why the smartest move for a neocloud might be to not build at all.
Path 1: The Ceph approach
On paper, Ceph is the obvious choice. It’s open-source and scalable. It unifies object, block, and file storage. And it runs on commodity hardware.
But people and complexity make Ceph more expensive than you’d think.
First of all, Ceph is not “set it and forget it.” It is notoriously complex to deploy, manage, and tune at petabyte scale. That means you can’t just assign Ceph to a junior sysadmin. You’ll have to hire and retain a dedicated team of expensive, hard-to-find Ceph specialists.
This creates a massive resource drain. Instead of allocating your engineering headcount to build your next great compute plan or AI feature, you’re burning the budget on operational overhead. And that overhead is significant because getting real-world performance out of Ceph depends on a deep, constant tuning of CRUSH maps, OSDs, and the perfect (and constantly evolving) co-design of the underlying hardware.
Ultimately, Ceph isn’t just a software choice; it’s a strategic commitment to building a storage operations division that works in tandem with all your other ops and engineering teams.
Path 2: The “purpose-built” approach
This approach gives you a lot of control. You can design hardware specifically for your cost model, data center layout, and performance needs.
But it comes with a catch: it means becoming a hardware R&D company.
Trust us, we know. It’s the path we took over 15 years ago. It worked for us, but looking back, it only made sense for two reasons:
The era. The operational realities of cloud storage were different back in 2007 when we launched the company. We simply didn’t have the options—and therefore, the competition around pricing and features—that we have now.
The pivot. We very quickly shifted our focus from being a single-product, consumer-focused company to a cloud storage provider whose first customer was Backblaze Computer Backup. We chose to double down on the infrastructure investment we’d made to support that scale.
The reality of the R&D treadmill
Our original Storage Pod, which we open sourced in the Petabytes on a Budget blog, required deep R&D to design a custom chassis, source specific components, and solve physics problems such as mass drive vibration and power draw.
However, solving those physics problems once was just the beginning. To stay competitive, we had to keep innovating. In fact, we’ve gone through seven major versions of our Storage Pods (1.0, 2.0, 3.0, 4.0, 4.5, 5.0, 6.0). After all that R&D, we eventually found that the build/buy incentives had flipped and commodification had finally caught up.
But to even make the decision to stop building custom chassis, we had to perform the same kind of testing we did for every previous version. Each iteration required new engineering to solve for higher drive densities, extended chassis lengths, changing cooling needs, and updated networking.
The operational reality
You’re not just “one and done” on drives or servers. Data centers are in constant flux. You are continually replacing old drives with new ones in existing chassis. Each time a new drive model enters the fleet, it must go through extensive testing to ensure it improves (or at least maintains) operations within the data center environment.
And drives are just one part of the equation. You also have to get files into and out of the data center efficiently. This requires constant, forward-thinking improvements in areas such as:
In other words, choosing the purpose-built path is a strategic commitment to becoming a full-time hardware and software engineering, supply chain, cybersecurity, and logistics company. If that sounds exhausting, that’s because it is.
Choose your distraction
The ultimate choice you need to make isn’t Ceph vs. purpose-built. The choice is, which resource-draining specialty do you want your product, engineering, and ops teams to be distracted by?
Do you want your best (and most expensive) engineers spending their days troubleshooting esoteric Ceph tuning parameters? Or would you rather have them re-designing a server chassis to introduce new CPU and GPU hardware and figuring out how to add essential security features with minimal overhead?
The answer is neither.
You want them focused on your specialty—building a better compute service, a faster AI model, or a more resilient database. Every hour they spend fighting with storage infrastructure is an hour they aren’t spending on the product your customers actually pay for.
The ideal solution: Storage as a specialty partner
This is why we exist.
Backblaze was built on the fundamental belief that storage is a specialty. We’ve spent 15+ years solving these hardware and operational problems so that you don’t have to.
A symbiotic relationship
The neocloud ecosystem thrives on interoperability. It functions best not when every provider tries to build the full stack, but when they connect with independent, open, and easy-to-use layers.
When neoclouds partner with Backblaze, the dynamic shifts from building to enabling. You gain a petabyte-scale storage layer that is:
Instantly available. No lead times, no hardware sourcing, no build-out.
S3 compatible. It fits seamlessly into your existing tools and your customers’ workflows.
Zero overhead. None of the R&D distraction and operational weight we outlined above.
We provide the foundational storage that enables the entire open cloud ecosystem to compete on equal footing against the “Big Three” cloud providers.
Focus on what makes your neocloud great. Let Backblaze handle the storage. Learn more about Powered by Backblaze, or reach out to our storage experts to start a conversation.
Every week, young people around the world gather in libraries, classrooms, community centres, and makerspaces to create with code. From Gujarat to Glasgow, Nairobi to New Jersey, the settings may differ, but the energy is unmistakable.
Code Club meeting at the shared hub at AEF Reuben in Kenya
We set out to learn from the wealth of experiences within the global Code Club community: how clubs adapt to local needs, and which practices consistently support young people’s learning. Although we found small differences in how clubs make Code Club work locally, what stands out far more are the shared principles that make it work everywhere. Here we share stories from across our network that collectively paint a vibrant picture of what makes this movement work.
Inspiration from the people who make Code Club thrive
One theme runs through every story: Code Club is powered by people who really understand the needs of their community.
During a visit to a set of Code Clubs in India, our team met a group of girls who once faced barriers to attending school. Now, they are confidently creating Scratch projects and exploring new technologies. Their club leader explained how a simple change — allowing girls to attend school wearing traditional attire — opened doors for families. Seeing these young creators code with pride is a vivid reminder of how opportunity can reshape futures.
Welspun Vapi Code Club in Gujarat
A club leader at Better Juniors Digital Club at Better Life Primary and Junior School in Kenya described how their programme began with just one laptop. Rather than letting that limit what learners could do, he found solutions everywhere: applying for grants, borrowing digital space from a nearby hub, and setting up equipment so children could work on projects together. His determination effectively created a bridge between schools and resources, opening up real opportunities for every child to learn.
We also heard from educators whose clubs have become long-standing pillars of their communities. At Rhiwbina Library in Wales, leaders have been running Code Club for over a decade, creating a space where older creators naturally guide new ones. When asked about club rules, one child replied: “There’s only one and that’s ‘respect’.” That simple principle continues to shape a joyful, collaborative atmosphere.
Fiona Lindsay and pupils at Hillside Primary School in Scotland
And sometimes the inspiration comes from the young people themselves. At Hillside Primary in Scotland, an enthusiastic creator took it upon himself to run taster sessions and codealongs for new members, helping them discover whether Code Club was right for them. His enthusiasm and leadership were infectious, and that spirit of young people lifting up their peers is something we’ve seen in clubs all over the world.
Moments of joyful learning capture the spirit of Code Club
In one club that meets across three different venues in Pennsylvania, USA — a creative arts centre, a coffee shop, and a library — the excitement became contagious. The librarian, Miss Sandy, was so inspired by the learners’ projects — including the moment they added “Shredder Cat”, the library’s pet mascot, into their digital creations — that she has begun learning to code alongside them.
Miss Sandy, Ruth, and her Code Club in Pennsylvania
At one showcase event in India, learners proudly demonstrated text-to-speech and video-sensing projects — remarkable achievements for many of them in their first year of coding. They explained their ideas with confidence and clarity, sharing the logic behind their work as parents, teachers, and mentors looked on with pride.
In the UK, we experienced a beautiful moment at Fakenham Academy where the room filled with a chorus of squarks, clicking, tapping, and squeaking sounds as learners adapted the Grow a Dragonfly project in their own creative ways.
From applause erupting whenever a project is finished to a room buzzing with micro:bits or young people debugging together on a shared laptop, these snapshots show Code Club at its best.
All about community: Belonging, identity, and a resourceful spirit
Whatever the context or setting, Code Club leaders are resourceful. They are not waiting for others to solve their challenges; together with their creators, they are finding local solutions that work. Communities share equipment, mentor each other, offer space, and build continuity for learners in imaginative ways. Young people are gaining far more than digital skills: they are developing belonging, confidence, and a clear sense that they are part of something bigger.
For example, in a club in Kenya, two groups learnt side by side in a shared space. Younger learners were welcomed by older peers who acted as mentors, creating a real sense of community — collaborative, vibrant, and full of pride in each other’s achievements.
We also saw how deeply this work is woven into people’s identities. One team member visiting three clubs in Malvern, UK, wrote about Bob Bilsland, a Code Club champion for 13 years, describing how naturally he connected with learners at their own level and how fully he embodied the role of mentor and champion:
“Seeing his versatility as he mentored, sparked excitement and connected with creators at their own level was a thing to behold… he truly walks the talk.”
The impact can sometimes show up in unexpected ways. Holly, a leader from Illinois, USA, shared that her learners wear their Code Club t-shirts to school as “spirit gear” on Fridays. She told us how much this meant to her students:
“They absolutely love their shirts and are thrilled to be able to wear them… It makes them feel like a team.” A reminder that belonging matters just as much as skills.
Across every example, we saw resilience, creativity, and generosity in action because Code Clubs grow from strong communities.
Locally rooted but globally informed
These stories underscore something essential: Code Club grows not because of any one model, but because communities everywhere make it their own. Sandra Keeru, Programme Coordinator in Kenya, put it beautifully when she reflected that:
“Code Club is locally rooted but globally informed.”
Each club reflects the needs, culture, and creativity of its community, yet everywhere the same shared values shine through: curiosity, inclusion, and the belief that young people can achieve remarkable things.
Code Club is more than just learning to code; it’s about creating opportunities, encouraging confidence, and building a global network of digital creators. Whether you’re a mentor, educator, or young digital maker, there’s a place for you in the community. Start your Code Club journey today and join a global community of digital creators.
You bet your ass we’re all alike… we’ve been spoon-fed baby food at school when we hungered for steak… the bits of meat that you did let slip through were pre-chewed and tasteless. We’ve been dominated by sadists, or ignored by the apathetic. The few that had something to teach found us willing pupils, but those few are like drops of water in the desert.
This is our world now… the world of the electron and the switch, the beauty of the baud. We make use of a service already existing without paying for what could be dirt-cheap if it wasn’t run by profiteering gluttons, and you call us criminals. We explore… and you call us criminals. We seek after knowledge… and you call us criminals. We exist without skin color, without nationality, without religious bias… and you call us criminals. You build atomic bombs, you wage wars, you murder, cheat, and lie to us and try to make us believe it’s for our own good, yet we’re the criminals.
Yes, I am a criminal. My crime is that of curiosity. My crime is that of judging people by what they say and think, not what they look like. My crime is that of outsmarting you, something that you will never forgive me for.
Ready to scale faster and grow smarter in 2026? If so, it might just be time to take a fresh look at the Zabbix Partner Program.
As we’ve mentioned before on this blog, the Partner Program is a lot more than just an extra layer on top of the software. It’s an invitation to be part of a community, a chance to gain access to advanced Zabbix training and certifications, an opportunity to dramatically expand your business reach, and a way to level up from skilled Zabbix practitioners to globally recognized experts.
What’s new in the Zabbix Partner Program?
In 2025, we made a good thing even better by revising and updating our Partner Program in order to bring Zabbix services to new users, in more locations, and in additional languages. Some of these changes include:
Granting more freedom to the Premium partners and supporting their business in cases outside their original territory.
Giving outstanding partners the visibility they deserve.
Engaging partners in a wider variety of activities and leveraging their expertise.
Sharing more business with partners.
Communicating better with partners about expectations and how to work with Zabbix.
As evidence of how becoming a Zabbix Partner can give your business a boost, here are a few success stories from five of our top partners.
Somone
Specialists in IT supervision and observability, Somone offers their clients strategic management of services and business indicators. Their experience with a variety of monitoring tools and track record of success with major accounts makes them a key player in the surveillance and observability spheres.
When the Paris-based company became a Zabbix Partner in 2023, they immediately took advantage of Zabbix Certified training and got all their employees certified as Zabbix users. This gave every employee a personal stake in the partnership and quickly brought everyone from the sales team to technical experts up to speed on Zabbix.
The company also notably encouraged its employees to speak at Zabbix events, which strengthened the relationship with Zabbix and encouraged a free exchange of ideas, which in turn helped Somone’s employees bring new ideas to life and improve processes.
As a result, Somone’s own Zabbix training offer has brought a host of new customers and increased their credibility in a crowded and competitive marketplace. In addition, having their own Zabbix team has been an enormous benefit – they have gone from having no real Zabbix strategy to building a dedicated team with their own sales and project leads.
Metricio
A Zabbix Partner for 5 years, Metricio is a Swedish IT services and consulting company that provides professional services and monitoring solutions for the Nordic market. They rely on Zabbix to deliver a cost-efficient and reliable monitoring solution that strengthens their portfolio and helps their customers achieve a higher level of efficiency.
When Metricio first became a Zabbix Partner, they found the Zabbix Partner team’s guidance to be invaluable. They have since advised other partners that they should not hesitate to reach out in the event of questions or concerns. In addition, they have organized several Zabbix-related meetings and events in Sweden, all of which have been more effective thanks to the presence of Zabbix team members.
Metricio worked hard during their first year to strengthen their position and set a goal to become Zabbix Premium Partners by their second year. Today, more than 70% of their leads come from the Zabbix Partner page or joint events. Teaming up with Zabbix has become the foundation of Metricio’s growth, their brand, and their success in the Swedish market.
ASPL Info
ASPL Info is a technology enterprise that aims to revolutionize businesses with best-in-class IT services and digital transformations. They boast a track record of success with global enterprises from a wide variety of sectors and geographies, including HP, Titan, Karnataka Bank, Trust Bank, Tata Sky, William Penn, Bajaj, Maruti, Emirates, Marico, Lupin, Dhanalaxmi Bank, HPE (Ministry of Home Affairs – MHA), Alstom, Birlasoft, Sify Digital, Tamilnad Mercantile Bank, Bank of Baroda, Union Bank of India, Tata-Elxsi, and Airtel.
As a Zabbix Premium Zabbix Partner, ASPL Info has broadened their horizons by working closely with the Zabbix OEM team, engaging early in joint opportunity planning, solution design and roadmap discussions. This has enabled faster project delivery and stronger customer outcomes. Meanwhile, building a highly skilled and certified team via Zabbix Certified trainings has guaranteed consistent delivery quality, better customer confidence, and a deeper understanding of enterprise-scale deployments.
The team at ASPL Info has also greatly benefited from active participation in Zabbix community and partner initiatives, including webinars, regional events, and marketing collaborations. These interactions not only enhance technical expertise but create valuable networking opportunities and visibility, while leveraging Zabbix’s cobranding, marketing and joint engagement initiatives have amplified ASPL Info’s credibility and supported business growth across new markets.
By getting the most out of their Zabbix partnership, ASPL Info has been able to rapidly streamline solution design, accelerate deployment, and strengthen customer confidence, while encouraging their engineers to pursue Zabbix certifications and continuous learning has built deeper in-house expertise, which in turn has allowed them to deliver more value and greater flexibility to their clients.
Since teaming up with Zabbix, ASPL Info have grown around 30–35% in service delivery engagements, maintained a consistent 97% resolution rate for customer queries and complaints related to Zabbix, and achieved 98% customer retention by continuing high-quality services and support throughout and even after the renewal process.
The ATS Group
As the sole Premium Zabbix Partner in North America, the ATS Group has seen strong growth by combining deep technical expertise with Zabbix’s proven monitoring platform and building a practical, collaborative relationship that’s based on accessibility, responsiveness, and mutual trust.
Their team has benefited greatly by engaging directly with the Zabbix team, finding out time and time again that open communication and quick access to the right people make a meaningful difference when delivering results for clients. Another key takeaway they have noted is the benefits of aligning Zabbix with a broader service conversation. Rather than leading with product features, they instead focus on outcomes, highlighting the ways in which Zabbix supports automation, observability, and operational excellence within modern IT environments.
The ATS Group’s partnership with Zabbix has led to new enterprise engagements and an increased awareness of Zabbix in North America through joint marketing activities and technical enablement efforts. The flexibility and support provided by the Zabbix team have been instrumental in helping their team tailor solutions to client needs while growing their monitoring and automation services.
OpenSource ICT Solutions
With a truly global footprint and Zabbix Premium Partner status, OpenSource ICT Solutions serves as a great example of how far being a Zabbix partner can take a business. In many regions they hold the status of “Certified Partner” or “Premium Delivery Partner.” They also operate as an official Zabbix reseller — meaning they can sell Zabbix support and services to customers while offering Zabbix consultancy and implementation services, providing official Zabbix training, and supplying support and managed services.
Their Premium Partner status and close relationship with the Zabbix team has taught them a few very important lessons – first and foremost of which is that Zabbix simply isn’t for everyone. It’s a great fit for many customers, but not all. Their policy is to always be honest about that and to never try to sell something that doesn’t make sense for the client.
They also recommend solving a problem rather than selling a product – in their view, the focus should be on delivering solutions that address real customer needs instead of pushing features. When in doubt, they always reach out to Partners team at Zabbix, who are always there to support them and who have access to valuable internal resources that let them come up with insights and materials that can make a real difference.
The team at OpenSource ICT Solutions also stresses the merits of participating in Zabbix meetings consistently, with every event they’ve attended leading to new customer relationships and strengthening existing ones. Additionally, they have found that blog posts and webinars have been highly effective, as they build visibility and showcase expertise, ultimately strengthening the team’s reputation and customer trust.
Lessons learned
The companies mentioned above are all very different and have used their status as members of the Zabbix Partner Program to achieve different goals, but there are a number of strategies for success and common best practices that apply to all of them, including:
Participation. Every partner mentioned above gained significant advantages by participating in Zabbix conferences and meetings, including better visibility and brand recognition, direct access to new leads and business opportunities, early insights Into the Zabbix roadmap and upcoming features, the opportunity to share expertise, and a strengthened relationship with the Zabbix team.
Engagement. Working more closely with the Zabbix Partner team provides tangible benefits in the form of more leads and co-selling opportunities, access to exclusive resources (like sales kits, marketing materials, and campaign support), direct access to Zabbix engineers (for architectural guidance, complex deployments, and troubleshooting), and strategic influence via feedback on product roadmaps and participation in advisory discussions.
Knowledge acquisition. Up-skilling and cross-skilling teams via Zabbix Certified trainings and our new Zabbix Academy courses benefits partners by ensuring that partner teams have the best possible understanding of Zabbix architecture, deployment, tuning, and troubleshooting, resulting in more efficient and stable deployments, faster problem resolution, and the ability to handle more complex customer environments. Partners can also use Zabbix certification to demonstrate competence during pre-sales, differentiate themselves from non-certified competitors, and build overall customer confidence in their services.
In conclusion
If your company provides IT services, system integration, managed services, or consulting, joining the Zabbix Partner Program is not just an extra benefit or a nice-to-have — it can become a core pillar of business growth. Reach out to our team and get started on the road to greater opportunity today!
Amazon Web Services (AWS) is pleased to announce that two additional AWS services and one additional AWS Region have been added to the scope of our Payment Card Industry Data Security Standard (PCI DSS) certification:
This certification allows customers to use these services while maintaining PCI DSS compliance, enabling innovation without compromising security. The full list of services can be found on the AWS Services in Scope by Compliance Program. The PCI DSS compliance package includes two key components:
Attestation of Compliance (AOC) demonstrating that AWS was successfully validated against the PCI DSS standard.
AWS Responsibility Summary provides guidance to help AWS customers understand their responsibility in developing and operating a highly secure environment on AWS for handling payment card data.
AWS was evaluated by Coalfire, a third-party Qualified Security Assessor (QSA).
This refreshed PCI certification offers customers greater flexibility in deploying regulated workloads while reducing compliance overhead. Customers can access the PCI DSS certification through AWS Artifact. This self-service portal provides on-demand access to AWS compliance reports, streamlining audit processes.
AWS is excited to be the first cloud service provider to offer compliance reports to customers in NIST’s Open Security Controls Assessment Language (OSCAL), an open source, machine-readable (JSON) format for security information. The PCI DSS report package (which includes both the PCI DSS AOC and the AWS Responsibility Summary) in OSCAL format is now available separately in AWS Artifact, marking a milestone towards open, standards-based compliance automation. This machine-readable version of the PCI DSS report package enables workflow automation to reduce manual processing time and modernize security and compliance processes. Your use cases for this content are innovative and we want to hear about them through the contact information found in the OSCAL report package.
To learn more about our PCI programs and other compliance and security programs, see the AWS Compliance Programs page. As always, we value your feedback and questions; reach out to the AWS Compliance team through the Compliance Support page.
If you have feedback about this post, submit comments in the Comments section below. If you have questions about this post, contact AWS Support.
In the last few days, Internet traffic from Iran has effectively dropped to zero. This is evident in the data available in Cloudflare Radar, as we’ll describe in this post.
Background
The Iranian government has a history of cutting off Internet connectivity when such protests take place. In November 2019, protests erupted following the announcement of a significant increase in fuel prices. In response, the Iranian government implemented an Internet shutdown for more than five days. In September 2022, protests and demonstrations erupted across Iran in response to the death in police custody of Mahsa/Zhina Amini, a 22-year-old woman from the Kurdistan Province of Iran. Internet services were disrupted across multiple network providers in the following days.
Amid the current protests, lower traffic volumes were already observed at the start of the year, indicating potential connectivity issues leading into the more dramatic shutdown that has followed.
Internet connectivity in Iran plummeted on January 8
Some traffic anomalies were seen in the first few days of 2026 (described in further detail below), though peak traffic levels recovered by January 5, and exceeded expected levels during the following days.
However, this strong recovery proved to be short-lived. IPv6-related shifts observed on January 8 provided the first indication of the changes to come. At 11:50 UTC (15:20 local time), the amount of IPv6 address space announced by Iranian networks dropped by 98.5%, falling from over 48 million /48s (blocks of 2^80 IPv6 addresses) to just over 737,000 /48s. A drop in announced IP address space (whether IPv6 or IPv4) means that the announcing networks are no longer telling the world how to reach those addresses. A major drop like this one can signal an intentional disruption to Internet connectivity, as there is no longer a path to the clients or servers using those IP addresses.
This drop in announced IPv6 address space served to reduce IPv6’s share of human-generated traffic from around 12% to around 2%.
As seen in the graph below, this drop in IPv6 traffic stayed at a relatively consistent level for approximately 100 minutes, before falling further just before 13:30 UTC (17:00 local time). This second drop resulted in IPv6 traffic from Iran all but disappearing.
Around 18:45 UTC, Internet traffic from Iran dropped to effectively zero, signaling a complete shutdown in the country and disconnection from the global Internet.
Brief windows of connectivity on January 9 — but they don’t last
After the shutdown took hold the previous day, internal traffic data showed an extremely low volume of traffic from Iran, amounting to less than 0.01% of pre-shutdown peaks, starting around 10:00 UTC (13:30 local time) on January 9. It appears that access to Cloudflare’s public DNS resolver, 1.1.1.1, also became available again around 10:00 UTC (13:30 local time), leading request traffic to briefly spike well above the expected range. However, after spiking, only a small amount of request traffic to 1.1.1.1 remained visible.
Changes in HTTP traffic preceded the Internet shutdown
Alongside the lower traffic levels observed at the start of the year, as discussed above, a clear shift in HTTP version usage from human-generated traffic was also observed across leading network providers, as seen in the graphs below. Prior to that point, as much as 40% of HTTP requests on IranCell (AS44244) used HTTP/3, but that figure fell to just 5% at 20:00 UTC (23:30 local time) on December 31, and continued to decline over the following days. Usage of QUIC from the network followed a similar pattern, as it relies on HTTP/3.
On TCI (AS58224), HTTP/3 also accounted for as much as 40% of requests at peak, but gradually declined starting on January 1 before falling below 5% starting around 07:00 UTC (10:30 local time) on January 3. QUIC usage on this network followed a similar pattern as well. MahsaNet, an organization that fights against Internet censorship in Iran, suggested that these shifts could indicate that “Severe filtering and layered, upgraded whitelisting are clearly evident and being implemented” (translation via X).
The shutdown continues
As we noted in social media posts (X, Mastodon, Bluesky), no significant changes have been observed in Iran’s Internet traffic since January 10. The country remains almost entirely cut off from the global Internet, with internal data showing traffic volumes remaining at a fraction of a percent of previous levels.
We will continue to monitor the state of Internet connectivity in Iran, and will continue to post updates on our social media accounts. Use Cloudflare Radar’s Traffic and Routing pages for Iran and the top networks within the country for near-real time insights into these metrics.
Organizations today are using data more than ever to drive decision-making and innovation. Because they work with petabytes of information, they have traditionally gravitated towards two distinct paradigms—data lakes and data warehouses. While each paradigm excels at specific use cases, they often create unintended barriers between the data assets.
Data lakes are often built on object storage such as Amazon Simple Storage Service (Amazon S3), which provide flexibility by supporting diverse data formats and schema-on-read capabilities. This enables multi-engine access where various processing frameworks (such as Apache Spark, Trino, and Presto) can query the same data. On the other hand, data warehouses (such as Amazon Redshift) excel in areas such as ACID (atomicity, consistency, isolation and durability) compliance, performance optimization, and straightforward deployment, making them suitable for structured and complex queries. As data volumes grow and analytics needs become more complex, organizations seek to bridge these silos and use the strengths of both paradigms. This is where the concept of lakehouse architecture is applied, offering a unified approach to data management and analytics.
Over time, several distinct lakehouse approaches have emerged. In this post, we show you how to evaluate and choose the right lakehouse pattern for your needs.
The data lake centric lakehouse approach begins with the scalability, cost-effectiveness, and flexibility of a traditional data lake built on object storage. The goal is to add a layer of transactional capabilities and data management traditionally found in databases, primarily through open table formats (such as Apache Hudi, Delta Lake, or Apache Iceberg). While open table formats have made significant strides by introducing ACID guarantees for single-table operations in data lakes, implementing multi-table transactions with complex referential integrity constraints and joins remains challenging. The fundamental nature of querying petabytes of files on object storage, often through distributed query engines, can result in slow interactive queries at high concurrency when compared to a highly optimized, indexed, and materialized data warehouse. Open table formats introduce compaction and indexing, but the full suite of intelligent storage optimizations found in highly mature, proprietary data warehouses is still evolving in data lake-centric architecture.
The data warehouse centric lakehouse approach offers robust analytical capabilities but has significant interoperability challenges. Though data warehouses provide JAVA Database Connectivity (JDBC) and Open Database Connectivity (ODBC) drivers for external access, the underlying data remains in proprietary formats, making it difficult for external tools or services to directly access it without complex extract, transform, and load (ETL) or API layers. This can lead to data duplication and latency. A data warehouse architecture might support reading open table formats, but its ability to write to them or participate in their transactional layers can be limited. This restricts true interoperability and can create shadow data silos.
On AWS, you can build a modern, open lakehouse architecture to achieve unified access to both data warehouses and data lakes. By using this approach, you can build sophisticated analytics, machine learning (ML), and generative AI applications while maintaining a single source of truth for their data. You don’t have to choose between a data lake or data warehouse. You can use existing investments and preserve the strengths of both paradigms while eliminating their respective weaknesses. The lakehouse architecture on AWS embraces open table formats such as Apache Hudi, Delta Lake, and Apache Iceberg.
You can accelerate your lakehouse journey with the next generation of Amazon SageMaker, which delivers an integrated experience for analytics and AI with unified access to data. SageMaker is built on an open lakehouse architecture that is fully compatible with Apache Iceberg. By extending support for Apache Iceberg REST APIs, SageMaker significantly adds interoperability and accessibility across various Apache Iceberg-compatible query engines and tools. At the core of this architecture is a metadata management layer built on AWS Glue Data Catalog and AWS Lake Formation, which provide unified governance and centralized access control.
Foundations of the Amazon SageMaker lakehouse architecture
The lakehouse architecture of Amazon SageMaker has four main components that work together to create a unified data platform.
Flexible storage to adapt to the workload patterns and requirements
Technical catalog that serves as a single source of truth for all metadata
Integrated permission management with fine-grained access control across all data assets
Open access framework built on Apache Iceberg REST APIs for universal compatibility
Catalogs and permissions
When building an open lakehouse, the catalog—your central repository of metadata—is a critical component for data discovery and governance. There are two types of catalogs in the lakehouse architecture of Amazon SageMaker: managed catalogs and federated catalogs.
Managed catalog refers to when the metadata is managed by the lakehouse, and the data is stored in a general purpose S3 bucket.
You can use an AWS Glue crawler to automatically discover and register this metadata in Data Catalog. Data Catalog stores the schema and table metadata of your data assets, effectively turning files into logical tables. After your data is cataloged, the next challenge is controlling who can access it. While you could use complex S3 bucket policies for every folder, this approach is difficult to manage and scale. Lake Formation provides a centralized database-style permissions model on the Data Catalog, giving you the flexibility to grant or revoke fine-grained access at row, column, and cell levels for individual users or roles.
Open access with Apache Iceberg REST APIs
The lakehouse architecture described in the preceding section and shown in the following figure also uses the AWS Glue Iceberg REST catalog through the service endpoint, which provides OSS compatibility, enabling increased interoperability for managing Iceberg table metadata across Spark and other open source analytics engines. You can choose the appropriate API based on table format and use case requirements.
In this post, we explore various lakehouse architecture patterns, focusing on how to optimally use data lake and data warehouse to create robust, scalable, and performance-driven data solutions.
Bringing data into your lakehouse on AWS
When building a lakehouse architecture, you can choose from three distinct patterns to access and integrate your data, each offering unique advantages for different use cases.
Traditional ETL is the classic method of extracting data, transforming it and loading it into your lakehouse.
When to use it:
You need complex transformations and require highly curated and optimized data sets for downstream applications for better performance
You need to perform historical data migrations
You need data quality enforcement and standardization at scale
You need highly governed curated data in a lakehouse
Zero-ETL is a modern architectural pattern where data automatically and continuously replicates from a source system to lakehouse with minimal or no manual intervention or custom code. Behind the scenes, the pattern uses change data capture (CDC) to automatically stream all new inserts, updates, and deletes from the source to the target. This architectural pattern is effective when the source system maintains a high degree of data cleanliness and structure, minimizing the need for heavy pre-load transformations, or when data refinement and aggregation can occur at the target end within lakehouse. Zero-ETL replicates data with minimal delay, and the transformation logic is performed on the target end closer to where the insights are generated by shifting it to a more efficient, post-load phase.
When to use it:
You need to reduce operational complexity and gain flexible control over data replication for both near real-time and batch use cases.
You need limited customization. While zero-ETL implies minimal work, some light transformations might still be required on the replicated data.
You need to minimize the need for specialized ETL expertise.
You need to maintain data freshness without processing delays and reduce risk of data inconsistencies. Zero-ETL facilitates faster time-to-insight.
Data federation (no-movement approach) is a method that enables querying and combining data from multiple disparate sources without physically moving or copying it into a single centralized location. This query-in-place approach allows the query engine to connect directly to the external source systems, delegate and execute queries, and combine results on the fly for presentation to the user. The effectiveness of this architecture pattern depends on three key factors: network latency between systems, source system performance capabilities, and the query engine’s ability to push down predicates to optimize query execution. This no-movement approach can significantly reduce data duplication and storage costs while providing real-time access to source data.
When to use it:
You need to query the source system directly to use operational analytics.
You don’t want to duplicate data to save on storage space and associated costs within your Lakehouse.
You’re willing to trade some query performance and governance for immediate data availability and one-time analysis of live data.
You don’t need to frequently query the data.
Understanding the storage layer of your lakehouse on AWS
Now that you’ve seen different ways to get data into a lakehouse, the next question is where to store the data. As shown in the following figure, you can architect a modern open lakehouse on AWS by storing the data in a data lake (Amazon S3 or Amazon S3 Tables) or data warehouse (Redshift Managed Storage), so you can optimize for both flexibility and performance based on your specific workload requirements.
A modern lakehouse isn’t a single storage technology but a strategic combination of them. The decision of where and how to store your data impacts everything from the speed of your dashboards to the efficiency of your ML models. You must consider not only the initial cost of storage but also the long-term costs of data retrieval, the latency required by your users, and the governance necessary to maintain a single source of truth. In this section, we delve into architectural patterns for the data lake and the data warehouse and provide a clear framework for when to use each storage pattern. While they have historically been seen as competing architectures, the modern and open lakehouse approach uses both to create a single, powerful data platform.
General purpose S3
A general purpose S3 bucket in Amazon S3 is the standard, foundational bucket type used for storing objects. It provides flexibility so that you can store your data in its native format without a rigid upfront schema. Because of the ability of an S3 bucket to decouple storage from compute, you can store the data in a highly scalable location, while a variety of query engines can access and process it independently. This means that you can choose the right tool for the job without having to move or duplicate the data. You can store petabytes of data without ever having to provision or manage storage capacity, and its tiered storage classes provide significant cost savings by automatically moving less-frequently accessed data to more affordable storage.
The existing Data Catalog functions as a managed catalog. It’s identified by the AWS account number, which means there is no migration needed for existing Data Catalogs; they’re already available in the lakehouse and become the default catalog for the new data, as shown in the following figure.
A foundational data lake on general purpose S3 is highly efficient for append-only workloads. However, its file-based nature lacks the transactional guarantees of a traditional database. This is where you can use the support of open-source transactional table formats such as Apache Hudi, Delta Lake, and Apache Iceberg. With these table formats, you can implement multi-version concurrency control, allowing multiple readers and writers to operate simultaneously without conflicts. They provide snapshot isolation, so that readers see consistent views of data even during write operations. A typical medallion architecture pattern with Apache Iceberg is depicted in the following figure. When building a lakehouse on AWS with Apache Iceberg, customers can choose between two primary approaches for storing their data on Amazon S3: General purpose S3 buckets with self-managed Iceberg or using the fully managed S3 Tables. Each path has distinct advantages, and the right choice depends on your specific needs for control, performance, and operational overhead.
General purpose S3 with Self-managed Iceberg
Using general purpose S3 buckets with self-managed Iceberg is a traditional approach where you store both data and Iceberg metadata files in standard S3 buckets. With this option, you maintain full control but are responsible for managing the complete Iceberg table lifecycle, including essential maintenance tasks such as compaction and garbage collection.
When to use it:
Maximum control: This approach provides complete control over the entire data life cycle. You can fine-tune every aspect of table maintenance, such as defining your own compaction schedules and strategies, which can be crucial for specific high-performance workloads or to optimize costs.
Flexibility and customization: It is ideal for organizations with strong in-house data engineering expertise that need to integrate with a wider range of open-source tools and custom scripts. You can use Amazon EMR or Apache Spark to manage the table operations.
Lower upfront costs: You pay only for Amazon S3 storage, API requests, and the compute resources you use for maintenance. This can be more cost-effective for smaller or less-frequent workloads where continuous, automated optimization isn’t necessary.
Note: The query performance depends entirely on your optimization strategy. Without continuous, scheduled jobs for compaction, performance can degrade over time as data gets fragmented. You must monitor these jobs to ensure efficient querying.
S3 Tables
S3 Tables provides S3 storage that’s optimized for analytic workloads and provides Apache Iceberg compatibility to store tabular data at scale. You can integrate S3 table buckets and tables with Data Catalog and register the catalog as a Lake Formation data location from the Lake Formation console or using service APIs, as shown in the following figure. This catalog will be registered and mounted as a federated lakehouse catalog.
When to use it:
Simplified operations: S3 Tables automatically handles table maintenance tasks such as compaction, snapshot management and orphan file cleanup in the background. This automation eliminates the need to build and manage custom maintenance jobs, significantly reducing your operational overhead.
Automated optimization: S3 Tables provides built-in automatic optimizations that improve query performance. These optimizations include background processes such as file compaction to address the small files problem and data layout optimizations specific to tabular data. However, this automation trades flexibility for convenience. Because you can’t control the timing or method of compaction operations, workloads with specific performance requirements might experience varying query performance.
Focus on data usage: S3 Tables reduces the engineering overhead and shifts the focus to data consumption, data governance and value creation.
Simplified entry to open table formats: It’s suitable for teams who are new to the concept of Apache Iceberg but want to use transactional capabilities on data lake.
No external catalog: Suitable for smaller teams who don’t want to manage an external catalog.
Redshift managed storage
While the data lake serves as the central source of truth for all your data, it’s not the most suitable data store for every job. For the most demanding business intelligence and reporting workloads, the data lake’s open and flexible nature can introduce performance unpredictability. To help ensure the desired performance, consider transitioning a curated subset of your data from the data lake to a data warehouse for the following reasons:
High concurrency BI and reporting: When hundreds of business users are concurrently running complex queries on live dashboards, a data warehouse is specifically optimized to handle these workloads with predictable, sub-second query latency.
Predictable performance SLAs:– For critical business processes that require data to be delivered at a guaranteed speed, such as financial reporting or end-of-day sales analysis, a data warehouse provides consistent performance.
Complex SQL workloads: While data lakes are powerful, they can struggle with highly complex queries involving numerous joins and massive aggregations. A data warehouse is purpose-built to run these relational workloads efficiently.
The lakehouse architecture on AWS supports Redshift Managed Storage (RMS), a storage option provided by Amazon Redshift, a fully managed, petabyte-scale data warehouse service in the cloud. RMS storage supports the automatic table optimization offered in Amazon Redshift such as built-in query optimizations for data warehousing workloads, automated materialized views, and AI-driven optimizations and scaling for frequently running workloads.
Federated RMS catalog: Onboard existing Amazon Redshift data warehouses to lakehouse
Implementing a federated catalog with existing Amazon Redshift data warehouses creates a metadata-only integration that requires no data movement. This approach lets you extend your established Amazon Redshift investments into a modern open lakehouse framework while maintaining compatibility with existing workflows. Amazon Redshift uses a hierarchical data organization structure:
Cluster level: Starts with a namespace
Database level: Contains multiple databases
Schema level: Organizes tables within databases
When you register your existing Amazon Redshift provisioned or serverless namespaces as a federated catalog in Data Catalog, this hierarchy maps directly into the lakehouse metadata layer. The lakehouse implementation on AWS supports multiple catalogs using a dynamic hierarchy to organize and map the underlying storage metadata.
After you register a namespace, the federated catalog automatically mounts across all Amazon Redshift data warehouses in your AWS Region and account. During this process, Amazon Redshift internally creates external databases that correspond to data shares. This mechanism remains completely abstracted from end users. By using federated catalogs, you can create and use immediate visibility and accessibility across your data ecosystem. Permissions on the federated catalogs can be managed by Lake Formation for both same account and cross account access.
The real capability of federated catalogs emerges when accessing Amazon Redshift-managed storage from external AWS engines such as Amazon Athena, Amazon EMR, or open source Spark. Because Amazon Redshift uses proprietary block-based storage that only Amazon Redshift engines can read natively, AWS automatically provisions a service-managed Amazon Redshift Serverless instance in the background. This service-managed instance acts as a translation layer between external engines and Amazon Redshift managed storage. AWS establishes automatic data shares between your registered federated catalog and the service-managed Amazon Redshift Serverless instance to enable secure, efficient data access. AWS also creates a service-managed Amazon S3 bucket in the background for data transfer.
When an external engine such as Athena submits queries against Amazon Redshift federated catalog, Lake Formation handles the credential vending by providing the temporary credentials to the requesting service. The query executes through the service-managed Amazon Redshift Serverless, which accesses data through automatically established data shares, processes results, offloads them to a service-managed Amazon S3 staging area, and then returns results to the original requesting engine.
To track the compute cost of the federated catalog of existing Amazon Redshift warehouse, use the following tag.
To activate the AWS generated cost allocation tags for billing insight, follow the activation instructions. You can also view the computational cost of the resources in AWS Billing.
When to use it:
Existing Amazon Redshift investments: Federated catalogs are designed for organizations with existing Amazon Redshift deployments who want to use their data across multiple services without migration.
Cross-service data sharing:– Implement so teams can share existing data in an Amazon Redshift data warehouse across different warehouses and centralize their permissions.
Enterprise integration requirements: This approach is suitable for organizations that need to integrate with established data governance. It also maintains compatibility with current workflows while adding lakehouse capabilities.
Infrastructure control and pricing:– You can retain full control over compute capacity for their existing warehouses for predictable workloads. You can optimize compute capacity, choose between on-demand and reserved capacity pricing, and fine-tune performance parameters. This provides cost predictability and performance control for consistent workloads.
When implementing lakehouse architecture with multiple catalog types, selecting the appropriate query engine is crucial for both performance and cost optimization. This post focuses on the storage foundation of lakehouse, however for critical workloads involving extensive Amazon Redshift data operations, consider executing queries within Amazon Redshift or using Spark when possible. Complex joins spanning multiple Amazon Redshift tables through external engines might result in higher compute costs if the engines don’t support full predicate push-down.
Other use-cases
Build a multi-warehouse architecture
Amazon Redshift supports data sharing, which you can use to share live data between source and target Amazon Redshift clusters. By using data sharing, you can share live data without creating copies or moving data, enabling uses cases such as workload isolation (hub and spoke architecture) and cross group collaboration (data mesh architecture). Without a lakehouse architecture, you must create an explicit data share between source and target Amazon Redshift clusters. While managing these data shares in small deployments is relatively straightforward, it becomes complex in data mesh architectures.
The lakehouse architecture addresses this challenge so customers can publish their existing Amazon Redshift warehouses as federated catalogs. These federated catalogs are automatically mounted and made available as external databases in other consumer Amazon Redshift warehouses within the same account and Region. By using this approach, you can maintain a single copy of data and use multiple data warehouses to query it, eliminating the need to create and manage multiple data shares and scale with workload isolation. The permission management becomes centralized through Lake Formation, streamlining governance across the entire multi-warehouse environment.
Near real-time analytics on petabytes of transactional data with no pipeline management:
Zero-ETL integrations seamlessly replicate transactional data from OLTP data sources to Amazon Redshift, general purpose S3 (with self-managed Iceberg) or S3 Tables. This approach eliminates the need to maintain complex ETL pipelines, reducing the number of moving parts in your data architecture and potential points of failure. Business users can analyze fresh operational data immediately rather than working with stale data from the last ETL run.
See Aurora zero-ETL integrations for a list of OLTP data sources that can be replicated to an existing Amazon Redshift warehouse.
See Zero-ETL integrations for information about other supported data sources that can be replicated to an existing Amazon Redshift warehouse, general purpose S3 with self-managed Iceberg, and S3 Tables.
Conclusion
A lakehouse architecture isn’t about choosing between a data lake and a data warehouse. Instead, it’s an approach to interoperability where both frameworks coexist and serve different purposes within a unified data architecture. By understanding fundamental storage patterns, implementing effective catalog strategies, and using native storage capabilities, you can build scalable, high-performance data architectures that support both your current analytics needs and future innovation. For more information, see The lakehouse architecture of Amazon SageMaker.
AWS has launched the catalog federation capability, enabling direct access to Apache Iceberg tables managed in Databricks Unity Catalog through the AWS Glue Data Catalog. With this integration, you can discover and query Unity Catalog data in Iceberg format using an Iceberg REST API endpoint, while maintaining granular access controls through AWS Lake Formation. This approach significantly reduces operational overhead for managing catalog synchronization and associated costs by alleviating the need to replicate or duplicate datasets between platforms.
In this post, we demonstrate how to set up catalog federation between the Glue Data Catalog and Databricks Unity Catalog, enabling data querying using AWS analytics services.
Use cases and key benefits
This federation capability is particularly valuable if you run multiple data platforms, because you can maintain your existing Iceberg catalog investments while using AWS analytics services. Catalog federation supports read operations and provides the following benefits:
Interoperability – You can enable interoperability across different data platforms and tools through Iceberg REST APIs while preserving the value of your established technology investments.
Cross-platform analytics – You can connect AWS analytics tools (Amazon Athena, Amazon Redshift, Apache Spark) to query Iceberg and UniForm tables stored in Databricks Unity Catalog. It supports Databricks on AWS integration with the AWS Glue Iceberg REST Catalog for metadata retrieval, while using Lake Formation for permission management.
Metadata management – The solution avoids manual catalog synchronization by making Databricks Unity Catalog databases and tables discoverable within the Data Catalog. You can implement unified governance through Lake Formation for fine-grained access control across federated catalog resources.
Solution overview
The solution uses catalog federation in the Data Catalog to integrate with Databricks Unity Catalog. The federated catalog created in AWS Glue mirrors the catalog objects in Databricks Unity Catalog and supports OAuth-based authentication. The solution is represented in the following diagram.
The integration involves three high-level steps:
Set up an integration principal in Databricks Unity Catalog and provide required read access on catalog resources to this principal. Enable OAuth-based authentication for the integration principal.
Set up catalog federation to Databricks Unity Catalog in the Glue Data Catalog:
Create a federated catalog in the Data Catalog using an AWS Glue connection.
Create an AWS Glue connection that uses the credentials of the integration principal (in Step 1) to connect to Databricks Unity Catalog. Configure an AWS Identity and Access Management (IAM) role with permission to Amazon Simple Storage Service (Amazon S3) locations where the Iceberg table data resides. In a cross-account scenario, make sure the bucket policy grants required access to this IAM role.
Discover Iceberg tables in federated catalogs using Lake Formation or AWS Glue APIs. During query operations, Lake Formation manages fine-grained permissions on federated resources and credential vending for access to the underlying data.
In the following sections, we walk through the steps to integrate the Glue Data Catalog with Databricks Unity Catalog on AWS.
Prerequisites
To follow along with the solution presented in this post, you must have the following prerequisites:
Databricks Workspace (on AWS) with Databricks Unity Catalog configured.
An IAM role that is a Lake Formation data lake administrator in your AWS account. A data lake administrator is an IAM principal that can register S3 locations, access the Data Catalog, grant Lake Formation permissions to other users, and view AWS CloudTrail logs. See Create a data lake administrator for more information.
Configure Databricks Unity Catalog for external access
Catalog federation to a Databricks Unity Catalog uses the OAuth2 credentials of a Databricks service principal configured in the workspace admin settings. This authentication mechanism allows the Data Catalog to access the metadata of various objects (such as catalogs, databases, and tables) within Databricks Unity Catalog, based on the privileges associated with the service principal. For proper functionality, grant the service principal with the necessary permissions (read permission on catalog, schema, and tables) to read the metadata of these objects and allow access from external engines.
Next, catalog federation enables discovery and query of Iceberg tables in your Databricks Unity Catalog. For reading delta tables, enable UniForm on a Delta Lake table in Databricks to generate Iceberg metadata. For more information, refer to Read Delta tables with Iceberg clients.
Follow the Databricks tutorial and documentation to create the service principal and associated privileges in your Databricks workspace. For this post, we use a service principal named integrationprincipal that is configured with required permissions (SELECT, USE CATALOG, USE SCHEMA) on Databricks Unity Catalog objects and will be used for authentication to catalog instance.
Catalog federation supports OAuth2 authentication, so enable OAuth for the service principal and note down the client_id and client_secret for later use.
Set up Data Catalog federation with Databricks Unity Catalog
Now that you have service principal access for Databricks Unity Catalog, you can set up catalog federation in the Data Catalog. To do so, you create an AWS Secrets Manager secret and create an IAM role for catalog federation.
Enter a name for your secret (for this post, we use dbx).
Choose Store.
Create IAM role for catalog federation
As the catalog owner of a federated catalog in the Data Catalog, you can use Lake Formation to implement comprehensive access controls, including table filters, column filters, and row filters, as well as tag-based access for your data teams.
Lake Formation requires an IAM role with permissions to access the underlying S3 locations of your external catalog.
In this step, you create an IAM role that enables the AWS Glue connection to access Secrets Manager, optional virtual private cloud (VPC) configurations, and Lake Formation to manage credential vending for the S3 bucket and prefix:
Secrets Manager access – The AWS Glue connection requires permissions to retrieve secret values from Secrets Manager for OAuth tokens stored for your Databricks Unity service connection.
VPC access (optional) – When using VPC endpoints to restrict connectivity to your Databricks Unity account, the AWS Glue connection needs permissions to describe and utilize VPC network interfaces. This configuration provides secure, controlled access to both your stored credentials and network resources while maintaining proper isolation through VPC endpoints.
S3 bucket and AWS KMS key permission – The AWS Glue connection requires Amazon S3 permissions to read certificates if used in the connection setup. Additionally, Lake Formation requires read permissions on the bucket and prefix where the remote catalog table data resides. If the data is encrypted using an AWS Key Management Service (AWS KMS) key, additional AWS KMS permissions are required.
Complete the following steps:
Create an IAM role called LFDataAccessRole with the following policies:
AWS Glue supports the DATABRICKSICEBERGRESTCATALOG connection type for connecting the Data Catalog with managed Databricks Unity Catalog. This AWS Glue connector supports OAuth2 authentication for discovering metadata in Databricks Unity Catalog.
Complete the following steps to create the federated catalog:
Sign in to the console as a data lake admin.
On the Lake Formation console, choose Catalogs in the navigation pane.
Choose Create catalog.
For Name, enter a name for your catalog.
For Catalog name in Databricks, enter the name of a catalog existing in Databricks Unity Catalog.
For Connection name, enter a name for the AWS Glue connection.
For Workspace URL, enter the Unity Iceberg REST API URL (in format https://<workspace-url>/cloud.databricks.com).
For Authentication, provide the following information:
For Authentication type,choose OAuth2. Alternatively, you can choose Custom authentication. For Custom authentication, an access token is created, refreshed, and managed by the customer’s application or system and stored using Secrets Manager.
For Token URL, enter the token authentication server URL.
For OAuth Client ID, enter the client_id for integrationprincipal.
For OAuth Secret, enter the secret ARN that you created in the previous step. Alternatively, you can provide the client_secret directly.
For Token URL parameter map scope, provide the API scope supported.
If you have AWS PrivateLink set up or a proxy set up, you can provide network details under Settings for network configurations.
For Register Glue connection with Lake Formation, choose the IAM role (LFDataAccessRole) created earlier to manage data access using Lake Formation.
When the setup is done using AWS Command Line Interface (AWS CLI) commands, you have options to create two separate IAM roles:
IAM role with policies to access network and secrets, which AWS Glue assumes to manage authentication
IAM role with access to the S3 bucket, which Lake Formation assumes to manage credential vending for data access
On the console, this setup is simplified with a single role having combined policies. For more details, refer to Federate to Databricks Unity Catalog.
To test the connection, choose Run test.
You can proceed to create the catalog.
After you create the catalog, you can see the databases and tables in Databricks Unity Catalog listed under the federated catalog. You can implement fine-grained access control on the tables by applying row and column filters using Lake Formation. The following video shows the catalog federation setup with Databricks Unity Catalog.
Discover and query the data using Athena
In this post, we show how to use the Athena query editor to discover and query the Databricks Unity Catalog tables. On the Athena console, run the following query to access the federated table:SELECT * FROM "customerschema"."person" limit 10;The following video demonstrates querying the federated table from Athena.
If you use the Amazon Redshift query engine, you must create a resource link on the federated database and grant permission on the resource link to the user or role. This database resource link is automounted under awsdatacatalog based on the permission granted for the user or role and available for querying. For instructions, refer to Creating resource links.
Clean up
To clean up your resources, complete the following steps:
Delete the catalog and namespace in Databricks Unity Catalog for this post.
Drop the resources in the Data Catalog and Lake Formation created for this post.
Delete the IAM roles and S3 buckets used for this post.
Delete any VPC and KMS keys if used for this post.
Conclusion
In this post, we explored the key elements of catalog federation and its architectural design, illustrating the interaction between the AWS Glue Data Catalog and Databricks Unity Catalog through centralized authorization and credential distribution for protected data access. By removing the requirement for complicated synchronization workflows, catalog federation makes it possible to query Iceberg data on Amazon S3 directly at its source using AWS analytics services with data governance across multi-catalog platforms. Try out the solution for your own use case, and share your feedback and questions in the comments.
As organizations scale their Kubernetes deployments, Kubernetes cluster scaling has traditionally been complex and slow, requiring careful management of node groups and auto scaling configurations. Karpenter, an open source node provisioning project for Kubernetes, can help transform this approach by directly provisioning right-sized nodes based on real-time workload demands. A recent Datadog report reveals that the percentage of nodes provisioned by Karpenter rose by 22% in the last 2 years as organizations migrate from traditional auto scaling approaches. This growth underscores Amazon Web Services (AWS) leadership in cloud-based innovation and the container ecosystem’s recognition of Karpenter’s strong performance and cost efficiency benefits. The following post examines how Salesforce, operating one of the world’s largest Kubernetes deployments, successfully migrated from Cluster Autoscaler to Karpenter across their fleet of 1,000 plus Amazon Elastic Kubernetes Service (Amazon EKS) clusters.
Salesforce operates one of the world’s most complex Kubernetes platforms, managing over 1,000 EKS clusters that serve thousands of internal tenants across the company. These clusters power a wide range of applications, from mission-critical services to experimental projects, and demand a high degree of scalability, reliability, and operational efficiency.
As the platform grew, Salesforce’s Kubernetes platform team began to face major hurdles with its traditional auto scaling approach based on AWS Auto Scaling groups and the Kubernetes Cluster Autoscaler. These limitations hampered the team’s ability to respond to application demands quickly, optimize compute resources, and empower internal developers to self-serve infrastructure needs.
To address these challenges, Salesforce undertook a large-scale migration to Karpenter, an open source Kubernetes [1] auto scaler built by AWS. This blog post details the motivation behind the transition, the implementation strategy, the challenges encountered along the way, and the impact it had on cost, performance, and operational complexity.
Opportunity for operational transformation
At Salesforce’s massive scale, the traditional Kubernetes infrastructure faced several critical challenges. The need to accommodate diverse workload requirements led to a proliferation of thousands of node groups and Auto Scaling groups, creating operational bottlenecks and slowing innovation. This architectural complexity was compounded by significant scaling performance issues, where the Auto Scaling group-dependent Cluster Autoscaler struggled to handle dynamic workloads, often resulting in multi-minute delays during demand spikes and degraded user experience. Resource utilization suffered as well, with inefficient bin-packing and conservative scale-down strategies leading to stranded resources and underutilized infrastructure—a particular concern given Salesforce’s focus on cost-to-serve and sustainability goals. These challenges were further exacerbated by structural limitations in the Auto Scaling group–based architecture, including poor Availability Zone balance and performance bottlenecks in large clusters, particularly for memory-intensive workloads. The combination of these factors made it clear that a more modern, flexible auto scaling solution was essential for maintaining Salesforce’s competitive edge and operational efficiency.
Solution overview
To migrate over 1,000 production clusters, without disruption, Salesforce engineered a highly automated, risk-mitigated transition process centered on Karpenter. Here’s how the migration was executed.
At this scale, a manual migration was infeasible. The team developed an in-house Karpenter transition tool to orchestrate the switch-over safely and consistently, and a Karpenter patching check tool. Karpenter transition tool and Karpenter patching check tool provide a comprehensive solution for migrating Kubernetes clusters to and from Karpenter node management while maintaining operational continuity through automated node rotation, Amazon Machine Image (AMI) validation, and graceful pod eviction handling.
Key design principles included:
Zero disruption – The tool cordoned and drained legacy nodes with full respect for pod disruption budgets (PDBs), maintaining workload safety
Rollback support – A reverse transition capability allowed fast recovery to Auto Scaling group–based auto scaling if needed
Continuous integration and continuous delivery (CI/CD) integration – The tool was embedded in the core infrastructure provisioning pipeline, standardizing the migration across services.
This foundation enabled repeatability across thousands of clusters and node pools, inspiring confidence in Salesforce developers.
Automated configuration mapping
To convert existing Auto Scaling group configurations to Karpenter-based definitions, the team automated the mapping logic between legacy and modern configurations. For example:
Auto Scaling group instance types → EC2NodeClass instance types
Root volume sizes → Storage parameters in Karpenter config
Node labels → Applied in both NodePool and EC2NodeClass
With over 1,180 node pools containing highly diverse configurations, automation was essential to minimize errors and reduce manual toil.
A deliberate, phased rollout strategy was adopted:
Mid-2025 to Early 2026 – A multistage migration across internal environments with soak times between stages
Start with lower-risk environments – Less critical workloads were migrated first to validate tooling and operational processes
Risk-based sequencing – High-stakes production environments continue to be migrated last after testing the process
By using this approach Salesforce, continuously learned and adapted, avoiding large-scale regressions.
Key insights from the migration
During this migration journey, the Salesforce team gained valuable insights and best practices that we’ll share to help guide your own transformation initiatives.
Managing application availability during nude Updates
PDBs emerged as a critical consideration during the migration because several services had overly restrictive or misconfigured PDBs that blocked node replacements. The team addressed this by identifying problematic configurations, partnering with application owners on remediation, and implementing Open Policy Agent (OPA) policies for proactive PDB validation. This experience highlighted how proper PDB configuration is essential for safe auto scaling and helped establish stronger governance practices.
Optimizing node maintenance workflows
The initial migration approach of cordoning Karpenter nodes in parallel led to unexpected cluster health issues. To address this, the team refined their strategy by implementing sequential node cordoning, adding manual verification checkpoints with rollback capabilities, and deploying enhanced monitoring for early detection of cluster instability. This experience reinforced that even with modern infrastructure tooling, careful orchestration of node maintenance remains crucial for system reliability.
Understanding Kubernetes label constraints
During the migration, the team discovered that Salesforce’s human-friendly legacy naming conventions often exceeded Kubernetes’s 63-character label length limit, creating challenges with Karpenter’s label-dependent operations. The team resolved this by refactoring naming conventions across node pools to comply with Kubernetes standards. This experience highlighted how seemingly minor technical constraints, such as label length limits, can become significant blockers in automated infrastructure management if not properly addressed early in the migration process.
For example, the following name is 67 characters long:
error: metadata.labels: Invalid value: must be no more than 63 characters
Protecting single-instance applications
The team discovered that Karpenter’s efficient bin-packing and consolidation features could unexpectedly impact applications running single-replica pods, leading to service disruptions in critical scenarios. To address this, we began implementing guaranteed pod lifetime features and workload-aware disruption policies to safeguard these singleton workloads. This experience demonstrated that effective auto scaling solutions must balance infrastructure efficiency with application availability requirements, particularly for mission-critical services.
Managing storage requirements in node migrations
The migration revealed that certain workloads failed to schedule due to incomplete ephemeral storage configurations. The team resolved this by implementing precise 1:1 mappings between the original Auto Scaling group–defined volume settings and Karpenter’s EC2NodeClass parameters. This experience emphasized the importance of carefully translating storage requirements during infrastructure migrations, particularly for I/O-intensive applications.
Realized value
The transition to Karpenter delivered measurable impact across multiple dimensions—performance, cost, and developer experience.
Operational efficiency
Salesforce eliminated thousands of node groups, significantly simplifying infrastructure management across its Kubernetes platform. Manual operational overhead was reduced by 80% through automation and the introduction of self-service capabilities. Developers can now define their own node pool requirements without waiting for centralized approvals, resulting in faster onboarding and greater agility.
Performance gains
With Karpenter, scaling latency was reduced from minutes to seconds by provisioning nodes based on actual pending pods, effectively bypassing delays associated with Auto Scaling groups. Node utilization improved significantly due to advanced bin-packing algorithms, resulting in fewer stranded resources and better efficiency. The migration eliminated Auto Scaling group thrashing, leading to more stable workloads and fewer scaling events during traffic spikes.
Cost optimization
Salesforce achieved 5% in cost savings in FY2026 by improving bin-packing efficiency and reducing idle capacity across its Kubernetes clusters. With the Karpenter rollout still in progress, an additional 5–10% in savings is projected for FY2027. The migration also lowered the overall cost-to-serve (CTS) by reducing the number of required nodes and improving multi-instance handling.
Enhanced developer and customer experience
The migration to Karpenter introduced true self-service infrastructure, allowing developers to define their capacity needs through straightforward node pool declarations. It also enabled greater flexibility by supporting heterogeneous instance types, including GPU, ARM, and x86, within a single node pool. Karpenter further improved IP efficiency by decoupling node provisioning from specific subnets, helping reduce IP fragmentation and exhaustion across the platform.
Conclusion
The migration to Karpenter represents a fundamental shift in how Salesforce manages Kubernetes infrastructure at scale. By addressing the limitations of traditional auto scaling approaches, we’ve achieved significant improvements in operational efficiency, cost optimization, and customer experience.
The key to our success was a combination of careful planning, custom tooling, and a phased approach that prioritized stability and zero-disruption migration. The results demonstrate that modern Kubernetes auto scaling solutions like Karpenter can transform platform operations while maintaining the reliability required for enterprise-scale deployments.
Salesforce’s success with Amazon EKS and Karpenter demonstrates how AWS continues to innovate alongside its largest enterprise customers, delivering solutions that scale from hundreds to thousands of clusters while reducing costs and operational complexity. This partnership showcases the power of combining AWS managed Kubernetes service with open source innovations like Karpenter to solve real-world challenges at unprecedented scale. To learn more, refer to the Karpenter Best Practices Guide in the Amazon EKS documentation.
At the beginning of January, I tend to set my top resolutions for the year, a way to focus on what I want to achieve. If AI and cloud computing are on your resolution list, consider creating an AWS Free Tier account to receive up to $200 in credits and have 6 months of risk-free experimentation with AWS services.
During this period, you can explore essential services across compute, storage, databases, and AI/ML, plus access to over 30 always-free services with monthly usage limits. After 6 months, you can decide whether to upgrade to a standard AWS account.
Whether you’re a student exploring career options, a developer expanding your skill set, or a professional building with cloud technologies, this hands-on approach lets you focus on what matters most: developing real expertise in the areas you’re passionate about.
Last week’s launches Here are the launches that got my attention this week:
AWS Config – Can now discover, assess, audit, and remediate additional AWS resource types across key services including Amazon EC2, Amazon SageMaker, and Amazon S3 Tables.
Crossmodal search with Amazon Nova Multimodal Embeddings – How to implement a crossmodal search system by generating embeddings, handling queries, and measuring performance with working code examples and tips to add these capabilities to your applications.
Upcoming AWS events Join us January 28 or 29 (depending on your time zone) for Best of AWS re:Invent, a free virtual event where we bring you the most impactful announcements and top sessions from AWS re:Invent. Jeff Barr, AWS VP and Chief Evangelist, will share his highlights during the opening session.
There is still time until January 21 to compete for $250,000 in prizes and AWS credits in the Global 10,000 AIdeas Competition (yes, the second letter is an I as in Idea, not an L as in like). No code required yet: simply submit your idea, and if you’re selected as a semifinalist, you’ll build your app using Kiro within AWS Free Tier limits. Beyond the cash prizes and potential featured placement at AWS re:Invent 2026, you’ll gain hands-on experience with next-generation AI tools and connect with innovators globally.
If you’re interested in these opportunities, join the AWS Builder Center to learn with builders in the AWS community.
That’s all for this week. Check back next Monday for another Weekly Roundup!
In open-source circles there are many situations, such as bug
reports, demos, and tutorials, when one might want to provide a
play-by-play of a session in one’s terminal. The asciinema project provides a set of
tools to do just that. Its tools let users record, edit, and share
terminal sessions in a text-based format that has quite a few
advantages compared to making and sharing videos of terminal sessions. For
example, it is easy to use, offers the ability to search text from
recorded sessions, and allows users to copy and paste directly from
the recording.
To provide the best experiences, we use technologies like cookies to store and/or access device information. Consenting to these technologies will allow us to process data such as browsing behavior or unique IDs on this site. Not consenting or withdrawing consent, may adversely affect certain features and functions.
Functional
Always active
The technical storage or access is strictly necessary for the legitimate purpose of enabling the use of a specific service explicitly requested by the subscriber or user, or for the sole purpose of carrying out the transmission of a communication over an electronic communications network.
Preferences
The technical storage or access is necessary for the legitimate purpose of storing preferences that are not requested by the subscriber or user.
Statistics
The technical storage or access that is used exclusively for statistical purposes.The technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance on the part of your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you.
Marketing
The technical storage or access is required to create user profiles to send advertising, or to track the user on a website or across several websites for similar marketing purposes.