Delta’s GoCool-150 Goes Big To Enable 150kW Liquid-To-Air Cooling for ASRock Rack’s NVIDIA VR NVL72

Post Syndicated from Ryan Smith original https://www.servethehome.com/deltas-gocool-150-goes-big-to-enable-150kw-liquid-to-air-cooling-for-asrock-racks-vr-nvl72/

How do you cool a giant rack of AI hardware? With an even bigger heat exchanger. Delta’s GoCool-150 is a liquid-to-air CDU that is designed to dissipate 150kW of heat from NVL72 and other high-density liquid cooled racks

The post Delta’s GoCool-150 Goes Big To Enable 150kW Liquid-To-Air Cooling for ASRock Rack’s NVIDIA VR NVL72 appeared first on ServeTheHome.

A decade of enterprise identity in the cloud with AWS Managed Microsoft AD

Post Syndicated from Vladimir Provorov original https://aws.amazon.com/blogs/security/a-decade-of-enterprise-identity-in-the-cloud-with-aws-managed-microsoft-ad/

Ten years ago, we launched AWS Directory Service for Microsoft Active Directory, a fully managed Microsoft Active Directory in the AWS Cloud. In that original announcement, Jeff Barr described a straightforward promise: “You will spend less time administering and more time working on your applications and your business.”

A decade later, AWS Managed Microsoft AD has become the identity backbone for thousands of enterprises worldwide. What started as a way to run directory-aware workloads in the cloud now powers SQL Server authentication, Amazon WorkSpaces virtual desktops, and Amazon FSx for Windows File Server for thousands of enterprises worldwide.

The beginning: Solving a real customer problem

In 2015, customers migrating Windows workloads to Amazon Web Services (AWS) faced a familiar challenge. Microsoft Active Directory (AD) had become the dominant standard for enterprise identity, by some estimates commanding 90% market share for directory services in the Fortune 1000. Running SharePoint, SQL Server, .NET applications, or virtually any Windows workload meant running AD.

However, running AD well comes with significant operational overhead. It requires careful capacity planning, high availability design across multiple sites, ongoing patching and maintenance, backup and disaster recovery procedures, and deep expertise that’s increasingly difficult to find and retain. Customers told us they wanted to focus on their applications, not on managing domain controllers.

So we built AWS Managed Microsoft AD. Powered by actual Windows Server, it delivered real Microsoft AD (not a compatible alternative, but the genuine article) as a fully managed service. We handled the domain controller deployment, the multi-AZ high availability, the automated backups, the patching, the monitoring, and many more features including scalability and multi-Region replication. Customers got a directory they could provision in 25–30 minutes and start using immediately.

From that original What’s New announcement by Bryan Nairn:

“AWS Directory Service now lets you run a Microsoft Active Directory (AD) as a managed service… Host monitoring and recovery, data replication, snapshots, and software updates are automatically configured and managed for you.”

The first decade of innovation

Looking back at the past 10 years, we’re struck by how much AWS Managed Microsoft AD has evolved in response to customer feedback. Here are some of the highlights:

2015: Launch of AWS Managed Microsoft AD (Enterprise Edition) in five AWS Regions, powered by Windows Server 2012 R2. Support for trust relationships with on-premises AD, seamless domain join for Amazon Elastic Compute Cloud (Amazon EC2) instances, and integration with Amazon WorkSpaces.

2017: Introduction of Standard Edition, optimized for small and midsize businesses. This gave customers a cost-effective option for resource forest deployments and smaller workloads.

2018: Added support for schema extensions, enabling customers to extend their directory schema for applications that require custom attributes. Support for Group Managed Service Accounts (gMSA) with Windows containers and other services.

2019: Launched multi-Region replication for Enterprise Edition, allowing customers to automatically replicate their directory across AWS Regions for improved performance and disaster recovery. Added directory sharing across AWS accounts and integration with AWS Organizations.

2020: Introduced fine-grained directory settings for security and compliance, enabling customers to configure secure channel settings for protocols and ciphers. Enhanced compliance support—with the service now HIPAA eligible—included as an in-scope service under PCI DSS, and achieving FedRAMP authorization.

2021: Added CloudWatch metrics for domain controllers, helping customers optimize scaling decisions based on CPU, memory, disk, and AD-specific metrics like DNS and directory read/write operations. Launched integration with AWS Transfer Family for SFTP/FTPS/FTP authentication.

2022: Windows Server 2019 upgrade became available, with customer-initiated updates and automatic migration for all directories beginning in 2023.

2023: AWS Private CA Connector for Active Directory launched, allowing customers to replace self-managed enterprise certificate authorities with AWS Private CA for automatic certificate enrollment to domain-joined objects, with no local agents or proxy servers required.

2024: Launched CRUD APIs for users and groups, enabling IT administrators to manage AD users and groups directly from the AWS Management Console, AWS Command Line Interface (AWS CLI), and APIs, without deploying bastion hosts or opening network ports.

2025: General availability of AWS Managed Microsoft AD (Hybrid Edition), allowing customers to extend their existing AD domain to AWS while retaining administrative control. Introduced self-service edition upgrades through the UpdateDirectorySetup API, eliminating the need for support tickets when scaling from Standard to Enterprise Edition.

2026 and beyond: As we enter our second decade, our roadmap continues to be shaped by the customers who depend on AWS Managed Microsoft AD every day. We’re working on new capabilities driven directly by your feedback, and we look forward to sharing more soon.

Powering identity across AWS

Over the past decade, more than 20 AWS services have added native integration with AWS Managed Microsoft AD. What started with WorkSpaces and EC2 domain join has expanded to more than 20 AWS services, making AWS Managed Microsoft AD foundational for many enterprise customers’ workloads on AWS.

Database services

For many customers, database authentication is a primary driver for adopting AWS Managed Microsoft AD. By pairing Amazon Relational Database Service (Amazon RDS) for SQL Server with AWS Managed Microsoft AD, they gain the benefits of fully managed services while achieving straightforward integration and reduced management overhead. This combination lets developers and DBAs use their existing AD credentials to access SQL Server databases, so they don’t need to manage separate database accounts.

Beyond SQL Server, AWS Managed Microsoft AD enables Windows authentication across the Amazon RDS family:

  • Amazon RDS for Oracle
  • Amazon RDS for PostgreSQL
  • Amazon RDS for MySQL
  • Amazon RDS for DB2
  • Amazon Aurora MySQL
  • Amazon Aurora PostgreSQL

File storage services

Amazon FSx for Windows File Server provides fully managed Windows file shares that integrate natively with AWS Managed Microsoft AD. Customers use AD users and groups to control access to file shares, apply Windows ACLs, and use features like DFS namespaces, all with the same management experience they use on premises.

AWS Storage Gateway supports AD authentication for SMB file shares, enabling hybrid storage architectures where on-premises applications access cloud storage using familiar AD credentials.

AWS Transfer Family added AD integration in 2021, allowing customers to authenticate SFTP, FTPS, and FTP users against their AWS Managed Microsoft AD. This allows customers to migrate file transfer workflows without changing end-user credentials.

End user computing

Amazon end-user computing services were among the first to integrate with AWS Managed Microsoft AD:

Security and identity

AWS IAM Identity Center (formerly AWS Single Sign-On) uses AWS Managed Microsoft AD as an identity source, synchronizing users and groups to provide single sign-on access across AWS accounts and applications. This provides centralized identity management while using your existing AD infrastructure.

AWS Client VPN authenticates users against AWS Managed Microsoft AD, providing secure remote access using corporate credentials.

AWS Management Console access can be federated through AWS Managed Microsoft AD, so AD users can assume AWS Identity and Access Management (IAM) roles and manage AWS resources with their existing credentials.

Compute services

Amazon EC2 instances (both Windows and Linux) support seamless domain join at launch. Windows instances can be managed using Group Policy, and Linux instances can authenticate users through SSSD or Realm integration.

Amazon Elastic Container Service (Amazon ECS) supports AD authentication for Windows containers through Group Managed Service Accounts (gMSA), enabling containerized applications to authenticate to AD-integrated resources.

Business applications

This breadth of integration means customers can standardize on a single directory for their entire AWS environment, from databases to desktops to file servers to analytics.

Choosing the right edition

Over the years, we’ve learned that customers have different needs when it comes to managed AD. Today, AWS Managed Microsoft AD is available in three editions, each designed for specific use cases.

Standard Edition: Basic, cost-effective identity

Standard Edition is optimized for small and midsize businesses, or for enterprises deploying a resource forest model in a single AWS Region. With 1 GB of directory object storage supporting up to 30,000 objects (approximately 5,000 users), Standard Edition provides everything needed to run directory-aware workloads without the overhead of managing domain controllers.

Common use cases:

  • Resource forest deployments – Many customers use Standard Edition as a resource forest, establishing a trust relationship with their on-premises AD. User identities remain in the customer’s existing domain, while the resource forest manages AWS resources like Amazon RDS for SQL Server and FSx for Windows File Server.
  • Development and test environments – Cost-effective option for non-production workloads
  • Single-Region applications – Workloads that don’t require global presence

Standard Edition is a great starting point, and customers aren’t locked in. With our new self-service upgrade capability (launched October 2025), you can upgrade to Enterprise Edition programmatically through the UpdateDirectorySetup API, no support tickets or maintenance window coordination required.

Enterprise Edition: Built for global scale

Enterprise Edition is designed for organizations with larger user populations, complex deployments, or global footprints. With 17 GB of storage supporting up to 500,000 directory objects, Enterprise Edition provides the capacity and capabilities that large enterprises require.

Key capabilities:

  • Multi-Region replication – Automatically replicate your directory across AWS Regions. Users and applications connect to local domain controllers, reducing latency and providing disaster recovery capabilities.
  • Extended directory sharing – Share your directory with up to 500 AWS accounts, enabling centralized identity across large organizations using AWS Organizations.
  • Higher compute capacity – Larger domain controller instances with more CPU and memory for demanding workloads

If you have users and applications in multiple geographic regions, or anticipate significant growth in directory objects, Enterprise Edition is the right choice.

Hybrid Edition: Extend your existing domain

Launched earlier this year, Hybrid Edition takes a fundamentally different approach. Instead of creating a new AD domain in AWS, Hybrid Edition extends your existing AD domain into the cloud.

What makes Hybrid Edition unique:

  • Same domain – AWS Managed Microsoft AD domain controllers join your existing AD. No new domain name, no trust relationships to configure.
  • Retain administrative control – Unlike Standard and Enterprise where you receive delegated OU permissions, Hybrid Edition preserves your existing administrative rights. Your AD administrators continue using familiar tools while changes replicate to AWS in real time.
  • Preserve existing investments – Security principals, group policies, and permissions transfer seamlessly. No migration of identities required.

Hybrid Edition is ideal for customers who want the operational benefits of AWS-managed domain controller infrastructure without changing their AD architecture or giving up administrative control.

Which edition should you choose?

Use the following table to determine which edition best fits your use case.

Use case Edition
A new AD domain for AWS workloads in a single Region Standard Edition
A resource forest with trust to on-premises AD Standard Edition
Multi-Region replication for global deployments Enterprise Edition
Support for more than 30,000 directory objects Enterprise Edition
To extend your existing AD domain to AWS Hybrid Edition
To retain full administrative control over your AD Hybrid Edition

What we’ve learned: Design decisions that stood the test of time

Looking back at the decisions we made in 2015, several have proven foundational to the service’s success:

  • High availability by default – Every AWS Managed Microsoft AD directory deploys with a minimum of two domain controllers across separate Availability Zones. Customers don’t need to design high availability (HA) architecture, it’s built in.
  • Real Microsoft AD – We chose to run actual Windows Server AD, not a compatible alternative. This means standard AD administration tools work, existing scripts and automation work, and applications that depend on specific AD behaviors typically work without modification.
  • Seamless integration with AWS services – By building native integrations between AWS Managed Microsoft AD and other AWS services, we’ve made it possible for customers to use a single directory across their entire AWS environment.
  • Customer retains control – While AWS manages the infrastructure, customers manage their directory content. You control your users, groups, OUs, and policies using familiar tools.
  • Room to grow – The edition model (and now self-service upgrades) means customers can start with what they need today and scale as requirements evolve.

Looking ahead: The next chapter

As we celebrate 10 years of AWS Managed Microsoft AD, we’re excited about what’s ahead. The launch of Hybrid Edition earlier this year represents a significant expansion of what’s possible, giving customers new flexibility in how they architect their identity infrastructure for hybrid and multi-cloud environments.

We continue to listen to customer feedback and invest in capabilities that reduce operational burden while expanding what you can build. Whether you’re running your first SQL Server database in the cloud, deploying virtual desktops to a global workforce, or modernizing legacy applications that depend on AD, AWS Managed Microsoft AD is here to help.

Thank you to all the customers who have trusted us with their identity infrastructure over the past decade. Your feedback has shaped this service, and we’re committed to continuing to earn that trust for the next 10 years and beyond.

Resources

Ready to get started or learn more? Here are some resources:

If you have feedback about this post, submit comments in the Comments section below.


Vladimir Provorov

Vladimir is a Product Solutions Architect from AWS Identity focused on Workforce Identity and Directory Service. He works on developing new features to make Enterprise Identity simpler and more scalable. He is excited to travel and explore the world with his family.

Rodney Underkoffler

Rodney Underkoffler

Rodney is a Senior Solutions Architect at Amazon Web Services, focused on guiding enterprise customers on their cloud journey. He has a background in infrastructure, security, and IT business practices. He is passionate about technology and enjoys building and exploring new solutions and methodologies.

Author

Tekena Orugbani

Tekena is a Sr. Specialist Solutions Architect at Amazon Web Services and a technologist of over 20 years, specializing in Microsoft technologies. At AWS, Tekena is focused on helping customers architect, migrate and modernize their Microsoft workloads on the AWS Cloud. Outside work, he enjoys hanging out with his family and watching soccer.

Securing your Amazon S3 buckets: Identifying and remediating over-permissioned access

Post Syndicated from Hetal Kolekar original https://aws.amazon.com/blogs/security/securing-your-amazon-s3-buckets-identifying-and-remediating-over-permissioned-access/

Misconfigured Amazon Simple Storage Service (Amazon S3) buckets can expose your data to unauthorized access. Without proactive review, S3 bucket policies or Access Control Lists (ACLs) configured with broad access may go unnoticed in your environment. In this post, you learn how to identify and fix over-permissioned S3 buckets across your AWS environment, along with best practice recommendations and automation opportunities to help you prevent security gaps. This post provides a workflow framework and methodology recommendations for your security team to adapt. The focus of this post is on the what and why rather than a prescriptive implementation. You will need to customize the approach based on your organization’s requirements and existing security tooling.

This solution is intended for security engineers, cloud architects, and DevOps teams managing single- or multiple-account AWS environments with Amazon S3 workloads that require access management.

Prerequisites

Before you begin, make sure you have the following in place:

Solution overview

This solution uses a five-phase workflow diagram to detect, remediate, and continuously monitor over-permissioned S3 buckets across your AWS accounts. The following workflow diagram illustrates the high-level end-to-end process for identifying and remediating over-permissioned S3 buckets across your Amazon Web Services (AWS) environment.

Figure 1: Amazon S3 over-permissive access – Detection, remediation, monitoring and cleanup workflow

Figure 1: Amazon S3 over-permissive access – Detection, remediation, monitoring and cleanup workflow

The diagram in Figure 1 consists of five phases:

  1. Setup and prerequisites – Configure AWS Organizations or multi-account access, designate a central security account, deploy AWS Config across all accounts, and enable AWS Security Hub with a central administrator.
  2. Detection and identification – Deploy AWS Config rules (such as s3-bucket-public-read-prohibited and s3-bucket-public-write-prohibited) and run an audit Lambda function that scans each S3 bucket. The function checks three areas: Public Access Block configuration, bucket policy status, and bucket ACL grants. Buckets with issues are added to a risky buckets list. The function then generates a report in CSV and JSON format, uploads it to an output S3 bucket, and sends an SNS alert.
  3. Remediation – Address findings using one or more approaches – Apply restrictive bucket policies to deny public read/write access and restrict access to specific IAM principals; deploy a remediation Lambda function to automatically update bucket policies and disable public access settings; or use CloudFormation StackSets to deploy standardized policies across multiple accounts.
  4. Continuous monitoring – Schedule the audit Lambda function for recurring scans (daily or weekly) using Amazon EventBridge. Use EventBridge to detect policy changes, configure automated notifications for new violations, enable IAM Access Analyzer for S3 to identify external access, and run regular compliance scans.
  5. Resource cleanup – Review and delete resources created during the audit that are no longer needed, including Lambda functions and IAM roles, EventBridge rules, SNS topics and subscriptions, audit output S3 buckets, AWS Config rules, and Security Hub (if enabled only for this audit).

Cost considerations

This section covers the AWS services used in this solution and their associated costs so you can estimate spend before deployment. The primary cost drivers are AWS Config and Security Hub, which scale with the number of accounts and resources you monitor. Lambda, Amazon EventBridge, Amazon SNS, and Amazon S3 typically add minimal costs for most environments. Start with a pilot in one or two accounts to validate costs before scaling.

  • AWS Config – Charges per configuration item recorded and per rule evaluation. Costs scale with the number of accounts and resources tracked.
  • Security Hub – Charges per account per AWS Region for security checks and finding ingestion.
  • Lambda – Charges per request and per GB-second of compute time.
  • EventBridge – Scheduled rules are free. Custom event bus usage might incur charges.
  • Amazon SNS – Charges per notification delivered.
  • Amazon S3 – Storage costs for audit report output files. Minimal for most environments.
  • AWS IAM Access Analyzer – Check the AWS IAM Access Analyzer pricing page to understand which features have costs associated with them.

Check the service pricing pages for current rates. Use the AWS Pricing Calculator to estimate costs for your specific environment before enabling services across all accounts. Consider starting with a pilot in one or two accounts to validate costs before scaling.

Detect and report over-permissioned buckets

This section walks you through setting up the audit environment, deploying the Lambda-based scanner, and generating reports of over-permissioned S3 buckets across your accounts. Follow these steps to identify over-permissioned S3 buckets in your multi-account environment, starting with preparing your environment for an Amazon S3 audit.

To set up the multi-account audit environment:

  1. Set up AWS Organizations or multi-account access. Set up centralized management of your AWS accounts using AWS Organizations or configure cross-account IAM roles.
  2. Choose a central security account. Choose one account as your security/audit account. This account will run the audit Lambda function and collect results from member accounts.
  3. Create an Amazon SNS topic for alerts. Subscribe your security team to receive notifications when over-permissioned buckets are detected. Note the topic Amazon Resource Name (ARN) from the output—you will need it when creating the Lambda execution role (step 6) and the Lambda function (step 9). Confirm the email subscription before testing; Amazon SNS doesn’t deliver alerts until the subscription is confirmed. Learn more in the Amazon SNS Developer Guide.
  4. (Optional): Create an S3 bucket for audit reports. If you plan to use Script v2 for historical reporting and trend analysis, create a dedicated bucket now. Skip this step if you only need real-time alerts using Script v1.
  5. Plan cross-account IAM roles. The central security account needs permission to scan member accounts. Design cross-account roles that:
    1. Grant minimum Amazon S3 read permissions (list buckets, read policies, ACLs, public access configurations).
    2. Include an external ID condition to mitigate the confused deputy problem.
    3. Can be deployed consistently using AWS CloudFormation StackSets.
    4. See the IAM documentation on creating cross-account roles, The confused deputy problem, and IAM security best practices for additional guidance on role configuration and trust policies.

      Note: The specific trust policy and permissions policy for your cross-account roles will depend on organizational requirements. Work with your IAM administrators to grant minimum necessary access for the audit function.

  6. Create the Lambda execution role. Create an IAM role for your Lambda function with the permissions it needs to scan buckets, publish alerts, and write logs. Apply the principle of least privilege—grant only the minimum Amazon S3 read permissions required for the audit (such as, listing buckets, reading bucket policies, ACLs, and public access block configurations), Amazon SNS publish permission for the alert topic created in step 3, Amazon S3 write permission for the output bucket created in step 4 (Script v2), and Amazon CloudWatch Logs permissions. For multi-account scanning, also include sts:AssumeRolepermission for the cross-account role ARNs created in step 5. The AWS Lambda execution role documentation has instructions on creating and configuring execution roles.
  7. To deploy the S3 audit solution Deploy the audit components
    1. Enable AWS Config in member accounts. AWS Config provides compliance monitoring and can detect when S3 buckets are created or modified with public access settings. This will enable the Lambda-based audit to receive real-time detection between scheduled scans. The AWS Config Developer Guide has setup instructions. Deploy pre-defined AWS Config rules to identify overly permissive settings. These managed rules provide automated compliance checking. When AWS Config detects violations, it sends findings to Security Hub (configured in step 8) for centralized visibility alongside the Lambda audit results.
      • s3-bucket-public-read-prohibited
      • s3-bucket-public-write-prohibited
      • Create AWS Config rules for specific permission patterns. For the full list of available rules, see the AWS Config managed rules reference
  8. Enable Security Hub for centralized visibility. Enable AWS Security Hub in member accounts and configure the central security account as the administrator. Security Hub aggregates findings from AWS Config rules (step 7), IAM Access Analyzer (enabled later), and can receive custom findings from your Lambda audit function, providing a single dashboard for Amazon S3 security issues across your organization. See the Security Hub User Guide for setup details.
  9. Deploy the audit Lambda function. Deploy a Python Lambda function using the Boto3 library to list S3 buckets, check their policies, ACLs, and IAM permissions, and identify over-permissioned buckets. See the example scripts that follow.

Important: These code examples aren’t production ready. Adapt them to meet your organization’s requirements and test them in a non-production environment before deployment.

Choose your approach:

  • Script v1 – Best for immediate SNS alerts when issues are detected.
  • Script v2 – Best for historical reports, trend analysis using BI tools.
  • Both scripts – Best for different schedules and ongoing needs.

Audit Lambda function – Example script v1 (Scan and alert)

The following is an example of a Lambda function script for reference purposes. Review, adapt, and test before use in your environment, it scans all S3 buckets in the current account and checks for:

  • Public Access block configuration gaps
  • Bucket policies that allow public access
  • ACL grants to AllUsers

Note: Replace placeholder values with actual values before deployment:

  • <REGION>– Your AWS Region (for example, us-east-1)
  • <ACCOUNT_ID>– Your 12-digit AWS account ID
  • <TOPIC_NAME>– The name of your SNS topic created in step 3
import boto3
import json

def lambda_handler(event, context):
    s3 = boto3.client('s3')
    sns = boto3.client('sns')
    risky_buckets = []
    errors = []

    try:
        buckets = s3.list_buckets()['Buckets']
    except Exception as e:
        return {'statusCode': 500, 'body': f'Failed to list buckets: {str(e)}'}

    for bucket in buckets:
        bucket_name = bucket['Name']
        issues = []

        try:
            # Check Public Access Block — all four settings should be enabled
            try:
                pab = s3.get_public_access_block(Bucket=bucket_name)
                config = pab['PublicAccessBlockConfiguration']
                if not all([
                    config.get('BlockPublicAcls'),      # Block new public ACLs
                    config.get('BlockPublicPolicy'),     # Block new public bucket policies
                    config.get('IgnorePublicAcls'),      # Ignore existing public ACLs
                    config.get('RestrictPublicBuckets')   # Restrict access to public buckets
                ]):
                    issues.append('Public Access Block not fully enabled')
            except s3.exceptions.NoSuchPublicAccessBlockConfiguration:
                issues.append('No Public Access Block configured')

            # Check bucket policy — flag if policy status is public
            try:
                policy_status = s3.get_bucket_policy_status(Bucket=bucket_name)
                if policy_status['PolicyStatus']['IsPublic']:
                    issues.append('Bucket policy allows public access')
            except s3.exceptions.NoSuchBucketPolicy:
                pass  # No bucket policy is acceptable

            # Check bucket ACL
            acl = s3.get_bucket_acl(Bucket=bucket_name)
            for grant in acl.get('Grants', []):
                grantee = grant.get('Grantee', {})
                uri = grantee.get('URI', '')
                # 'AllUsers' = anonymous public access
                # 'AuthenticatedUsers' = any AWS account (still overly permissive)
                if grantee.get('Type') == 'Group' and ('AllUsers' in uri or 'AuthenticatedUsers' in uri):
                    issues.append('Bucket ACL grants public access')
                    break

            if issues:
                risky_buckets.append({'bucket': bucket_name, 'issues': issues})

        except Exception as e:
            errors.append(f'{bucket_name}: {str(e)}')

    # Send alert if risky buckets found
    if risky_buckets:
        message = f'Found {len(risky_buckets)} buckets with public access:\n\n'
        for item in risky_buckets:
            message += f"  {item['bucket']}: {', '.join(item['issues'])}\n"

        sns.publish(
            TopicArn='arn:aws:sns:<REGION>:<ACCOUNT_ID>:<TOPIC_NAME>',
            Subject='S3 Public Access Alert',
            Message=message
        )

    return {
        'statusCode': 200,
        'body': json.dumps({
            'risky_buckets': risky_buckets,
            'errors': errors,
            'total_checked': len(buckets)
        })
    }

Multi-account scanning: This script scans the current account only. To scan across member accounts, see the Multi-account extension section later in this post.

Audit Lambda function – Example script v2 (CSV and JSON report)

The following is an example Lambda function script for reference purposes. Before deploying any script, review error handling, logging, output structure, and permissions. This script generates CSV and JSON output files and uploads them to an S3 bucket for reporting and business intelligence (BI) dashboard integration.

You can deploy both functions with different EventBridge schedules, for example, Script v1 daily for alerts and Script v2 weekly for reports.

Note: Before you deploy this script, replace <OUTPUT_BUCKET_NAME> with the S3 bucket you created for audit reports in step 4.

import boto3
import csv
import json
import os

def lambda_handler(event, context):
    s3 = boto3.client('s3')
    buckets = s3.list_buckets()['Buckets']

    full_access_buckets = []
    for bucket in buckets:
        bucket_name = bucket['Name']
        try:
            bucket_policy = s3.get_bucket_policy(Bucket=bucket_name)['Policy']
            policy = json.loads(bucket_policy)
            for statement in policy['Statement']:
                if (statement['Effect'] == 'Allow'
                    and statement['Principal'] == '*'
                    and 'Action' in statement
                    and 's3:*' in statement['Action']):
                    full_access_buckets.append({'BucketName': bucket_name})
                    break
        except s3.exceptions.ClientError as e:
            if e.response['Error']['Code'] != 'NoSuchBucketPolicy':
                print(f'Error checking bucket policy for {bucket_name}: {e}')

    # Output CSV
    csv_output = os.path.join('/tmp', 'full_access_buckets.csv')
    with open(csv_output, 'w', newline='') as csvfile:
        writer = csv.DictWriter(csvfile, fieldnames=['BucketName'])
        writer.writeheader()
        writer.writerows(full_access_buckets)

    # Output JSON
    json_output = os.path.join('/tmp', 'full_access_buckets.json')
    with open(json_output, 'w') as jsonfile:
        json.dump(full_access_buckets, jsonfile, indent=2)

    # Upload to Amazon S3
    output_bucket = '<OUTPUT_BUCKET_NAME>'
    s3.upload_file(csv_output, output_bucket, 'full_access_buckets.csv')
    s3.upload_file(json_output, output_bucket, 'full_access_buckets.json')

    return {
        'statusCode': 200,
        'body': json.dumps(f'CSV and JSON files uploaded to {output_bucket}')
    }

Important: If this function runs on a schedule, consider implementing a file naming strategy with timestamps to prevent overwriting previous reports or establish a lifecycle policy to manage retention. Include the output bucket in your cleanup procedures when the auditing process is no longer needed.

What if no over-permissioned buckets are found?

If the audit scan returns zero risky buckets, document the clean baseline for future comparison and move to the verification and monitoring phase to so new buckets or policy changes don’t introduce risk over time.

Multi-account extension

The preceding example scripts scan buckets in the current account only. To scan across member accounts in your organization, add the following AssumeRole logic. This function assumes the cross-account IAM role you created during setup, then returns an Amazon S3 client with temporary credentials for each member account.

Note: Before you deploy, configure the following Lambda environment variables:

  • <MEMBER_ACCOUNTS> – Comma-separated list of 12-digit account IDs to scan (for example, 111111111111,222222222222)
  • <CROSS_ACCOUNT_ROLE_NAME> – The IAM role name created in each member account (for example, S3AuditRole)
  • <EXTERNAL_ID> – The external ID configured in the trust policy (for example, s3-audit-external-id)
import boto3
import os

def get_member_s3_clients():
    """
    Assumes the cross-account audit role in each member account
    and returns a list of (account_id, s3_client) tuples.
    """
    sts = boto3.client('sts')
    member_accounts = os.environ.get('<MEMBER_ACCOUNTS>', '').split(',')
    cross_account_role_name = os.environ.get('<CROSS_ACCOUNT_ROLE_NAME>')
    external_id = os.environ.get('<EXTERNAL_ID>')

    clients = []
    for account_id in member_accounts:
        account_id = account_id.strip()
        if not account_id:
            continue

        try:
            assumed_role = sts.assume_role(
                RoleArn=f'arn:aws:iam::{account_id}:role/{cross_account_role_name}',
                RoleSessionName='S3AuditSession',
                ExternalId=external_id
            )

            # Create S3 client with assumed credentials
            s3_client = boto3.client(
                's3',
                aws_access_key_id=assumed_role['Credentials']['AccessKeyId'],
                aws_secret_access_key=assumed_role['Credentials']['SecretAccessKey'],
                aws_session_token=assumed_role['Credentials']['SessionToken']
            )
            clients.append((account_id, s3_client))

        except Exception as e:
            print(f'Failed to assume role in account {account_id}: {e}')

    return clients

To scan each member account, replace the single-account s3.list_buckets() call with a loop over member accounts:

def lambda_handler(event, context):
    all_risky_buckets = []
    all_errors = []

    # Scan each member account
    for account_id, s3_client in get_member_s3_clients():
        try:
            buckets = s3_client.list_buckets()['Buckets']
            for bucket in buckets:
                # ... same scanning logic as the single-account scripts ...
                # Use s3_client instead of s3 for each API call
                pass
        except Exception as e:
            all_errors.append(f'Account {account_id}: {e}')

    # ... same alerting/reporting logic ...

The Lambda execution role in the central security account needs sts:AssumeRole permission for the cross-account role ARNs. Add this to the execution role policy you created in step 5.

Remediate elevated access

This section describes how to fix over-permissioned buckets using account-level controls, bucket policies, and optional automation. Any elevated access that you find needs to be remediated.

Enable Amazon S3 Block Public Access (account level)

Before applying individual bucket policies, enable Amazon S3 Block Public Access at the account level. This prevents buckets in the account from being made public, regardless of individual bucket policies or ACLs. See theS3 Block Public Access documentation for configuration details. See the following example AWS CLI command; replace <ACCOUNT_ID> with the ID of the account you’re using to manage resource access:

aws s3control put-public-access-block \
  --account-id <ACCOUNT_ID> \
  --public-access-block-configuration \
BlockPublicAcls=true,IgnorePublicAcls=true,BlockPublicPolicy=true,RestrictPublicBuckets=true

For multi-account environments, deploy this setting across member accounts using AWS CloudFormation StackSets or AWS Organizations service control policies (SCPs).

Important: Before enabling account-level S3 Block Public Access, check whether any workloads need public bucket access (for example, static website hosting, public dataset sharing). Coordinate with your application teams to identify any exceptions.

Remediate using bucket policies

Implement bucket policies that restrict access to specific IAM users, roles, or accounts. When crafting policies, apply the principle of least privilege and include only the actions and principals required for your use case.

Example S3 bucket policy: deny public read/write access. Modify the resource ARN, actions, and conditions to match your requirements:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Deny",
      "Principal": "*",
      "Action": [
        "s3:PutObject", "s3:PutObjectAcl",
        "s3:GetObject", "s3:GetObjectAcl",
        "s3:DeleteObject"
      ],
      "Resource": "arn:aws:s3:::<BUCKET_NAME>/*",
      "Condition": {
        "StringEquals": {
          "s3:x-amz-acl": ["public-read", "public-read-write"]
        }
      }
    }
  ]
}

Example S3 bucket policy: restrict access to specific IAM principals. Replace <ACCOUNT_ID>, <USERNAME>, and <ROLE_NAME>:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "AllowObjectAccess",
      "Effect": "Allow",
      "Principal": {
        "AWS": [
          "arn:aws:iam::<ACCOUNT_ID>:user/<USERNAME>",
          "arn:aws:iam::<ACCOUNT_ID>:role/<ROLE_NAME>"
        ]
      },
      "Action": ["s3:GetObject", "s3:PutObject", "s3:DeleteObject"],
      "Resource": "arn:aws:s3:::<BUCKET_NAME>/*"
    },
    {
      "Sid": "AllowBucketAccess",
      "Effect": "Allow",
      "Principal": {
        "AWS": [
          "arn:aws:iam::<ACCOUNT_ID>:user/<USERNAME>",
          "arn:aws:iam::<ACCOUNT_ID>:role/<ROLE_NAME>"
        ]
      },
      "Action": ["s3:ListBucket", "s3:GetBucketLocation"],
      "Resource": "arn:aws:s3:::<BUCKET_NAME>"
    }
  ]
}

See the Amazon S3 bucket policy documentation for additional examples and guidance.

Automate remediation with Lambda or CloudFormation StackSets (optional):

You can also remediate using Lambda or CloudFormation Stacksets:

  • Create Lambda functions to automatically update bucket policies or disable public access settings for flagged buckets
  • Use CloudFormation StackSets to deploy standardized bucket policies and S3 Block Public Access settings across multiple accounts

Verify your remediation

This section explains how to confirm that your fixes are effective before moving to ongoing monitoring. After applying remediation, verify the fix is effective before setting up ongoing monitoring:

  1. Re-run the audit Lambda function – Confirm the previously flagged buckets no longer appear in the risky buckets list.
  2. Check Security Hub compliance – Verify the compliance status has changed from FAILED to PASSED for Amazon S3-related controls.
  3. Validate with IAM Access Analyzer – Review findings for the remediated S3 buckets. Active findings should resolve automatically after public access is removed.
  4. Test application functionality – Confirm that legitimate workloads continue to function correctly.

Document the verification results for your auditing needs. If any S3 buckets still show issues, investigate whether the policy was applied correctly or if there are conflicting permissions.

Automation opportunities

This section covers optional strategies to automate ongoing detection and maintain your security posture without manual intervention.

  1. (Optional) Schedule recurring scans with Amazon EventBridge
    • Regular security scans help identify new issues arising from configuration changes or newly created S3 buckets. When new security risks are detected, Amazon SNS sends an alert and automatically initiates the remediation phase (Workflow 2 in Figure 1). To avoid repeated alerts, you can configure the audit Lambda function to run on a schedule and compare current results with the previous baseline to generate notifications when new findings are discovered.
    • For ongoing monitoring, you can schedule the audit Lambda function to run on a recurring basis using EventBridge. Create a scheduled rule with a cron expression (for example, daily at 6:00 AM UTC or weekly on Mondays), add the Lambda function as the target, and grant EventBridge permission to invoke it. See Amazon EventBridge scheduling documentation for instructions on creating scheduled rules and configuring targets.
  2. Enable IAM Access Analyzer for Amazon S3
    • IAM Access Analyzer monitors bucket policies, ACLs, and access points to identify buckets accessible from outside your account or organization. Create an analyzer scoped to your organization or individual account, then review findings to identify unintended external access. Findings automatically flow into Security Hub when both services are enabled, giving you a dashboard view for Amazon S3 security findings. See the IAM Access Analyzer documentation for setup and usage instructions.
  3. Automate notifications for policy drift
    • Recurring scans might surface new findings from policy drift or newly created buckets. When new risks are detected, Amazon SNS alert triggers and the remediation cycle repeat (as shown in Workflow 2 in Figure 1) sends email notifications. Configure the audit Lambda function to compare current scan results against the previous baseline and alert on new findings for ongoing reviews.

Clean up

This section lists the resources created during this walkthrough that you should review and remove when they are no longer needed. If the following services were not previously active in your account, leaving them enabled might result in additional ongoing charges. See the Cost considerations section for details. Review and remove unused resources to optimize costs.

Delete or disable the following script-generated resources if they’re not required after outputs are generated. Focus first on Lambda functions and EventBridge rules if you’re not running recurring scans. If you enabled AWS Config or Security Hub specifically for this audit, evaluate whether you need them for other compliance requirements before disabling.

  • Lambda – Functions, IAM roles, and policies created for auditing
  • Amazon EventBridge – Scheduled rules created for recurring audit triggers
  • Amazon SNS – Topics and subscriptions created for notifications
  • Amazon S3 – Buckets containing script-generated audit output files
  • AWS Config – Rules and recorders if no longer needed for compliance
  • Security Hub – Disable if enabled solely for this audit
  • IAM Access Analyzer – Delete the analyzer if no longer needed for ongoing monitoring

Note: Be careful when deleting data and consider temporarily disabling services first to check for dependencies. Only delete resources generated as part of your audit outputs. Verify you have retained any necessary results before proceeding. Verify resources are not used by other workloads before deletion.

Best practices

This section provides recommendations to maintain secure Amazon S3 configurations long-term. To learn more about maintaining secure Amazon S3 configurations, review the AWS documentation links provided in the conclusion. The following recommendations aren’t exhaustive. Adapt and extend them based on your organization’s evolving security requirements and AWS best practices guidance. After you’ve fixed existing issues, these practices help you maintain secure Amazon S3 configurations.

  • Start with account-level controls – Enable S3 Block Public Access at the account level. This prevents buckets from becoming public even if someone misconfigures an individual bucket policy. For multi-account environments, enforce this through AWS Organizations SCPs.
  • Automate detection – Use IAM Access Analyzer to detect external access. Schedule your audit Lambda function with EventBridge to catch new issues weekly or daily, depending on your change frequency. Compare scan results against previous baselines to identify drift.
  • Standardize across accounts – Use CloudFormation StackSets to deploy the same secure configuration to all accounts in your organization, reducing the chance of configuration drift. Use StackSets for IAM roles, AWS Config rules, and S3 Block Public Access settings.

Additional security measures

  • Regularly review and rotate cross-account IAM role credentials and external IDs
  • Implement Amazon S3 server-side encryption (SSE-S3 or SSE-KMS) for data at rest
  • Enable S3 access logging and AWS CloudTrail data events for audit trails

Conclusion

This section summarizes what you accomplished and suggests next steps to maintain your S3 security posture. By implementing the detection, remediation, and monitoring workflow outlined in this post, you can proactively identify and secure over-permissioned S3 buckets across your AWS environment. To maintain your ongoing security posture, enable IAM Access Analyzer for continuous monitoring and schedule recurring audits with EventBridge. To learn more about Amazon S3 security best practices, see Security best practices for Amazon S3

For more information:

If you have feedback about this post, submit comments in the Comments section below.


Hetal Kolekar

Hetal Kolekar

Hetal is a Sr. Technical Account Manager at AWS with more than 21 years of experience in Infrastructure Architecture, Security, Systems Engineering, and Consulting. He excels in leading teams to strengthen their cloud security posture and helps customers scale up their security using AWS services. Hetal is a guitarist and loves playing at church.

Manomayi Vedam

Manonmayi Vedam

Manonmayi is a Senior TAM and Product Owner at AWS, specializing in AI-driven cloud enablement, security, and generative AI risk across Healthcare, Financial Services, Energy, and Public Sector. She co-leads global security programs for Fortune 500 clients, contributes to the NIST Cyber AI Profile RMF and NCCoE, and is a Fellow at SCRS with recognition from GlobeeAwards and IEEE.

Fernando Freitas

Fernando Freitas

Fernando is a Sr. Technical Account Manager at AWS in Salt Lake City, focused on helping customers achieve their desired outcomes with the AWS Cloud. Fernando is passionate about Identity and Security, Training and Education.

Amazon OpenSearch Service extends version lifecycle support timelines

Post Syndicated from Kuldeep Yadav original https://aws.amazon.com/blogs/big-data/amazon-opensearch-service-extends-version-lifecycle-support-timelines/

In November 2024, we announced Standard and Extended Support dates for legacy Elasticsearch versions (1.5 through 7.8) and OpenSearch versions (1.0 through 1.2, and 2.3 through 2.9) running on Amazon OpenSearch Service. At that time, Extended Support for these versions was set to end on November 7, 2026 (except Elasticsearch 5.6, for which Extended Support ends on November 7, 2028), after which domains would no longer receive security fixes or operating system patches.

Since that announcement, many customers have upgraded to newer versions. However, some customers need more time to plan and complete their migrations. To provide this flexibility, we are continuing security and operating system patch coverage for these versions for an additional 12 months, through November 7, 2027, at an updated support rate.

Today, we’re announcing two updates: Extended Support extension for present versions and End of Standard Support and Extended Support for additional versions.

Extended Support extension for present versions

We are continuing security and operating system patch coverage for Elasticsearch versions 1.5 through 7.8, OpenSearch versions 1.0 through 1.2, and OpenSearch versions 2.3 through 2.9 for an additional 12 months. Coverage will now continue through November 7, 2027, giving customers additional time to plan and execute their migrations to the latest OpenSearch versions.

From November 7, 2026, the Extended Support surcharge for these versions will effectively double your instance pricing for the extension period. Storage costs are not affected. Elasticsearch 5.6, for which existing Extended Support rates end on November 7, 2028, will continue at the standard Extended Support cost of $0.0065 per Normalized Instance Hour (NIH). During this period, these versions will continue to receive critical security patches and operating system updates.

See the following table for the updated Extended Support end dates.

Software version End of Standard Support Original End of Extended Support date Updated End of Extended Support date
Elasticsearch versions 1.5 and 2.3 November 7, 2025 November 7, 2026 November 7, 2027
Elasticsearch versions 5.1 to 5.5 November 7, 2025 November 7, 2026 November 7, 2027
Elasticsearch version 5.6 November 7, 2025 November 7, 2028 No change
Elasticsearch versions 6.0 to 6.7 November 7, 2025 November 7, 2026 November 7, 2027
Elasticsearch versions 7.1 to 7.8 November 7, 2025 November 7, 2026 November 7, 2027
OpenSearch versions 1.0 to 1.2 November 7, 2025 November 7, 2026 November 7, 2027
OpenSearch versions 2.3 to 2.9 November 7, 2025 November 7, 2026 November 7, 2027

We recommend that you upgrade to the latest available OpenSearch version.

End of Standard Support and Extended Support for additional versions

Today we are announcing end of Standard and Extended Support dates for Elasticsearch versions 6.8, 7.9, and 7.10, OpenSearch version 1.3, and OpenSearch versions 2.11 to 2.19. For future updates on versions in Standard Support and Extended Support, follow supported versions.

For OpenSearch versions running on Amazon OpenSearch Service, we provide at least 12 months of Standard Support after the end-of-support date for the corresponding upstream open source OpenSearch version. Alternatively, we provide 12 months of Standard Support after the release of the next minor version on Amazon OpenSearch Service, whichever is longer. This aligns with the open source OpenSearch maintenance policy.

We categorize these versions into two groups:

  • The last versions of each major version family (ES 6.8, ES 7.10, OS 1.3, OS 2.19) will receive 3 years of Extended Support at the standard Extended Support charge of $0.0065 per NIH.
  • All other minor versions with clear upgrade paths within the same major family (ES 7.9, OS 2.11, OS 2.13, OS 2.15, OS 2.17) will receive 1 year of Extended Support at the same standard rate of $0.0065 per NIH.

After Extended Support ends for a version, domains running that version will not receive bug fixes or security updates. The following table shows the end of Standard Support and Extended Support dates for Elasticsearch and OpenSearch versions.

Elasticsearch versions

Software version End of Standard Support End of Extended Support
Elasticsearch version 6.8 November 7, 2027 November 7, 2030
Elasticsearch version 7.9 November 7, 2027 November 7, 2028
Elasticsearch version 7.10 November 7, 2027 November 7, 2030

OpenSearch versions

Software version End of Standard Support End of Extended Support
OpenSearch version 1.3 November 7, 2027 November 7, 2030
OpenSearch version 2.11 November 7, 2027 November 7, 2028
OpenSearch version 2.13 November 7, 2027 November 7, 2028
OpenSearch version 2.15 November 7, 2027 November 7, 2028
OpenSearch version 2.17 November 7, 2027 November 7, 2028
OpenSearch version 2.19 November 7, 2027 November 7, 2030
OpenSearch version 3.1 and above Not announced Not announced

Upgrading OpenSearch Service domains: We recommend that you upgrade your domains to the latest available OpenSearch version to derive maximum value out of Amazon OpenSearch Service. Minor version upgrades on OpenSearch don’t contain breaking changes. These version upgrades are typically non-disruptive. We recommend moving to the latest minor version. See Upgrading OpenSearch Service domains for detailed instructions. You can also use the Migration Assistant for Amazon OpenSearch Service for upgrading to newer versions.

New domain creation: New domain creation will be blocked after Extended Support ends for each version.

Calculating Extended Support charges

Amazon OpenSearch Service domains running versions under Extended Support will be charged a flat additional fee per NIH. NIH is computed as a factor of the instance size (for example, medium or large), and the number of instance hours.

Depending on which version you are on, the Extended Support charges are as follows:

  • Versions that have been on Extended Support (ES 1.5–7.8 (other than ES 5.6), OS 1.0–1.2, OS 2.3–2.9) from November 7, 2026: Extended Support cost will be equal to your instance price. This will be in addition to the standard instance pricing, effectively doubling the instance cost. Storage costs are not affected. Example (for versions that have been on Extended Support from November 7, 2026): If you are running an m7g.medium.search instance priced at $0.068/hr (on-demand) in US East (N. Virginia) for 24 hours, your standard instance cost is $1.632/day ($0.068×24). The Extended Support surcharge for the extension period will be equal to your instance cost ($1.632/day), doubling your instance pricing to ~$3.264/day. Storage costs remain unchanged.
  • New versions coming under Extended Support (ES 6.8, 7.9, 7.10, OS 1.3, OS 2.11–2.19): $0.0065 per NIH (standard Extended Support rate). See the pricing page for exact pricing by Region. Example (new versions — standard rate): If you are running an m7g.medium.search instance for 24 hours in the US East (N. Virginia) Region, priced at $0.068 per instance hour (on-demand), you will typically pay $1.632 ($0.068×24). If you are running a version that is in Extended Support, you will pay an additional $0.0065 per NIH. This is computed as $0.0065 × 24 (instance hours) × 2 (normalization factor for medium) = $0.312 for Extended Support for 24 hours. The total amount you will pay for 24 hours is $1.944 ($1.632 + $0.312, excluding storage cost).

The following table shows the normalization factor for various instance sizes in OpenSearch Service.

Instance size Normalization Factor
nano 0.25
micro 0.5
small 1
medium 2
large 4
xlarge 8
2xlarge 16
4xlarge 32
8xlarge 64
9xlarge 72
10xlarge 80
12xlarge 96
16xlarge 128
18xlarge 144
24xlarge 192
32xlarge 256

Summary

The latest OpenSearch versions include new features, performance and resiliency improvements, and security enhancements. With today’s announcement, we are:

  • Continuing security and operating system patch coverage for present versions through November 7, 2027, giving customers additional time to complete their upgrades at a new Extended Support rate.
  • Announcing Standard and Extended Support timelines for the next set of versions (ES 6.8, 7.9, 7.10, OS 1.3, 2.11–2.19) with predictable cost visibility.

We recommend that you upgrade to the latest OpenSearch versions to get the most benefit out of OpenSearch Service. For any questions on Standard and Extended Support options, see the FAQs. For further questions, contact AWS Support.


About the authors

Kuldeep Yadav

Kuldeep Yadav

Kuldeep is a Principal Technical Program Manager at AWS. He’s passionate about driving innovation and complex problem solving. He works closely with teams and customers to ensure operational excellence and achieve more with less.

Arvind Mahesh

Arvind Mahesh

Arvind is a Senior Manager-Product at AWS (Amazon OpenSearch Service). With close to two decades of technology experience, he brings deep expertise across Analytics, Search, Cloud, Network Security, and Telecom.

Jon Handler

Jon Handler

Jon is a Senior Principal Solutions Architect at AWS. He works closely with OpenSearch and Amazon OpenSearch Service, guiding a broad range of customers looking to move their search and log analytics workloads to the AWS Cloud.

How and Why Netflix Built a Real-Time Distributed Graph: Part 3 — Querying the graph with gRPC…

Post Syndicated from Netflix Technology Blog original https://netflixtechblog.com/how-and-why-netflix-built-a-real-time-distributed-graph-part-3-querying-the-graph-with-grpc-0f3468349607

How and Why Netflix Built a Real-Time Distributed Graph: Part 3 — Querying the graph with gRPC execution API

Authors: Nilesh Mishra and Ajit Koti

This is the third entry of a multi-part blog series describing how we built a Real-Time Distributed Graph (RDG). In Part 1, we discussed the motivation for creating the RDG and the architecture of the data processing pipeline that populates it. In Part 2, we discussed how we designed the storage layer to handle billions of nodes and edges while maintaining single-digit-millisecond latency. In Part 3, we will explore how we designed a fast, flexible serving layer to efficiently query the graph.

Introduction

In Part 1 of this series, we described why Netflix needed a Real-Time Distributed Graph (RDG) and how we used Apache Flink to build an ingestion and processing pipeline that turns streaming events into graph primitives. In Part 2, we explored how we designed a storage layer capable of handling billions of nodes and edges while still delivering single-digit-millisecond latency.

In this post, we focus on the next challenge: querying the graph efficiently to power real-time insights for our internal partners. All of the work on ingestion and storage only matters if we can actually ask complex questions and get answers back quickly. As we optimized for lower latency, we found that the serving layer posed its own set of challenges, distinct from those of ingestion and storage. How do we turn a constantly evolving, billion-edge graph into sub-100ms responses across a wide variety of workloads? This is the problem we tackle in this post.

The Real World Needs

As we integrated the RDG into Netflix’s ecosystem, we realized that “querying the graph” is not a one-size-fits-all operation. We needed to handle a wide range of access patterns: from high-volume security lookups to deep, exploratory personalization traces.

Let’s revisit our example from Part 1 and expand on it slightly. In the earlier posts, we focused on accounts, devices and content. In practice, the graph is richer: each account has multiple profiles.

A member journey often looks like this:

  1. Alex logs in to their Netflix profile on a smartphone and starts watching Stranger Things.
  2. They later switch to a smart TV in the living room to continue the episode.
  3. The next morning, they use a tablet to play the game Stranger Things: 1984.

In the RDG, this journey creates the following graph structure:

Graph queries vary along two axes: how wide they fan out at each hop, and how deep they chain across hops. To see this range, let’s look at two scenarios from opposite ends:

1. Shallow and wide: “Which devices has this account used?”

Consider a “shallow, wide” query: “Which devices has this account used to stream in the last 30 days?”

Using the graph structure above, this translates to:

  • Starting Point: A specific Account Node.
  • Hop 1 Edge Traversed: The streamed_from edge.
  • Hop 1 Destination: Device Nodes.

While this is only a “single hop,” it presents a significant scaling challenge. For a highly active account, the fan-out can be massive. The query layer must fetch hundreds of streamed_from edges, apply temporal filters on each edge’s last_watch_timestamp property to capture only those within the last 30 days, and aggregate the results, all while maintaining sub-100ms latency.

2. Deep & Narrow: What has this profile watched?

Consider a scenario where personalization teams need to understand a member’s viewing journey. They might ask: “For Account X, show me the Stranger Things viewing history across all profiles: which profiles watched it, what they watched, and when”.

This path unfolds as follows:

  • Starting Point: A specific Account Node.
  • Hop 1 Edge Traversed: has_profile
  • Hop 1 Destination: Profile Nodes
  • Hop 2 Edge Traversed: started_watching (filtered for title_name = “Stranger Things”)
  • Hop 2 Destination: Content Nodes

The core challenge in this scenario is sequential dependency: we cannot fetch a profile’s viewing history until Hop 1 has identified which profiles exist. In a distributed environment, the client has to wait for Hop 1 to finish before sending Hop 2. If each hop takes 10ms of network time, that’s 20ms of overhead before we’ve processed a single byte. To hit our sub-100ms goal, we needed a way to package this multi-step logic into a single request.

This example is a 2-hop traversal, but queries can chain 3–4 hops across different entity types, and the latency penalty of sequential execution only grows with depth.

Balancing Depth and Breadth

These two scenarios pull the system in opposite directions. Shallow-wide queries stress I/O throughput: can we handle massive fan-out without slowing down? Deep-narrow queries stress execution efficiency: can we chain multiple hops without the network overhead adding up? Supporting both on the same system is what shaped the design that follows.

Design Constraints and Key Choices

The two scenarios above sit at opposite ends of the spectrum, but they are not unusual. In practice, the RDG serves tens of thousands of queries per second, each potentially different, all needing sub-100ms responses while the underlying graph continues to grow. Scale, latency, query diversity, and the need for extensibility pulled the design in different directions at once, and every choice came with a trade-off we had to live with.

Why breadth-first, not depth-first? The most intuitive way to traverse a graph is depth-first: pick a path, follow it to the end, backtrack, try another path. But in a distributed system where every hop is a network call, depth-first can lead to high latency. If Account X has 5 profiles and each profile has watched hundreds of titles, depth-first would trace all of one profile’s watched titles before moving to the next, missing the opportunity to batch lookups across profiles. Breadth-first flips this by working one level at a time across all nodes, rather than one path at a time through each node. We fetch all profiles for the account at once, then fetch the started_watching edges for all profiles, and finally fetch content details for all matching titles. Three rounds of parallel calls instead of sequential chains. With breadth-first, there is a clear trade-off in memory, because we hold each level of the graph in memory at once, so the cost scales with how wide a level fans out rather than how deep the query goes. We keep this comfortable by bounding each hop with the per-edge-type limits described in Step 5 below, so even a high fan-out level stays a manageable frontier. We’ll walk through how this works, level by level, in Step 3 below.

Why async-first, not thread-per-request? Latency in the RDG is dominated by I/O, reading from the storage layer, calling enrichment services, and waiting on caches. A traditional thread-per-request model would pin a thread to each in-flight query, and most of the time, the thread would be idle, waiting for a network response. With thousands of concurrent queries, we’d need thousands of threads, most of which would be doing nothing. Instead, we decided to build the entire execution pipeline around asynchronous composition. A small set of dedicated thread pools (16–24 threads total) handles thousands of concurrent requests because no thread ever blocks on I/O. While a storage call is in flight, the thread continues with other work and picks up the result when it arrives. This is the foundational design decision on which everything else rests. We’ll see this in action in Step 4 below, where we cover parallel execution.

Why cache selectively, not everything? Not all data in the graph changes at the same rate. Some properties, such as account plan type and content metadata, are relatively stable: they change on the order of hours or days. Edges like who watched what and when change constantly. For stable data that many queries touch, we use a distributed cache (EVCache) with TTLs tuned to data volatility. Getting the caching strategy right took iteration. We started by caching aggressively and measured the impact: tracking hit rates, monitoring stale-data incidents, and adjusting TTLs based on how quickly different node types actually changed in production. The result: 70–80% hit rates on node lookups, achieved by narrowing the cache to nodes that are both frequently accessed and slow to change, while skipping data that would expire before the TTL ran out. Step 6 below covers how this works in practice.

Why opt-in enrichments, not automatic? Clients know what they need. A query checking account relationships doesn’t care about title artwork; a personalization service building a viewing timeline does. Rather than fetching metadata from external services by default and penalizing every query, we make enrichments opt-in: clients specify exactly which external data they want per request. Also, enrichment is fail-open: if a service is slow or unavailable, we return the graph data without it.

Why eventual consistency, not strong? Most of our queries ask “What has this member done recently?”, not “What happened in the last millisecond?” By defaulting to eventual consistency, we read from the nearest replica and avoid coordination overhead. While the RDG is used to power in-the-moment experiences, it is not set up as the source of truth for the data it holds.

Architecture Overview

The above choices lead to the following three-layer architecture:

The Graph Query Service is the entry point. It accepts gRPC requests, validates the traversal specification, and hands it to the query execution engine. The execution engine orchestrates breadth-first traversal: expanding one level at a time, applying filters and limits at each hop, and composing all I/O asynchronously.

The Storage Abstraction Layer sits between the execution engine and the underlying KVDAL storage. It provides a clean interface for node lookups and edge retrieval, handles streaming for large adjacency lists, and manages node caching (EVCache).

The Enrichment Layer fetches additional metadata from external Netflix services on demand. It batches requests, runs them in parallel with graph data assembly, and degrades gracefully when an enrichment source is unavailable.

When a client sends a query, the request flows through these layers in sequence: the Query Service parses the request into an execution plan, the execution engine walks the graph level by level through the Storage Abstraction Layer, and if enrichments are requested, the Enrichment Layer fetches and merges external data before the response is serialized back to the client.

Now, with that mental model in place, let’s follow a query through this system and see how these choices play out in practice.

Executing Queries Efficiently: Following a Query’s Journey

To see how the RDG query layer works in practice, let’s follow a single query end-to-end and focus on one question: how do we make every step fast?

We’ll reuse the deep-narrow example from above:

For Account X, show me the Stranger Things viewing history across all profiles: which profiles watched it, what they watched, and when.

In graph terms, this becomes a 2‑hop traversal:

  1. Account X → has_profile → Profiles
  2. Profiles → started_watching → Content (filtered for “Stranger Things”)

We’ll walk through how this query moves through the layers we described above:

  1. Reading and interpreting the request
  2. Reading from storage efficiently
  3. Executing traversal with breadth‑first levels
  4. Running many operations in parallel, but safely
  5. Filtering smartly to keep only what matters
  6. Making repeat queries faster with caching

By the end, we’ll see how a 2-hop query like our Stranger Things example, with streaming, filtering, and parallel execution, can complete in under 100ms.

Step 1: Reading the Request: Deciding What the Query Really Wants

Every query starts as a gRPC request. Before we touch storage or walk a single edge, the engine needs to understand what the caller actually wants.

For our running example below:

For Account X, show me the Stranger Things viewing history across all profiles

The engine creates a traversal plan with a set of levers: how many hops, how many edges per hop, how much history to consider, and whether to favor recent activity.

We resolve these upfront by merging a hierarchy of filters and limits, from application-level defaults down to per-edge-type overrides, into a concrete execution plan. By the time we read from storage, every hop has clear rules. We’ll see how this hierarchy works in detail in Step 5, but the key insight is simple: interpreting the request up front prevents over-fetching from the downstream storage layer.

Step 2: Reading from Storage: Direct Lookups and Streaming Fan‑Out

Once we’ve parsed the request and decided what the query should do, the next step is to actually touch the graph. For our running example:

For Account X, show me the Stranger Things viewing history across all profiles…

The first concrete question the engine has to answer is very simple:

Which profiles does Account X have?

Under the covers, that really means: how do we find all relevant edges for Account X without scanning the entire graph every time?

Finding Edges Fast with Adjacency Lists

If we stored every edge in one massive table, the naive approach would be to scan for rows where source = Account X. Even with indexing, doing that across billions of edges for every request would be slow.

Instead, we organize edges as adjacency lists. For each node, we keep a compact list of “who it’s connected to” by edge type. For Account X, a simplified view might look like:

Account_X: has_profile → [Profile_Alex, Profile_Kids, …,]

Now “get all profiles for Account X” is no longer a global search; it’s a direct lookup into Account X’s stored adjacency. The storage layer can usually pull that list back in a few milliseconds because it’s reading a small, well‑indexed slice of data instead of hunting through everything.

For our query, the first hop is quick: Account X has just two profiles. The engine fetches those edges with has_profile and moves on. For more information on Storage, refer to our previous post.

When One Node Has A Lot of Neighbors

The first hop was small, but the second is where things get interesting. Each profile can have a large number of started_watchingedges. Loading the entire adjacency list at once would spike latency and memory usage.

To avoid this, we treat adjacency lists as streams rather than blobs.

When the engine requests Profile_Alex’s started_watching edges, the storage layer streams them in batches of 100. As each batch arrives, we apply filters (e.g., “last 30 days”) and decide whether to continue.

If we’ve collected enough edges to satisfy the query’s limits ( max_edge_cnt, lookback window, etc.), we stop reading. Otherwise, we pull the next batch.

In our Stranger Things example:

  • Storage streams the started_watching adjacency for Profile_Alex.
  • There are about 500 edges total: months of viewing history
  • As each batch arrives, we filter for Stranger Things and drop anything older than 30 days.
  • After a few batches, we’ve found what we need: a handful of Stranger Things sessions.
  • We never materialize more data than needed. Filtering happens at the source.

Why This Matters Later

These two choices, the adjacency‑list lookups and streaming fan‑out, enable everything that follows:

  • Small fan‑outs (like Account → Profiles) yield predictable, low‑millisecond lookups.
  • Large fan‑outs (like Profile → Content) stay efficient by reading only what’s needed.
  • Traversal logic treats “neighbors of this node” as a cheap, bounded operation.

By Step 3, we’re working with concise frontiers like “Profile_Alex and Profile_Kids,” ready for the next hop into their viewing histories.

Step 3: Traversal Execution: Walking the Graph Level by Level

We’ve completed the first hop. From Account X, we pulled the has_profile edges and found two profiles: Profile_Alex and Profile_Kids.

But we’re not done. The query was:

For Account X, show me the Stranger Things viewing history across all profiles: which profiles watched it, what they watched, and when

So we still need to fetch each profile’s history and filter it down to Stranger Things sessions. As we covered in our design choices, we use breadth-first traversal: expanding all nodes at the current level in parallel before moving to the next.

Querying, Level by Level

Let’s walk through the Stranger Things query level by level.

Level 1: Account → Profiles

Starting at Account X, the engine pulls has_profile edges, discovering two profiles:

  • Profile_Alex, Profile_Kids

These become the frontier for Level 2, a single small lookup that takes a few milliseconds.

Level 2: Profiles → Content (Stranger Things)

From those two profiles, we fetch started_watchingedges and filter for Stranger Things. Instead of exhausting Profile_Alex’s entire viewing history before touching Profile_Kids, we treat this as one logical step:

  • For each profile, fetch started_watching edges in parallel.
  • Filter for title_name = “Stranger Things” as edges stream in.
  • Each profile might have hundreds of content edges, but filtering at the source keeps the result set small.

We discover that Profile_Alex watched Season 1 and Season 2, while Profile_Kids watched Season 4. Level 2 turns “2 profiles” into “a handful of Stranger Things sessions” in roughly one storage round trip.

The traversal completes: two levels, two frontiers.

Why This Matters for Latency

We parallelize within each phase, then regroup. This provides:

  1. Predictable resource usage: known requests per level
  2. Maximum parallelism: all frontier nodes processed together
  3. Far fewer round trips: one per level, not per path

For a 2-hop query: two rounds of parallel lookups instead of hundreds of sequential ones. That’s why our Stranger Things query completes in under 100ms.

Step 4: Parallel Execution: Doing Many Things at Once, Safely

Breadth-first traversal enables parallel work at each level, which is the key to low latency.

At Level 2 of our Stranger Things query, we fetch started_watching edges for each profile. With two profiles, this is trivial, but in production queries fan out across many profiles, each with hundreds of edges to stream and filter. So do we process them sequentially or in parallel? Sequential means waiting for each profile before starting the next, and the delays stack up. Parallel finishes in the time of the single slowest profile, but hundreds of queries doing this at once could overwhelm storage with unbounded concurrency.

The goal: parallel speed without unbounded chaos.

A Kitchen, Not a Single Queue

We structured the query engine like a professional kitchen, with specialized stations for appetizers, mains, and desserts, each with its own capacity. If one station is slammed, the others keep flowing. In practice, that means dedicated thread pools for different work types: fetching nodes, reading adjacency lists, and performing enrichments. When the Stranger Things query reaches Level 2, calls route to the adjacency-list pool, where 8 workers stream and filter each profile’s edges in parallel.

Knowing When to Back Off

Thread pools give us local control, but we also need a global view of total capacity, so we use adaptive concurrency limiting. When things are healthy, we raise the limit gradually (100 in-flight, then 101, 102, and so on); when timeouts or errors spike, we back off by a larger step (say, 100 down to 70). Combined with per-pool limits, the engine constantly tunes parallelism, fanning out within each level while staying inside safe storage and network limits.

Fetching Extra Metadata Along the Way

If the client opted into enrichments (say, maturity ratings for the matched content), the Enrichment Layer fetches them in parallel on its own thread pool and merges them into the response. Enrichment is fail-open: a slow or unavailable source never blocks the query, and we just return the graph data without it.

​​Step 5: Smart Filtering: Keeping Only What Matters

We’ve traversed from Account X to profiles, then to their viewing histories. But raw edges aren’t what our partners need. They care about recent, relevant activity, not every started_watching edge accumulated over the years. This is where filtering decides which parts of the story make the final cut.

From “All Activity” to “The Last 30 Days”

Go back to the original question:

For Account X, show me the Stranger Things viewing history across all profiles: which profiles watched it, what they watched, and when.

The phrase “viewing history” is deceptively simple. Under the hood, it means we need to:

  • Ignore older viewing activity, even if it exists in the graph
  • Avoid pulling more edges than we actually need
  • Let different teams choose their version of “recent enough.”

We handle this with a filtering hierarchy. The system starts with conservative defaults (e.g., 100-day lookback, 300 edges per hop), and requests can override them globally, per-hop, or down to specific edge types. In our query, the 100-day default applies broadly, but the caller sets 30 days for started_watching edges, and the narrower rule wins. Older sessions are discarded. The same engine can just as easily provide a tight recent window on one edge type and full history on another, all in a single query.

Choosing Which Edges to Keep: LATEST vs ANY

Sometimes there are still more edges than we want to return after time filtering. If Profile_Alex watched the same episode several times last month, pausing and resuming, we don’t want to send all those edges back. So we offer two selection modes.

LATEST sorts edges by timestamp and keeps the newest ones up to the limit, ideal for “what has this profile watched recently?” where teams want the current state, not every play event. ANY grabs whichever edges it encounters first, no sorting, which is faster and fine for “has this profile ever watched Stranger Things?” where timing doesn’t matter. Teams default to LATEST and switch specific edge types to ANY when “any proof” is enough.

Bringing It Back to Our Story

So what happens for our running query?

We start with all the started_watching edges for each profile. The time filter narrows this to 30 days. Edge-count limits prevent response flooding. LATEST mode selects the most recent viewing session per title. The result: a concise answer distilled from a verbose history:

  • Profiles that watched Stranger Things in the last month.
  • Which seasons and episodes they watched.
  • The most recent session for each, tying it all together.

This filtering turns raw history into a focused answer.

Step 6: Making It Even Faster: Caching the Things We Keep Seeing

By now, we’ve walked the full path of our query: we’ve traversed from account to profiles, filtered viewing history by time, and focused on Stranger Things sessions.

Despite our optimizations, each storage call still costs a network round-trip. When the same nodes appear across thousands of queries per minute, those redundant calls add up: both in infrastructure cost and in tail latency at scale.

The key question: what can we avoid repeating?

The Things That Don’t Change Every Second

Look back at the entities in our Stranger Things journey:

  • The Account node (plan type, region, etc.)
  • The Profile nodes (“Alex”, “Kids”, whether it’s a kids profile)
  • The Content nodes (Stranger Things seasons and episodes)

These rarely change. Profiles don’t flip between “kids” and “non-kids” every minute. Title metadata is stable.

To improve efficiency, we keep a distributed cache of hot nodes (accounts, profiles, content) that are likely to reappear. When the same entity appears again, we answer “What is this node?” from memory, skipping storage.

Result: for high-traffic entities, we eliminate storage calls and noticeably reduce infrastructure cost and tail latency at scale.

A Quick Replay of Our Query With Caching Turned On

The first time the Stranger Things query runs for Account X, the cache is cold, so we pay the full cost: we fetch the account and its profiles, then the started_watching edges and matching content nodes, caching each node as we go. Minutes later, a different query arrives:

Show me everything Account X’s profiles have watched in the last 7 days, and flag anything rated TV-MA on the kids profile.

This time, many of those nodes are already in the distributed cache. Storage still handles the adjacency lists and edges, but node lookups are lighter and latency drops. At scale, that reuse gives us comfortable headroom for traffic spikes.

Not Everything Deserves a Spot in Cache

We can’t cache everything. The RDG prunes old activity after a set retention window, so caching a node that’s about to be deleted is wasteful.

To avoid polluting the cache, we consider:

  1. The node’s last activity timestamp
  2. The graph’s retention period (e.g., 100 days)
  3. The cache TTL (e.g., 30 days)

If a node was last active 99 days ago, it expires from the graph in a day, so a 30-day TTL makes no sense, and we skip it. We reserve cache space for active nodes like Account X. This “smart TTL” policy keeps the cache focused on live stories rather than archival ones, so repeat queries for the same part of the graph return faster.

Caching is integrated into the journey, not an afterthought. The engine reuses knowledge from previous queries, so repeated traversals over the same part of the graph keep getting cheaper

The Payoff

The serving layer sits in front of 8 billion nodes and 150 billion edges, serving mixed workloads, all of which need to feel interactive. Single-hop queries return at a P50 of 15–30ms with P99 under 100ms. Even 3-hop traversals, the kind that chain across accounts, profiles, and content, come back at P99 between 100–150ms. Breadth-first execution and parallelism within each level keep these numbers stable even as fan-out grows.

The async-first design is what enables the throughput. Thousands of concurrent requests flow through just 16–24 threads spread across dedicated pools because no thread ever blocks on I/O. When load spikes, our concurrency limiter lets work queue briefly: slowly increasing capacity when things are healthy, backing off aggressively when they’re not

Caching has the most visible impact on day-to-day efficiency. Popular entities like accounts, profiles, and content achieve 70–80% cache hit rates, resulting in roughly 3–4x fewer storage calls on common query paths. Smart TTLs keep the cache focused on active data, avoiding wasted memory on nodes that are near the end of their graph retention window.

These properties, together, make multi-hop graph queries over billions of entities feel, at query time, much closer to in-memory lookups than to remote calls.

What We Learned Along the Way

The biggest surprise wasn’t any single optimization: it was how much async composition changed the economics of our system. We expected it to help latency; we didn’t expect it to slash infrastructure cost. A serving layer that would have needed hundreds of threads per instance runs comfortably on 16–24, because no thread ever blocks on I/O. The tradeoff is debuggability: async stack traces are hard to read, and exceptions can get lost in future chains. We compensated with per-stage metrics, measuring each request at validation, storage, enrichment, and end-to-end, so when something is slow, we know exactly which stage to blame.

Caching took longer to get right than expected. Our first instinct was to cache everything in EVCache and let TTLs handle freshness, but that wastes memory on nodes about to expire from the graph anyway. The breakthrough was matching TTLs to data volatility: stable node properties get long TTLs, while nodes near the end of their retention window aren’t cached at all. The 70–80% hit rate we see today came from being selective, not aggressive.

The filtering hierarchy was born out of frustration. Early on, every new use case meant a code change: one team wanted a 7-day lookback, another 90 days, a third different limits at different depths. Instead of bespoke logic per team, we built a layered override system: application defaults, global overrides, per-depth limits, and per-edge-type limits. It took real effort, but it eliminated an entire class of feature requests and teams now tune their own queries without touching our code.

Closing: Principles for Distributed Systems

The lessons above are specific to the RDG, but the underlying principles apply to any distributed system built around I/O-heavy, fan-out workloads.

  • Think in terms of frontiers, not features. Design your APIs so callers describe what frontier to explore, then let the system decide how to walk it efficiently.
  • Filter early, not late. Every byte you fetch but don’t need is wasted I/O. Push filters and limits as close to the storage layer as possible: discard irrelevant data at each stage rather than fetching everything and trimming at the end.
  • Parallelize deliberately, not by default. Unbounded concurrency feels fast until it overwhelms the systems you depend on. Set explicit limits, monitor them, and adjust dynamically: treat concurrency as a dial, not a switch.
  • Treat caching as a first‑class design choice, not an afterthought. Decide what is worth remembering, for how long, and what should be allowed to fade out of memory. Match TTLs to data volatility, and don’t cache what’s about to expire.

Thanks for reading Part 3 of the RDG blog series. For us, getting these details right is what turns a constantly changing, billion-edge graph into something that, at query time, feels like a responsive, in-memory data structure.


How and Why Netflix Built a Real-Time Distributed Graph: Part 3 — Querying the graph with gRPC… was originally published in Netflix TechBlog on Medium, where people are continuing the conversation by highlighting and responding to this story.

[$] Changes in shadow-utils password-expiration features

Post Syndicated from jzb original https://lwn.net/Articles/1086949/

The shadow-utils
project provides the tools that handle /etc/shadow,
/etc/passwd, and other related databases; in
general, manages users and groups on many Linux systems. While most
software releases are notable for what is added, the recent shadow-utils 4.20.0
release is most noteworthy for what has been removed. Specifically,
several utilities and functionality related to periodic password
expiry, which were deprecated in the December 2025 4.19.0
release, have been removed as planned. It is still possible to manage
some aspects of password aging with shadow-utils, but organizations
that depend on such features should start planning for their complete
removal within a few years.

The collective thoughts of the interwebz