Introducing EC2 AMI Tag Sharing: Share EC2 tags across AWS accounts

Post Syndicated from Gena Gizzi original https://aws.amazon.com/blogs/compute/introducing-ec2-ami-tag-sharing-share-ec2-tags-across-aws-accounts/

Tags are an important and versatile tool on AWS. FinOps teams track costs with them, security teams enforce compliance, and governance teams gate access through attribute-based access control (ABAC).

Tags become especially important as infrastructure management grows more complex across multiple accounts in an organization. Because compute is a common shared resource, many organizations tag their AMIs to track approval status, OS versions, and team. However, when you share an AMI with another AWS account, the tags you attached to it stay behind. The other account sees the image, but tag metadata used to describe it could not be shared.

Until today, the only way to solve this was to build a custom workflow to copy tags into every account the AMI is shared with. Typically created with services such as AWS Simple Notification Service (Amazon SNS) and AWS Lambda, this caused operational burden.

Today, we’re introducing EC2 Tag Sharing, a new feature that AMI owners can use to share tags with any account that has access to the image. Whether an AMI is shared with an account, across an organization, or made public, a shared tag is automatically shared alongside the image by adding the ec2:SharedTag/ prefix to the tag key. Updates also propagate automatically without any need for custom replication workflows. In this post, we walk through how tag sharing works, demonstrate a common use case, and cover considerations and best practices.

How it works

Any tag whose key starts with ec2:SharedTag/ is visible to all AWS accounts the AMI is shared with, and also when the AMI is shared publicly. Tags without the prefix remain private to the account that created them, exactly as they do today.

Only the resource owner can create, modify, or delete tags with the ec2:SharedTag/ prefix. Accounts you share the AMI with can view these tags but cannot change them. They can continue to add their own private tags to a shared AMI.

Shared tags count against the resource owner’s 50-tag-per-resource quota for both shared and private tags. They do not count against the quota of accounts you share the image with. When you copy an AMI that has shared tags, AWS Elastic Compute Cloud (AWS EC2) copies the shared tags to the new image when the --copy-image-tags parameter is included. At launch, tag sharing is supported specifically for AMI resources.

Creating a shared tag (CLI)

To share a tag, create it with the ec2:SharedTag/ prefix on the key name. For example, to share a tag indicating the AMI’s approval status:

aws ec2 create-tags \
    --resources ami-0abcdef1234567890 \
    --tags Key=ec2:SharedTag/status,Value=approved

You can add multiple shared tags alongside your private tags. In the following example, status and os-version are shared with other accounts. The team tag remains private to the owner.

aws ec2 create-tags \
    --resources ami-0abcdef1234567890 \
    --tags \
    Key=ec2:SharedTag/status,Value=approved \
    Key=ec2:SharedTag/os-version,Value=al2023-2026.09 \
    Key=team,Value=platform-engineering

You can also apply shared tags when you create an AMI from the start, using the --tag-specifications parameter. For example:

aws ec2 create-image \
    --instance-id i-1234567890abcdef0 \
    --name "My server" \
    --tag-specifications "ResourceType=image,Tags=[{Key=ec2:SharedTag/publicTag,Value=text}]"

Listing all resources with shared tags (CLI)

To see which of your AMIs have shared tags, you can run the following CLI command:

aws ec2 describe-tags \
    --filters "Name=key,Values=ec2:SharedTag/*"

And get an output such as:

{
    "Tags": [
        {
            "Key": "ec2:SharedTag/status",
            "ResourceId": "ami-0abcdef1234567890",
            "ResourceType": "image",
            "Value": "approved"
        },
        {
            "Key": "ec2:SharedTag/os-version",
            "ResourceId": "ami-0abcdef1234567890",
            "ResourceType": "image",
            "Value": "al2023-2026.09"
        }
    ]
}

You can also receive a similar output using the describe-images command as shown in the following:

aws ec2 describe-images --filters "Name=tag-key,Values=ec2:SharedTag/*"

Stopping tag sharing

To stop sharing a specific tag, delete the tag with the ec2:SharedTag/ prefix:

# Remove the shared tag
aws ec2 delete-tags \
    --resources ami-0abcdef1234567890 \
    --tags Key=ec2:SharedTag/status

If you want to keep the tag private, recreate it without the prefix:

# Optionally, recreate as a private tag
aws ec2 create-tags \
    --resources ami-0abcdef1234567890 \
    --tags Key=status,Value=approved

Viewing shared tags as a recipient

When an AMI is shared with your account, you can see the shared tags alongside any tags you’ve added yourself. Figure 1 shows an example of an AMI shared from the original account with both private and shared tags. Figure 2 shows that AMI in the recipient account. The ec2:SharedTag/ prefix in the key name makes it clear which tags came from the owner.

Amazon EC2 console in the owner account showing an AMI with both private tags and ec2:SharedTag/ prefixed shared tags

Figure 1: AMI tags in the originating account, showing both private and shared tags

Amazon EC2 console in the recipient account showing only the shared tags on the same AMI, with the private tags absent

Figure 2: The same AMI in the recipient account, where only the shared tags are visible

As you can see in the preceding figures, the shared tags are visible across both accounts. However, the private tags created in the original account are only visible in the original account.

A common use case: Golden AMIs

Many organizations maintain a central build factory account that produces hardened, security-approved AMIs. These golden images can be shared with hundreds or thousands of workload accounts across an organization. Tags like status=approved or patch-date=2026-09-18 are commonly used by recipient accounts to enforce policies that only allow launches from approved images.

Before tag sharing, the build factory had to replicate tags to every account the AMI is shared with after each AMI publish. A typical workflow looked like this:

  1. The build factory creates a new AMI and tags it.
  2. An SNS notification triggers a Lambda function in each account.
  3. Each Lambda function reads the tag values from the AWS SNS message and calls CreateTags on the shared AMI in its own account.

Because this runs for every AMI, in every Region, across every account, the number of CreateTags API calls can grow rapidly in concentrated bursts. At that scale, teams risk throttling, failed Lambda invocations, and tag drift when calls silently fail.

With tag sharing, the build factory tags the AMI with the ec2:SharedTag/ prefix. Every account the AMI is shared with sees these tags, without the need for additional pipelines to replicate the tags. When the build factory updates a tag (for example, marking an older image as deprecated), the change is visible across all accounts without any additional API calls.

Considerations

Shared tags are visible to every account with access. We recommend that you do not include personally identifying, confidential, or sensitive information in these tags.

If consuming accounts have ABAC policies that evaluate tags, those policies will also apply to the shared AMI. For example, if a recipient account has a policy that allows ec2:RunInstances only when the AMI has the tag status=approved, and you share that tag by using ec2:SharedTag/status=approved, the policy will deny it.

When you adopt the shared tags model, any action that depends on those tags is controlled by the AMI owner, not by the recipient account. This includes tag-gated instance launches, AWS Identity and Access Management (IAM) and ABAC policy evaluations, and any downstream automation that reads tag values. Because only the owner can create, modify, or delete shared tags, the recipient account cannot change the values its own policies depend on. Before you rely on a shared tag in a policy or workflow, make sure you trust the owner account to set and maintain that value.

As a best practice, apply account-level governance so that you only consume AMIs from providers you trust. With Allowed AMIs, you can define criteria for which AMIs are allowed in your account, for example a list of trusted AMI provider account IDs. Only AMIs that meet the criteria are discoverable and available to launch in your account. To preview the impact before enforcing, start in audit mode.

Conclusion

EC2 Tag Sharing removes the need to build and maintain cross-account tag replication workflows. For organizations that share AMIs, this means less undifferentiated heavy lifting. To get started, add a tag with the ec2:SharedTag/ prefix to a shared AMI in the AWS EC2 console. To learn more, see Tag your AWS EC2 resources in the Amazon EC2 User Guide.

Restart EC2 and on-premises fleets faster with AWS CodeDeploy RESTART deployment mode

Post Syndicated from Iskandar Anvarov original https://aws.amazon.com/blogs/devops/restart-ec2-and-on-premises-fleets-faster-with-aws-codedeploy-restart-deployment-mode/

Operators often restart fleets when they need to pick up runtime configuration changes, recover unhealthy processes, or return hosts to a known state. Before RESTART, AWS CodeDeploy customers either redeployed the current revision, repeating completed work, or ran custom scripts outside of CodeDeploy production safeguards. Now there’s a purpose-built option: RESTART deployment mode. It keeps the operation inside CodeDeploy, so you get the same batch sizing, health checks, alarm monitoring, and rollback behavior as a standard deployment.

This new option reapplies the last successful revision to Amazon Elastic Compute Cloud (Amazon EC2) and on-premises in-place deployments. It retains your deployment configurations, lifecycle hooks, Amazon CloudWatch alarm monitoring, rollback settings, and deployment history.

In testing, a fleet restart completed up to 6.1x faster than a standard deployment, turning a multi-minute rollout into tens of seconds.

In this post, we explain how RESTART works, walk through starting a restart deployment from the command line and the CodeDeploy console, and review the safety controls that carry over from a standard deployment.

Why use a CodeDeploy restart?

A restart changes production even when the application revision stays the same. Hosts stop and start, failures reduce fleet capacity, and external configuration errors spread across restarted hosts. Fleet scripts must recreate batch sizing, minimum healthy capacity, validation, alarm monitoring, and audit history. One customer experienced this firsthand. Their restart script restarted hosts in batches with no health checks in between. When a bad configuration left the first batch unable to start, the script never noticed and continued. What should have been a routine restart took down their fleet.

This new deployment mode keeps the operation in CodeDeploy. You configure how the deployment runs with the familiar CreateDeployment API. CodeDeploy still runs your lifecycle hooks, validates the result, and stops after a health or alarm breach. A failed restart stays within the active batch instead of continuing through the fleet.

RESTART also works well as a building block for automated remediation. Point a CloudWatch alarm on a memory or resource-utilization metric at an Amazon EventBridge rule, and have that rule invoke an AWS Lambda function that calls CreateDeployment with deploymentMode: RESTART. Long-running or stateful workloads benefit from periodic recycling. Examples include self-managed Kafka brokers and JVM services with memory growth. This turns that recycling into a managed, self-healing loop instead of a cron job or a manual restart. The following diagram shows the flow:

CloudWatch alarm triggering an Amazon EventBridge rule that invokes a Lambda function to call CreateDeployment in RESTART mode

Figure 1: Automated remediation loop using an Amazon CloudWatch alarm, an Amazon EventBridge rule, and an AWS Lambda function to trigger a RESTART deployment

How RESTART works

Set deploymentMode to RESTART on CreateDeployment and provide no revision. CodeDeploy resolves the deployment group’s last successful revision and creates a deployment record.

Each selected host runs:

ApplicationStop → DownloadBundle → BeforeInstall → Install → AfterInstall → ApplicationStart → ValidateService

CodeDeploy agents that support local reuse (from version 2.1.0 onward) use the previous deployment’s archive. DownloadBundle remains in the signed workflow. If the archive is unavailable or invalid, the agent downloads the same pinned revision, covering added or replaced hosts.

Install reapplies the revision and corrects drift in managed files. BeforeInstall, AfterInstall, and ValidateService run as defined in the AppSpec file. Local reuse removes the network transfer, and end-to-end savings vary with revision size, agent version, hooks, fleet size, and deployment configuration.

Performance in feature testing

Feature tests peaked at 6.14x. Tests used m5.large instances, a 1 GB revision, and 10, 50, and 100-host fleets with and without an Application Load Balancer (ALB). Each scenario ran seven times. We discarded the minimum and maximum and report the median of five. Results combine local reuse with omitted traffic control and are not predictive of results for other applications.

At 75 percent minimum healthy, four or five waves saved ALB-backed fleets 397.1 to 424.8 seconds. The 50-host fleet reached 6.14x.

Host Count ALB in Front? Rollout Waves Standard Deployment Time Restart Deployment Time Speedup Time Saved
10 Yes 5 526.2s 129.1s 4.08x 397.1s
50 Yes 5 507.4s 82.6s 6.14x 424.8s
100 Yes 4 544.5s 144.7s 3.76x 399.8s
10 No 5 192.4s 167.8s 1.15x 24.6s
50 No 5 190.5s 117.2s 1.63x 73.3s
100 No 4 260.8s 139.6s 1.87x 121.2s

With two-wave HalfAtATime, ALB-backed fleets saved 168.6–185.7 seconds. No-ALB rows isolate local reuse.

Host Count ALB in Front? Standard Deployment Time Restart Deployment Time Speedup Time Saved
10 Yes 230.2s 61.6s 3.74x 168.6s
50 Yes 276.2s 90.5s 3.05x 185.7s
100 Yes 275.2s 103.9s 2.65x 171.3s
10 No 133.8s 37.0s 3.62x 96.8s
50 No 152.4s 86.7s 1.76x 65.7s
100 No 162.7s 109.1s 1.49x 53.6s

Multi-wave ALB deployments repeated traffic control and saved the most time. These results carry a couple of safety implications to keep in mind.

The safety controls remain familiar

A RESTART deployment uses the controls already configured for the deployment group:

  • Deployment configuration: Use CodeDeployDefault.OneAtATime, CodeDeployDefault.HalfAtATime, CodeDeployDefault.AllAtOnce, or a custom minimum healthy host setting to bound concurrent restarts.
  • Lifecycle validation: CodeDeploy runs ValidateService on each host before it considers that host healthy.
  • CloudWatch alarms: CodeDeploy polls the alarms configured on the deployment group and stops the deployment when an alarm enters ALARM.
  • Automatic rollback: CodeDeploy applies the deployment group’s automatic rollback configuration for qualifying failures. A rollback can’t reverse an external configuration change, so correct the underlying configuration before retrying.
  • Deployment history: GetDeployment and ListDeployments expose the restart’s status, timestamps, revision, and result for monitoring and audit.

Traffic-control behavior

RESTART doesn’t run the load balancer BlockTraffic and AllowTraffic steps. The host remains registered while its application stops and starts. Treat this as an operational constraint, not as the reason to use RESTART.

For request-serving fleets, make ApplicationStop stop accepting new work and drain in-flight work before the process exits. Select a deployment configuration that preserves enough healthy capacity for the expected restart duration. Pull-based workers stop receiving work when the process stops, but their hooks still need to handle in-flight jobs safely.

Walk through a restart deployment

This section walks through starting a restart deployment from the command line and the CodeDeploy console, and then monitoring its progress.

Prerequisites

  • An existing Amazon EC2 or on-premises application and deployment group with at least one successful deployment.
  • An AWS Command Line Interface (AWS CLI) or SDK version that supports the deploymentMode request field.
  • CodeDeploy agent version 2.1.0 or newer on target hosts to benefit from local revision reuse. Earlier agents work but fall back to downloading the revision.

Start a restart deployment with the AWS CLI

The following command restarts the fleet with the deployment group’s default deployment configuration:

aws deploy create-deployment \
    --application-name MyApp \
    --deployment-group-name MyDeploymentGroup \
    --deployment-mode RESTART \
    --description "Restart processes after a runtime configuration update"

CodeDeploy returns a deployment ID:

{
  "deploymentId": "d-EXAMPLE123"
}

To restart one host at a time, provide a deployment configuration in the request:

aws deploy create-deployment \
    --application-name MyApp \
    --deployment-group-name MyDeploymentGroup \
    --deployment-mode RESTART \
    --deployment-config-name CodeDeployDefault.OneAtATime \
    --description "Restart one host at a time"

Omit --deployment-config-name to use the deployment group’s configured default.

Start a restart deployment on the console

You can also use this feature in the CodeDeploy console. Under Applications, select the deployment group you want to restart and choose Create deployment.

CodeDeploy console Applications page with a deployment group selected and the Create deployment button

Figure 2: Choosing Create deployment for a deployment group in the CodeDeploy console

For Deployment mode, select Restart, and configure any other settings or overrides on the page (the same options you would set with the AWS CLI).

CodeDeploy Create deployment page with Restart selected as the deployment mode

Figure 3: Selecting Restart as the deployment mode on the Create deployment page

Monitor the restart

You can see the status of the deployment on the console, or use the deployment ID with the GetDeployment API or CLI command. For example:

aws deploy get-deployment --deployment-id d-EXAMPLE123

The deployment moves through the standard Created, InProgress, and terminal states. If a lifecycle hook fails, the deployment breaches its minimum healthy host requirement, or a configured alarm enters ALARM, CodeDeploy stops the restart and applies the configured failure behavior.

You can distinguish restart deployments by the deploymentMode field in the GetDeployment API, or visually on the console under Deployment details.

CodeDeploy console Deployment details showing the deploymentMode field set to RESTART

Figure 4: Deployment details showing the deploymentMode field set to RESTART

Handle alarms during incident recovery

If an alarm is already in ALARM state, CreateDeployment still creates the deployment. CodeDeploy then stops it when the deployment workflow observes that alarm during polling. The ignorePollAlarmFailure setting does not ignore an alarm in ALARM. It only controls behavior when CodeDeploy cannot retrieve alarm state.

If an approved incident runbook requires a restart despite the current alarm state, override alarm monitoring for that deployment:

aws deploy create-deployment \
    --application-name MyApp \
    --deployment-group-name MyDeploymentGroup \
    --deployment-mode RESTART \
    --override-alarm-configuration enabled=false \
    --description "Restart under an approved incident runbook"

This override disables all deployment-group alarms for that deployment. It requires codedeploy:UpdateDeploymentGroup in addition to the permission to create a deployment. Use it only when your incident process provides another health signal and explicitly authorizes the override. The deployment configuration and lifecycle validation continue to apply.

Request constraints

RESTART has the following constraints:

  1. It supports EC2 and on-premises in-place deployment groups. It doesn’t support Amazon Elastic Container Service (Amazon ECS) or AWS Lambda deployment groups.
  2. The deployment group must have a successful revision for CodeDeploy to reapply.
  3. Don’t provide revision, s3Location, gitHubLocation, or deploymentRevisions. CodeDeploy resolves the revision from deployment history.
  4. Don’t combine RESTART with updateOutdatedInstancesOnly. That option selects hosts that aren’t running the target revision, which conflicts with restarting hosts on the current successful revision.

For exact request validation and error types, see the CreateDeployment API reference.

Conclusion

Restarting a fleet is routine, but it still changes production availability and exposes configuration or process failures. RESTART gives the operation a first-class CodeDeploy path instead of requiring a separate fleet script.

Set deploymentMode to RESTART to reapply the deployment group’s last successful revision. CodeDeploy controls the batch size, runs the lifecycle hooks, validates each host, monitors configured alarms, records the result, and reuses the local revision archive when possible. The operation is faster when the archive is already present, while hosts that need to download it use the normal fallback path.

To get started, review the AWS CodeDeploy documentation and create-deployment CLI reference.

 


About the authors

Iskandar Anvarov

Iskandar Anvarov

Iskandar is a software engineer with over 10 years of experience across enterprise software and cloud computing. For the past 4 years, he’s been with Amazon Web Services, building developer tools and cloud infrastructure that support customers at scale. He holds a bachelor’s degree in Business Information Systems from Westminster University in Tashkent.

Filip Danić

Filip Danić

Filip is a software developer with 12 years of experience across service agencies, startups, and enterprise. For the past 5 years he’s been at Amazon, currently a Software Developer Engineer on AWS CodeDeploy. He previously led an internal finance systems team supporting compliance and reporting at scale. Filip holds a bachelor’s degree in Computer Science from the School of Computing (Računarski Fakultet) in Belgrade.

Monika Awasthi

Monika Awasthi

Monika is a Technical Product Manager with over 10 years of experience launching and scaling products across countries and functions. At Amazon Web Services, she builds developer tools that help customers ship software safely and efficiently, drawing on a technical background spanning networking, data analytics, and product management. She holds an engineering degree in Electronics and Communication and an MBA from INSEAD, France.

Daz Akbarov

Daz Akbarov

Daz is a Solutions Architect at Amazon Web Services, based in Munich. For the past 2 years at AWS, he has worked with customers across Central Asia and the Caucasus in Financial Services Industry, helping them migrate to and build on AWS. He focuses on modernization to cloud-native architectures, building serverless applications and enhancing developers’ experience. He has over 7 years of experience in enterprises and startups, including launching (and failing) his own one.

Run an automated operational review with the Amazon Redshift MCP server

Post Syndicated from Sidhanth Muralidhar original https://aws.amazon.com/blogs/big-data/run-an-automated-operational-review-with-the-amazon-redshift-mcp-server/

It’s the end of a strong quarter, and your Amazon Redshift workloads have grown with the business. Data volumes are up, new pipelines have shipped, and more teams are querying than when you first sized the cluster. Nothing is broken, but this is exactly when a periodic operational review pays off. It confirms the cluster is still tuned for how it’s used today, and surfaces ways to optimize cost and performance as you scale.

Amazon Redshift already automates significant operational complexity. Autonomics features such as automatic table optimization, automatic workload management, and automatic vacuum handle much of the complex work, so you can focus on writing queries rather than managing infrastructure. Amazon Redshift Serverless goes further, using AI-driven scaling that adapts compute to workload demand.

Even so, some decisions still benefit from human judgment. For example, table design choices may not have accounted for common join patterns, or query patterns may have shifted since the tables were built. Both are worth revisiting. Likewise, as workloads increase, it’s worth deciding whether the current deployment model is still correctly sized.

Many organizations have turned this into a recurring business process: monitoring dashboards and key performance indicators (KPIs), narrowing down long-running queries, collaborating across teams to resolve them, and coaching users on efficient query patterns. This isn’t unique to Amazon Redshift. It’s a best practice for any production data system.

AWS Enterprise Support helps customers with these reviews. But even with specialist assistance, the process typically consumes 4–8 hours of focused effort per cluster: assembling diagnostic queries, interpreting results, cross-referencing documentation, and compiling findings into a prioritized report. Most teams recognize the value, but consistently deprioritize it in favor of feature delivery and day-to-day operations.

In a previous post, we demonstrated how to query Amazon Redshift using natural language with Kiro and the Amazon Redshift Model Context Protocol (MCP) server. That approach replaced manual schema navigation and hand-written SQL with conversational analytics. This post takes the next step: from asking questions about your data to asking questions about your cluster’s operational health.

The review_cluster tool in the Amazon Redshift MCP server makes the entire diagnostic process available from a single natural-language request. It evaluates 12 diagnostic areas, identifies potential issues, and returns prioritized recommendations linked directly to AWS documentation, covering both provisioned clusters and serverless workgroups in one invocation. What previously required hours of specialist effort now completes in a few minutes. This makes it practical to review your cluster weekly, after significant schema changes, or before peak traffic events.

In this post, you learn how to:

  1. Run an automated operational review of your Amazon Redshift cluster using natural language.
  2. Interpret the structured findings and recommendations.
  3. Act on specific findings conversationally, turning diagnostics into remediation without leaving Kiro, Claude, or any MCP-compatible client.
  4. Incorporate periodic reviews into your operational workflow.

What is the review_cluster tool?

A single natural-language request triggers a comprehensive diagnostic assessment. It complements the server’s discovery and query tools: those give you conversational access to your data, while review_cluster gives you the same for your cluster’s health. The tool executes a curated set of queries against Amazon Redshift system views, evaluates the results against known best-practice thresholds, and returns structured findings with prioritized recommendations. Every recommendation includes direct links to the relevant AWS documentation, so you can move from identification to remediation without searching. The entire process is read-only. No data is modified, no configuration is changed, and no resources are created. The tool observes and reports. Remediation decisions remain with you.

The tool returns a structured result containing:

  • Signals evaluated: The total number of diagnostic checks executed.
  • Findings: A list of triggered conditions, each with a name, the number of affected objects (tables, queries, nodes), and linked recommendation IDs.
  • Recommendations: A deduplicated list of corrective actions, ordered by effort, with documentation links.

What it inspects

The review currently evaluates your cluster across 12 diagnostic areas:

  1. Automatic Table Optimization (ATO): Whether the automated tuning actions of Amazon Redshift (encoding, sort keys, distribution styles) are completing successfully or not.
  2. Table design recommendations: Amazon Redshift Advisor recommendations for encoding, sort keys, and distribution that have not yet been applied.
  3. COPY and data ingestion performance: File sizing, parallelism relative to slice count, and ingestion throughput patterns.
  4. External query (Spectrum) performance: Partition pruning effectiveness, file sizes, and scan efficiency for queries against external tables.
  5. Materialized view health: Staleness, auto-refresh status, and maintenance overhead of materialized views.
  6. Node and storage utilization: Disk usage, node type currency, and whether the cluster would benefit from migration to the latest instance generation.
  7. Table-level statistics: Vacuum status, stale statistics, sort key effectiveness, distribution skew, and compression encoding coverage.
  8. Query performance: The longest-running queries, nested loop joins, and disk spill patterns.
  9. Workload usage patterns: How intensively the cluster is used throughout the day and whether the workload suits the current deployment model.
  10. Workload Management (WLM) configuration: Queue setup, concurrency scaling, short query acceleration, priority settings, and query monitoring rules.
  11. Workload evaluation: Whether the cluster’s utilization pattern suggests it could benefit from a different deployment model such as serverless.
  12. Serverless scaling: Whether a serverless workgroup’s observed compute range falls within the AI-driven scaling window.

Not every area applies to every deployment, because some checks target configuration that only exists in one model. Four checks are provisioned-only: COPY parallelism relative to slice count, node and storage utilization, WLM configuration, and workload evaluation. These tune constructs that Amazon Redshift Serverless manages for you. Serverless has no nodes or slices to size, and it always uses automatic WLM rather than user-configured queues, so there is nothing for those checks to act on. The serverless-scaling check is the reverse: it evaluates AI-driven Redshift Processing Unit (RPU) scaling, which is specific to serverless and has no equivalent on a provisioned cluster. As a result, a provisioned cluster evaluates 11 of the areas and a serverless workgroup evaluates 8, and the tool reports the number actually run as signals evaluated.

Running your first review

The following steps cover what you need in place and how to start a review.

Prerequisites

Before running a review, make sure that you have:

  • An MCP-compatible client configured with the Amazon Redshift MCP server (for example, Kiro). See the Amazon Redshift MCP server README for installation and configuration steps.
  • Valid AWS credentials with the AWS Identity and Access Management (IAM) permissions the server requires.
  • The sys:monitor role granted to the connecting database user, the one additional database permission review_cluster needs.

To grant this access, a database administrator with superuser privileges runs:

GRANT ROLE sys:monitor TO "IAMR:YourRoleName";

The database user name for IAM identities follows the format IAMR:RoleName or IAM:UserName. Confirm yours with SELECT current_user in the Amazon Redshift Query Editor. The double quotes in the GRANT statement are required for IAM identity names.

With prerequisites in place, the review is a single prompt:

Run a review of analytics-cluster

The agent identifies the target cluster, connects to the database, and executes the diagnostic assessment.

The following example shows the output from a review of a serverless workgroup. Because four provisioned-only areas don’t apply, the tool evaluated 8 of the 12 diagnostic areas, identified 8 findings, and mapped them to 6 distinct recommendations:

Cluster Review: analytics-cluster (Serverless, eu-west-2 / London)
Ran the diagnostic --- result: 8 signals evaluated, 8 findings, 6 recommendations.
Findings
- Missing statistics --- 16 queries hit tables that have never been analyzed
- Dist/Broadcast alerts --- 10 queries reshuffle data across nodes
- Nested loop joins / cross-joins --- 5 long-running queries
- Small tables without ALL distribution --- 4 tables
- Small tables with a manual sort key --- 4 tables
- Large table with data skew --- 1 table (slice imbalance ≥4x)
- Large table distributed by date/datetime --- 1 table
- Variable serverless workload in the AI-scaling range --- 1 workgroup
Recommendations
1. REC_003 --- Analyze tables (fix missing stats). Effort: Low. Run or schedule ANALYZE on the 16 affected tables. Lowest-effort, highest-impact item; also improves the poor distribution/plan decisions the optimizer is making. Docs: https://docs.aws.amazon.com/redshift/latest/dg/t_Analyzing_tables.html
2. REC_007 --- Right-size sort keys. Effort: Low. The 4 small tables carry manual sort keys that add overhead for little benefit --- switch them to ALTER TABLE [table] ALTER SORTKEY AUTO. Docs: https://docs.aws.amazon.com/redshift/latest/dg/c_best-practices-sort-key.html
3. REC_008 --- Choose the right distribution style. Effort: Low--Medium. Small tables → DISTSTYLE AUTO/ALL; the skewed large table → pick a better DISTKEY (skew often from NULLs in the key); the date-distributed large table → reconsider the key or move to DISTSTYLE AUTO. This also resolves the 10 Dist/Broadcast query alerts. Docs: https://docs.aws.amazon.com/redshift/latest/dg/c_best-practices-best-dist-key.html
4. REC_030 --- Tune serverless AI-driven scaling. Effort: Low. Review your price-performance target (Optimizes for cost / Balanced / Optimizes for performance) and pair it with Max capacity + Max RPU-hours caps to keep cost predictable. Adapts over 1--3 days. Docs: https://docs.aws.amazon.com/redshift/latest/mgmt/serverless-capacity.html
5. REC_009 --- Remove nested loop joins (cross-joins). Effort: Medium. Add join conditions to the 5 cross-join queries. Docs: https://docs.aws.amazon.com/redshift/latest/dg/query-performance-improvement-opportunities.html#nested-loop
6. REC_019 --- Add QMR rules. Effort: Medium. Add query monitoring rules to catch nested loops and long runtimes going forward (log/hop/abort actions). Docs: https://docs.aws.amazon.com/redshift/latest/dg/cm-c-wlm-query-monitoring-rules.html
Suggested starting point
Start with REC_003 (ANALYZE) --- cheapest, highest-impact, and likely a root cause behind several slow-query findings. Then apply the AUTO table-optimization changes in REC_007 and REC_008, which are mostly one-line ALTER TABLE statements.
Want me to pull the specific table names behind each finding and draft the exact ANALYZE / ALTER TABLE remediation statements?

Each finding represents an independent diagnostic condition that was triggered. The affected_row_count indicates how many objects match that specific condition. These counts describe different dimensions of the cluster and are not additive across findings.

When the tool returns zero findings, the cluster is operating within best-practice thresholds across the evaluated areas.

From findings to fixes

The review output isn’t a static report. It’s a starting point for an interactive conversation. Each finding identifies a specific condition, and each recommendation provides a clear remediation path with documentation links. From here, you can continue working within the same Kiro session to plan and prepare your next steps.

Recommendations fall into four broad categories:

  • Quick configuration changes: Actions such as enabling concurrency scaling or short query acceleration require a single parameter change in the Amazon Redshift console or a brief API call. These are low-risk, high-impact adjustments that can often be applied immediately.
  • Batch table operations: Findings related to table design (distribution style, sort keys, compression encoding) typically affect multiple tables. You can ask Kiro to list the specific tables involved and generate the corresponding ALTER TABLE statements. Review the generated SQL, then apply it through the Amazon Redshift Query Editor or your preferred SQL client.
  • Data ingestion optimization: Findings related to COPY performance identify inefficiencies in how data is loaded. For example, source files might be too small, or file counts might not align with the cluster’s slice count. Addressing these requires changes upstream in your extract, transform, and load (ETL) pipeline or Amazon Simple Storage Service (Amazon S3) staging process rather than within Amazon Redshift itself.
  • Architectural decisions: Recommendations such as migrating to Amazon Redshift Serverless or resizing to RG instances require broader evaluation. These are not single-command fixes. They involve capacity planning, workload testing, and potentially migration steps. The recommendation text and linked documentation provide the context needed to begin that planning.

After presenting findings, Kiro proposes the next best step to act upon, typically starting with the lowest-effort, highest-impact recommendation. Alternatively, you can direct the conversation yourself. For example:

  • “Which tables are affected by the distribution style finding?”
  • “What would the ALTER TABLE statements look like for those tables?”
  • “Explain what concurrency scaling does and how to enable it.”
  • “What are the trade-offs of migrating this workload to serverless?”

Kiro retrieves the relevant details, generates SQL where applicable, and references the documentation. The tool diagnoses and recommends, but does not modify your cluster. The decision to apply changes remains with you.

Best practices

Tips for getting the most out of review_cluster.

  1. Embed reviews in your DataOps practice. Regular table and cluster maintenance is a foundational best practice for any production data warehouse, and its value comes from consistency. Operational reviews are one component of a broader DataOps discipline: the practice of maintaining data systems that are clean, reliable, governed, and always available. Define KPIs for your cluster, such as query latency percentiles, disk spill frequency, and WLM queue wait times. Then use periodic review_cluster runs to track your query and workload optimization progress against them. Over time, this creates a feedback loop: findings inform remediation, KPIs measure impact, and the next review validates improvement.
  2. Run reviews on a regular cadence. Treat operational reviews like testing: the more routinely you run them, the sooner you catch drift before it affects users. Consider running a review weekly, after significant schema changes, after major data loads, or before anticipated peak traffic events.
  3. Follow the documentation links. Every recommendation includes direct links to the relevant AWS documentation. These pages provide detailed guidance, edge cases, and configuration examples that go beyond what the recommendation text can cover. Use them as your primary reference when planning remediation.
  4. Start with low-effort wins. The recommendations are ordered by effort. Begin with quick configuration changes (concurrency scaling, short query acceleration, query monitoring rules) before moving to structural changes that require broader planning. Early wins build confidence and often improve cluster performance enough to create headroom for larger changes.
  5. Use the review as a baseline. Run a review before and after significant changes to measure their effect. For example, after applying table design recommendations, a follow-up review should show fewer table-related findings. This before-and-after pattern helps validate that your changes had the intended impact.

Conclusion

In this post, you learned how to run an automated operational review of your Amazon Redshift cluster using natural language with Kiro. The review_cluster tool evaluates 12 diagnostic areas, from table design and workload management to node utilization and data ingestion performance. It returns prioritized, actionable recommendations linked directly to AWS documentation.

What previously required a specialist to assemble diagnostic scripts, interpret system view outputs, and compile findings over several hours now completes in a single request. This shift makes it practical to incorporate operational reviews into your regular workflow rather than treating them as an infrequent, resource-intensive exercise.

To get started:

  1. Make sure the Amazon Redshift MCP server is configured in Kiro (see the setup post).
  2. Grant the sys:monitor role to your database user.
  3. Ask Kiro to review your cluster.

To go further:

  • Visit the Amazon Redshift MCP server documentation for the full tool reference.
  • Read the MCP protocol to understand how AI agents integrate with external tools.
  • Explore Kiro for additional capabilities including steering files, hooks, and agent automation.

About the authors

Sidhanth Muralidhar

Sidhanth Muralidhar

Sidhanth is a Principal Technical Account Manager at AWS, where he partners with enterprise customers to design, scale, and optimize cloud-focused systems. He specializes in guiding organizations through complex architectural decisions across cost efficiency, reliability, performance, and operational excellence. His work increasingly sits at the intersection of data systems and AI as well, helping customers operationalize modern data architectures and build intelligent, production-ready systems.

Sergey Konoplev

Sergey Konoplev

Sergey is a Senior Database Engineer on the Amazon Redshift team at AWS. For more than a decade he has focused on the automation and improvement of database and data operations, and he currently drives a range of initiatives spanning operations, observability, and AI tooling.

Vlad Siniavin

Vlad Siniavin

Vlad is a Sr. Technical Account Manager at AWS with over 15 years of experience in building innovative solutions, products and services. He is driven by delivering measurable outcomes for his customers – whether that’s reducing operational risk, optimising costs, or accelerating cloud adoption. He believes the best technical guidance starts with deeply understanding what matters most to the customer and acting in their best interest.

AWS Weekly Roundup: Amazon Bedrock Managed Agents powered by OpenAI, Q3 service availability updates, Kiro workflows, and more (October 5, 2026)

Post Syndicated from Channy Yun (윤석찬) original https://aws.amazon.com/blogs/aws/aws-weekly-roundup-amazon-bedrock-managed-agents-powered-by-openai-q3-service-availability-updates-kiro-workflows-and-more-october-5-2026/

Last week, we announced a public preview of Amazon Bedrock Managed Agents powered by OpenAI, built on a customized version of OpenAI’s Agents API engineered to be AWS-native and integrated with AWS resources. You can now build agents optimized for OpenAI models that run entirely inside AWS with the identities, permissions, and governance controls you already use.

You can choose an execution environment: self-hosted compute to use an existing development machine, container, or compute environment or Amazon Bedrock AgentCore Runtime, for managed runtime sessions and configurable storage in your AWS account. To learn more, visit the Amazon Bedrock documentation.

In addition, we are adding new frontier models on Amazon Bedrock to expand your model choices:

  • OpenAI GPT-6.1 Sol: An upgrade to GPT-6 Sol, GPT-6.1 Sol delivers exceptionally strong performance on agentic coding, computer use, and professional work. According to OpenAI, it approaches GPT-6 Astra across demanding evaluations at roughly one-fifth of the cost, giving developers more room to build and run capable agents at scale. To learn more, visit the GPT-6.1 Sol model card.
  • OpenAI GPT-6 Astra UltraFast mode: Ultrafast is a premium speed tier for GPT-6 Astra, built for workloads where speed matters most. According to OpenAI, Ultrafast delivers up to 6x faster inference in the API, with up to 300 tokens per second. The Amazon Bedrock inference engine delivers the performance, security, and reliability required for production workloads. To learn more, visit the GPT-6 Astra model card.
  • Anthropic Claude Sonnet 5.5: Claude Sonnet 5.5 is a smarter, more efficient Sonnet and a step up from Sonnet 5, making it a natural upgrade for teams already building on Sonnet. It’s stronger for coding, completing well-scoped tasks as part of a larger coding strategy such as building and fixing features with Claude in the same session or verifying output against requirements. To learn more, visit the Claude Sonnet 5.5 model card.
  • SpaceXAI Grok 4.7: Grok 4.7 builds on Grok 4.6 with better mixed-document handling, more dependable repo-scale coding with planning and error recovery, and enhanced browser-use agents for form fills and portal navigation. To learn more, visit the Grok 4.7 model card.

Last week’s launches

Here are some launches that got my attention:

  • AWS Well-Architected Agent (preview): You can use an AI-powered agent service that analyzes your AWS environment to deliver targeted, contextual recommendations for improving your applications’ cost, security, performance, and resilience. The agent analyzes your infrastructure, understands your unique business goals, and delivers contextual recommendations.
  • Amazon Aurora PostgreSQL supports direct querying of Apache Iceberg and Parquet data: You can directly query operational data together with data stored in data lakes in Apache Iceberg and Parquet formats using your existing PostgreSQL applications and tools, without extract, transform, and load (ETL) pipelines or data duplication.
  • Amazon S3 Tables support all Apache Iceberg V3 data types: Amazon S3 Tables add support for geometry, geography, unknown, and nanosecond timestamp data types, along with column default values, as defined in the Iceberg V3 specification. You can now store geospatial coordinates and nanosecond-precision event times natively instead of encoding them in strings or integers.

For a full list of AWS announcements, be sure to keep an eye on the What’s New with AWS page.

AWS service availability updates

When the availability of an AWS service or feature changes, we provide customers guidance in AWS Product Lifecycle Changes on available alternatives and support for migration so that disruptions to your operations are minimized. The following lifecycle changes were updated on September 29, 2026.

Services moving to Maintenance (no longer accessible to new customers starting October 29, 2026):

Services entering Sunset:

Services reaching End of Support (as of September 29, 2026):

  • Amazon Mechanical Turk

We understand that changes in availability can impact your operations. For specific guidance, consult the relevant service documentation or contact AWS Support.

Other AWS news

Here are some additional projects and news items you may find interesting:

  • Introducing Kiro workflows: Kiro workflows enable you to carry out complex tasks from start to finish with multiple agents and less supervision. We’ve been building Kiro itself with workflows, including the new cloud configuration, cloud sessions, and most of the workflow experience.
  • Introducing Strands Decider: Strands Decider is one of a new class of decision models or system one models, a type of model that has been gaining significant attention since TypeSafe AI’s launch of Jev earlier this month. Strands Decider 2B is a small, open source, decision model optimized for fast experimentation, local development, and innovation.
  • New FDE pathways for AWS Partners: On June 30, AWS announced the Forward Deployed Engineering (FDE) organization, backed by a $1 billion investment, and extended this hands-on delivery approach to AWS Partners through the Partner-Led FDE motion. Now, AWS Partners have a structured way to build and validate that depth with three new Partner FDE pathways and credentials that recognize the applied proficiency required to deliver production agentic AI.

For a full list of AWS blog posts, be sure to keep an eye on the AWS Blogs page.

Learn more about AWS, browse and join upcoming AWS-led in-person and virtual events, startup events, and developer-focused events, including upcoming AWS re:Invent and AWS Community Days. Join the AWS Builder Center to connect with builders, share solutions, and access content that supports your development.

That is all for this week. Check back next Monday for another Weekly Roundup!

— Channy

[$] An update on the Sashiko patch-review system

Post Syndicated from corbet original https://lwn.net/Articles/1096963/

Patch review has long been one of the limiting constraints for the kernel
project (and most others); there just aren’t enough people to properly
review all of the code that is submitted for inclusion. The Sashiko
system, which uses a large language model (LLM) to generate reviews
automatically, offers the prospect of some relief, and has already become
an important part of the kernel’s development process. At the 2026 edition
of Kernel Recipes, Roman
Gushchin, the maintainer of Sashiko, provided an overview of how
the system works and what is being done to improve it.

Security updates for Monday

Post Syndicated from jzb original https://lwn.net/Articles/1098599/

Security updates have been issued by AlmaLinux (ghostscript, libvirt, and osbuild-composer), Debian (freecad, linux-6.12, node-lodash, pcre2, perl, php8.2, ruby-rack-session, wireshark, and xen), Fedora (assimp, budgie-control-center, budgie-desktop, budgie-desktop-services, budgie-desktop-view, chromium, cri-o1.36, curl, flatpak, lemonldap-ng, libX11, nagios-plugins, nanosvg, noctalia, openssl, pgbouncer, pocillo-gtk-theme, prometheus, python-streamlink, python-urllib3, python-uv-build, python3.12, ruff, rust-libcst, rust-libcst_derive, rust-salsa, rust-salsa-macro-rules, rust-salsa-macros, ty, and uv), Red Hat (rhc-worker-script), SUSE (amazon-ecs-init, binaryen-133, bind, chromium, distribution, firefox, firefox-esr, glib2, gnumeric, helm, jline3, kubevirt-1.6, libparted-fs-resize0, libslirp-devel, libtcnative-1-0, libtcnative-2-0, tomcat, tomcat10, tomcat11, libwireshark19, openai-codex, python313-litellm, rpcbind, rustup, sccache, suseconnect-ng, valkey, wget, and xdg-dbus-proxy), and Ubuntu (ceph).

One year later: the power of 1.1.1.1 interns

Post Syndicated from Kelly Russell original https://blog.cloudflare.com/one-year-later-1111-interns/

A year ago, many companies were cutting intern and new-graduate hiring. We went the other way. We announced a goal to hire as many as 1,111 interns in 2026, a number that’s a nod to 1.1.1.1, our public DNS resolver.

The bet was that AI makes early-career talent more valuable and able to make an impact faster. The best AI tools help people learn a system faster, try more ideas, and take on harder problems. They don’t supply the energy, curiosity and fresh eyes a new person brings to a team.

A year in, and our interns are shipping to our internal teams and to millions of customers.

If you’re reading this on the Cloudflare Blog, you’re already using some of their work. The blog runs on EmDash, and EmDash’s second maintainer started at Cloudflare as an intern this past summer. 

A new generation of builders

We’re still working toward 1,111. So far, we’ve hosted 750 internships across 48 teams in nine offices: Austin, San Francisco, London, Lisbon, New York, Singapore, Bengaluru, Washington DC, and Sydney. And we’re still hiring.

From their first day, interns joined active teams and worked on real problems. Each was expected to leave something behind: a shipped product improvement, a better process, a new piece of infrastructure, or an insight that changes how a team approaches its work.

That work reached far beyond engineering. An internal audit intern built an AI-assisted pipeline to automate ISO compliance control testing and documentation. A product manager intern worked on an API, dashboard, and migration tooling to modernize credit management for our Startup Program. A people team software engineering intern improved background-verification workflows for new hires. A customer support intern explored ways to use AI to trigger troubleshooting commands from case descriptions.

AI-native interns, real-world results

AI was a key theme through the whole program, but it wasn’t a shortcut around learning or accountability. Interns used AI to understand unfamiliar codebases and systems, compare product requirements with implementation, prototype ideas, automate repetitive work, and accelerate the path from a question to a working solution. Managers and mentors remained essential in setting direction, reviewing decisions, and ensuring that what shipped met Cloudflare’s standards.

For many interns, AI moved the starting line. They could map a new project in days instead of weeks, create a working scaffold quickly, and spend more time on the hard questions: What should we build? Who will use it? How do we know it is correct? What happens when it reaches production?

One intern put it more bluntly: “How did previous interns ship anything in 12 weeks without AI?”

Interns ship at Cloudflare

Last year’s announcement had a section with that title. This year, the interns wrote the posts:

More cache from the same hardware. RAM and disk prices have climbed sharply over the past year, so our intern Aashi asked whether Cloudflare could get more cache capacity out of the servers it already has. Aashi built Cache Transcoding, which compresses eligible assets with Zstandard inside our primary proxy before they’re written to disk. In initial testing, it shrank them to a third of their original on-disk size on average. That points toward petabytes of effective cache capacity and less data moving between our data centers.

The CMS behind this blog. Noah joined as an intern and became EmDash’s second maintainer, contributing changes across the media library, content editor, and admin interface. EmDash 1.0 shipped during Birthday Week.

Getting ready for post-quantum. Cloudflare is targeting 2029 for post-quantum security, and you can’t migrate what you can’t see. Tiago helped build CryptoLabe, an internal AI tool that finds cryptography across our codebase and maps what depends on it. Sophie helped bring per-connection post-quantum visibility to Logpush, Log Explorer, and HTTP Traffic Analytics, so customers can check their own progress too.

Fixing the Internet’s plumbing. Some intern work goes beyond our products. Iliana measured how often networks rewrite BGP’s ORIGIN attribute to pull traffic their way, and found it on roughly 70% of observed paths. Then she publicly advocated for taking ORIGIN out of route selection altogether. She also tracked the adoption of RFC 9234, which lets routers reject route leaks on their own. Helping build a better Internet includes work like this: measuring a problem no single network owns, publishing the data, and proposing a fix.

Our interns didn't just help with Birthday Week, they shipped impactful work! Several of this year's Birthday Week launches included intern-authored and intern-engineered work:

See more on the Cloudflare Blog.

Building community together

An internship is also about the people you meet. Across our offices, interns shared meals, volunteered together, and made friends. Our CEO, Matthew Prince, hosted dinners with interns in offices around the world to hear directly about what they were working on and learning.

Executives also joined Q&A sessions where interns could ask candid questions about Cloudflare, strategy, technology, and careers. These sessions complemented the day-to-day support of managers and mentors, and gave interns a view of how the whole company works and where it’s going.

What comes next

We’ll keep growing the program next year, with a new cohort of interns, more teams taking part, and more learning about how early career people do their best work with AI and with the people around them.

And many of this year’s interns are coming back to Cloudflare full-time. A good summer is nice, but the outcome we want is a career and a new generation of people helping build a better Internet.

If you’re a student or early-career technologist, and you want your first project to ship to millions of people, apply to our internship program. We’ll keep opening up roles throughout the year.

Everything we launched during Birthday Week 2026

Post Syndicated from Carlos Armada original https://blog.cloudflare.com/birthday-week-2026-wrap-up/

We celebrated our 16th birthday last week by sharing how we’re building a better Internet for today’s world. As Matthew and Michelle reflected in this year’s Founders’ Letter, this year saw some of the most consequential changes in the history of the Internet.

For the first time, automated traffic surpassed human activity. AI is empowering people to build like never before, leading the Internet to grow massively in scale and unlocking more ambition and creativity. As we witnessed the influence that agent-driven recommendations have on consumer choices, we identified the need for a new approach that creates space for new businesses to succeed.

Each day of Birthday Week explored a different way we are helping to build the future of the Internet. We began on Monday by strengthening our commitment to open source. Tuesday focused on application security and the post-quantum transition. On Wednesday, we explored new economic models for the agentic Internet. Thursday, we expanded the Developer Platform with new tools for data analysis, storage, AI, and agent development. Finally, we closed out the week by launching features that make Cloudflare faster, easier to operate, and more accessible to everyone. As a special Birthday Week follow-up, we shared an update on our intern program, one year after announcing our goal to hire 1,111 interns. Interns directly contributed to many of the projects launched this week, including EmDash, post-quantum visibility, CryptoLabe, and Protected Quick Tunnels.

We shipped 46 announcements this week. In case you missed any, here’s the full list of everything we announced during Birthday Week 2026.

Monday, September 28 – Commitment to open source

With the announcement of our new CLI, which we released alongside the pipeline we use to generate it and our SDKs and docs, we shared how we’re building to support agents and developers as they use Cloudflare — and supporting the projects that you rely on, too.

What

In a sentence…

Introducing cf: the agentic CLI for the entire Cloudflare API

The new cf CLI mirrors the Cloudflare API, uses JSON-first output and typed configuration, and gives people and agents one consistent command-line interface.

Introducing Forge: the open source pipeline for generating SDKs, CLIs, docs, and more

Forge is a pluggable, open-source pipeline that runs in CI to generate SDKs, CLIs, documentation, and other interfaces directly from API definitions.

Introducing EmDash – the spiritual successor to WordPress that solves plugin security

EmDash is an open-source, Astro-based serverless CMS that runs plugins in isolated Worker sandboxes with explicitly approved capabilities.

Four months of VoidZero at Cloudflare: making the open-source JavaScript toolchain faster for all humans and agents

Since joining Cloudflare, VoidZero has delivered more than 80 releases across the Vite ecosystem, and its previously commercial Void platform will become fully open source.

Next.js applications, powered by Vite: introducing Vinext 1.0

Vinext 1.0 turns an AI-built experiment into a production-ready, portable way to run Next.js applications on Vite.

The road to the agentic browser: A Kitesurf update

Kitesurf, our Workers-based browser for agents, adds WebMCP support, faster DOM operations, broader web compatibility, and terminal-based rendering.

How fast is the web? Explore billions of real-user measurements with BEACON

BEACON makes billions of anonymized real-user performance measurements from 10,000 major websites available as a public BigQuery dataset.

Supporting native Rust in Workers with the new Emscripten target for wasm-bindgen

Experimental Emscripten target support lets developers bring more native Rust libraries and applications, including progress toward Tokio support, to Workers.

Introducing The Cold Start: pitch your startup live at Cloudflare Connect

The Cold Start gives five early-stage companies the opportunity to pitch live at Cloudflare Connect and compete for resources to help them grow.

Tuesday, September 29 – Helping secure the agentic Internet

Technological progress is rapidly changing how we think about application security. We announced our intention to become a certificate authority, as well as how we’re preparing foundational Internet cryptography for the post-quantum era and adapting application security to counter AI-driven attacks.

What

In a sentence…

Building a certificate authority for the whole Internet

Twelve years after launching Universal SSL, Cloudflare announced its intention to become a public certificate authority (CA) and add resilience to free, automated certificate issuance.

Building a post-quantum certificate authority with Merkle Tree Certificates

Our planned CA will issue free Merkle Tree Certificates designed to make post-quantum authentication practical without imposing large certificate and handshake costs.

Using AI to chart a course for our post-quantum migration

CryptoLabe uses AI to find and classify cryptography across our codebase as Cloudflare works toward completing its post-quantum migration by 2029.

Preventing quantum downgrade attacks against IPsec

Cloudflare helped develop an IETF extension that authenticates the full IKEv2 transcript and prevents attackers from downgrading post-quantum IPsec tunnels.

Is your domain using post-quantum encryption? Now you can see for yourself

HTTP Analytics, Log Explorer, and Logpush now show whether requests negotiated post-quantum key exchange, giving customers evidence they can inspect and report.

Enforce positive security with Cloudflare Application Profiles

Application Profiles learns the expected structure of HTTP requests so customers can identify deviations and enforce what valid application traffic should look like.

We tested our own WAF with frontier AI models. Here's what we found

An adaptive AI red-team system found WAF detection gaps across six attack categories, helping us improve normalization and managed rules for customers.

Introducing Threat Signals: agentic skills for open-source threat intelligence, free for every Cloudflare account

Threat Signals turns open-source reporting into structured indicators and connects the context to WAF rules, while the Threat Events Platform expands to every account.

Adaptive application security for the AI era: how Cloudflare connects code, traffic, and intelligence to stop attacks

Our application-security framework connects discovery, governance, runtime protection, investigation, and response in a continuous learning loop.

Wednesday, September 30 – Powering the agent economy

With our announcements of Pay Per Use and the release of our Monetization Gateway in beta, we shared how we’re building support for a new economic model that empowers creators to monetize their content and services.

What

In a sentence…

The Internet has a second audience

AI agent requests have grown rapidly, and our strategy helps creators see agents, set terms for access, and get paid when agents use their work.

Cloudflare Containers, rebuilt to scale agent sandboxes

Containers now has faster startup, flexible image and instance selection, new scheduling controls, and filesystem snapshots for persistent agent workspaces.

Monetization Gateway beta: charge AI agents for consumption with HTTP 402

Monetization Gateway lets sellers put a price on resources behind Cloudflare and collect agent payments using HTTP 402 and x402.

Pay Per Use: when AI uses your work, you should get paid

Pay Per Use gives enrolled publishers usage reports, billing, and payouts when verified AI buyers use their content.

Simplifying domains for people and agents

A new domain-search experience and expanded Registrar APIs make it easier for both people and agents to search, register, transfer, and manage domains.

Identify AI model overuse with User Insights

AI Gateway User Insights identifies tasks, model fit, and overuse, so teams can understand where a smaller or less expensive model may work.

Detect and send production issues straight to your agent

Issues groups Workers errors and sends the relevant stack traces, logs, and traces to coding agents or any webhook for faster investigation.

Cut your AI spend with AI Gateway's Auto Router

Auto Router classifies each request at the edge and sends it to a suitable model, reducing cost while preserving response quality.

Cloudflare Impact reaches $100 million in donations

Initiatives including Project Galileo, the Athenian Project, and Cloudflare for Campaigns have now delivered more than $100 million in donated services.

Thursday, October 1 – Bringing more of the developer stack to Cloudflare

We expanded what is possible to achieve on Cloudflare’s platform with the general availability launch of Cloudflare Basin, our data analytics platform, the launch of K2, a durable serverless event stream, and the announcement of our new contest — inviting developers to build a Git platform designed for agentic development.

What

In a sentence…

Introducing Cloudflare Basin: an open, serverless data platform, now generally available

Basin is now generally available, giving developers a serverless platform built on Apache Iceberg and R2 for ingesting, managing, and querying large datasets.

Support for modern cryptographic algorithms in Workers

Workers adds opt-in native Web Crypto support for ML-KEM and ML-DSA, giving developers post-quantum primitives without bundling their own implementations.

AI Search is now generally available

AI Search reaches general availability with visual search, OCR for scanned PDFs, larger files, and support for any chat model.

We want you to build the next Git platform on Cloudflare

Artifacts enters open beta and a new competition invites developers to build a Git platform designed for the era of AI agents.

Announcing Cloudflare K2: serverless event streams

K2 provides durable, ordered event streams on R2, separating producers and consumers without the operational overhead of managing broker clusters.

Cloudflare OS: your company's agent workspace, managed for you

Cloudflare OS provides an agent workspace connected to an organization’s data and systems, with a waitlist open for fully managed deployments.

Introducing Workers KV Instant – powered by Quicksilver

Workers KV Instant delivers sub-two-millisecond p99 reads and fast global replication across more than 300 locations using the familiar Workers KV API.

One year later: Sovereign AI and the fight for choice

We are expanding local open-source model choice and model-agnostic security tools, so nations can pursue AI sovereignty without isolation.

Introducing Clef: our open-source decision models, and new RL fine-tuning platform

Clef and Clef-flash are open-source decision models for fast classification and agent workflows, accompanied by a platform for reinforcement-learning fine-tuning.

Friday, October 2 – Delivering a faster, simpler Internet for everyone

We wrapped up the week with major updates to Cloudflare Observability, alongside adding Cloudflare Traces, network performance improvements that make Cloudflare faster, and an announcement on how we’re supporting civil society organizations.

What

In a sentence…

8 major updates to Cloudflare Observability

Eight updates bring logs, traces, analytics, alerts, dashboards, querying, and telemetry export into one observability platform with simpler pricing.

Introducing Cloudflare Traces: follow requests through our entire platform

Cloudflare Traces provides request-level visibility across security rules, transformations, cache, Workers, services, and origins without requiring an agent or SDK.

Updates on our pledge to make Cloudflare features accessible to everyone

One year after our pledge, Logpush, multi-account governance, higher platform limits, and other capabilities are available to more customers across plans.

Announcing Cloudflare OHTTP Gateway – expanding access to Cloudflare's privacy-preserving infrastructure

A self-serve OHTTP Gateway enters closed beta, while Privacy Gateway becomes Cloudflare OHTTP Relay to distinguish the two roles.

Follow the thread: a new dashboard to investigate account abuse

Account Abuse Protection uses stateful analysis and privacy-preserving Hashed User IDs to help teams investigate credential stuffing and fake-account creation.

Protected Quick Tunnels: simple accountless authentication for your next dev project

Quick Tunnels now support email authentication, letting developers share a local application with selected people or domains without requiring Cloudflare accounts.

Building for good: How civil society organizations are automating on Cloudflare

Civil society organizations are using Cloudflare’s developer platform to automate and scale work that protects human rights and the public interest.

2026 Birthday week: network performance update

Using an expanded real-user measurement methodology, Cloudflare now ranks as the fastest provider across 74% of the top 1,000 networks.

Introducing Web Search API via AI Gateway

AI Gateway’s Web Search API brings current web context from multiple providers into model calls through REST APIs, Workers bindings, or customer-managed keys.

Streamline: custom video pipelines with Cloudflare Stream and Workers

Streamline is an open-source example for building continuous video pipelines by combining Workers, Durable Objects, and a containerized media engine.

Building the Internet’s next chapter together

Across this week’s announcements, we kept returning to a consistent theme: the Internet should continue to open up more opportunities for people to create, contribute, and succeed. That means open tools developers can shape, security that keeps pace with new threats, a fairer exchange between agents and the people whose work they use, and infrastructure designed for the agentic Internet.

For 16 years, we have been building alongside developers, creators, researchers, customers, partners, and open-source communities. Your ideas, feedback, and willingness to challenge us have shaped Cloudflare, and that collaboration matters now more than ever.

Response Overview and Colonel Clustered – Grouping Burp Responses by Content

Post Syndicated from Darknet original https://www.darknet.org.uk/2026/10/response-overview-colonel-clustered-burp-response-grouping/

Two Burp Suite extensions that group responses by content: Response Overview, a BApp with a threshold you set, and Colonel Clustered, which picks its own.

Kernel prepatch 7.3-rc6

Post Syndicated from corbet original https://lwn.net/Articles/1098476/

The 7.3-rc6 kernel prepatch is out for
testing. Linus said:

Next week might look a bit different: we’ve got the annual
maintainer summit and the Linux plumbers conference going on , so
I’ll be on the road, as will a number of other maintainers. That
may or may not end up changing the stats for -rc7. But it’s
unlikely to affect the release schedule, although the fact that I
have my yearly family vacation the week after that might make the
next merge window a bit wonky.

The collective thoughts of the interwebz