Post Syndicated from The Atlantic original https://www.youtube.com/watch?v=GfXIxpD_1vY
How Property Finder automated incident management with AWS DevOps Agent
Post Syndicated from Nada Tlohi original https://aws.amazon.com/blogs/devops/how-property-finder-automated-incident-management-with-aws-devops-agent/
When a production service starts saturating the CPU at 1 AM, every minute counts for incident management. For Property Finder, a production incident could mean failed searches, frustrated users, and direct revenue impact. Property Finder is the leading property portal in the Middle East and North Africa (MENA), serving millions of property seekers across five markets.
Before adopting AWS DevOps Agent, incident response followed a familiar pattern: an alert fires, an on-call engineer wakes up, spends 20–40 minutes correlating metrics across tools, manually documents findings, and opens a fix. Mean Time to Resolution stretched to 2–3 days for non-critical issues.
Today, that entire workflow runs autonomously. From alert to root cause analysis, Slack notification, Jira ticket, on-call phone call with context, and auto-remediation pull request (PR), the full lifecycle completes in 14 minutes. This post walks through the implementation and shows how a separate custom agent that automatically generates code fixes is the key differentiator.
The business problem
Property Finder runs a distributed microservices architecture on Amazon Elastic Container Service (Amazon ECS) fronted by Application Load Balancers (ALBs). When infrastructure issues occur, the impact is immediate: users see failed searches, agents cannot update listings, and revenue is directly impacted during peak hours.
The traditional workflow had three gaps:
- Detection lag. Non-critical anomalies could go undetected for days.
- Context switching. Engineers bounced between five or more tools per incident.
- Knowledge silos. Runbooks lived in people’s heads, not automation.
Solution architecture
Property Finder’s implementation connects AWS DevOps Agent at the center of a three-tier pipeline: Detection and Trigger, Autonomous Investigation, and Event-Driven Output.
The numbered steps correspond to the data flow in Figure 1:
- ECS CPU spike triggers an Amazon CloudWatch Alarm. CloudWatch Metrics Insights monitors service health across all ECS clusters. When sustained CPU exceeds 98%, the alarm transitions to ALARM state.
- AWS Lambda formats and HMAC-signs the payload. Triggered directly by the CloudWatch alarm action (which fires only on ALARM state transitions), AWS Lambda enriches the payload with service metadata, signs it with HMAC-SHA256 using credentials from AWS Secrets Manager, and POSTs to the webhook.
- The agent begins autonomous investigation. Parallel subagents query ECS metrics, AWS CloudTrail, ALB traffic patterns, and Grafana telemetry (Prometheus, Loki, Pyroscope). The agent reads relevant source code from GitHub for correlation.
- Findings post to Slack in real time. The native Slack integration posts investigation progress to #incidents. The full root cause analysis, impact assessment, and mitigation plan appear at the end of the thread.
- Investigation Completed event fires to Amazon EventBridge. Amazon EventBridge triggers an orchestrator Lambda that fans out to three independent targets simultaneously.
- Lambda creates a Jira ticket with the full root cause analysis. The Lambda retrieves the investigation summary from journal records and creates a prioritized ticket with root cause, severity, and affected service.
- Grafana IRM pages the on-call engineer by phone. A Lambda posts a Grafana Alerting-compatible payload to the IRM webhook. The escalation chain calls the engineer with full investigation context: what broke, why, and the recommended fix.
- The remediation agent opens a GitHub PR with the auto-fix. It receives the root cause, generates a Terraform or code fix, and opens a Draft PR through a GitHub Model Context Protocol (MCP) server. Engineers review before merging.
A real incident
The example-service, Property Finder’s core property search microservice serving millions of queries per day across five MENA markets, experienced CPU saturation at 99.11%. The pipeline resolved it end-to-end in 14 minutes.
1:21 AM │ Alarm fires (ECS CPU > 98%)
1:22 AM │ Investigation starts + Slack posted
1:22 AM │ 4 parallel subagents launched
1:32 AM │ Root cause identified
1:33 AM │ Jira ticket [redacted] created
1:34 AM │ On-call paged via phone call
1:35 AM │ GitHub PR [redacted] opened with fix
The detection Lambda handles three tasks: (1) retrieves the webhook secret from AWS Secrets Manager, (2) enriches the CloudWatch alarm event with ECS service metadata (cluster name, service name, task count), and (3) HMAC-signs the payload before POSTing to the webhook. The key authentication pattern:
Figure 2: Four parallel subagents investigating ECS, CloudTrail, ALB, and Grafana data sources simultaneously
Root cause: Conflicting CPU and memory target-tracking autoscaling policies combined with an insufficient capacity floor. The service had both a CPU policy (target 70%) and a memory policy (target 75%). Actual memory usage sat at 3–8%, creating a persistent conflict between the two policies.
With MinCapacity set too low, the service could not sustain the task count needed to absorb CPU load. The resulting instability (22+ scaling flips observed) prevented stable scale-out, leaving the service effectively pinned at two tasks with no CPU headroom.
This is a common organizational issue: teams configure both scaling dimensions without realizing the interaction, especially when the capacity floor is not sized for baseline traffic. The agent identified the pattern in 10 minutes, a task that typically requires senior engineers with deep scaling expertise and hours of CloudWatch metric correlation.
At 1:34 AM, the on-call engineer received a phone call through Grafana IRM with the complete investigation context. No need to wake up and hunt for root cause across dashboards.
Mitigation plan generated: (1) Remove the memory-based scaling policy, (2) raise MinCapacity to handle baseline traffic, (3) implement CPU-only target tracking at 70%. This plan was passed to a separate custom agent for remediation.
Remediation
Remediation is the key differentiator in this pipeline. It is a dedicated remediation agent (pr-creation-agent) invoked only after investigation completes. AWS DevOps Agent enforces read-only access to infrastructure through a per-session permission guardrail. Effective permissions are the intersection of the execution role’s IAM policy and the guardrail, and write actions are excluded.
The split separates concerns: investigation stays within that read-only envelope, whereas the remediation agent is scoped to a GitHub MCP server as its only external integration. Safety at the remediation layer does not rely on the agent’s built-in directed actions approval mechanism. Instead, two controls enforce the boundary. First, the remediation agent is a separate, narrowly scoped agent with access limited to GitHub MCP. Second, every output is a Draft pull request that requires human review and merge before taking effect. The GitHub MCP connection is authenticated with a fine-grained personal access token scoped to the specific infrastructure repositories, with an expiration and rotation policy. No elevated IAM role or additional agent permissions are required.
How it works
When the “Investigation Completed” Amazon EventBridge event fires, a Lambda orchestrator invokes the remediation agent with the investigation ID. The agent then:
- Reads findings from journal records to understand the root cause and recommended fix.
- Maps the AWS account to the correct repository. Property Finder has six infrastructure repos for different teams (B2B, B2C, core-platform, growth, data-engineering, shared-infra). The agent extracts the account ID from resource ARNs and routes to the right repo. This mapping is validated through automated tests and updated as new accounts or repositories are onboarded.
- Checks for duplicate PRs by searching existing PR titles and bodies for the investigation ID. If a matching PR exists, it reports the URL and exits without creating a duplicate.
- Reads the relevant Terraform files through GitHub MCP (GITHUB-MCP_get_file_contents), identifies the exact changes required, and plans the fix.
- Creates a feature branch (fix/{investigation_id}), commits the changes, and opens a Draft PR with a structured template including problem summary, root cause, changes made, and a testing checklist.
AWS also supports remediation through Kiro CLI with AWS CodeBuild or Kiro-ready prompts. Property Finder chose an approach that fits their multi-team repository structure: the remediation agent runs entirely within the Agent Space (the managed environment where custom agents execute), uses GitHub MCP for repository access, and maps multiple repositories to different teams automatically.
The orchestrator Lambda is triggered by the “Investigation Completed” Amazon EventBridge event. It first retrieves the investigation findings from journal records, then fans out to three targets simultaneously. Target one creates a Jira ticket with the full root cause analysis, severity, and affected service. Target two posts a Grafana Alerting-compatible payload to the Grafana IRM webhook to trigger phone call escalation. Target three invokes the remediation agent through the CreateChat and SendMessage API, passing the investigation ID and root cause context so it can generate the appropriate code fix.
Figure 8: GitHub PR [redacted] generated by the remediation agent with a structured problem, root cause, and changes template
Figure 9: Terraform diff showing the new CPU-only scaling policy replacing the conflicting memory configuration
The PR is always opened as Draft. Engineers review, run terraform plan, validate in staging, and merge. The agent never auto-merges.
Results
| Metric | Before | After | Improvement |
| End-to-end time | Hours to days | 14 minutes | >88% reduction |
| Investigation | 20 to 40 min (manual) | 10 min (autonomous) | 50–75% reduction |
| Documentation | Manual, incomplete | Auto-generated root cause analysis + Jira | 100% documented |
| Remediation | Manual PR by engineer | Auto-fix PR + review | Minutes to code fix |
Cost considerations: Each incident invokes two agent sessions (investigation + remediation) with up to four parallel subagents. Billing is based on agent minutes. For detailed pricing, see the AWS DevOps Agent pricing page. We recommend reviewing pricing for all services used in this architecture.
“We now rely fully on AWS DevOps Agent to identify infrastructure-related issues. It has helped us identify multiple complex issues without even opening a support ticket. Even if we had raised tickets, it would likely have taken support engineers hours to find the root cause, whereas we resolved these issues in minutes.”
— Yasitha Bogamuwa, Cloud Engineering Manager, Property Finder
Getting started
Prerequisites:
- An Agent Space configured in your account.
- Amazon CloudWatch and AWS CloudTrail enabled for observability.
- Slack, Grafana, and GitHub connected as capabilities.
- Infrastructure resources tagged for topology mapping.
Step 1: Configure the webhook trigger. Set up CloudWatch Alarm action to invoke a Lambda function. The Lambda enriches the payload, HMAC-signs it, and POSTs to your Agent Space webhook endpoint.
Step 2: Set up event-driven outputs. Create an Amazon EventBridge rule for “Investigation Completed” events (source: aws.aidevops). Add Lambda targets for Jira, Grafana IRM, and optionally a remediation custom agent.
Step 3: Test end-to-end. Trigger a test alarm and verify the full pipeline: investigation starts, Slack posts, Jira ticket created, on-call paged, and PR opened.
For a similar integration pattern with Salesforce, see Automating Incident Investigation with AWS DevOps Agent and Salesforce MCP Server on the AWS DevOps Blog.
Clean up
This post describes an architecture pattern implemented by Property Finder. If you deployed test resources while following along, remember to delete any CloudWatch Alarms, Lambda functions, Amazon EventBridge rules, and Agent Space configurations to avoid ongoing charges. For a full list of resources and associated costs, review the pricing pages for each AWS service used in this architecture.
Conclusion
Property Finder’s implementation shows that autonomous incident management works in production today, with their pipeline running since early 2026. The agent never auto-merges. Human review remains in the loop by design: the agent accelerates, the engineer decides. The on-call engineer wakes up to a phone call with the root cause already identified, a Jira ticket filed, and a PR ready for review.
Explore the AWS DevOps Agent documentation to get started with your own autonomous pipeline.
Related resources
- Getting Started with AWS DevOps Agent.
- Automating Incident Investigation with Salesforce MCP.
- Building an End-to-End Agentic SRE.
- Amazon EventBridge User Guide.
- Grafana IRM Documentation.
About the authors
Git v2.56.0 released
Post Syndicated from jake original https://lwn.net/Articles/1097213/
Version 2.56 of the Git distributed
version-control system has been released. It has 748 non-merge commits
since Git 2.55 was released back in
June; those commits came from 104 developers, 39 of whom are first-time
contributors. New features include a safer workflow for conflict
resolution, smaller path-walk repacks, a new git history drop
sub-command, and much more. LWN looked at Git
2.56 recently and the GitHub blog has a lengthy
look at 2.56 as well.
AWS European Sovereign Cloud: Demonstrating an independent operation
Post Syndicated from Stéphane Israël original https://aws.amazon.com/blogs/security/aws-european-sovereign-cloud-demonstrating-an-independent-operation/
On Saturday, October 24, 2026 we will conduct an exercise, demonstrating that the AWS European Sovereign Cloud can operate without depending on any infrastructure outside of the European Union (EU).
For several hours, the AWS European Sovereign Cloud will operate without a connection to the AWS Global Network backbone. The backbone is the private network that moves authorized AWS operational data between AWS locations without using the public internet. During the exercise, this traffic will securely reroute over the public internet.
The exercise will not affect service availability within the AWS European Sovereign Cloud, other AWS Regions, or private connectivity through AWS Direct Connect. Customers may experience brief connectivity disruptions as traffic moves onto a separate network route at the beginning or the end of the exercise, after which normal connectivity resumes.
The operational team, composed entirely of EU residents within the EU, will execute the exercise using only the hardware and software resources of the AWS European Sovereign Cloud. The AWS European Sovereign Cloud Managing Directors called for this exercise to showcase its operational independence.
An independent cloud for Europe
The AWS European Sovereign Cloud is a new, independent cloud for Europe. Located in Brandenburg, Germany, its data centers are physically and logically separate from other AWS Regions, with a local in-EU copy of the source code. All customer content and customer-created metadata stay in the EU. It has no critical dependencies on non-EU infrastructure and is operated exclusively by EU residents.
In standard operations, the AWS European Sovereign Cloud uses two global systems. The first is the AWS Global Network backbone. The second is a dedicated system that the local EU team controls and supervises to securely manage limited, controlled transfers of operational AWS data.
Neither is operation-critical, and neither affects the sovereignty assurance of the AWS European Sovereign Cloud. The AWS European Sovereign Cloud can operate independently at any time without a connection to these global systems, and on October 24 that’s what the team will demonstrate.
Built to meet regulatory standards
This exercise will produce verifiable technical and operational evidence that the AWS European Sovereign Cloud can operate independently within the EU. The exercise is designed to be consistent with the objectives of the European Commission’s EU Cloud Sovereignty Framework (CSF) and the criteria of the C3A framework from Germany’s Federal Office for Information Security (BSI). These frameworks set out objectives and criteria for assessing whether cloud services can be provided independently and autonomously.
We designed the AWS European Sovereign Cloud for regulated customers and the public sector across the EU. The AWS European Sovereign Cloud: Sovereign Reference Framework (ESC-SRF) gives our customers and partners a comprehensive set of evidence points, maps to controls, artifacts, and other elements regulators and compliance authorities need to accelerate their adoption of the AWS European Sovereign Cloud. The results of this exercise will provide additional evidence for their compliance and assurance packages.
Standalone and fully secure
The AWS European Sovereign Cloud runs connected to the AWS Global Network backbone because it delivers superior performance, capacity, reliability, security, and cost savings to customers. That includes always-on encryption and distributed denial of service (DDoS) defenses; and the backbone can’t decrypt or see the encrypted data that AWS European Sovereign Cloud customers send and receive. While the backbone delivers these benefits day-to-day, the AWS European Sovereign Cloud can continue to operate independently, with the appropriate security controls in place.
During the exercise, instead of using the AWS Global Network backbone, the AWS European Sovereign Cloud will exclusively use its dedicated internet connectivity from European internet service providers. This will provide connectivity to the worldwide internet.
Whenever traffic moves between internet links, there’s a small window of limited disruption called convergence, a short time when other non-AWS networks change their routing information to reflect the change. This could happen at the beginning of the exercise, when traffic moves to dedicated AWS European Sovereign Cloud internet providers, and at the end of the exercise, when traffic moves back to the AWS Global Network backbone.
Customer data stays in the EU
AWS has committed to not moving AWS European Sovereign Cloud customer content and customer-created metadata outside of the EU. Only certain data, which is neither customer content nor customer-created metadata, such as AWS operational data, leaves the EU. We use a dedicated system to securely manage these limited, controlled transfers under the control and supervision of the local EU team. We’re rigorous about what the system transfers. It accepts vetted source code mirroring and software updates, and transfers out very limited and approved routine information. During the exercise, the AWS European Sovereign Cloud team will disable the system entirely, confirming that the AWS European Sovereign Cloud continues to operate independently without it.
Learn more
AWS will share an update after the exercise with regulators and customers. To learn more about the AWS European Sovereign Cloud’s design and digital sovereignty controls, visit aws.eu. If you have questions about this exercise or would like to discuss how it may impact your workloads, reach out to AWS Support or contact your AWS Account team.
AWS Weekly Roundup: GPT-6 Sol and Luna, Claude Opus 5.5 on Amazon Bedrock, Strands harness, and more (September 28, 2026)
Post Syndicated from Daniel Abib original https://aws.amazon.com/blogs/aws/aws-weekly-roundup-gpt-6-sol-and-luna-claude-opus-5-5-on-amazon-bedrock-strands-harness-and-more-september-28-2026/
If there’s one theme that defined last week, it’s choice. The frontier models keep arriving, and the interesting question is no longer just “how smart is it?” but “which model fits this step, at this cost, at this latency?” That’s exactly what landed on Amazon Bedrock over the past few days: GPT-6 Sol and GPT-6 Luna from OpenAI, giving you two new points on the intelligence-versus-efficiency curve, and Claude Opus 5.5 from Anthropic, the first of the Claude 5.5 family.

GPT-6 Sol is built for the demanding, recurring work of development and operations, while GPT-6 Luna makes focused, repeatable tasks practical at high volume, and both ship at significantly lower pricing than their GPT-5.6 predecessors. Claude Opus 5.5, meanwhile, does more with fewer tokens than Opus 5 and is tuned for agentic coding and long-running tasks. What I like about all three is that they push toward the same idea: match the model to the job instead of reaching for the biggest one every time. The other thread was observability catching up to this agentic world, including a launch I had the pleasure of writing about myself.
Now, let’s get into this week’s AWS news…
Last week’s launches
Here are some launches and updates from this past week that caught my attention:
- Introducing Amazon CloudWatch Omni – You can now observe your applications and AI agents together in a single, collaborative experience. Amazon CloudWatch Omni is built on OpenTelemetry, so your existing telemetry shows up with nothing to reconfigure, and your whole team reaches it through one URL with enterprise SSO — no console access required. It auto-discovers your services, maps dependencies, and brings AWS DevOps Agent into investigation sessions to correlate signals and trace root causes. There’s a companion post on the agent-observability side, a deeper dive on the AWS Cloud Operations blog on what observability for the AI era looks like, and the announcement on What’s New with the specifics. If you want the bigger picture, Matt Wood’s Wrong, not broken is a great read on why correctness now has to be measured at the level of the run.
- Enhanced custom event buses in Amazon EventBridge – Amazon EventBridge now offers an enhanced custom event bus purpose-built for organizations scaling event-driven applications across teams and accounts. You can now deploy a single centralized bus shared across every account in your organization through AWS RAM, with optional event ordering, a simplified Subscriber resource that bundles filtering, targets, and retries, content-based deduplication, and synchronous invocation for targets like AWS Lambda. A new ingress/egress pricing model replaces the compounding cross-account routing charges of multi-bus setups, and your existing buses keep working unchanged as “classic.”
- Amazon SageMaker HyperPod Inference Gateway – You can now front your LLM inference on Amazon SageMaker HyperPod with a Kubernetes-native, GPU-aware routing layer that deploys as a single Amazon EKS managed add-on with zero application changes. Instead of round-robin load balancing, it routes on real-time inference signals — KV cache utilization, queue depth, prefix cache hits, predicted latency, and more — cutting first-token latency by up to 82% in mixed-hardware and bursty scenarios. It works with any OpenAI-compatible model server, including vLLM and SGLang.
- AI agent skills for AWS End User Messaging and Amazon SES – You can now build and send messages by asking your AI coding agent in plain language. Amazon SES and AWS End User Messaging publish AI agent skills for the AWS MCP Server, giving your agent step-by-step, validated guidance for tasks like verifying a sending identity, sending a production email, or building a branded RCS agent with cards and buttons. The skills work with Claude Code, Codex, Cursor, and Kiro, so you can complete messaging workflows without hopping between docs and console screens.
For a full list of AWS announcements, be sure to keep an eye on the What’s New with AWS page.
Other AWS news
Here are some additional posts and resources that you might find interesting:
- Introducing Strands harness – The Strands Agents team released Strands harness, a fully assembled, general-purpose agent harness you can run locally or deploy anywhere, under Apache 2.0. It takes one line of Python or TypeScript to wire up your model of choice across Amazon Bedrock, Anthropic, OpenAI, Google, or a local Ollama model, and it ships with sensible defaults for prompt caching and context management (truncating bulky tool results, compacting when the context window fills up, and keeping memory across runs). The team reports it costs about 28% less than comparable harnesses on the same models while holding accuracy steady.
- AWS named a Leader in the 2026 Gartner Magic Quadrant for Container Management – Gartner recognized AWS as a Leader for the fourth consecutive year. The post is a nice tour of where containers are heading, from Amazon ECS Express Mode and Amazon EKS Auto Mode to the 99.99% availability SLA on the EKS Provisioned Control Plane — with containers increasingly becoming the default substrate for how AI agents are built and run.
- Announcing the new AWS Reimagine report on AI – The AWS Executive in Residence team spent nine months interviewing 154 leaders across 27 countries about what separates organizations that turn AI into value from those that don’t. The report is candid (including where AI hasn’t worked at Amazon), and the recurring insight is that once building gets fast, the bottleneck moves to deciding, funding, and governing the work. Well worth a read if you’re thinking about how your teams adopt AI in practice.
For a full list of AWS blog posts, be sure to keep an eye on the AWS Blogs page.
Upcoming AWS events
Check your calendar and sign up for upcoming AWS events:
- AWS re:Invent – AWS re:Invent returns to Las Vegas from November 30 to December 4, and session times, locations, and speakers are live. Reserved seating opens October 6, so register now and be ready to claim your spot in chalk talks, workshops, and builders’ sessions.
- AWS Summits – With re:Invent on the horizon, the Summits are wrapping up for the year. The last stop is Dubai (September 30) at the Dubai World Trade Center, with 60+ sessions, an AWS Village, and hands-on workshops.
- AWS Community Days – Community-led conferences planned and delivered by community leaders. Upcoming events include ComSum Manchester, UK (October 1) and Rome, Italy (October 2).
Join the AWS Builder Center to connect with builders, share solutions, and access content that supports your development. Browse here for upcoming AWS-led in-person and virtual events and developer-focused events. That’s all for this week. Check back next Monday for another Weekly Roundup!
— Daniel Abib
Audit trails for autonomous agents with AWS DevOps Agent
Post Syndicated from Ben Peterson original https://aws.amazon.com/blogs/devops/audit-trails-for-autonomous-agents-with-aws-devops-agent/
Autonomous agents need audit trails. AWS DevOps Agent (DevOps Agent) investigates production incidents and proposes or applies fixes on your behalf. Every operation and security review then raises the same two questions: what did the agent do, and how do you understand its impact?
AWS DevOps Agent maintains an immutable, step-by-step record of its own reasoning and actions. This post shows how to capture the agent’s full operational trail using the agent journal, recommendations, Amazon EventBridge lifecycle events, and AWS CloudTrail. We then wire them into an audit pipeline built on Amazon EventBridge, AWS Lambda, and Amazon Simple Storage Service (Amazon S3).
By the end, you will have deployable audit patterns that show, for any investigation the agent runs, what it concluded, what it recommended, when it ran, and whether the fix landed.
Why auditing an autonomous agent is different
CloudTrail records the API calls made in your account, but an autonomous agent adds reasoning that CloudTrail doesn’t capture. “The agent ran a metric query” is far less valuable than “the agent concluded the Lambda was timing out because its security group blocks egress to the database.” The latter is a decision, and that’s what an agent audit needs to capture.
The four surfaces
AWS DevOps Agent exposes four surfaces. Two capture the agent’s output, what it found and what it advises, and two capture context: when it ran, and who configured the agent and its permissions.
The agent journal
The agent journal (API) is the heart of the audit trail. For every execution, AWS DevOps Agent records an ordered, immutable log of its reasoning step, sub-agent it dispatches, observations, findings, and root-cause summary. Journal entries cannot be modified once written, making them resistant to prompt injection and trustworthy as an audit record.
Each record carries a recordType: symptom, observation, finding, and investigation_summary / investigation_summary_md are what matters for audit. This is the surface you archive per investigation.
Recommendations polling
Recommendations (API) are cross-incident preventative advice. The agent generates these on a schedule through a goal, and each recommendation carries a status and a version. Each evaluation run writes new records rather than updating the previous run’s, so advice that persists week over week appears as a series of records. The superseded ones remain at whatever status they last held. “The agent recommended X, the same failure recurred Y weeks later, and here is every version of that advice in between” is something you reconstruct from the archived snapshots, because the API returns current and superseded records together. Recommendations have no Amazon EventBridge event. You capture them by polling on a schedule.
Amazon EventBridge lifecycle events
Amazon EventBridge is how you capture lifecycle transitions in real time. A successful investigation produces Created, In Progress, and Completed events. Each carries the execution_id you need to fetch the journal and a summary_record_id pointing at the root-cause summary. Investigations can also end as Failed, Timed Out, or Canceled, and mitigations emit their own parallel set.
AWS CloudTrail
CloudTrail records API calls made to the AWS DevOps Agent service and stamps agent-initiated service calls: invokedBy: aidevops.amazonaws.com. It doesn’t capture the agent’s investigation reads, the metric and log queries it runs while diagnosing an incident in your account’s trail. Use CloudTrail for control-plane accountability, and the journal for behavioral audit.
IAM: Action boundary
As with anything in AWS, the agent can only do what its AWS Identity and Access Management (IAM) role permits. During an investigation, AWS DevOps Agent assumes an Agent Space role. That role’s policies are the hard ceiling on its capabilities. You can inspect it directly:
The AWS-managed AIOpsAssistantPolicy is attached to the default role. As of policy version 15, 848 of its actions are reads except 6 read-oriented query lifecycle operations. The only actions that change anything come from a companion policy: support:CreateCase and a service-linked-role creation scoped to the Amazon Resource Name (ARN) of a single role.
Keep that role least-privilege, and your audit surface stays small by construction. If you enable agent actions, a later section covers the write path which uses a separate actions role.
The reference architecture
The agent produces output that arrives two different ways, and this shapes how you capture each:
| Agent output | Delivery | How you capture it | Latency |
| Investigation lifecycle | Push: Amazon EventBridge events | React to events (rules + targets) | Seconds |
| Recommendations | Pull: no event emitted. Generated on goal cadence | Poll list-recommendations on a schedule |
depends on your poll frequency |
The journal itself has no dedicated event, but the terminal lifecycle event carries the execution_id you need to fetch it. The journal is push-triggered, pull-retrieved: the event tells you when to look, and the API gives you what to archive.
Layer 1: Lifecycle capture. One Amazon EventBridge rule matching {"source":["aws.aidevops"]}, targeting an Amazon CloudWatch Logs (CloudWatch Logs) group directly. This durably records every lifecycle transition. Start here for operational visibility. If your primary goal is behavioral audit rather than operational visibility, deploy layer 2 alongside it.
Layer 2: Behavior capture. A second rule matches only terminal events and invokes a Lambda function. The function reads the execution_id from the event, calls list-journal-records, and writes the journal to Amazon S3. Subscribe to each terminal investigation and mitigation type. This is the layer that captures the agent’s decisions for the long term, including agent-based mitigations.
Layer 3: Recommendations snapshot. Because recommendations are generated on a schedule and have no event, capture them with an Amazon EventBridge Scheduler rule that invokes a Lambda function on a cadence (start daily). The function calls list-recommendations and writes each to Amazon S3, keyed on recommendation ID and version. It also calls list-goals in the same invocation, because a recommendation carries no field saying whether it is still current and the owning goal is the only thing that does. The journal captures what the agent found, and this layer captures what it advised and what you did about it.
Layer 4: Control-plane alerting. On your existing organization trail, alert on mutating aidevops.amazonaws.com events including UpdateApprovalAction, which is produced on elevated actions. This is your tripwire for changes to the agent itself.
Layer 5: Query. AWS Glue Data Catalog tables and an Amazon Athena (Athena) workgroup over the archived journals, recommendations, and goals.
Querying the archive: AWS Glue and Athena
The sample implementation overlays an AWS Glue Data Catalog and an Athena workgroup on the Amazon S3 archive. Three external tables cover the full archive. The journals table uses Athena partition projection, and Hive-partitioned by agent space and date:
Volume of recommendations is low (tens to hundreds of objects), so a flat external table over the recommendations/ prefix is sufficient. Athena recurses subdirectories by default, picking up every versioned snapshot.
The result bucket has Amazon S3 Object Lock but Object Lock prevents Athena from managing its own query-result objects. The query layer deploys a dedicated results bucket with a seven-day lifecycle rule for ephemeral query outputs.
Access control
Use IAM to control access. Investigation journals contain the agent’s full reasoning about your infrastructure. Scope your IAM permissions on the Athena workgroup, AWS Glue database, and on the archive bucket itself since bucket read access bypasses Athena entirely. Scope all three to your audit and operations teams.
To find all findings from the past 7 days for a specific resource:
The Athena workgroup integrates with Amazon Quick or any business intelligence tool that speaks JDBC/ODBC. Additional examples are available in the sample repository.
Closing the loop: Correlating findings to actual changes
The capture layers record what the agent found and what it recommended. But did the recommended fix actually land? This requires connecting the agent’s output to the real infrastructure change that followed.
The sample implementation includes correlate.py, an on-demand operator CLI that takes an archived finding or recommendation, resolves the resource it references, and reports what changed, when, and who did it. The correlation is heuristic by looking at resource identity and a tight time window in minutes to produce reliable attribution. This is why the sample implementation pairs it with a deterministic engine for agent-initiated actions.
It works by pivoting through two services:
- AWS Config resolves the resource identity by using
select-resource-config, then pulls its configuration timeline fromget-resource-config-history. This shows the before/after state of the resource around the time of the agent’s finding. - CloudTrail looks up the write event that caused the change: who called what API, from where, and when. This attributes the change to a principal.
The output is a correlated record: the agent found X, the resource changed from state A to state B, and that change was made by principal Y at time T.
Because CloudTrail indexes resources by different identifiers depending on the service, you require a strategy registry. Examples are in the following table:
| Resource type | How CloudTrail indexes it | Lookup strategy |
| S3 bucket | Bucket name | By name |
| Lambda function | Function name | By name |
| Amazon Relational Database Service (Amazon RDS) instance/cluster | Full ARN (not the DB ID) | Build ARN from template |
| Amazon Elastic Compute Cloud (Amazon EC2) security group | Group ID as ResourceName |
By name, with a resource-type scan as fallback |
A naive “look up by resource name” works for Amazon S3 and Lambda but returns zero results for Amazon RDS (RDS). The strategy registry encodes the right ID per resource type.
Correlating agent actions
When an operator approves an elevated action, the service stamps the approval ID into the credential it mints, so the executed call carries that ID inside its own principal ARN (op.system.apr.<approvalId>). The sample implementation includes correlate_agent.py that uses this. Because the ID is present on both sides, the correlation is a join. The engine checks the executed call against the argumentPins the operator was shown at approval time, so you can prove the agent’s behavior.
| . | correlate.py |
correlate_agent.py |
| Pivots on | A resource the agent named | Agent’s approval ID |
| Correlation | heuristic | deterministic |
| Answers | Who changed? | Who approved, and did it match? |
| Dependency | CloudTrail and AWS Config |
CloudTrail |
Production considerations
Understand the data volume. Journal size scales with investigation complexity. As an example:
| Scenario | Journal size | API calls (pagination) | Notes |
| Minimal (single-service, shallow investigation) | ~65 KB | 2–3 pages | Quick symptom to finding arc |
| Typical (multi-signal, 1–2 findings) | 250–340 KB | 65–106 calls | Typical investigations |
| Exhaustive (account-wide, high-priority) | ~428 KB | 150+ calls | Full cross-service correlation |
At 100 investigations/month at 300 KB average, you are storing roughly 30 MB/month of journal data.
Concurrency per agent space. By default, you can run three concurrent investigations per agent space. Additional requests queue as PENDING_START and start when a slot opens. The archival pipeline is unaffected because each terminal event triggers its own Lambda invocation. Refer to the AWS DevOps Agent Quotas page for future updates.
Paginate the journal. The journal API is server-paginated: pass limit, follow nextToken until it’s empty. A real incident’s journal can span several pages. Always loop.
Design for at-least-once delivery. Amazon EventBridge can deliver an event more than once. Key the Amazon S3 object on execution_id so a redelivery overwrites rather than duplicates, and attach an Amazon Simple Queue Service (Amazon SQS) dead-letter queue (DLQ) so a dropped terminal event is not lost.
Deploy per AWS Region and per account. Events land on the default bus in each Agent Space’s hosting account and Region. If you run agent spaces in multiple accounts, you must aggregate events to a central monitoring account for unified visibility. Refer to Amazon EventBridge cross-account document for further details.
Make the archive immutable. Enable Amazon S3 Object Lock and versioning. The sample implementation defaults to GOVERNANCE mode but for stronger compliance posture, use COMPLIANCE mode.
Warning: COMPLIANCE mode is irreversible. After it’s set, no principal (including the account root user) can delete or modify locked objects before their retention period expires. The only way out is closing the AWS account, and Object Lock itself can’t be disabled once enabled. Choose COMPLIANCE mode deliberately. If you use GOVERNANCE mode, enable CloudTrail data events on the bucket.
Encrypt your data. The sample implementation uses SSE-S3. If your compliance framework requires you to control and audit decryption events, use SSE-KMS with customer managed key.
The full loop
Here’s what a complete audit trail looks like for a single incident through resolution.
Step 1: Investigation. The agent investigates a failing Lambda function, concludes its security group restricts necessary egress, and writes the finding to the journal. Layer 2 archives the journal to Amazon S3.
Step 2: Recommendation. On its goal cadence, the agent generates a recommendation: “Update the security group egress rules to allow…” Layer 3 polls and captures it as recommendations/rec-a1b2c3.../v1.json with status PROPOSED. A later poll captures v2.json as the status changes.
Step 3: Engineer applies the fix. An engineer runs the suggested command. AWS Config records the new configuration item, and CloudTrail records the API call with principal, source IP, and timestamp.
Step 4: Correlation.
The agent found the problem, recommended the fix, and you can prove who applied it and when.
Step 5: A new investigation. A later investigation examines the same Lambda function, still erroring. The agent compares new advice against advice it has already given, and that comparison is semantic. But it compares against the recommendations currently attached to the goal, not against everything it has ever advised, and when the comparison is uncertain it keeps the two separate. Older advice drops out of that comparison set over time. Because you archived every recommendation and every finding with their resource identifiers, you can now compare across the full history:
The archive diagnosed the cause of the cause. Two recommendations, raised separately, on one resource, in one view. The agent’s own comparison covers the advice currently attached to the goal. The archive covers all of it. That is the feedback loop the audit trail adds.
Agent Actions changes Step 3’s actor, and the agent applies the fix directly. In the recommendation path, the human runs the command. In the elevated-action path, the human approves a specific call, and the agent executes it under a single-use session. correlate_agent.py uses a single-use session named for the approval (op.system.apr.<approvalId>), with invokedBy: aidevops.amazonaws.com rather than a time-window heuristic.
Operating the pipeline: Common failures
Always design for failure. Here are some common failures and how to detect and recover.
| Failure | Symptom | Detection | Recovery |
| Lambda timeout | No archive in Amazon S3. Event in DLQ | DLQ ApproximateNumberOfMessagesVisible alarm |
Increase timeout above the 2-minute default. Replay DLQ message which is idempotent on the execution_id key |
| Missed recommendation poll | Gap in recommendations/ prefix with a version number skipped |
Periodic reconciliation: compare Amazon S3 keys against list-recommendations response |
Re-run poll Lambda manually (idempotent) |
| Amazon EventBridge delivery failure | Missing lifecycle event in Layer 1 logs | Layer 2 archive exists without matching Layer 1 log entry | No data loss since journal already archived. Gap is in lifecycle visibility only |
| Amazon S3 write failure | Lambda errors spike. DLQ grows | Lambda error rate metric and DLQ alarm | Fix IAM/bucket policy. Replay DLQ (all messages are idempotent) |
| AWS Config recorder stopped | correlate.py returns no configuration history |
AWS Config recorder status alarm | Re-enable recorder. Note: historical gap is permanent for the stopped period |
| Journal API throttled | Partial archive. Lambda retries exhaust timeout | Lambda error logs showing throttling exceptions | Implement exponential backoff in the pagination loop. Increase timeout |
| Approval recorded but not executing | Approval exists in CloudTrail with no corresponding write | Join approvals to execution on the approval ID | None needed |
The highest value alarm is on the DLQ message count. A non-empty DLQ means a terminal event triggered, but the journal was not archived. Terminal events aren’t re-emitted, and the DLQ retains messages for 14 days. After that, the record is lost. The sample implementation ships this alarm at a threshold of 1, wired to an Amazon Simple Notification Service topic.
Run a reconciliation check weekly or monthly. Compare the execution_id values in the Layer 1 lifecycle log against the set of keys in the Amazon S3 journals/ prefix. Any ID in the logs but not in Amazon S3 represents a missed archive.
Limitations
Automated correlation – The current design requires a human to run correlate.py. Extend to a Lambda function that triggers on each new journal archive, cross-references the finding’s resource identifiers against the recommendations table, and alerts when a new finding touches a resource that was the subject of a prior recommendation.
Schema evolution – The Athena table definitions depend on the journal’s recordType values and content structure. If new record types appear, queries can return incomplete results without raising an error. Monitor for unknown recordType values. A query that returns zero findings for a week of active investigations is a signal that the schema moved.
Conclusion
Adopting an autonomous agent is a trust decision, and trust needs evidence. AWS DevOps Agent gives you the raw material: a journal of its reasoning, a real-time lifecycle event stream, a control-plane audit in CloudTrail, and an action boundary you can read straight from IAM. The pattern in this post assembles those into a durable, low-maintenance audit trail using services you already run.
The archive is more than compliance paperwork. With a persistent record of every finding and every recommendation, you can correlate across investigations and recommendations the agent no longer has in view, and against the present state of your infrastructure. That feedback loop is the difference between trusting the agent and understanding it.
Start with Layer 1. A single Amazon EventBridge rule to a log group gives you visibility into every investigation within minutes. Add the journal-archiving Lambda when you are ready to retain the agent’s decisions for the long term. Add the correlation layer when you want to prove that recommendations were acted on and catch the ones that created new problems.
Clone the sample repo to get started. It covers prerequisites, deploy steps, codebases, and teardown instruction. If you want the agent’s mitigations to become code, Automated incident remediation with AWS DevOps Agent and Kiro CLI builds a pipeline.
About the authors
How Meta Glasses Are Fueling the Anti-Surveillance Sentiment
Post Syndicated from The Atlantic original https://www.youtube.com/shorts/674xQ_hmkCs
Isolate email reputation in Amazon SES Mail Manager with tenant management
Post Syndicated from Abilashkumar P C original https://aws.amazon.com/blogs/messaging-and-targeting/isolate-email-reputation-in-amazon-ses-mail-manager-with-tenant-management/
When AnyCompany’s new IT outsource team misconfigured the email settings on 200 of the company’s multifunction printer/scanners, it had two bad outcomes. First, nobody received their scanned documents in their inboxes. Somewhat predictably, many users rescanned the same documents multiple times before creating support tickets. Second, the misconfiguration along with the multiple failed attempts resulted in a “bounce storm” that was quickly reported by a major email service provider, but unfortunately ignored by the IT team.
Within 48 hours, the bounce rate crossed the provider’s threshold. The company’s entire Amazon Simple Email Service (Amazon SES) account lost its sending reputation. Password resets, order confirmations, and service notifications from every business unit on the account started landing in spam or failing to deliver. The damage spread because every sender on the account, from the mission-critical billing system to the misconfigured printers, shared the same reputation score.
This is a preventable problem. With Amazon SES tenant management you can isolate email reputation per tenant inside a single account so one misbehaving sender cannot affect the rest. If you use Amazon SES Mail Manager for Simple Mail Transfer Protocol (SMTP) filtering, routing, archiving, or relay, you can activate tenant isolation. To do so, tag each message with the X-SES-TENANT header in your Mail Manager rule set. In this post, you will compare five architectural patterns for applying the X-SES-TENANT header, from static per-tenant endpoints to AWS Lambda driven runtime resolution.
This post complements Isolate email suppression per tenant with Amazon SES. That post explains how tenant-level suppression lists prevent cross-tenant bounce and complaint contamination, which is the “what happens after the message is tagged” story. This post focuses on the upstream problem: how to get the X-SES-TENANT tag onto messages when your senders are legacy appliances, printers, or applications that can’t set custom MIME headers. Together, the two posts cover the full tenant isolation pipeline, from tagging through delivery and suppression.
This post provides architectural guidance. For step-by-step implementation, refer to the Amazon SES documentation.
How SES tenant isolation works
Amazon SES tenant management isolates reputation per tenant inside a single Amazon SES account. Each tenant acts as a container organized around sending identities, configuration sets, and the resulting reputation metrics. Amazon SES attributes bounces, complaints, and Trust and Safety signals to the tenant, not the account, so a deliverability issue in one tenant doesn’t affect the others.
A critical benefit of tenant isolation: when one tenant’s reputation degrades beyond a threshold, Amazon SES can pause sending for that tenant only. Other tenants continue delivering normally. Without tenant isolation, a reputation issue affects the entire account. This pause-and-contain mechanism is one of the strongest reasons to adopt tenant management, especially for accounts with diverse sender types.
You associate a message with a tenant by passing the TenantName parameter on the Amazon SES API v2 SendEmail operation, or by adding an X-SES-TENANT Multipurpose Internet Mail Extensions (MIME) header to an SMTP message. For a detailed walkthrough of tenant management concepts, including identity ownership, the ses:TenantName AWS Identity and Access Management (IAM) condition key, and tenant-level suppression lists, see Improve email deliverability with tenant management in Amazon SES.
How Mail Manager works
Mail Manager processes inbound and outbound SMTP traffic through a pipeline of three components:
- Ingress endpoint: an authenticated SMTP endpoint that accepts connections from your senders. Mail Manager ingress endpoints handle SMTP only, not the Amazon SES API.
- Traffic policy: filters connections based on sender attributes (IP, TLS version, authentication) before messages reach rule processing.
- Rule set: an ordered list of rules. Each rule has conditions (match on envelope sender, recipient, source IP, or header values) and actions (Add header, Write to S3, Invoke Lambda, Send to internet, SMTP relay, Drop).
The “Add header” rule action is what makes tenant isolation possible for legacy senders: it injects the X-SES-TENANT SMTP header before the “Send to internet” action hands the message to Amazon SES for delivery.
With the “Add header” rule action inserted before the “Send to internet” action in the same rule, Mail Manager effectively tags the message with the SMTP header that defines the tenant. When Amazon SES processes the send, it reads the X-SES-TENANT header and attributes the message to the corresponding tenant.
Amazon SES performs tenant attribution only during send processing. A Send to internet action, or a Lambda function that calls SendEmail with the TenantName parameter or X-SES-TENANT header, activates tenant management. An SMTP relay action forwards to a third-party SMTP server (Google Workspace, Microsoft 365, or on-premises mail), so Amazon SES doesn’t process the send and tenant attribution doesn’t apply. Write to S3 and Drop don’t hand messages to Amazon SES, so they don’t activate tenant management either. This post describes flows that include a Send to internet action or a Lambda function calling SendEmail.
Understanding the outbound email flow
An outbound message flows from the SMTP client to the Mail Manager ingress endpoint, passes through the traffic policy and rule set, then routes through Amazon SES to the internet.
Figure 1: Outbound email flow from an SMTP client through Mail Manager to Amazon SES
Compare the patterns
Before diving into each pattern, use this table to identify which one fits your workload. You can then read only the pattern section that applies, or read all five for the full picture.
| Consideration | Pattern 1 | Pattern 2 | Pattern 3 | Pattern 4 | Pattern 5 |
| Works for legacy and appliance senders | — | Yes | Yes | Yes | Yes |
| Retrieve tenant from static value | Yes | Yes | Yes | Yes | Yes |
| Retrieve tenant from source IP or sender condition | — | Yes | Yes | Yes | Yes |
| Retrieve tenant from runtime lookup or body inspection | — | — | — | Yes | Yes |
| Records Send in Mail Manager log | Yes | Yes | Yes | — | — |
| Tenants per Region | Up to 10,000 | ~50 | 400 (per-tenant Send) or 1,560 (chained) | Up to 10,000 | Up to 10,000 |
One difference cuts across the patterns: where the tenant mapping lives determines what it takes to change it. Patterns 2 and 3 hold the mapping in rule-set configuration, so adding or removing a tenant is a rule-set edit and deployment (a control-plane change, not a data change). Patterns 4 and 5 resolve the tenant from a runtime source such as a database, so onboarding or offboarding a tenant is a data update that takes effect without a deployment. In Pattern 1, the sender supplies the tenant, so there’s no mapping to maintain in Mail Manager at all.
Pattern 1: The SMTP sender sets the header before Mail Manager
If the SMTP sender (a backend service, internal tool, or any application that can add a custom MIME header) sets X-SES-TENANT on the message before connecting to the Mail Manager ingress endpoint, the message arrives pre-tagged. The rule set only needs a Send to internet action.
Pattern 1 fits customers who already use Mail Manager for filtering, archiving, or compliance and whose sending applications can add one header at send time. You keep Mail Manager gateway capabilities without adding Add header or conditional logic to the rule set.
Pattern 2: Mail Manager adds a static header with Add header
A rule with Add header followed by Send to internet attaches a fixed tenant value to each message. This pattern fits a one-tenant-per-endpoint model: provision one authenticated ingress endpoint per tenant, give each tenant its own SMTP credentials, and attach a rule set that injects the tenant value.
For example, an enterprise provisions one endpoint for facilities-printer notifications and a second for corporate alerts. The Send to internet action’s IAM role grants permission only to that tenant’s Amazon SES identities, preventing a misrouted client from sending as another tenant.
You can group tenants behind one endpoint when they share a sending configuration. The header value and IAM scope live in the rule-set configuration, and no code runs at send time.
Pattern 3: Mail Manager derives the header from rule conditions
If multiple tenants share an endpoint but have stable distinguishing attributes (like source IP), one rule set handles each of them. Rule conditions match on envelope properties, and matching rules run an Add header action that sets X-SES-TENANT to the correct value.
Mail Manager rule sets allow 40 rules with up to 10 conditions and 10 actions per rule, but caps Send to internet and SMTP relay actions at 10 per rule set (counting every occurrence). One Send to internet per tenant rule tops out at 10 tenants.
To support more tenants, separate header-setting from delivery:
- Rules 1 to 39: Each matches a distinguishing condition and runs a single Add header action.
- Rule 40: A catch-all with no conditions and a single Send to internet action.
Each message matches at most one header-setting rule, picks up its tenant header, and passes through the catch-all. The effective ceiling is now 39 tenants per rule set with one Send to internet action and one IAM role.
Scale limits of Pattern 3
Pattern 3’s ceiling depends on how you structure the rule set. Two cases:
Case A: Chained structure (39 Add header rules + 1 Send to internet rule): Each rule set uses one Send to internet action, so the 10-action cap isn’t binding. Capacity is 39 tenants per rule set × 40 rule sets per Region = 1,560 tenants per Region.
Case B: Per-tenant Send to internet (each tenant rule has its own Send action): The 10-action cap binds at 10 tenants per rule set. Capacity is 10 tenants per rule set × 40 rule sets per Region = 400 tenants per Region.
The two cases trade off scale against IAM scoping. Case A shares one IAM role across all tenants in the rule set. Case B gives each tenant its own IAM role at the cost of 4× fewer tenants.
Amazon SES supports up to 10,000 tenants per account (adjustable). Workloads that exceed a few hundred tenants, or need runtime tenant changes, can use Pattern 4 or Pattern 5.
Pattern 4: Mail Manager calls Lambda for runtime tenant resolution
Some tenant values require runtime logic, such as a database lookup on the sender IP, an external policy service, or content inspection. For these cases, the Mail Manager Invoke Lambda action runs a Lambda function inside the rule chain.
The Lambda event carries only metadata (headers, envelope sender, recipients, verdicts), not the MIME body. The function also can’t modify the message for downstream actions. Lambda must therefore handle delivery.
The rule writes the raw MIME to Amazon S3 with Write to S3, then invokes the Lambda function with the message ID. The function fetches the object and determines the tenant through the runtime logic your workload requires. That logic might be a database lookup (for example, an Amazon DynamoDB query), a call to an external policy service, or inspection of the message body. It then calls the Amazon SES API v2 SendEmail operation, passing the resolved tenant in the TenantName parameter. Delivery permissions live on the function’s execution role, which carries the ses:TenantName condition key.
The Lambda function is yours to build and maintain. This gives you full control over the tenant resolution logic and everything downstream (retries, dead-letter queues, observability), but it also means you own the operational overhead: code updates, monitoring, and cost management.
Mail Manager can invoke the function synchronously or asynchronously. Synchronous invocation (REQUEST_RESPONSE) keeps Lambda in Mail Manager’s critical path: Mail Manager waits up to 30 seconds for the function to return, and retries on failure. Asynchronous invocation (EVENT) hands control to Lambda instead, so Mail Manager invokes the function and moves on. There are no additional Mail Manager charges for the Lambda invocation beyond standard Lambda pricing.
Pattern 5: Mail Manager stages to Amazon S3, Lambda delivers asynchronously
In Pattern 4, Mail Manager invokes the function directly through the Invoke Lambda rule action. Pattern 5 removes that direct invocation: the Mail Manager rule ends at Write to S3, and an Amazon S3 event notification triggers the Lambda function instead. Mail Manager’s work finishes at the write, and delivery becomes fully event-driven.
The rule set has two actions: write the raw MIME to Amazon S3, followed by an explicit Drop action. The Drop action prevents accidental duplicate delivery if a Send to internet action is inadvertently added to the rule later. The Lambda function handles delivery through the Amazon SES API, so Mail Manager’s job ends at writing the MIME to Amazon S3. The Amazon S3 event routes to the function directly or through Amazon Simple Queue Service (Amazon SQS) or Amazon EventBridge for fan-out and back-pressure.
The function reads the object, performs the tenant lookup, and calls SendEmail with the TenantName parameter. The Mail Manager critical path is minimal, and Lambda retries use the Lambda retry model with dead-letter queue support. The same Amazon S3 object fans out to multiple consumers (delivery, analytics) without changing the Mail Manager rule.
The Lambda function is yours to build and maintain. The upside is full control over the function and everything after it: tenant resolution logic, retries, dead-letter queues, and observability. The tradeoff is cost and upkeep, since you own code updates, monitoring, and operational overhead.
The other tradeoff is less visibility. After Write to S3, the Mail Manager log no longer records the delivery outcome.
Secure tenant attribution with IAM
Regardless of which pattern sets the X-SES-TENANT header, the Send to internet action’s IAM role should enforce tenant boundaries. Scope the IAM role with a Condition element that includes the ses:TenantName condition key.
Example IAM policy:
This policy allows the role to send email only when the message is attributed to the facilities-printers tenant. Messages tagged with any other tenant value, or messages with no tenant header, are denied.
In Pattern 1, the sender sets the header, the IAM role on the Send to internet action validates that the claimed tenant matches the role’s permissions. In Pattern 2, the Add header action sets a fixed value, and the IAM role confirms the header matches the expected tenant for that endpoint. In Pattern 3 with a chained structure, a single Send to internet action services all tenants. Scope its role to the set of valid tenant names so untagged messages (those matching no Add header rule) fail authorization. For Patterns 4 and 5, the Lambda function’s execution role carries the ses:TenantName condition key, providing the same enforcement at the API call level.
Paused tenants
Each of the five patterns handles paused tenants the same way. When Amazon SES pauses a tenant (through a reputation policy or manually), sends for that tenant fail with a rejection error. Other tenants keep delivering. The failure surfaces depending on the pattern:
- Patterns 1 to 3: Mail Manager records the rejection in the rule set log.
- Patterns 4 and 5: The rejection surfaces in the Lambda function’s Amazon CloudWatch Logs.
- Patterns 1 to 5: Amazon SES publishes tenant status changes to Amazon EventBridge (such as Sending Status Disabled).
Mail Manager won’t re-route or retry a paused tenant send. Graceful handling (queueing, failover, notification) belongs in the Lambda function in Patterns 4 and 5.
Observability
Observability for these patterns draws on three sources, each answering a different question:
Mail Manager vended log: which rule actions ran, and whether Amazon SES accepted the message from a Send to internet action. Mail Manager delivers this log to a destination you configure: Amazon CloudWatch Logs, Amazon S3, or Amazon Data Firehose. Query CloudWatch Logs with CloudWatch Logs Insights, or query Amazon S3 with Amazon Athena to surface IAM denials, configuration errors, and throttling.
Amazon SES event publishing: the final delivery outcome (delivered, bounced, or complaint), routed through a configuration set. This applies to every pattern.
Lambda Amazon CloudWatch Logs: for Patterns 4 and 5, where delivery runs inside the Lambda function, the acceptance result and any application errors.
To trace a message end to end, correlate these sources. For Patterns 1 to 3, the Mail Manager log and Amazon SES event publishing cover the flow. For Patterns 4 and 5, add the Lambda function’s CloudWatch Logs, since the Mail Manager log ends at Invoke Lambda (Pattern 4) or Write to S3 (Pattern 5).
Limits that shape the architecture
Review the Amazon SES Mail Manager service quotas before committing to a pattern. These quotas most often drive your pattern choice:
| Resource | Default | Where it matters |
| Maximum message size (SMTP ingress) | 40 MB | Patterns 1 to 5 |
| Authenticated ingress endpoints per Region | 50 | Pattern 2 per-tenant endpoints |
| Rule sets per Region | 40 | Pattern 2, Pattern 3 partitioning |
| Rules per rule set | 40 | Pattern 3 |
| Send to internet action per rule set | 10 | Pattern 3 tightest constraint |
| Actions per rule | 10 | |
| Conditions per rule | 10 | |
| Addresses per address list | 100,000 | Pattern 3 consolidation |
| Tenants per account (Amazon SES) | 10,000 (adjustable) | Patterns 4 and 5 ceiling |
| Lambda concurrent executions per Region | 1,000 (adjustable) | Patterns 4 and 5 throughput ceiling |
| Lambda timeout (Mail Manager InvokeLambda) | 30 seconds | Pattern 4 synchronous path |
| S3 event notification destinations per prefix | 1 (use Amazon EventBridge for fan-out) | Pattern 5 fan-out design |
| Lambda invocation payload (synchronous) | 6 MB | Pattern 4 metadata-only (body in S3) |
| Sending quota per 24 hours (Amazon SES) | 200 in sandbox (adjustable in production) | Patterns 1 to 5 |
| Maximum send rate (Amazon SES) | 1 message/second in sandbox (adjustable in production) | Patterns 1 to 5 |
Conclusion
The five patterns in this post show how to architect tenant tagging, whether through static endpoints, rule-set headers, or runtime resolution, so you can choose the approach that fits your workload.
Next steps
- Create your first ingress endpoint and rule set in the Mail Manager console.
- Create your first tenant.
- Review the tenant management page.
- Review the Amazon SES product detail page for pricing and Regional availability.
- Read the Amazon SES Mail Manager documentation.
- Learn how tenant-level suppression lists prevent cross-tenant contamination in Isolate email suppression per tenant with Amazon SES.
About the authors
Getting started with Apache Iceberg write support in Amazon Redshift – Part 3
Post Syndicated from Raghu Kuppala original https://aws.amazon.com/blogs/big-data/getting-started-with-apache-iceberg-write-support-in-amazon-redshift-part-3/
Production data is always evolving. Tables gain and lose columns, outgrow their data types, and get re-partitioned as query patterns shift. Multiple engines often need to read the same data. These changes used to mean expensive data rewrites or rebuilt pipelines. Apache Iceberg makes them metadata-only operations, and Amazon Redshift now supports evolving schemas and partitioning layouts through ALTER statements, with no data rewrites and no pipeline rebuilds. You can also create AWS Lake Formation resource links in the catalog of Amazon S3 Tables, a capability of Amazon Simple Storage Service (Amazon S3), for centralized cross-engine governance.
In Part 1, you created Apache Iceberg tables and wrote data directly from Amazon Redshift to your data lake, setting up external schemas, creating tables in both Amazon Simple Storage Service (Amazon S3) and Amazon S3 Tables, and performing INSERT operations with full ACID (Atomicity, Consistency, Isolation, Durability) compliance. In Part 2, you performed DELETE, UPDATE, and MERGE operations to modify data at the row level and synchronize staging and production tables.
In this post, you use the customer and orders datasets from the previous posts to evolve Iceberg table schemas and partitioning with ALTER operations. You also create an AWS Lake Formation resource link in the S3 Tables catalog to share tables with other analytics engines under a single, centralized permission model.
Solution overview
This solution demonstrates ALTER operations for Apache Iceberg tables in Amazon Redshift and Lake Formation resource link creation for the S3 Tables catalog. The walkthrough includes the following key operations:
- ALTER TABLE RENAME COLUMN – Rename existing columns without changing data types or partition specs.
- ALTER TABLE ADD/DROP COLUMN – Add new columns or remove existing columns as metadata-only operations.
- ALTER TABLE ALTER COLUMN – Widen column data types (for example, INT to BIGINT) without rewriting data.
- ALTER TABLE SET TABLE PROPERTIES – Change compression type for future writes.
- ALTER TABLE ADD/DROP/REPLACE PARTITION FIELD – Evolve partition specs without re-partitioning existing data.
- Lake Formation resource link – Create a resource link in the S3 Tables catalog for centralized access governance.
The following diagram shows the end-to-end architecture:
Figure 1: Architecture showing Amazon Redshift performing ALTER operations on Iceberg tables in S3 Tables, with Lake Formation resource links providing access from Amazon Athena and other engines
Prerequisites
Complete the setup from Part 1 and Part 2, including:
- An Amazon Redshift data warehouse (provisioned or Serverless) on patch 201 or higher.
- The AWS Identity and Access Management (IAM) role (
RedshifticebergRole) with permissions for Amazon S3, AWS Glue Data Catalog, and Lake Formation. - The
customertable in a standard Amazon S3 bucket (AWS Glue catalog:customer_db). - The
orderstable in an Amazon S3 table bucket (iceberg-write-blog@s3tablescatalog). - Access to an IAM role that is a Lake Formation data lake administrator.
- AWS Glue Data Catalog integrated with S3 Tables (
s3tablescatalogexists).
Schema evolution with ALTER TABLE
With ALTER TABLE, you can change Iceberg table definitions, including schema, partition specs, and properties, without rewriting stored data. Each operation updates only metadata. The table structure changes instantly while existing data files remain untouched. This helps make schema evolution, partition adjustments, and property updates safe to run on production tables.
Add a column
You can add a new column to an Iceberg table using ALTER TABLE. Each new column is added with a unique field ID that Iceberg uses for column tracking across schema evolution. Existing rows return NULL for the newly added column.
Verify the current schema:
Add the column:
Verify the schema change:
The following output shows the new loyalty_tier column as NULL for existing rows:
Populate the new column by aggregating order totals from the orders table in S3 Tables:
The following output shows customer loyalty tiers after the update:
Note: Customer IDs 11, 13, and 15 show NULL for loyalty_tier because they have no matching orders in the orders table.
Drop a column
Remove columns that are no longer needed. The column is removed from the current schema, but data in existing files remains untouched and simply becomes invisible to queries.
Verify the current schema:
Drop the column:
Verify the schema change:
Verify the column is dropped:
Note: To drop a column used in the current partition spec, first drop or replace the partition field, then drop the column.
Rename a column
Rename a column without affecting data types or partition specs:
The following output confirms the column has been renamed to location:
Widen a column type
Widen a column’s data type without rewriting data. This is useful when your data outgrows the original precision, for example when order amounts exceed the original decimal range.
Verify the current column type:
Now run the ALTER to widen the column:
Verify the updated column type:
Note: Amazon Redshift supports safe type promotions (for example, INT to BIGINT, FLOAT to DOUBLE, DECIMAL(10,2) to DECIMAL(18,2)). Plan column types accordingly for future growth.
Set table properties
Change the compression type for future writes:
Verify the current compression type:
Now run the ALTER to change the compression type:
The following SHOW TABLE output confirms the updated compression setting:
Note: This affects only future writes. Existing data files retain their original compression.
Partition evolution
A powerful feature of Iceberg is partition evolution, the ability to change how a table is partitioned without rewriting existing data. Amazon Redshift writes new data with the updated partition scheme, while existing data remains in the old layout. Query engines handle both layouts transparently.
Adding a partition field
The orders table from Part 1 is partitioned by DAY(order_date). Add an additional bucket partition to distribute data across hash buckets:
Verify the current partition spec:
Add the partition field:
After this change, new data is partitioned by both DAY(order_date) and bucket(16, customer_id), while existing data remains in the original day-only layout.
Verify the updated spec:
Figure 15: SHOW TABLE output showing the updated partition spec with DAY(order_date) and bucket(16, customer_id)
Replacing a partition field
Instead of separately dropping and adding, use REPLACE PARTITION FIELD as a single atomic operation. This is the recommended approach when swapping one transform for another on the same source column, because it makes the intent explicit and avoids a transient state where the table is unpartitioned between operations.
Verify the current partition spec:
Figure 16: SHOW TABLE output showing the current partition spec with DAY(order_date) and bucket(16, customer_id)
Replace the partition field:
After this change:
- Existing data remains in day-based partition folders.
- Amazon Redshift writes new data into month-based partition folders.
- The query engine reads both layouts transparently.
Confirm the new partition spec:
Insert new data and verify that both partition layouts are queryable:
Converting to a multi-level partition
Iceberg supports multi-level (composite) partition specs, where data is organized by more than one partition field. You can evolve an existing single-level spec into a multi-level spec by adding partition fields one at a time. Each ADD PARTITION FIELD is a lightweight metadata operation, and no data is rewritten.
The orders table is currently partitioned by MONTH(order_date) and bucket(16, customer_id). Add one more partition field to create a three-level spec:
Verify the current partition spec:
Figure 19: SHOW TABLE output showing the current two-level partition spec of MONTH(order_date) and bucket(16, customer_id)
Add partition field to build the three-level spec:
Verify the new multi-level partition spec:
Figure 20: SHOW TABLE output showing the three-level partition spec of MONTH(order_date), bucket(16, customer_id), and day(order_created_at_tz)
After these changes:
- Existing data remains in the original single-level layout (month-based folders).
- Amazon Redshift writes new data into the multi-level layout (month, then bucket, then day folders).
- The query engine reads both layouts transparently.
Dropping partition fields from a multi-level partition
You can also evolve in the other direction by removing partition fields from a multi-level spec to simplify the partition layout. Like adding fields, dropping a partition field is a metadata-only operation and removes one field per statement.
Verify the current multi-level partition spec:
Drop the partition fields one at a time:
Verify the table is back to its original single-level spec:
Figure 22: SHOW TABLE output confirming the table is back to a single-level MONTH(order_date) partition spec
After dropping a partition field:
- Data written under the dropped field’s layout stays in place and remains queryable.
- Amazon Redshift writes new data using only the remaining partition fields.
- Queries that filtered on the dropped field still work, but they no longer benefit from partition pruning on that field for newly written data.
Supported partition transforms
The following table lists the partition transforms available for Iceberg tables in Amazon Redshift:
| Partition transform | Syntax example | What it does |
| Year | year(order_date) |
Groups data into yearly partitions based on a date or timestamp column. |
| Month | month(order_date) |
Groups data into monthly partitions based on a date or timestamp column. |
| Day | day(order_date) |
Groups data into daily partitions based on a date or timestamp column. |
| Hour | hour(event_ts) |
Groups data into hourly partitions based on a timestamp column. |
| Bucket | bucket(16, customer_id) |
Distributes data across N hash buckets for even distribution on high-cardinality columns. |
| Truncate | truncate(3, zip_code) |
Truncates column values to a fixed width W for grouping similar values together. |
| Identity | identity(region) |
Partitions by the exact column value with no transformation applied. |
Note: A column that is already part of an existing partition field can’t be used in a new partition field. Drop or replace the existing field first.
Accessing S3 Tables with external schemas
Lake Formation resource links provide cross-engine access to your S3 Tables through centralized governance. You create a resource link in the default AWS Glue Data Catalog that points to your S3 Tables database. Amazon Redshift, Amazon Athena, Amazon EMR, and other engines can then discover and query the tables using a single permission model.
For the complete setup walkthrough, including Lake Formation prerequisites, resource link creation, and permission grants, see Optimize Amazon S3 Tables queries with Amazon Redshift. For conceptual details on resource links and S3 Tables catalog integration, see About resource links and Creating an S3 Tables catalog.
The following steps show how to query S3 Tables through a resource link after completing the setup from the referenced blog.
Grant access to the resource link
In the Lake Formation console, the resource link appears as a database named iceberg_write_blog_rl (type: Resource link). To grant access to the resource link:
- In the Lake Formation console, choose Databases.
- Locate
iceberg_write_blog_rl(type: Resource link). - Choose Actions, then Grant.
- Grant DESCRIBE permission to RedshiftIcebergRole.
Create an external schema
With the resource link in place, create an external schema in Amazon Redshift for two-part notation access.
For IAM federated users:
For database users and business intelligence (BI) tools:
Grant access to specific users or roles:
Query with two-part notation
With the external schema created, query S3 Tables using two-part notation:
Access methods comparison
The following table compares the available methods for accessing Iceberg tables in Amazon Redshift:
| Access method | Query syntax | Authentication | Best for |
| S3 Tables three-part notation | "bucket@s3tablescatalog".namespace.table |
IAM federated identity only | Interactive queries in Query Editor v2 with direct catalog access. |
| External schema through resource link | schema_name.table |
Any (IAM role defined in schema) | BI tools, Data API, JDBC/ODBC applications, and shared team access. |
| awsdatacatalog | awsdatacatalog.database.table |
IAM federated identity only | Multi-database access in a single session without creating external schemas. |
Bringing it together
Combine schema evolution with cross-engine access in a single workflow. The following example adds a column to the orders table and immediately queries it through the external schema:
Figure 26: Cross-catalog join showing the evolved schema immediately visible through the external schema
The new column is visible through both the three-part notation and the external schema without any additional configuration, because the schema evolution in Iceberg propagates automatically.
Best practices
- Test ALTER operations in non-production first. While metadata-only, schema changes affect all readers immediately.
- Use REPLACE PARTITION FIELD instead of DROP + ADD. The atomic operation avoids a transient unpartitioned state.
- Monitor partition spec changes with SHOW TABLE. Verify the current spec after any partition evolution.
- Choose partition transforms based on query patterns. Use
month()orday()for time-range filters. Usebucket()for high-cardinality join keys. - Set table properties before bulk loads. Change compression type (
zstdfor better ratios,snappyfor speed) before large INSERT operations. - Run table maintenance after mutations. After performing multiple UPDATE, DELETE, or MERGE operations, run AWS Glue table optimizers to compact deletion files and improve read performance.
- Use Lake Formation for fine-grained access. Column-level and row-level security can be applied through Lake Formation on tables accessed through resource links.
- Grant schema access to specific users or roles. Avoid granting to PUBLIC. Use named IAM roles or database users for least-privilege access.
- Monitor query performance. Use Amazon Redshift query monitoring features to track performance of write operations and optimize partitioning strategies as needed.
Considerations
Keep the following in mind when working with ALTER TABLE and partition evolution on Iceberg tables:
- Plan for metadata-only behavior. ALTER TABLE operations update metadata instantly, and existing data files remain unchanged. All readers see the new schema immediately after the operation completes.
- Drop partition fields before dropping partitioned columns. To remove a column used in the current partition spec, first drop or replace the partition field, then drop the column.
- Use safe type promotions for ALTER COLUMN TYPE. Amazon Redshift supports widening within compatible families (INT to BIGINT, FLOAT to DOUBLE, DECIMAL(10,2) to DECIMAL(18,2)). Plan column types with future growth in mind.
- Account for mixed partition layouts after evolution. Partition evolution doesn’t re-partition existing data. Old files remain in their original layout, and the query engine reads both layouts transparently.
- Use external schemas for database user access. The auto-mounted three-part notation (
"bucket@s3tablescatalog") requires IAM federated authentication. For database users and BI tools, create an external schema with an explicit IAM role. - Use full three-part notation with awsdatacatalog. The USE statement isn’t supported with awsdatacatalog, so always specify the full path.
- Clean up S3 data separately after dropping tables. Dropping an Iceberg table removes only the catalog entry from AWS Glue Data Catalog. Delete the underlying S3 data files separately, or use AWS Glue table optimizers to remove orphaned files.
Clean up
To avoid ongoing charges, run the following:
Conclusion
In this post, you evolved Apache Iceberg table schemas using ALTER TABLE operations. You added, dropped, and renamed columns, widened data types, changed compression, and evolved partition specs, all as metadata-only operations without rewriting data. You also created Lake Formation resource links to provide governed cross-engine access to S3 Tables, and simplified query syntax with external schemas.
This concludes the three-part series on getting started with Apache Iceberg write support in Amazon Redshift:
- Part 1: Create Iceberg tables and perform INSERT operations.
- Part 2: Run DELETE, UPDATE, and MERGE for row-level modifications.
- Part 3: Evolve schemas with ALTER TABLE and add cross-engine access with Lake Formation resource links.
If you have questions or feedback about this series, leave a comment on this post.
Additional resources
- Amazon Redshift Iceberg integration – Complete syntax reference.
- Writing to Apache Iceberg tables – Detailed examples.
- ALTER TABLE for Iceberg – Full ALTER reference.
- Amazon S3 Tables – Managed Iceberg storage.
- AWS Lake Formation – Centralized data governance.
- Optimize S3 Tables queries with Amazon Redshift – Resource links and performance tuning.
About the authors
Enforce IAM permissions boundaries for Amazon SageMaker Unified Studio Tooling blueprints
Post Syndicated from Sanjana Sekar original https://aws.amazon.com/blogs/big-data/enforce-iam-permissions-boundaries-for-amazon-sagemaker-unified-studio-tooling-blueprints/
Amazon SageMaker Unified Studio now supports custom permissions boundaries for IAM roles created by the Tooling blueprint. Organizations that enforce Service Control Policies (SCPs) requiring permissions boundaries on all AWS Identity and Access Management (IAM) roles can now adopt Amazon SageMaker Unified Studio without modifying their security posture.
Amazon SageMaker Unified Studio is a unified development environment that brings together data engineering, machine learning, and analytics tools into a single workspace. In Amazon SageMaker Unified Studio, a project is a collaborative workspace that bundles people, tools, and access permissions together. It builds every project from a project profile, which defines a list of blueprints. Blueprints are pre-configured infrastructure templates that provision AWS resources at project creation time or on demand, along with their default parameters. The Tooling blueprint is the only mandatory one. Amazon SageMaker Unified Studio deploys it with every project, creating foundational resources such as the project IAM role and security groups.
In this post, you learn how to create a permissions boundary that restricts AI agent capabilities. You then configure it on the Tooling blueprint using the AWS Command Line Interface (AWS CLI). Finally, you validate that the boundary is enforced on all provisioned roles.
The problem
Enterprises in regulated industries use SCPs to require that every IAM role in an account carries a permissions boundary. A well-scoped boundary prevents privilege escalation and verifies no role exceeds the maximum permissions defined by the organization’s security team. Before this feature, Amazon SageMaker Unified Studio Tooling blueprints created IAM roles without permissions boundaries. When an SCP enforced permissions boundaries, project creation failed with an explicit deny:
Amazon SageMaker Unified Studio surfaces the blocked role creation as a Tooling environment provisioning failure, as shown in Figure 1.
The project is marked as failed because its Tooling environment couldn’t deploy in the US East (N. Virginia) AWS Region (us-east-1). The details show a 403 permissions error, while the preceding IAM message identifies the underlying iam:CreateRole SCP denial. This blocked adoption for any organization with SCP-enforced permissions boundaries. The AWS CloudFormation event for the Tooling stack exposes the IAM failure behind the project-level error, as shown in Figure 2.
The AmazonBedrockServiceRole resource entered CREATE_FAILED because iam:CreateRole was explicitly denied by the SCP, even though AWS CloudFormation surfaced the wrapper error as UnauthorizedTaggingOperation.
Granular control using a permissions boundary: Example use case
Beyond satisfying SCP requirements, permissions boundaries give administrators granular control over what the Tooling blueprint roles can do. For instance, some organizations have SecOps policies that require disabling Data Agent and Data Notebook capabilities across their accounts. These organizations want project members to access data connections and run SQL queries directly, but must block conversational AI, code generation, and notebook cell execution through the agent.
When PermissionsBoundaryArn is configured on the Tooling blueprint, SageMaker Unified Studio attaches the specified customer-managed permissions boundary to all IAM roles provisioned by that blueprint. If your governance requires a boundary, configure it explicitly and verify the resulting roles.
The following permissions boundary policy scopes the roles to the AWS services that Amazon SageMaker Unified Studio uses and then explicitly denies the Amazon DataZone actions that power the AI agent. This permissions boundary is provided for illustrative purposes only and isn’t a recommendation or reference for environment configuration. You should tailor your permissions boundaries to your specific workloads in accordance with the principle of least privilege.
Warning: validate before using in production. This example scopes the roles to the service namespaces Amazon SageMaker Unified Studio uses, but it is still coarse (it allows each listed service in full) and is provided only for illustration. Because a permissions boundary is a ceiling, it must remain a superset of everything the three Tooling roles (datazone_usr_role, AmazonBedrockServiceRole, and AmazonBedrockLambdaExecutionRole) actually need. If Amazon SageMaker Unified Studio adds a dependency that isn’t listed, provisioning or in-console actions will fail with an access denied error. Validate in a non-production domain first.
With this boundary attached, the Tooling blueprint provisions normally, project members can access data connections and run SQL queries. However, any attempt to invoke the AI assistant or execute notebook cells through the agent returns an access denied error. The boundary acts as a ceiling that no policy attached to the role can override.
How it works
The custom permissions boundary feature operates at the blueprint configuration level. An administrator sets a PermissionsBoundaryArn in the Tooling blueprint’s regional parameters. When a user creates a new project that includes the Tooling blueprint, Amazon SageMaker Unified Studio provisions an AWS CloudFormation stack that creates three IAM roles and attaches the specified boundary to each:
datazone_usr_role– the role that all project members assume to access data and resources in that project.AmazonBedrockServiceRole– for Amazon Bedrock operations.AmazonBedrockLambdaExecutionRole– for Amazon Bedrock-related AWS Lambda functions.
Because the boundary is set at the blueprint level, it applies to every project created under that blueprint. No per-project configuration is needed.
Prerequisites
Before you begin, make sure that you have:
- An AWS account with a Amazon SageMaker Unified Studio Identity Center-based domain created.
- The Tooling blueprint enabled in the domain.
- AWS CLI v2 configured with administrator access.
If your organization uses AWS Organizations with SCPs that require permissions boundaries, you will also need an organization with the target account as a member and permissions to create and attach SCPs in the management account.
Setting up the SCP (optional)
This section provides instructions to create an SCP and attach it to your AWS Organizations organizational unit or accounts. If your organization already enforces permissions boundaries through SCPs, skip this section. Otherwise, create an SCP in your AWS Organizations management account that denies IAM role creation unless an approved permissions boundary is attached. This also prevents the boundary from being removed, swapped, or weakened afterward:
This policy does three things:
DenyRoleWithoutApprovedBoundaryblocks creating a role, or attaching a boundary to an existing role, with anything other than the approved boundary ARN. Denyingiam:PutRolePermissionsBoundaryin addition toiam:CreateRolestops a privileged principal from swapping in a weaker boundary after the role exists.DenyRemovingBoundaryblocksiam:DeleteRolePermissionsBoundaryoutright, so the boundary cannot be stripped off. (This action doesn’t support theiam:PermissionsBoundarycondition key, so it must be denied unconditionally.)ProtectBoundaryPolicyprevents tampering with the boundary policy itself. Deleting it, or publishing and defaulting a new version that quietly widens what it allows.
Note: Scope these denies so you don’t lock yourself out. A broad deny on iam:PutRolePermissionsBoundary and iam:DeleteRolePermissionsBoundary also applies to your own administrators. Add an exception for a break-glass or IAM-admin role (for example, an aws:PrincipalArn StringNotLike condition) so a trusted principal can still manage boundaries.
To create the SCP, sign in to the AWS Organizations console with your management account and go to AWS Organizations → Policies → Service control policies. If SCPs aren’t enabled for your organization yet, choose Enable service control policies first. Choose Create policy, give it a name (for example, test_scp), and replace the default content in the policy editor with the JSON above substituting <account-id> with your account ID. Choose Create policy to save it.
After creating the SCP in the management account, verify its content before attaching it. Figure 3 shows the core create-role control. The full example above adds controls that prevent replacing or removing the boundary and modifying the protected policy.
Figure 3: Service Control Policy defined in the AWS Organizations management account
The AWS Organizations Content tab displays the customer-managed test_scp policy. Its visible statement denies iam:CreateRole unless the request uses the SMUSToolingBoundary policy.
Attach this SCP to the organizational unit or account where your Amazon SageMaker Unified Studio domain and domain-associated accounts reside. To do so, open the test_scp service control policy, choose the Targets tab, and choose Attach. The AWS organization structure appears; select the OU or account where the SCP should apply, then choose Attach policy.
Figure 4 identifies the member account that must inherit the SCP in this example organization. The target member account, datazone-account2, resides under OU2, while datazone-account1 is the organization’s management account. Attaching the SCP to the target account or a parent organizational unit enforces it there.
Figure 4: AWS Organizations account structure showing the management account and the target member account
After attaching the policy, verify the association on the SCP’s Targets tab, as shown in Figure 5.
The Targets tab lists datazone-account2 as an ACCOUNT target, confirming that test_scp is enforced directly on the intended member account.
Configuring the permissions boundary
In this section you will execute the required steps to create the permissions boundary and enable it in the Tooling blueprint. The example in this walkthrough uses us-east-1. Change it to the Region where your Amazon SageMaker Unified Studio domain is deployed. You must execute the configuration in the account where you plan to create your project. This can be your Amazon SageMaker Unified Studio domain account or accounts associated to your Amazon SageMaker Unified Studio domain.
Step 1: Create the permissions boundary policy
If you haven’t already created the boundary policy, save the following JSON document as a boundary-policy.json file on your workstation:
As noted previously, this illustrative policy is scoped to the services Amazon SageMaker Unified Studio uses but is still coarse, and its allow list must stay a superset of what all three Tooling roles need.
Then create the policy using the following command:
Note the policy ARN from the output, because it will be used later in the procedure.
Step 2: Retrieve the ID of your domain
Retrieve the ID of your domain by running the following command. Replace <YOUR_DOMAIN_NAME> with the name of your SageMaker Unified Studio domain.
Note the returned ID, because it will be used later in the procedure.
Step 3: Identify the Tooling blueprint
Retrieve the Tooling blueprint ID by running the following command. Replace <domain-id> with the ID you noted in Step 2.
Note the returned ID, because it will be used later in the procedure.
Step 4: Read the current configuration
Retrieve the current Tooling blueprint configuration by executing the following command. Replace <domain-id> with the ID from Step 2 and <tooling-bp-id> with the ID from Step 3.
Important: Back up the output of get-environment-blueprint-configuration before making any changes. The command above pipes the response to tooling-bp-config-backup.json so you have a restore point if you need to revert.
Note the values of provisioningRoleArn, manageAccessRoleArn, enabledRegions, and all fields inside regionalParameters (AZs, S3Location, Subnets, VpcId). You will need all of these in the next step.
Step 5: Set the permissions boundary
Update the blueprint configuration to include PermissionsBoundaryArn in the regional parameters using the following command.
Important: The put-environment-blueprint-configuration API operates in overwrite mode, it replaces the entire configuration with what you provide. You must include all existing values from the previous step’s output. The only new addition is PermissionsBoundaryArn inside the regional parameters. Omitting any existing parameter removes it.
Make sure to replace <domain-id> with the ID you noted in Step 2, <tooling-bp-id> with the ID you noted in Step 3, and all other <placeholder> values with the corresponding values from Step 4’s output.
The following anonymized example is based on an existing Tooling blueprint configuration. Its S3Location reflects the bucket naming pattern used in that environment. Copy the exact S3Location returned in Step 4. Don’t use the following illustrative value. Here’s an example:
Step 6: Verify the configuration was applied
Confirm the permissions boundary ARN is now set in the blueprint configuration using the following command. Make sure to replace <domain-id> with the ID you noted in Step 2 and <tooling-bp-id> with the ID you noted in Step 3.
The output should return your boundary policy ARN:
Validating the configuration
After configuring the permissions boundary, in this section you will get instructions to create a new project to verify it works end to end and that the IAM roles created with the project actually include the permissions boundary.
Step 1: Select a project profile in enabled state
Use the following command to list project profiles configured in your domain. Make sure to replace <domain-id> with the ID you noted in Step 2 of the “Configuring the permissions boundary” section.
Choose a project profile that has "status": "ENABLED". Note the id of any project profile returned in the previous command.
Step 2: Create a test project
Create a new project using the following command. Make sure to replace <domain-id> with the ID you noted in Step 2 of the “Configuring the permissions boundary” section and to replace <profile-id> with the project profile ID noted in Step 1 of this section.
Note the id (project ID) returned in the response. Wait for the Tooling blueprint to provision. This typically takes a minute or two. After provisioning completes, confirm that the validation project reaches the Active state, as shown in Figure 6.
The Projects list shows the timestamped PB-Validation-* project with an Active status, confirming that project creation succeeded with the custom boundary configured.
Step 3: Verify the roles have the boundary attached
In this section you check that the IAM roles created with the project have the permissions boundary attached. Use the following commands to get the configuration for the IAM roles created with the project you just created. Replace <domain-id> and <project-id> with the values from the previous steps.
All three roles should return a response showing the permissions boundary ARN:
You can also verify each role in the IAM console. Figure 7 shows the permissions boundary for the project user role.
The datazone_usr_role Permissions tab displays SMUSToolingBoundary as its customer-managed permissions boundary.
Figure 8 confirms that the same boundary is attached to the Amazon Bedrock service role.
The AmazonBedrockServiceRole also displays SMUSToolingBoundary as its customer-managed permissions boundary.
Figure 9 verifies the boundary on the third Tooling role, the Bedrock Lambda execution role.
Figure 9: IAM console showing the permissions boundary attached to the AmazonBedrockLambdaExecutionRole
The AmazonBedrockLambdaExecutionRole likewise displays SMUSToolingBoundary, confirming that all three provisioned roles carry the boundary.
Step 4: Verify the boundary denies AI agent actions
In this section you verify the boundary actually denies AI agent actions. If you configured the boundary from the use case section earlier, the boundary blocks Data Notebooks and messages to the Data Agent, such as the Query Editor assistant. Any such attempt returns an access denied error. The project user role has the boundary attached, so even if the role’s identity policies grant the relevant APIs, the boundary’s explicit deny takes precedence.
To confirm, navigate to your project in SageMaker Unified Studio and test the following actions:
- Attempt to create a notebook – In the left sidebar, select Notebooks. Select Create notebook. The operation will fail because the permissions boundary prevents the
datazone:CreateNotebookaction (Figure 10).
Figure 10: Permissions boundary preventing creation of Data Notebooks
After the create action, Amazon SageMaker Unified Studio reports Failed to create notebook and identifies datazone:CreateNotebook as explicitly denied by SMUSToolingBoundary.
- Attempt to use Data Agent in the Query Editor – In the left sidebar, select Query Editor, then select the Chat with AI icon. The agent chat will fail to load because the permissions boundary blocks the APIs required by Data Agent (Figure 11).
Figure 11: Permissions boundary preventing using Data Agent on Query Editor
The Query Editor remains available, but the Agent panel reports “You don’t have access to Data Agent“. In this configured test, that message is the user-visible result of denying the Data Agent APIs. The screenshot itself doesn’t display the denied API or boundary ARN.
Important considerations
- Immutable after project creation – The permissions boundary is set at provisioning time. Changing the boundary ARN on the blueprint configuration only affects new projects. Existing projects retain their original boundary.
- Applies to all Tooling-provisioned roles – When
PermissionsBoundaryArnis configured on the Tooling blueprint, SageMaker Unified Studio attaches the specified customer-managed permissions boundary to all three IAM roles created by that blueprint. It’s applied uniformly — you can’t selectively apply it to individual roles. No boundary is attached unless you configure one, so if your governance requires a boundary, set it explicitly and verify the resulting roles rather than assuming one is present by default. - Policy must exist – The IAM policy referenced by
PermissionsBoundaryArnmust exist in the account before project creation. If the policy is deleted or the ARN is invalid, provisioning will fail. - Tooling blueprint only – Among Amazon SageMaker Unified Studio provided blueprints, only the Tooling blueprint supports custom permissions boundaries. Other provided blueprints that create IAM roles (for example, the EmrOnEc2 blueprint) don’t currently support this feature. If your organization requires permissions boundaries on roles created by additional blueprints, you can build custom blueprints that include a permissions boundary configuration so you can extend this security control across your entire project infrastructure.
Clean up
To remove test resources, delete the test project from the SageMaker Unified Studio UI. On the project’s Overview page, choose the ⋮ (more actions) menu in the top-right and choose Delete project.
Figure 12: Deleting the test project from the project Overview page.
In the Delete project dialog, type confirm in the text box to acknowledge that the action is final, then choose Delete project. This permanently deletes the project and its underlying resources, and triggers an asynchronous AWS CloudFormation stack deletion.
Figure 13: Confirming project deletion.
To remove the boundary from future projects, re-run the put-environment-blueprint-configuration command from Step 5: Set the permissions boundary, but omit the PermissionsBoundaryArn field from the regional parameters. Because you backed up the original configuration in Step 4: Read the current configuration (tooling-bp-config-backup.json), you can reuse the exact same provisioningRoleArn, manageAccessRoleArn, enabledRegions, and regionalParameters values (AZs, S3Location, Subnets, VpcId) — just without PermissionsBoundaryArn — so the blueprint returns to provisioning roles with no permissions boundary.
Conclusion
With the custom permissions boundary feature for Amazon SageMaker Unified Studio, organizations can adopt Amazon SageMaker Unified Studio Tooling blueprints without compromising their IAM governance posture. By configuring a single parameter on the Tooling blueprint, all IAM roles provisioned by future projects automatically carry the specified permissions boundary. This satisfies SCPs that mandate a boundary on every role and gives administrators granular control over what the Tooling roles can do, for example disabling AI agent and notebook capabilities. Remember that the example boundary in this post is illustrative, because it scopes to the services SageMaker Unified Studio uses but is still coarse.
“I just updated the EnvironmentBlueprintConfiguration for the Tooling blueprint to include the new PermissionsBoundaryArn param. After that the blueprint provisioned successfully with the required permissions boundary attached to all the IAM roles, in line with our security policies. In the end it was a one-line change.”
— Nat Noordanus, Data Tech Lead at Nexthink
To get started, create your permissions boundary policy, configure it on the Tooling blueprint using the CLI, and create a project to verify the boundary is attached.
For more information, see the documentation for Amazon SageMaker Unified Studio, IAM permissions boundaries, and Service Control Policies.
About the authors
[$] Reducing undefined behavior in the C language
Post Syndicated from corbet original https://lwn.net/Articles/1095811/
As a professor of biomedical engineering, Martin Uecker perhaps does not
fit the profile of a typical presenter at Kernel Recipes. He is,
however, a longtime Linux user, and works on free software for controlling
magnetic resonance imaging (MRI) scanners. He was at the conference to
talk about the C programming language, the specific problem of undefined
behavior in C, and whether it can eventually be made into a memory-safe
language.
Security updates for Monday
Post Syndicated from jake original https://lwn.net/Articles/1097191/
Security updates have been issued by AlmaLinux (firefox, ipa, kernel, libxml2, perl-DBI, python-cryptography, thunderbird, and unbound), Debian (chromium, evolution-data-server, exim4, ghostscript, incus, lemonldap-ng, libheif, nodejs, php8.4, ruby-oj, swift, and vlc), Fedora (chromium, cinnamon, cinnamon-desktop, cinnamon-session, cinnamon-settings-daemon, ckermit, dnf5, forgejo, goose, gssntlmssp, libheif, libpcap, librsvg2, mingw-gstreamer1, mingw-gstreamer1-plugins-bad-free, mingw-gstreamer1-plugins-base, mingw-gstreamer1-plugins-good, mingw-python3, mongo-c-driver, muffin, nemo, nemo-extensions, nextcloud, pgadmin4, postgresql16-postgis, postgresql17-postgis, postgresql18-postgis, rust-librsvg, rust-xml5ever, sipp, suricata, tesseract, and xreader), Mageia (erlang, gpsd, libreswan, python3 & python-pip, and udisks2), Oracle (abrt, buildah, cockpit-image-builder, corosync, ipa, kernel, libxml2, openexr, perl-DBI, perl-DBI:1.641, postgresql, thunderbird, unbound, and yelp), SUSE (389-ds, ansible-lint, cups, firefox, flatpak-builder, forgejo-longterm, gdb, gimp, gitoxide, glib2, gnome-shell, google-guest-agent, google-osconfig-agent, haveged, helm, ImageMagick, kbd, libsoup, libtpms, obs-service-cargo, openai-codex, opensuse-signkey-cert, osmo-iuh, perl-mojolicious, poppler, python-WebOb, python-WebOb-doc, python313-vllm, python314, sdbootutil, suseconnect-ng, and swtpm), and Ubuntu (exim4, freerdp3, libvirt, libvirt-hwe, libwebsockets, lxc, pyjwt, and requests).
Next.js applications, powered by Vite: introducing Vinext 1.0
Post Syndicated from James Anderson original https://blog.cloudflare.com/vinext-nextjs-on-vite/
When we launched Vinext in February, it was the result of an audacious week-long AI-driven experiment to see how far one engineer, and a stack of tokens, could get to replicating the NextJS framework backed by Vite.
In the seven months since that experiment, Vinext has grown into a framework that our customers trust and run in production for high-traffic, dynamic applications.
Today we are announcing the release of Vinext 1.0, the latest step on our journey to make it possible to deploy Next.js apps anywhere. Vinext lets you take any Next.js application, whether it was built for the Pages or App Router, and make it portable to be deployed to any web platform, including the Cloudflare Workers free plan, Netlify, or AWS Lambda.
Vinext 1.0 brings with it sweeping improvements to compatibility, stability, and caching behaviors, and sets the project up for the long term. There’s never been a better time to take your Next.js project and convert it to Vinext; just run npx vinext check and npx vinext init.
Graduation to 1.0
On release Vinext was promising, but it was incomplete. Since then, we’ve spent a lot of time both improving App Router compatibility and expanding that to Pages Router apps — which we’ve learned many customers are longtime fans of, with large applications that are complex to migrate. We didn’t want Vinext to be a tool that only worked for people using the latest App Router features.
Our focus has been on adopting both these routers, and watching our test compatibility closely, which for most important customer-requested features now surpasses 99%.
This improvement has been fueled through the community around our GitHub project. As soon as Vinext launched, that community threw it at a wide variety of applications to find the gaps. With their scrutiny, we found challenges not immediately obvious in the test coverage. Vinext needs to act exactly as Next.js behaves. It is not good enough to imitate functions with the same name. Building an alternative import { revalidatePath } is simple enough; the difficulty is in making sure it correctly affects the rendered pages, cache entry, and future requests.
Tracing requests through the application to make sure Vinext responds in the way expected — and replicating not just the API, but the behavior of this machine — was by far the more challenging aspect.
Once we’ve patched problems and brought new features forward, it’s important that we don’t regress, especially if Next.js makes a change. That’s why we’ve also built out our test suite: thousands of focused tests covering core framework behavior across both routers, the development and production server, and the deployment targets of Nodejs and Cloudflare Workers. We also run the Next.js end-to-end test suite against Vinext nightly, giving us a continually moving window on our compatibility, and making sure we immediately become aware of regressions coming from merged changes. Alongside the automated testing, we’ve been working directly with large customers that have Vinext in production to make sure they are not facing issues.
What’s in 1.0
The clearest messages we got from customers using Vinext is that certain Next.js features carry the framework and Vinext didn’t actually need to do everything that Next.js has launched in recent versions to be incredibly useful to them. So we focused on better support where you need it:
- App Router, Pages Router, and Hybrid applications: We heard from customers that Pages Router was still important, and migrations are not a one-step process. Vinext therefore has support for both routing paths, including React Server Components, Server Actions, API routes, route handlers, middleware, and client-side navigation.
- The complete page lifecycle: Pages can be rendered in many different ways: on the server, pre-rendered in the build, exported as static assets, or cached with page-level Incremental Static Regeneration (ISR). We’ve made sure that Background and on-demand revalidation work with any output.
- Caching: Vinext has a shared set of caching functions across the App and Pages Router and the supported runtimes. We have further support for using Cloudflare’s Workers Cache.
- Observability: Vinext provides Next.js-compatible tracing across both routers, so existing OpenTelemetry and Sentry setups continue to work. On Cloudflare Workers, traces also integrate with native Workers Observability.
- Next.js ecosystem compatibility: Vinext implements the public
next/*surface and supports common Next patterns for use of authentication, MDX, image optimization, fonts, metadata, environment variables, and more. - First-class runtime support for Workers: While Vinext can run anywhere, server code can run in the Cloudflare workerd runtime during development and production, with direct access to bindings such as image optimization and hyperdrive.
We’ve also made migration part of the framework: it takes two commands to verify that your Next.js install and any modification you have made is compatible, and set up the Vite and deployment configuration while keeping all your previous Next.js project structure.
When we talked to teams about what features were important for them, something stood out. Next.js 16 took a stance that Cache Components were an important part of the future of the framework, and yet most teams that we talked to were not using them and did not consider support a prerequisite to move. Therefore, Vinext today has limited support for the “use cache” directive that drives Cache Components, and though we will continue to improve compatibility there, we’re much more focused on the core priorities above.
Pre-rendering and cache warming
When we first announced Vinext, it supported Incremental Static Regeneration (ISR) after the first request, but it did not yet render pages during the build. Applications use generateStaticParams() and getStaticPaths() to identify pages that should be rendered when building, and they expect page-level ISR to connect those initial responses to background and on-demand revalidation.
Vinext 1.0 supports that lifecycle for both routers. It can prerender App Router and Pages Router routes during the build, serve those responses through page-level ISR, and invalidate them by path or tag. It also supports output: "export" when the result you want is a fully static site.
But this led us to question something: Why should this rendering happen during the build at all?
A site with tens or hundreds of thousands of possible URLs can spend a seriously long time rendering pages that receive little traffic. The build process cannot evaluate the long tail of traffic that most sites experience and therefore cannot focus compute time on the smaller number of more critical pages. Instead you waste hours of time waiting for sequential builds working their way through thousands of pages, long after the most important routes are done.
Cache warming is our solution to this, moving page prerendering from the build machine to Cloudflare’s network. Developers can continue to use Next.js primitives to identify the pages for prerendering, and Vinext can additionally identify high-traffic pages to add to this list. This happens in the background before your site is deployed to production, so that the moment it is, it is ready to serve rapid responses from the Cloudflare cache.
Inside the deployment process, this works by uploading a new Worker version and deploying it to 0% of production traffic, before then requesting pages specifically from that version. This allows the rendering pipeline to work before any real users hit the new deployment. Once the caches have been populated, the deployment can be promoted safely.
What we’re doing next
If the original experiment invented the one-off slopfork, the more consequential part has been how we can keep that process of self-improvement running indefinitely.
The project is now focused on keeping the framework up to date with everything happening upstream. Next.js canary receives new commits every day. Each morning, an agent reviews the changes, fetches diffs, and opens tracking issues for anything that could affect Vinext. Every night, the compatibility matrix is regenerated as we run the Next.js test suite against Vinext.
When one of these tests or issues reveals a gap, agents are now in the position where they can identify the change across both codebases, build a reproduction, port any relevant tests, and propose a fix.
This review has been catching missing cases, unsafe caching behaviors, and differences in the development vs. production servers.
Automation has helped us narrow the stream of activity into a focused set of changes that deserve attention, allowing the maintainers of the project to focus on only the issues that need context of how a process should map onto Vite from the Next.js implementation.
We’re building a software factory for open source at Cloudflare, and you can see what we’re up to on GitHub.
Try it out
Vinext is available for new applications, and existing Next.js projects.
Start a new application today:
Or migrate an existing application:
And then deploy it to Cloudflare Workers, with our cache warming:
Visit vinext.dev for documentation, examples, and the current compatibility matrix. Vinext is open source at github.com/cloudflare/vinext. Issues, pull requests, application reproductions, and feedback are welcome.
Introducing cf: the agentic CLI for the entire Cloudflare API
Post Syndicated from Matt “TK” Taylor original https://blog.cloudflare.com/cloudflare-cf-cli-launch/
Over the last year, agent use of Wrangler has skyrocketed.
In March 2026, agents were responsible for a quarter of Wrangler use, up from single-digit percentages the year prior. Last week, agent usage reached 48%.
Agents are more prolific users, using almost twice as many distinct commands per day, and are almost four times as likely to use six or more commands.
Agents love CLIs. But Wrangler only provides commands for around 280 operations, and Cloudflare offers thousands.
Earlier in the year we teased how we were planning to solve this and today, we’re enabling agents to use every Cloudflare product by introducing a new CLI: cf.
cf is a CLI that is built for the next generation of software development:
- Agents can find the command they need to do anything they want to do with bespoke search and steering.
- JSON is the default interface, pretty printed for humans and condensed for agents for maximum context savings.
- cloudflare.config.ts is the new configuration format for the whole of Cloudflare, starting with Workers, and bringing the safety and accuracy of TypeScript to you and your agent’s language server protocol (LSP)
- Vite becomes default, bringing with it the best local development server, and a plugin suite for developers and framework authors.
Install the open beta today globally and run it from anywhere:
cf gives your agent access to the entire Cloudflare API
What if your agent could do everything Cloudflare can do? That’s the question that sparked our interest earlier this year: agents were getting ever more powerful, but what they were able to do with Cloudflare’s CLI was still limited.
Wrangler was hand-built with each product team contributing and taking their own approach to their command developer experience. Enforcing patterns across teams was virtually impossible, even across our ~280 command paths. We had inconsistent terminology across d1 info, hyperdrive get, workflows describe as each team came up with their own practices at different times. Some teams built entirely custom experiences across thousands of lines of code that turned out to be used extremely rarely, and teams came up with different approaches to solve the same problems.
We wanted to both standardize what we had and make a massive expansion, all at once. Forge — Cloudflare’s new unified API generation pipeline — enabled us to do this, building on the idea of generating our CLI commands directly from the API schema that powers our API documentation and SDK generation. Everything we provide has an OpenAPI schema, and if we annotate this with just a little more information, we can use it as the source for Forge to make a CLI.
This enables us to expand cf from the ~280 functions that Wrangler had built up over time, to cover the entirety of the Cloudflare API surface of over 3,000 operations.
Now it’s simple to give your agent cf and ask it to go set up a worker, deploy it, monitor and observe it, protect it with Cloudflare Access, buy a domain, and front it with Cloudflare WAF, all from a single tool.
Building for an agent that has never used cf
cf is built for the trajectory of software engineering, where agentic development is drastically changing how software is built and deployed. This year we’ve been focused on providing tools to support this shift, culminating in cf. cf has been built from the ground up with agents in mind, and includes novel tools for agentic command discovery that we think will become standard in more CLIs in the near future.
Wrangler came with the advantage that years of documentation, blogs, and third-party guides have been absorbed into the training process of LLMs. It also came with the same disadvantage: changing how Wrangler works now goes against learned behavior, and significant change would be inevitable given the scale of improvement we want to make.
Introducing a new CLI that agents have never seen sounds like a big disruptive change — but actually it’s the cleanest thing we can do. Because of the design decisions we have made, the context injections we can make, and the AGENTS.md files we can append, making a switch in this way is actually less confusing than having an agent contextualize the major differences between two versions of a tool it is familiar with. We’re launching with a couple of these agent-focused features built in, with more to come.
Agents need to filter JSON, not look at tables
When agents use Wrangler, they append --json to every command they run, and then often filter the output with jq to extract a subset of fields. But only some commands in Wrangler supported --json ; many commands returned unicode tables, designed for humans looking at output in their terminal. Agents can figure these out, but it costs them more time and tokens than a jq filter.
In cf we’re taking the opposite stance: agents just need JSON, and if agents are the future primary user of this tool, it should be the default. For the vast majority of commands that will rarely be accessed by humans, this is obviously the right call.
You as the human customer of this CLI are, in reality, one step removed from using it. Agents being able to easily filter their results and then return that filtered list in whatever format you request is preferable to supplying tables you will never likely read directly.
But what if you’re looking to do something that might require real personal input, like searching for a domain to buy?
For commands that your agent can access through chaining named parameters in a long and unwieldy sequence, you can simply fill in a form. Cf deconstructs the requirements of the API into a series of validated inputs, so buying a domain, even one with complex requirements, is simple to follow.
Or, if you insist, just ask your agent to do it.
Your agent can find the right command itself
With 3,000 possible routes through a CLI, how can your agent find the right operation it needs quickly without bloating your context? For this reason we have also added cf cli search.
This command allows your agent to ask in natural language what it needs to do, and a small search index will provide a list of appropriate commands, based on their API description and parameters. We automatically tell your agent about this command when it runs --help for the first time.
Configuration that type-checks your agent
Our new configuration format is based on TypeScript, which is easy for humans and agents to parse, and allows you to write your configuration programmatically.
Typed configuration is enormously helpful for agents. We’ve found that even with no prior context of the programmatic configuration format, agents are able to easily identify and edit the configuration on demand, even across elements like env which have dramatically changed from the same named feature in Wrangler. All agents that use LSP plugins, such as Claude Code and Codex, benefit from being able to interpret more about the configuration file format in context, and make much more accurate suggestions as a result.
Compare this to TOML, which had no accessible schema, or JSONC, which had a linked schema that agents rarely used.
Some Wrangler configuration files inside Cloudflare have been condensed by 40% from over 5,000 lines, with many custom environments per developer, to factory files that build each developer’s configuration more efficiently.
This is achieved through programmatically defining each environment from the same universal base, instead of copying env blocks as was typical in Wrangler. A simple Worker with multiple environments simply switches on the Vite-native mode argument to swap between one set of configuration and another.
A simple configuration that does this now looks like:
You can migrate your Cloudflare Worker to this new format through cf migrate.
We’re also providing a few helper functions to make building your Worker a breeze.
bindings gives you a simple place for your agent to discover all the developer platform has to offer. Everything — from environment variables to storage, database, and queues — can be auto-completed and explained by your editor.
Similarly, we have included a helper for triggers, which is the new way to define routes, queues, schedules, and email triggers for your Worker. Rather than having these scattered through your configuration file, it’s now simple to find, in a single block, the actions that could trigger your Worker to run.
defineConfig.worker is just the start here. Our intention with cloudflare.config.ts is that this is how you manage Cloudflare as a whole. Every product you need — along with its API being available to your agent through cf — will be able to be expressed through typesafe configuration. Soon you will be able to configure entire policies, set up zones, configure DNS and more, all through this configuration file.
A best in class development experience
When Wrangler first started building JavaScript Workers, Vite didn’t exist. Instead, we used esbuild in Wrangler to bundle your Workers. The dev server that Wrangler made available on :8787 was something that the Wrangler team built, and modifying any of this meant reaching into the internals of Cloudflare-specific local tooling like Miniflare.
Vite is a huge improvement on this, and comes with a large ecosystem of plugins you can use, as well as providing a best in class dev server with HMR (hot module replacement), and builds that use the Rust-based library Rolldown for tree-shaking. Anything you can do with Vite, you can do with the Cloudflare Vite Plugin.
The Cloudflare Vite Plugin is the recommended way we suggest you build Workers, whatever you are building: whether that’s a frontend-focused project or a backend API. Together with our Vitest plugin it provides a cohesive development and testing environment that matches the Workers runtime and gives you direct access to bindings and platform APIs.
cf is built on Vite as default. Most of your Workers will migrate simply with agents. Others may take more time, which is why cf will continue to delegate to Wrangler for dev and deployment for JavaScript Workers that need to continue to use esbuild and Rust and Python Workers.
Migrating from Wrangler
Migrating a Worker from Wrangler is as simple as running
Workers that already build with Vite will be converted to cloudflare.config.ts for you. If your Worker relies on Wrangler for esbuild, then cf will continue to delegate builds to Wrangler.
When the open beta ends we will release a final major version of Wrangler that directs you and your agent to use cf. We’ll continue to provide maintenance support for Wrangler for 18 months after the beta ends, to give you time to migrate.
You can also take new projects and automatically configure them for Cloudflare by running cf init/deploy, which will install the Cloudflare Vite Plugin for you and create a configuration file.
Static sites still don’t require a configuration file to start, and deploying them is as simple as running cf deploy in your project.
To start a new Hello World project with cf, use cf init.
cf is open source and issues can be reported to our GitHub repository.
How fast is the web? Explore billions of real-user measurements with BEACON
Post Syndicated from Ryan Townsend original https://blog.cloudflare.com/how-fast-is-the-web/
If you work in technology, you’re probably reading this on a powerful laptop or flagship mobile on robust, lightning-fast Wi-Fi. This is fantastic for building software, but often is far removed from the reality facing many who are using that software.
End users might be nursing a four-year-old budget phone, running low on battery, on a data plan that throttles after 2 GB, living with under-invested public infrastructure, or even just walking into that corner of the gym where the Wi-Fi never seems to work. Multiply this by billions of people around the world and the gap between “works on my machine” and “works for everyone” starts to widen, distorting critical decisions regarding our technology choices and priorities.
Closing the perception gap with objective data is central to our mission of helping build a better Internet, one that's fast and accessible to everyone, not just those of us using the best hardware.
That’s why today, we’re sharing a view that offers insight on how real people experience the web, by publishing the Cloudflare BEACON dataset — Browser Experience Across Cloudflare's Observed Network. Cloudflare has collected this kind of telemetry for years on behalf of our customers, giving them a customer-specific, detailed understanding of how real users experience their sites. Today, we're opening that view up to everyone.
BEACON is an anonymized dataset built from billions of real-world performance measurements across 10,000 of the largest websites on our network. It covers every major browser engine, is updated daily in Google BigQuery, and uses standards defined by the community-led RUM Archive, a publicly available Real User Monitoring (RUM) database. By expanding the footprint of that project 100-fold, BEACON gives researchers an unprecedented view of how the web performs across browsers, devices, and countries.
What BEACON reveals
The Core Web Vitals have long been the de facto standard for measuring perceived performance on the web, and BEACON reports all three:
- Largest Contentful Paint (LCP): load time
- Cumulative Layout Shift (CLS): visual stability
- Interaction to Next Paint (INP): interaction responsiveness
Because we’re publishing these as full histograms rather than single averages, you can derive any percentile you like. Instead of stopping at P75 (the 75th percentile), you can examine the long tail and see where the industry still struggles to deliver fast experiences for everyone.
Who experiences a slower web?
WebKit, currently the only browser engine on iOS, performs best on these metrics overall, but that advantage is not universal. In 46 countries where WebKit represents more than 10% of traffic, its LCP, INP, or both are at least 10% worse than those of Blink-based browsers such as Chrome, Edge, and Opera. In Cambodia, for example, WebKit accounts for 17.5% of page views, but its LCP is 50% worse than Blink’s.
BEACON also includes domain industry classifications from our Intel API. Government and Politics, Health, and Safe for Kids stand out as high-performing categories, while Ads, Religion, and Weather typically perform worst:
The additional percentiles expose differences hidden by a single P75 result. In several industries, the slowest experiences fall sharply in the long tail, particularly for visual stability as measured by Cumulative Layout Shift.
What makes pages feel slow?
BEACON extends the RUM Archive standard with LCP and INP sub-parts that separate the stages of loading and responding to an interaction. We’ll also add these metrics to our Real User Monitoring (RUM) dashboard in the coming weeks. Aggregating them into the suggested ‘Good’, ‘Needs Improvement’, and ‘Poor’ thresholds helps narrow down what needs to be optimized:
|
LCP Sub-part |
Document TTFB Nothing can be rendered until we have our HTML document. |
Load Delay Is JavaScript dependence slowing discovery of our LCP candidates? |
Load Duration Is bandwidth an issue, with the LCP image/video/webfont taking a long time to download? |
Render Delay When all is ready, is there something blocking the LCP from rendering? |
|
Good |
|
|
|
|
|
Needs Improvement |
|
|
|
|
|
Poor |
|
|
|
|
The query for the table above and others for every data is stored in BigQuery alongside the dataset as ‘Global LCP Sub-parts’ so you can customize it as you see fit.
The results challenge a common assumption: downloading the resource itself, such as an image, font, or video, typically contributes the least to perceived loading time. For most page views that breach the ‘Good’ threshold, the larger opportunities are discovering the LCP (load time) candidate and unblocking its render. Cloudflare customers can address some resource-discovery delays with Smart Hints.
The same workflow breaks down Interaction to Next Paint (INP), our measure of responsiveness, into input delay, processing time, and presentation delay:
|
INP Sub-part |
Input Delay Are our interactions waiting on the main thread becoming available? |
Processing Time Does processing the interactions themselves block the main thread? |
Presentation Delay How long does it take to paint any update to the screen? |
|
Good |
|
|
|
|
Needs Improvement |
|
|
|
|
Poor |
|
|
|
Query for above table stored as ‘Global INP Sub-parts’ in BigQuery
For the slowest interactions, JavaScript execution time covers the longest period but we also see a significant rise in presentation time, which is typically dominated by complex CSS layout recalculations. Cloudflare customers can use tools such as Zaraz to reduce the performance impact of third-party JavaScript.
How application architecture changes the picture
Speaking of JavaScript, we recently added support for Google Chrome’s new Soft Navigations API, which provides accurate LCP reporting for client-side navigations commonly used in single-page applications built with frameworks such as React, Vue, Angular, and Svelte.
|
LCP Percentile |
P50 |
P75 |
P90 |
P95 |
|
Hard Navigations |
|
|
|
|
|
Soft Navigations |
|
|
|
|
Query for above table stored as ‘Blink – Hard vs Soft Navigations’ in BigQuery
Soft navigations render two to three times faster than hard navigations at every percentile. But they do not eliminate the cost of the initial landing page, which is often considerably heavier:
|
LCP Percentile |
P50 |
P75 |
P90 |
P95 |
|
Landing Page |
|
|
|
|
Query for above table stored as ‘Blink – Landing Pages’ in BigQuery
For teams choosing this architecture, it’s important to be mindful of tradeoffs: faster subsequent navigations must offset a slower first experience. If users rarely progress beyond the landing page, a heavier initial load may never pay for itself.
What can researchers discover by combining data?
Because BEACON is an open dataset, its value is not limited to the fields it contains. Researchers can join it with other sources to explore new questions. For example, combining BEACON with the World Bank Group’s measure of GDP per capita reveals how economic conditions correlate with web performance by country:
Cloudflare Radar, our public data insights and visualizations platform, will start using this approach in a new Web Performance section on Radar, featuring correlations of their Internet Quality Index (IQI) data with BEACON data. IQI is an aggregation of the performance benchmarking data that powers our bi-annual network performance updates.
Pairing BEACON and IQI is particularly compelling because it splits the user experience into its two most influential components: the performance of the website a user is visiting, and the performance of the eyeball network getting them there. Both affect how quickly the page will load, and together they determine whether a visit to a website feels painful or seamless.
Early analysis shows the two tend to move together: good web performance usually coincides with good network quality, and vice versa. In the above graphs we see that the higher the bandwidth, the faster the perceived load time (LCP). For example in Europe, the bandwidth is higher relative to other continents, while the LCP is higher overall with the lower end.
The relationship between bandwidth and LCP was expected, but when comparing IQI to Transfer Size we saw something surprising. We would expect that transfer size would be uniform across continents — after all, the content of the sites are typically the same. However, see that Africa has a noticeably smaller transfer size, suggesting less content being downloaded by these users. Although we can’t be sure of the cause, we can observe that in the IQI data Africa has lower bandwidth, which suggests that businesses on the continent are adapting their websites to optimize towards the constraints of network conditions. High-quality eyeball networks are also potentially more forgiving to poorly-optimized websites, while a slow one exacerbates bottlenecks. Each of these hypotheses are potential directions for future analysis.
The new Web Performance section will explore how common these patterns are, and in particular how often one half of the experience cancels out the gains made by the other. Follow our progress on Radar.
We’ll continue to introduce more metrics and more dimensions over time, and we welcome requests for data you’d like to see next.
How we process all this data and make it useful
Anonymization
Publicizing a real-world dataset at this scale brings with it responsibility for privacy. Our Real User Monitoring (RUM) product is already built to be privacy-first, and for BEACON, we also strip out any potential customer website identifiers such as the domain name and URL paths. Ultimately, the community gains valuable insight without compromising the trust of the people and businesses behind it.
Normalization
For the primary table, including all websites from our data would lead to one of two problems:
- The largest sites would dominate the data by traffic volume, skewing performance metrics towards their architecture, visitor profiles etc., or:
- If we instead normalize every site down to the traffic volume smallest website, that would drastically reduce the overall number of records in the data.
These two extremes made it necessary to scope the dataset to the greatest number of the largest websites on our network to provide maximum diversity across architectures, technologies, geography, and more, all while retaining the total beacon count after normalizing. We found 10,000 to be a good balance: a globally representative sample with enough volume in the 10,000th that when we normalize the data down to their level, the overall dataset still represents billions of daily records collectively.
Aggregation
Finally, we aggregate records together where they share dimensions such as country, operating system, browser, and connection protocol, and discard any records with fewer than five data points to further guarantee no individual or specific site can be identified.
How to get access and contribute
BEACON is publicly available on Google BigQuery. We’ve included queries for all the data included in this post as examples you can adapt for your own analysis, and the RUM Archive website includes further documentation too.
We’d love to hear what you discover. Share your findings in our community forum or on Discord.
BEACON makes it possible to study web performance at a scale and level of geographic and browser diversity that has not previously been publicly available. We hope researchers, browser vendors, developers, and standards groups use it to identify where the web falls short and help make fast experiences available to everyone.
A special shout out to Cloudflare’s 1,111 intern project. This couldn’t have happened without the hard work of two of our wonderful summer interns, taking the initial idea through to what you see today. Their contributions were instrumental in launching this project. Thanks Chisara Duru and Tong Zhou!
Texas Midterm Elections #lastweektonight
Post Syndicated from LastWeekTonight original https://www.youtube.com/shorts/9vflU7z90Yk
Four months of VoidZero at Cloudflare: making the open-source JavaScript toolchain faster for all humans and agents
Post Syndicated from Evan You original https://blog.cloudflare.com/voidzero-update/
When VoidZero joined Cloudflare four months ago, we made a commitment to open source, promising that Vite, Vitest, Rolldown, Oxc, and Vite+ will stay open source, vendor-agnostic, and community-driven. As part of Cloudflare’s Birthday Week, where Cloudflare gives back to the Internet, we thought it’d be a good time to check in on how we’ve been doing against this commitment.
In the four months since VoidZero joined Cloudflare, we’ve shipped more than 80 releases, closed over 1,200 issues, and landed some big performance improvements, including:
- Oxc React Compiler, released in August, compiles React.js apps 10x faster
- Vitest 5, released in September, is up to 50% faster than Vitest 4
- tsgolint is now stable, and is up to 18x faster than ESLint in large codebases
- Oxfmt’s JSON, CSS, SCSS, Less, GraphQL, and YAML formatters are now written in Rust, making it 7x faster than Prettier
On top of this, Vite+, which unifies the entire toolchain with a set of great defaults, is now 1.0. And we’re making progress towards shipping “Bundled Dev” (f.k.a. Full Bundle Mode) — a new development mode that has been shaped by working with customers with massive web applications, including Cloudflare’s own dashboard. By being part of Cloudflare, our engineering team gets a much better view into scenarios that only occur in massive codebases.
VoidZero’s mission is to make the next generation of JavaScript developers more productive than ever before. And in 2026, that means making developers’ agents faster.
A faster developer experience for humans and agents
Developer experience and performance were always about improving the feedback loop while creating software. But now when we optimize for developer experience, we no longer just optimize for humans — we optimize for agents. And as inference gets faster, the speed of type-checking, linting, or building the code becomes the bottleneck again. The longer those processes take, the longer an agent has to wait before making progress and completing its goals.
VoidZero has always been focused on performance, but this new era of software has given us even more reason and motivation to make the entire toolchain faster for both agents and humans. And since the tools in the VoidZero toolchain are built on each other, from the compiler (Oxc) to the bundler (Rolldown) to the build tool (Vite), the linter (Oxlint) and the test runner (Vitest), any optimization at one layer automatically benefits everything built on top.
Oxc React Compiler — 10x faster compile times for React.js apps
We’ve recently shipped the Oxc React Compiler, a rewrite of the React Compiler, based on the React team’s rewrite in Rust. It is 10x faster than the original Babel implementation, uses less memory, and has a more complete implementation with better error handling. If you use Vite, you can enable the Oxc React Compiler by installing the oxc-transform-react package and enabling the compiler flag:
Vitest 5 — up to 50% faster than Vitest 4
Vitest 5, released in September, cuts test times by double-digit percentages across many common scenarios.
- Faster test runs. Vitest shares transformed files across projects, caches modules on disk, and sends less data between its main process and workers.
- vitest doctor. It breaks down setup, import, transform, and test time, then tests other configurations and recommends faster settings.
- Trace View. It records browser interactions, assertions, and DOM snapshots, so you can replay failures or inspect them in an HTML report.
- Conditional mocks with vi.when. Map arguments to responses without writing a manual
mockImplementation. - Better benchmarks. Benchmarks now work like regular tests, with fixtures, hooks, retries, filters, and assertions.
- Fewer false passes. Vitest fails unawaited async assertions, clears mock calls before each test, and adds the
--repeatsflag to help find intermittent failures.
tsgolint — now stable, up to 18x faster than ESLint in large codebases
tsgolint, the type-aware linting engine behind Oxlint, is now stable. It catches bugs that require TypeScript type data while running 12 to 18 times faster than ESLint with typescript-eslint.
- 59 type-aware rules. tsgolint now supports 59 of typescript-eslint's 61 type-aware rules.
- One pass for linting and type-checking. Oxlint can share one TypeScript program across both tasks instead of analyzing the project twice.
Enable type-aware linting and TypeScript diagnostics in your Oxlint config:
Oxfmt — 7x faster than Prettier with formatters now written in Rust
Oxfmt brings the speed of the Oxc toolchain to formatting. We rewrote its JSON, CSS, SCSS, Less, GraphQL, and YAML formatters in Rust, making Oxfmt many times faster than Prettier, while keeping Prettier-compatible output and ergonomics.
Bundled Dev — faster dev server for larger apps
We’ve made progress towards shipping “Bundled Dev” (f.k.a. Full Bundle Mode), which uses Vite’s production bundler during development. This should lead to significant dev server speed-ups in larger applications and reduce network overhead when working with remote sandboxes.
We’re looking forward to bringing Bundled Dev out of experimental status soon. Cloudflare’s Dashboard already uses Bundled Dev for all internal developers.
Vite+ is now 1.0 — a unified toolchain
Speeding up tools is one way of making humans and agents ship software faster. Another way is by reducing decision fatigue (“Which linter shall I use?”) and providing great defaults. We are excited to announce that Vite+ is now 1.0. Vite+ bundles all of VoidZero’s tools together into a single unified and fast toolchain.
Vite+ ships with Vite 8, Vitest 5, Rolldown, Oxlint, Oxfmt and task caching built in, and comes with great defaults. Check out the Getting Started guide to try it out today.
Cloudflare’s Open Source Investment
VoidZero was born in open source, and we believe a more sustainable open-source ecosystem benefits everyone. VoidZero is a proud member of the Open Source Pledge and as part of Cloudflare, we have the opportunity to expand that impact.
Cloudflare committed $1M to a Vite ecosystem fund to support maintainers and contributors in the original announcement. Since then, Cloudflare has committed an additional $1M to open source. Together, we’re doubling down on open source.
The open-source toolchain for the entire JavaScript community
There is more to come! Features we plan to ship in the next few months include major improvements to Oxc’s parser with up to 3x potential speedup, a re-designed chunking algorithm in Rolldown, and an open-source, self-hostable version of Void, the Vite-native deployment platform built on top of Cloudflare. Cloudflare and VoidZero both recognize our responsibility that we have to developers, and we do not take the community’s trust in us for granted. We’re in this for the long haul.
We are excited to keep shipping faster tools, and will continue to make every decision with the community in mind, just like we did when raising venture capital, or when we open sourced Vite+, or when we joined Cloudflare. Thank you for continuing to trust us with your projects and apps, supporting us, and building with us.
Introducing Forge: the open source pipeline for generating SDKs, CLIs, docs, and more
Post Syndicated from Dimitri Mitropoulos original https://blog.cloudflare.com/forge-open-source-generation-pipeline/
Today we’re introducing Forge, a fresh approach to generating SDKs, CLIs, docs, and libraries. Forge is an open source, pluggable generation pipeline that anyone can deploy and run for free.
Forge is early in its life, but already generates the output required for the cf CLI, and over the next few months will power Cloudflare’s API documentation, SDKs, and much more.
We built Forge because we needed it ourselves in order to treat agents as our customers. Now, we’re open sourcing it because we think everyone should be able to generate all the surfaces that agents need. It used to be that only developer products needed CLIs, API SDKs, MCP servers, all with great corresponding docs. Now these are table stakes for every product.
Our API outgrew our generators
Cloudflare’s API has over 3,500 operations, and the hundreds of services that power these APIs are written in many languages, including Rust, Go, TypeScript, and Python. As we embarked on building a CLI for the entire Cloudflare API, including our SDKs and API docs, we needed a code generation pipeline that could handle this scale. That pipeline needs to be flexible enough to work across languages and the ways each of our engineering teams operate.
We needed a way to reduce coordination overhead between teams. When a Cloudflare product team makes an API change, they need to be able to use a preview build of the Cloudflare-wide CLI, SDK, and docs site that will be generated, before merging that change and shipping to customers. We needed a way to ensure they didn’t inadvertently break the generation pipeline. And we needed a system that we could extend to generate more than just an SDK, from Cap‘n Web to MCP and beyond.
We’ve tried several hosted products that attempt to solve this, and relied on some in production. None of them solved this problem for us, and some have shut down entirely. One team would merge a change that inadvertently would break the generation pipeline, another team would discover this at release time, and we spent too much time swimming upstream through hosted tools we couldn’t control, coordinating changes between teams and vendors.
That’s how we started building Forge.
Forge seeks to fix all these problems: it runs in CI, on each team’s API repos, just like our AI code reviewer and test pipelines. It lints every change, and then generates preview builds of the CLI, docs, and SDKs with just your changes highlighted that you can install to test. It’s the same premise as Workers Previews: a full preview build for every change, but applied to SDK generation at scale, including when the API surface is distributed across hundreds of services and repositories. That’s what Forge seeks to deliver.
Forge transformers can generate anything, including Cap’n Web
Cloudflare has more reasons than most to want a generator that can go well beyond the normal language targets. Cap’n Web is Cloudflare’s RPC system that lets TypeScript call a remote API as if it were calling a local method:
Forge makes it possible to take an OpenAPI spec and generate Cap’n Web directly. This opens the door to generating bindings from Workers to other APIs. After all, bindings in the Workers runtime are implemented as Workers that expose RPC methods.
This isn’t specific to Cap’n Web: other popular tools you may already rely on need the same thing. If you use TanStack Query, you’d ideally want to be able to generate TanStack Query bindings for your application, built directly from your API itself. Always up to date, always validated against your real API. The same is true of generating Zod or Valibot schemas, MCP servers, or anything else that makes it easier to consume your API.
This is possible because Forge code generators are flexible. They’re built for flowing information from one output to another.
Forge transformers can be chained: generate outputs from other outputs
We’ve designed Forge to be pluggable, and support many input and output types. Forge provides CLI, SDK and docs generators, but there’s nothing stopping you from adding a transformer that generates a library-specific package or even a full dashboard or application. Forge supports OpenAPI as an input type today, but we’ve designed it to allow AsyncAPI, GraphQL, Cap’n Proto, Protobuf, or other input formats in the future.
This is about more than just compatibility: it lets you chain targets, using one target output to produce others. This is common in other generators where the CLI and Terraform targets are produced from the Go SDK. But what’s missing, and what Forge provides, is a way for the user to control this chaining system themselves.
We needed a solution for this ourselves, because our own cf CLI is written in TypeScript, which other SDK generators don't generally chain from for CLIs. But our own situation made us recognize the deeper problem: why should any SDK generator tool make this decision for you? Maybe you’re a Python shop, and you want the CLI to be in Python.
If you’re thinking “Well, but who cares if it’s in Python or not? The code is automatically generated,” it’s because CLIs are different. CLIs often introduce local-only behaviors that wouldn’t make sense in an SDK. Behaviors that you write by hand since they’re inherently not something backed by any API call. For example, the cf CLI has commands like cf dev and cf build that are added on top of the rest of the generated output. These commands need to call TypeScript APIs from other packages like Vite.
Now let’s add docs to the mix. If you’re generating your CLI and your docs purely from your OpenAPI spec, how do you feed those handwritten commands back into your docs, so they can be documented alongside the rest?
We couldn’t find an existing tool that does this today, and yet this is exactly what we need for cf. So we’re building it into Forge.
Change your API without breaking users
Forge is also setting us up for better API versioning. Cloudflare’s v4 API has been the one major version of our API for 10 years. Since then, it appears like we haven’t launched any new major versions, but by SemVer definitions we’ve made quite a few changes worthy of a new major version. At the same time, several operations across our API feature internal ‘v2’ tags or ‘beta’ identifiers that have long outlived that part of the product’s lifecycle.
After so many years of our v4 API, we’re keenly aware that a big new v5 would leave a lot of our customers behind. That’s why, with Forge releasing artifacts along the way, we’re working on an API versioning approach that allows us to release new major API versions without breaking old clients or SDKs.
We’ll have more on our SDKs very soon, including TypeScript, Rust, Python, Go, PHP, and Terraform. Especially Terraform. We know that upgrading any Terraform provider comes with its own set of rigor, and we’re going to put extra special care into the Terraform transition.
Critical tools should be open to all
We believe that building tools for APIs is a core part of the Internet, and you should be able to do that without needing a SaaS product. You should own your SDKs, CLIs, and docs. And if you generate them, then you should be able to do whatever you want, wherever you want, for free.
That’s why we’re making Forge available open source under the permissive Apache 2.0 license. We want people to join in on this journey with us, and contribute.
Or not? Maybe you want to keep everything to yourself. Go for it! You can run Forge on your own for any purpose, with custom modifications, for free, in private.
Acknowledgements: This project was also made possible by the design and implementation efforts of Dan Carter, Steven Chong, Krishna Paritala, and Shelley Jones.
The road to the agentic browser: A Kitesurf update
Post Syndicated from Celso Martinho original https://blog.cloudflare.com/kitesurf-update/
In August, we introduced Kitesurf, a browser for the agentic age that runs entirely on Cloudflare Workers. We built it around what agents need from the web, rather than carrying all the features and bloat of a browser designed for humans. If this is the first you’re hearing about it, we highly recommend you read the blog post where we introduced Kitesurf, for all the juicy technical details of how we did it.
Since then, we’ve put Kitesurf through increasingly realistic tasks and used internal and external feedback from customers to make it more capable and more efficient. Here’s what has changed, how you can try it today, and where we’re going next.
WebMCP support
Websites were not built for agents to use. Browsing today is a messy process of clicking pixels and hoping the right element loads. In a programmatic world this is slow and fragile. WebMCP helps by allowing developers to expose site functionality directly to agents, where they can call functions like searchFlights() instead of simulating clicks.
Cloudflare has been supporting WebMCP since its early days; just a few weeks ago we announced that site owners can now turn on WebMCP with one switch, so browser agents can discover and use tools on their sites without changing the site’s code, and Browser Run has been supporting WebMCP when using Chrome beta for some time now.
Today we are announcing that Kitesurf now supports WebMCP.
You can test this by going to our public Kitesurf playground, opening Cloudflare Radar, and navigating to WebMCP on the Application tab in the DevTools panel. As you can see, Radar exposes a list of WebMCP tools like navigate-to or set-location which allow clients to interact with the page and explore Radar programmatically.
If you target Kitesurf with your AI Agent:
You can see that the AI model can interact with the exposed WebMCP tools which you can use to complete tasks more reliably.
You can read more about how to use WebMCP with Kitesurf and Browser Run here.
New APIs, better WPT coverage
Since the initial announcement, we’ve been adding more browser standards so that agents can render more sophisticated pages. The list of the APIs that Kitesurf supports has grown, and now includes:
We added URL-based module resolution, JSON modules, and import map handling—important for sites that load JavaScript in chunks. Additionally we are using the new Cloudflare Workers’ module registry to support imports from URLs.
Iframe behavior has improved as well; now they load at the right time, stay better isolated, and display text correctly across more languages and encodings.
As we said at launch, running tests is how we keep the quality of both code and results under control without losing velocity while improving Kitesurf. Web Platform Tests (WPT) is a shared, open-source test suite that checks whether browsers implement web standards consistently.
We now pass 730,000+ WPT subtests and are growing. That’s 500,000 more subtests than when we launched. Here you can see the evolution over time, up to the latest version since we started the project:
Efficiency optimized for agents
For an AI agent, efficiency isn’t so much about loading pages fast, but about the latency of the agentic loop. To make Kitesurf truly agentic, we’ve aggressively optimized the browser engine’s internals so that every DOM traversal, timer, and font fetch is as lightweight as possible, ensuring the agent spends its compute cycles on reasoning, not waiting for the browser to catch up.
These optimizations include:
- Improved JavaScript execution with less work crossing between Boa and the DOM, making the boundary more compatible with real Web frameworks. Common reads such as getAttribute, id, and parentNode can now be answered inside the Wasm DOM instead of making repeated Boa → JavaScript shim → Wasm trips.
- Kitesurf does less repeated work when running timers and loading scripts, and releases memory from objects it no longer needs. It also handles objects and classes more consistently when code moves between its two JavaScript engines, resulting in running busy pages more efficiently.
- Kitesurf now loads fonts when they’re needed, fetches fewer fonts a page won’t use, checks which characters appear on the page before fetching language-specific font files and renders synthetic italics more faithfully.
Together, these optimizations help keep Kitesurf efficient for agents. Despite adding support for more web standards—and bringing Kitesurf closer to the capabilities of full-featured browsers like Chrome—its wall-clock time and CPU usage remain roughly in line with our launch benchmarks, and in some cases they have improved.
Plays better with Browser Run
Browser Run is our developer platform product that lets you programmatically control and run headless browser instances. When you use this API, you can select from a list of browser flavors we support, including Kitesurf.
This means that we have to make sure that all of our browsers are supported across the API surface. Starting today, Kitesurf has full Browser Run API coverage. You can use Kitesurf with CDP, Playwright, Puppeteer, or MCP.
One of the most popular Browser Run features, Quick Actions, provide simple interfaces for common browser tasks like capturing screenshots, extracting HTML content, generating PDFs, and more. When we launched Kitesurf, you could use Quick Actions from our REST API. Now you can also use them from inside a Worker script using the env.BROWSER.quickAction() binding:
Kitesurf runs in the terminal now
As we detailed in the How we built it section of our announcement blog post, Kitesurf separates PageScript, the isolate that handles the page session and runs the page code, from PageRenderer, which is responsible for generating the actual pixels from the computed page objects.
This not only gives great isolation and flexibility, but it also allows us to decouple and move the rendering logic to outside Kitesurf (to the client, for example, or to another Worker), while keeping the security-critical parts server-side, running in our network.
If this model sounds familiar, it may be because Cloudflare has another SASE product called Cloudflare Browser Isolation, which runs all untrusted web code at the edge of our global network while it “streams” the rendering data back to the clients.
We can do something similar with Kitesurf. To prove it, we moved PageRenderer to our Playground Worker and patched this version so that instead of converting scene data into an image, it outputs to Kitty—a terminal graphics protocol supported by modern terminals like Kitty itself, Ghostty, WezTerm, and others. We even went a step further and added a pure ANSI text mode for environments where Kitty isn't available.
The result is that you can now quickly open and render a page using Kitesurf without leaving the comfort of your terminal application. This is super useful not only because you can now browse the modern Web at a glance without context-switching, but you can also use this tool to see how an agent using Kitesurf “sees” a page.
To install the terminal version of Kitesurf, do this:
From now on just type this in terminal:
Here’s a demo of it working.
The terminal also sends back scrolling and click events, so you use the keyboard, arrows, or the mouse normally, as if you were in a dedicated browser application window.
And here is our Silent Space Marine Doom demo running in Kitesurf inside the terminal:
Where we go from here
We continue to iterate rapidly on the road to the best agentic browser for our customers and developers. Expect ever-better performance benchmarks and for the list of supported Web standards and WPT test coverage to continue to rise quickly. In fact, we’ve decided to publish the results here and here, in the open, so that you track them as we move forward, in real time.
We are going to continue exploring scenarios where decoupling Kitesurf and moving PageRenderer away from PageScript is an advantage for agents, or where higher frame rates are important. We may or may not have a 30fps Doom version running in Kitesurf as we write this.
We also want to address the elephant in the room: While we are currently prioritizing rapid development, we remain committed to open-sourcing Kitesurf. This is coming soon, but we want to do this right, so we are set up to support it for the long term.
Kitesurf stays true to its initial design goal: It runs entirely on top of Workers just like any other customer application does; that means we only use our publicly available features and APIs and have no access to any special privileges. This is not only a great way to dogfood and prove our own platform, but also the only way to make Kitesurf very cheap and scale automatically across the Cloudflare global network.
Give Kitesurf a try in the refreshed kitesurf.dev playground, and use it in your projects via Browser Run. It’s available for free while in beta, behind per-account limits. Keep an eye on our changelog and come chat with the team on Discord. Share your experience and send us feedback—we’ll be listening.
Introducing The Cold Start: pitch your startup live at Cloudflare Connect
Post Syndicated from Fatima Yusuf original https://blog.cloudflare.com/introducing-the-cold-start/
Sixteen years ago, Cloudflare was one of more than 1,000 startups hoping for a spot on stage at TechCrunch Disrupt.
We were not, on the face of it, an obvious choice. Cloudflare was infrastructure: we made websites faster and protected them from attack, which was not well understood by the general market at the time. Infrastructure is often invisible right up until the moment it becomes important.
But on September 27, 2010, Matthew Prince and Michelle Zatlyn got on the Startup Battlefield stage and launched Cloudflare to the public. During the presentation, people started signing up. Then more people started signing up. By the time the judges had finished asking questions, hundreds of websites had joined Cloudflare, putting our initial five data centers to the test in real time. In the seven days following, traffic through our network increased almost 10x and Cloudflare jumped from the 1,000th largest site online to one of the top 50.
Cloudflare didn’t win the main trophy that day. At the awards ceremony, TechCrunch founder Mike Arrington described what we did as something akin to "muffler repair for the Internet" and honestly, he had a point. But then he named us the Most Innovative Company. As Matthew wrote later: "You may not win the trophy, but you'll receive something else far more important."
There are moments in the life of a company when somebody gives you a room, a microphone, and a small amount of time to explain the thing you have spent months or years obsessing over. Most of the time nothing magical happens. Sometimes, though, the right people hear it at the right moment and suddenly an idea that has mostly existed between a handful of people starts moving through the world.
This October, as Cloudflare turns 16, we want to give five early-stage startups a stage of their own.
Introducing The Cold Start
The Cold Start is a live startup competition taking place next month at Cloudflare Connect in San Francisco. We’ll select five early-stage companies and give each of them five minutes on stage to explain what they’re building, why it needs to exist, and why they are the people who should build it.
We are less interested in perfect pitch decks than in interesting ideas clearly explained. You do not need thirty slides, a suspiciously precise TAM calculation, or a rehearsed story about how your childhood prepared you to disrupt accounts receivable. What we want is to understand the vision: what changed in the world that makes it possible, what you see that other people have missed, and why you cannot stop thinking about it.
The five finalists will make their case in front of the Cloudflare Connect audience and three people who've spent a fair amount of their lives thinking about companies, infrastructure, and the Internet:
- Matthew Prince, co-founder and CEO of Cloudflare
- Michelle Zatlyn, co-founder and President of Cloudflare
- Dane Knecht, CTO of Cloudflare
The judges will select one startup to win $500,000 in Cloudflare credits, take over a billboard in San Francisco, and receive an invitation to our VIP speakers dinner that evening.
Five companies, five minutes each, and a room full of people paying attention.
What are we looking for?
The Cold Start is open to ambitious early-stage startups that have raised less than $10 million. Beyond that, we are deliberately keeping the definition broad because the most interesting companies rarely arrive neatly categorized.
We want to see ideas that seem obvious once somebody finally builds them, and ideas that initially sound slightly unreasonable. We want infrastructure that appears boring until you realize everyone is going to need it; products that could not have existed a few years ago; strange new interfaces; new ways of building software; things aimed at enormous existing markets and things aimed at markets nobody has bothered to name yet.
Most of all, we want to meet people who have noticed something about the world and decided to do something about it.
As part of the application, we’ll ask you to tell us who you are, give us your one-line pitch, explain what you’re building and why, tell us about your funding and revenue stage, show us how Cloudflare fits into your stack, and point us toward anything else that helps us understand you and your work.
The goal is simple: make us understand why the thing you’re building should exist.
Applications are open now and close Friday, October 2, 2026.
Five minutes in San Francisco
The Cold Start will take place Monday, October 19, from 4:00-5:00 p.m. PDT at Moscone West in San Francisco, as part of Cloudflare Connect. Cloudflare will pay for the five finalists to fly to San Francisco for the competition.
Connect brings together people building and thinking about what comes next for the Internet. This year’s lineup includes AI pioneer Dr. Fei-Fei Li; Idealab founder Bill Gross; organizational psychologist and author Adam Grant; AMD CTO Mark Papermaster; Vue.js and Vite creator Evan You; Lovable co-founder and CTO Fabian Hedin; and OpenAI member of technical staff and creator of OpenClaw, Peter Steinberger.
For five young companies, we are reserving part of that stage. Each startup will have 5 minutes to pitch, followed by 3–5 minutes of questions by the judges.
There is something we like about that symmetry. Sixteen years ago, Cloudflare needed someone to take a chance on an infrastructure company with a difficult story to tell and give us a few minutes in front of the right room. Today, we are fortunate enough to have a stage of our own, and we want to pass that same opportunity on to companies that are just getting started.
Start small. Build something enormous.
There is a practical reason Cloudflare spends so much time working with startups: very small groups of people have an uncanny ability to attempt very large things.
The problem is that ambitious software increasingly depends on infrastructure that, not very long ago, only the largest technology companies in the world could afford to build for themselves. Global compute, storage, networking, security, real-time systems, AI inference, and the ability to survive the possibility that the thing you made suddenly becomes popular should not require a company to first become enormous.
We think you should be able to reach for those capabilities on day one.
That is part of the idea behind Cloudflare for Startups, through which eligible early-stage companies can receive up to $350,000 in Cloudflare credits for one year. It is also part of the reason we continue expanding Cloudflare’s developer platform: a tiny team should be able to build something on Tuesday and if the Internet decides it likes it on Wednesday, spend Thursday focused on the product rather than hastily becoming experts in global infrastructure.
Cloudflare began with its own slightly unreasonable premise: that the performance, security, and global compute available to the largest companies on the Internet should be available to everyone on day one. In 2010, we got a chance on a stage to explain why that matters.
Sixteen years later, we have a much larger network, a somewhat larger team, and significantly better circuit breakers.
Now we want to hear what you’re building.




































