Post Syndicated from The Atlantic original https://www.youtube.com/shorts/zBydMyArSfg
[$] Analyzing Rust programs with Charon
Post Syndicated from daroc original https://lwn.net/Articles/1097198/
Nadrieril is a long-time Rust contributor, and the maintainer of the rustc
pattern-matching infrastructure. During his involvement with Rust, he has
noticed a problem with the usability of the language: it is difficult to
automatically extract information from a Rust crate for use with other tooling.
The Charon project aims to fix that by providing a stable API for accessing
internal information from the Rust compiler.
Home Assistant 2026.10: The Best New Features and Changes
Post Syndicated from BeardedTinker original https://www.youtube.com/watch?v=U7dWtoP22Go
Why American Cities Cost So Much and Deliver So Little | The David Frum Show
Post Syndicated from The Atlantic original https://www.youtube.com/watch?v=yRkz854fKx0
I Stopped Flashing ESPHome on Sonoff — OpenEdge Changes Everything
Post Syndicated from digiblur DIY original https://www.youtube.com/watch?v=xS1kNcHtUDg
[$] Evolving the LAVD scheduler from gaming to servers
Post Syndicated from corbet original https://lwn.net/Articles/1097209/
The extensible scheduler class, which
enables the creation of custom CPU schedulers with BPF, has led to a burst
of innovation in this area; the
LAVD scheduler has, perhaps, been one of the most noteworthy schedulers
to emerge. Though it was originally designed
for gaming applications, the LAVD scheduler has since grown to serve
other types of workloads as well. At the 2026 edition of Kernel Recipes, Changwoo Min
and Gavin Guo presented an overview of this scheduler and how it has
evolved over time.
A new Raspberry Pi Desktop release is finally available for x86-64
Post Syndicated from jzb original https://lwn.net/Articles/1099206/
Simon Long has announced
a long-awaited release of Raspberry Pi OS, based on Debian 13 (“trixie”),
for x86-64 systems.
We managed to find the time to update the Desktop for the Buster and Bullseye
releases of Debian, but then we all just got too busy with other things, and,
while we left the Bullseye version on the website for anyone who wanted it, we
simply didn’t have time to release any newer versions. But people kept on asking
for it – we get two or three emails every week asking when the PC Desktop will
be updated, and we haven’t had an answer, because we honestly didn’t know when
we might get a chance to do it. We’ve continually tried to allocate time to be
able to work on this, but it hasn’t been easy. […]Earlier this year, we (or rather Serge) finally got the latest version of the
Desktop running on top of a Debian Trixie image. It’s now based on 64-bit Debian
(the amd64 architecture) rather than the older 32-bit version, as Debian itself
has stopped supporting 32-bit for PC architectures. This shouldn’t be a major
problem – most PCs made in the last 15 years or so will quite happily run the
64-bit version of Debian, as will most Intel-based Macs. (Debian support for
Apple Silicon is still experimental, so unfortunately those of you with the
latest and greatest shiny fruit products will not be able to run this.)
Security updates for Wednesday
Post Syndicated from jzb original https://lwn.net/Articles/1099202/
Security updates have been issued by AlmaLinux (bind, dovecot, freerdp, kernel, mariadb-connector-c, mod_auth_openidc, nodejs22, nodejs:22, sudo, and vim), Debian (node-shell-quote, puma, rails, ruby-jwt, and suricata-update), Fedora (chromium, cockpit, flocq, freerdp, gappalib-coq, golang-x-mod, httpd, janus, libical, musescore, python3-docs, python3.14, python3.15, rocq, rocq-stdlib, tesseract, why3, and zenon), Mageia (srt and tor), Oracle (freerdp, kernel, libpcap, mariadb-connector-c, nodejs22, sudo, and vim), Red Hat (expat and grafana), Slackware (openssh), SUSE (chromium, docker-stable, firefox, jupyter-jupyterlab, libtcnative-1-0, tomcat, tomcat10, libtcnative-1-0, tomcat11, openexr, python310, python313-azure-storage-queue, python313-langchain-anthropic, python313-sglang, python313-Werkzeug, and valkey), and Ubuntu (fluidsynth, freerdp3, freetype, golang-1.18, golang-1.21, golang-1.24, gst-plugins-good1.0, libsoup2.4, libsoup3, libwebsockets, linux, linux-aws, linux-gcp, linux-gke, linux-ibm, linux-oracle, linux-realtime, linux-azure, linux-azure-fde, linux-nvidia-tegra, linux-oem-7.0, redis, sg3-utils, tesseract, and u-boot).
Sierra Screamin’ 3D Graphics Accelerator from 1996
Post Syndicated from LGR original https://www.youtube.com/watch?v=o0wDkCp3Eu4
CVE-2026-21589: Critical unauthenticated arbitrary file access in Atlassian products
Post Syndicated from Rapid7 original https://www.rapid7.com/blog/post/etr-cve-2026-21589-critical-unauthenticated-arbitrary-file-access-in-atlassian-products
Overview
On October 5, 2026, Atlassian published a security advisory for CVE-2026-21589, a critical arbitrary file access vulnerability affecting eight products: Bitbucket Data Center, Confluence Data Center, Jira Service Management Data Center, Jira Software Data Center, Bamboo Data Center, Crowd Data Center, Crucible, and Fisheye. Atlassian assigned the vulnerability a CVSSv4 score of 9.3. An unauthenticated remote attacker who knows a target file’s exact name and path can access it within the application’s web root; the vulnerability does not provide directory listing or enumeration.
Atlassian’s advisory treats all versions before the applicable fixed releases as affected, including unsupported versions. Affected Atlassian Cloud products have already been patched, and no action is required from Cloud customers.
Detailed technical analysis and file-read proof-of-concept scripts are public, so Rapid7 recommends patching on an emergency basis, outside of normal patch cycles, and reviewing access logs for attempted exploitation.
Technical overview
NVD lists files or directories accessible to external parties (CWE-552) as the weakness associated with CVE-2026-21589.
On October 6, watchTowr Labs published a technical analysis based on comparisons of vulnerable and patched Jira, Confluence, and Bitbucket packages. Their analysis identified a path traversal vulnerability in Atlassian’s web-resource handling: double-colon (::) sequences can become path separators during request processing, allowing traversal components to reach the resource-loading code, resulting in the contents of arbitrary file being read back to an attacker.
Their testing could not traverse outside the Tomcat context, but could read files throughout the application web root. In an Atlassian Crowd deployment that had Jira configured, reading WEB-INF/classes/crowd.properties exposed application credentials. With network access to Crowd, they used those credentials to create a user and add it to jira-administrators; Crowd’s IP allowlisting can block this direct route.
Mitigation guidance
Organizations should upgrade each affected installation to a listed fixed version or the latest available version. Atlassian’s October 5 advisory lists the following fixed versions:
|
Product |
Fixed versions |
|---|---|
|
Bitbucket Data Center |
9.4.26, 10.2.8, 10.5.1 |
|
Confluence Data Center |
9.2.26, 10.2.19 |
|
Jira Service Management Data Center |
5.12.40, 10.3.26, 11.3.12 |
|
Jira Software Data Center |
9.12.40, 10.3.26, 11.3.12 |
|
Bamboo Data Center |
10.2.24, 12.1.12 |
|
Crowd Data Center |
6.3.7, 7.0.3, 7.1.7, 7.2.4 |
|
Crucible |
4.9.15 |
|
Fisheye |
4.9.15 |
Organizations unable to patch immediately should remove affected instances from the internet or otherwise restrict them from external network access. Atlassian provides a Web Application Firewall or proxy rule for all affected products, a Tomcat RewriteValve mitigation for Confluence, Jira Service Management, Jira Software, Bamboo, and Crowd, and a separate urlrewrite.xml rule for Bitbucket. These mitigations are limited and are not replacements for patching.
Rapid7 strongly recommends looking for signs of compromise even after the patch has been applied. Atlassian recommends URL-decoding each access-log request line up to twice, then searching for .. immediately adjacent to /, \, or ::. Alternatively, search raw logs with the vendor-supplied regex:
(?is).*(?:/|\\|::|%(?:25)*(?:2f|5c)|(?::|%(?:25)*3a){2})(?:\.|%(?:25)*2e){2}(?:/|\\|::|%(?:25)*(?:2f|5c)|(?::|%(?:25)*3a){2}|;|%(?:25)*3b|$).*
Public testing artifacts for Jira, Confluence, and Bitbucket include a Python file-read PoC and a Nuclei template. If investigation identifies access to protected configuration files, organizations should rotate exposed credentials and other secrets after containing the affected systems.
For the latest mitigation and investigation guidance, please refer to the vendor security advisory.
Rapid7 customers
Exposure Command, Vulnerability Management, and Nexpose
Exposure Command, Vulnerability Management, and Nexpose customers can assess exposure to CVE-2026-21589 with unauthenticated vulnerability checks on Jira Software Data Center expected to be available in the October 8 content release.
Updates
-
October 7, 2026: Initial publication.
Capture of CSS Florida
Post Syndicated from The History Guy: History Deserves to Be Remembered original https://www.youtube.com/watch?v=M5i04SaPdHM
Apple’s Verified Photography System
Post Syndicated from Bruce Schneier original https://www.schneier.com/blog/archives/2026/10/apples-verified-photography-system.html
Apple just released a system called “Reference Image.” It can verify the image is exactly as taken by an iPhone—new models only—without tying it to a specific iPhone or photographer. It can also verify that multiple images came from the same iPhone.
Other industry solutions require a photographer or institution to vouch for an image using their own credentials. We are concerned this puts some photographers, such as those operating in conflict zones, in a difficult position; it should not be necessary to forgo anonymity in order to prove image authenticity. We built Apple Reference Image to avoid using an explicit, public credential for photographers, and to avoid even implicit public association between different photos taken by the same sensor. The final reference image is instead signed by Apple’s signing service, after validation by PCC. That signature is backed by Apple’s strongest technical guarantees.
Our implementation also protects the confidentiality of the image itself, including from Apple. Merely capturing a reference image should never expose the actual pixels to Apple or anyone else. We achieve this through the exceptional privacy properties of PCC the nodes themselves are architected so that not even Apple can access image data, just as Apple cannot see the information processed for Apple Intelligence in PCC. While the revocation service must maintain a private record of photo GUIDs and associated sensors to allow for revocation, it never has access to the image data, and does not allow for public access to this record. And as final revocation checks occur using on-device lists, a device never reveals to anyone which photo it’s looking at in order to find out whether it’s still valid.
The report makes for good reading; the details are interesting.
The ASOS incident: When attackers use the channels customers trust
Post Syndicated from Emma Burdett original https://www.rapid7.com/blog/post/it-asos-incident-attackers-using-channels-customers-trust
ASOS customers opened their phones to find a hostile push notification delivered through the retailer’s own app. The message claimed the company’s Snowflake environment had been compromised and directed ASOS to engage with the sender through Telegram. ASOS later confirmed to Sky News that an unauthorized customer notification had been sent and said it was investigating activity involving third-party platforms used to communicate with customers. The company also said basic personal information, including names and contact details, may have been accessed, while payment-card information and account passwords were not believed to be affected.
The attackers’ wider claims remain unverified, and Snowflake told Sky News that its investigation had found no compromise of the Snowflake platform at that point. Even without knowing the full route into ASOS’s environment, though, the notification raises a useful question for security teams: what happens when an attacker can communicate through a channel customers already trust?
When the message comes from the real app
Most security awareness advice assumes there will be something suspicious for the recipient to notice. The sender might be unfamiliar, the domain slightly wrong, or the request out of character. Those checks become much less useful when the message arrives through the genuine app sitting on someone’s phone.
Attackers have already been moving in this direction elsewhere. Rapid7 research into calendar-based phishing showed how malicious content can appear inside familiar workflows, while our earlier look at how social engineering is evolving explored the growing use of collaboration tools and other everyday platforms to make attacks feel routine.
The ASOS incident moves that problem into a customer-facing environment. Once an attacker has access to a system that can speak on behalf of a business, the trust built around that system can work in the attacker’s favor too.
“Let’s face it, an attacker would much rather borrow trust that already exists than spend time building their own. Our recent Zimbra research is a good example, because once you can impersonate a sender or edit a calendar from inside the platform, everything the victim checks lives in a system they have no reason to question. I can’t say how this one happened, but a notification coming out of a real app gives an attacker that same head start. There is no strange domain or unfamiliar sender to catch, so the activity can look a lot like a normal Tuesday afternoon.” Douglas McKee, Director, Vulnerability Intelligence at Rapid7
What suspicious activity looks like inside legitimate services
An attacker does not always need obviously malicious infrastructure to create damage. A legitimate account, integration, or SaaS platform used in an unexpected way can provide access to employees, customers, or partners while generating activity that may look relatively ordinary when viewed on its own.
If a customer communications service suddenly sends an unusual notification, the security team needs to understand what happened around it: who accessed the platform, whether credentials or permissions changed, which connected services were involved, and whether suspicious activity appeared elsewhere in the environment.
ASOS said the activity involved third-party platforms used for customer communications, while TechRadar reported that the claimed Snowflake connection could potentially have been indirect through services running on the platform rather than evidence of a compromise of Snowflake itself. That kind of environment can leave investigators working across several providers, identities, and systems before they have a complete picture of what happened.
MDR has to follow the activity across the environment
When attackers use legitimate identities, integrations, cloud services, or communication platforms, analysts need to connect behavior across systems rather than depend on a known-bad IP address or malware signature to tell the story. An unexpected authentication, a permission change, third-party access, or unusual activity from a customer-facing service may not be enough to raise the alarm independently, but the sequence can reveal a much clearer pattern.
A preemptive MDR approach brings those signals together across endpoints, identities, cloud environments, and other parts of the attack surface so analysts can investigate the activity in context. Businesses now rely on a growing number of SaaS services and external platforms that can act on their behalf, and although security teams may not operate every one of those systems directly, they still need to understand what access they hold, how they connect to the wider environment, and how misuse would surface.
That becomes particularly relevant when a third-party service can communicate externally in the organization’s name. Access to the platform is only part of the picture; teams also need visibility into how that access is being used and whether activity elsewhere suggests the account or integration has been compromised.
The first message can create a second wave of risk
A visible incident can give other attackers useful material. Once customers know something has happened, a phishing email or text offering an account update, refund, password reset, or security check immediately has a credible event behind it.
Rapid7 research into digital footprint exposure has shown how breached information can be combined with publicly available data to support more convincing phishing, impersonation, and fraud. Names and contact details may appear relatively limited compared with passwords or payment information, but they can still become valuable when combined with a real incident and a recognizable brand.
The investigation therefore has to support several decisions at once: understanding the technical scope, establishing which customer or business data may have been involved, working with third-party providers, assessing regulatory obligations, and preparing for the possibility that the incident will be reused in follow-on attacks.
As more customer communication moves through apps, SaaS platforms, automated workflows, and third-party services, security teams need visibility into how those channels are being used as well as who can access them. The earlier unusual activity can be connected across those systems, the more room analysts have to investigate and respond before a trusted channel becomes part of a much larger incident.
Comic for 2026.10.07 – Resuscitate
Post Syndicated from Explosm.net original https://explosm.net/comics/resuscitate
New Cyanide and Happiness Comic
How a Democratic Flip in Texas Could Alter the 2028 Political Map
Post Syndicated from The Atlantic original https://www.youtube.com/shorts/nErwwQAcBTA
64-Day Certificate Lifetimes Coming Feb 2027
Post Syndicated from Let's Encrypt original https://letsencrypt.org/2026/10/07/64-day-certs.html
On February 10, 2027, all Let’s Encrypt subscribers will move to certificates with 64 day lifetimes by default unless they select an even shorter lifetime (45 or 6 days, as previously announced). This means that any certificate we issue or renew on and after that date will have a 64 day validity period, and we expect the last 90-day certificate to expire on May 11, 2027. We will not revoke valid certificates as a part of this process.
We will switch to issuing 64 day certificates in our staging environment on October 14, 2026 to enable testing. We recommend testing in staging before the change takes effect in production.
If your renewals are automated and your client supports ACME Renewal Info (ARI), you should be all set since ARI allows Let’s Encrypt to tell your client when to renew (you can review your ACME client’s documentation to determine if ARI is implemented).
If your renewals are hard-coded to a date from expiration you should update them to renew at approximately ⅔ of the lifetime instead. Taking this step in preparation for 64 day lifetimes will lay the groundwork for default lifetimes of 45 days in 2028. Grep for common hardcoded numbers like 83, 80 or 60 in cron jobs, wrapper scripts and runbooks if you’re not sure.
We will also be reducing the authorization reuse period from 30 days to 10 days. In 2028, the reuse period will shrink to seven hours. We are making this change to comply with a 2029 reduction in maximum validation reuse periods, and to remove the need for “CAA rechecking”, where we have to repeat part of the validation process if the validation data is more than 7 hours old. Unless you have specifically designed your ACME client to rely on validation reuse, you will not need to make any changes.
This is also an opportunity to automate certificate management processes like reload and deployment and to add alerting for renewal failures.
Rate limits will not be impacted by this change; you can learn more in our previous blog post.
This change will not affect ACME endpoints or our issuance chains.
We are moving to shorter certificate lifetimes because this reduces the risk of key compromise and mis-issuance. As a nonprofit we see it as part of our mission to make this change to advance security for everyone using the Web globally. We anticipate a smooth transition, but if you experience issues, our community forum and documentation are good resources.
Juice
Post Syndicated from xkcd.com original https://xkcd.com/3308/

Building Git infrastructure for agent-scale development
Post Syndicated from Brian Celenza original https://github.blog/engineering/architecture-optimization/building-git-infrastructure-for-agent-scale-development/
Every day on GitHub, millions of developers build the products their customers rely on, contribute to open source, and pursue personal projects. GitHub’s architecture has changed steadily over the years to support that work and the growing demands of the developers and organizations who depend on it.
Agentic software development is driving the next architectural shift. With developers and agents working concurrently in repositories that receive millions of commits a day, these workloads demand a different Git architecture. We’re rebuilding GitHub’s Git infrastructure to support them. This post explores the demands shaping that work and the design principles behind it.
Today’s highest-volume workloads show the scale we’re building for. The gap between a typical repository and the busiest ones is wider than most people expect. Here’s a rough picture of the monthly repository activity distribution on GitHub from August 2026:
Repository activity climbs sharply at the far end of the distribution. The busiest repository on GitHub saw roughly a billion requests in August.
Beyond these highest-volume workloads, total Git activity on GitHub is also growing rapidly: between September 2025 and August 2026, it increased to more than 2x its previous level, from 218.2 billion events per month to 473.3 billion.
In September alone, developers and agents made 7.38 billion commits on GitHub, more than five times as many as a year earlier.
The repositories at the top of this curve show what agentic development looks like at its leading edge: large engineering teams running busy CI pipelines alongside growing fleets of agents. Supporting these teams means building Git infrastructure for sustained, concurrent reads and writes at a scale few repositories reach today. We’re investing deeply in Git infrastructure to meet the demands of agentic software development and give teams a foundation built for their most ambitious workloads.
Building for this scale means addressing several architectural challenges:
- Commit turnaround becomes a bottleneck per agent. An agent in a tight loop commits or checkpoints after nearly every action. Its speed is bounded by how fast a single push completes, so latency that a human would never notice becomes the limiting factor.
- Write throughput demand is increasing by orders of magnitude. Pushes grew 4.9x year over year, from 0.69 billion to 3.35 billion per month. Thousands of agents working on their own branches in one repository produce a sustained write rate that converges on a single point in our architecture.
- Merges contend on one reference. Trunk-based development, release trains, and merge queues funnel all that work onto a single ref that has to absorb every merge. Pull request merges on GitHub grew to nearly 4x their volume a year ago.
- Each push multiplies into thousands of reads. For example, CI and code scanning clone or fetch the same branch tip thousands of times per minute, and that fan-out has to be cheap. GitHub Actions alone ran 3.26 billion times in September, more than 4x as many as a year ago.
- Operations on a repository must continue to be fast. To keep them fast, we continually compact repository data and clean up objects that are no longer needed. Every new write adds to that work, and the cost compounds as volume climbs.
This is why fast clones only solve part of the problem. Reads are relatively easy to scale: add caches, add replicas, and serve the same bytes to more clients. Scaling reads is essential, but these workloads require more than that. Writes are way harder. Every push has to be stored durably and made visible consistently before the next agent or CI job can build on it.
Where today’s architecture meets new demands
The current architecture has served developers well for years. Every repository is stored by Spokes, which keeps a full copy on the local disks of several fileservers, five by default. Those fast local disks let Git operations read native repository data with low latency, and the extra copies provide redundancy while spreading reads across fileservers. When a push updates a reference, a three-phase commit protocol uses a quorum to ensure that CI, the web UI, and API clients see a consistent repository state. That pairing serves a billion repositories today.
However, the mechanism we use for durability is the same one we use for scale. The copies on disk are the source of truth, so adding read capacity means adding another durable replica. Every replica participates in every write, so a push is only as fast as the slowest replica in its set. The net effect: adding replicas to absorb read load makes writes slower.
For most repositories, this tradeoff works well. At the highest activity levels, it becomes a ceiling: adding read replicas adds overhead to writes, losing a replica reduces read capacity, and losing quorum stops writes entirely. To meet agent-first demands, we need to separate durability from scale without losing what teams rely on today.
Built for the busiest, better for everyone
We’re rebuilding the infrastructure while GitHub keeps running. There’s no maintenance window where the world’s code stops moving, and no version of this work where we ask people to change how they build software while we do it.
We’re building for the most demanding workloads on GitHub: an enterprise shipping under strict regulatory requirements, a team landing a change across a repo that builds an operating system, and an organization running thousands of agents against a single codebase. Engineering for that scale raises the floor for everyone. The maintainer reviewing contributions from volunteers across time zones and the student opening their first pull request get the same faster, more resilient foundation.
The new architecture also must preserve the controls teams already operate on. A maintainer needs branch protections and required reviews so an unreviewed change never reaches the default branch. A security team needs audit logs and repository visibility to investigate a suspicious access event. An on-call engineer needs dependable automation and enough observability to understand why a deployment failed.
For the platform to keep serving everyone here while it scales for the busiest workloads, these are our guiding principles:
- Build on the workflows developers already trust. Teams rely on workflows like branching, review, merge, and history to build, ship, and govern software at scale. Our new infrastructure is designed to support those same workflows at much higher volumes of activity.
- Put reliability first. We measure every decision against the reliability that developers and organizations require. Confidence in the platform is what lets an engineering organization build automation, ship on a schedule, meet compliance obligations, and understand the software it produces. This work will meaningfully improve throughput and scale, and those gains extend a foundation of trust that’s already there.
- Keep people in control of their code. If the system isn’t helping the people and organizations who use it, and isn’t under their control, it isn’t worth building. As agents take on more of the work, the people who own the code can still review, understand, and approve it.
The approach
We are building a new GitHub architecture that can scale much more effectively. Our approach centers around core distributed systems design tenets, applied to the concurrency and scale of agentic software development. Our goal is to continue the forward momentum for open-source communities and enterprises around the world who have built their projects with Git and GitHub, while preserving and adapting the features and controls around it to meet the new needs of the agentic era.
Minimize coordination
A repository that receives many pushes must accept and publish updates quickly. Coordination is valuable when it protects correctness, but too much of it limits write throughput and can turn a busy repository into a bottleneck. Our current architecture is tightly coupled in places it doesn’t need to be, which limits our ability to scale across reads and writes without tough trade-offs. We’re redesigning the system to preserve the coordination that Git semantics require and let everything else proceed independently.
- Coordinate only what needs agreement. The part of a push that truly needs agreement is the reference update itself. Storing the underlying objects, validating object connectivity, and secret scanning are much more work, but most of it can happen in parallel to other writes. That shrinks the critical path of a push to the small step that needs coordination, so the rest of the work no longer delays the acknowledgment.
- Move maintenance off the serving path. Compaction and garbage collection are among the heaviest work a repository does, and today they run on the same hosts that answer live Git requests. In the new architecture, separate workers handle maintenance directly against durable storage. A busy repository can be optimized continuously in the background without slowing pushes and fetches.
Decouple storage from compute
Today, complete repository copies on local disks serve both as durable storage and as the layer that answers Git requests. Separating the two lets us scale each one independently.
- Scale reads without adding durable copies. In the new architecture, read capacity comes from lightweight workers that cache data to serve requests. The authoritative copy of the repository lives in a durable storage layer underneath. That way, the platform can absorb large read spikes from CI fan-out, agent fleets, and large clones without adding work to every push.
- Let each layer do one job. Authoritative repository data lives in Azure Blob Storage, which already provides durability and replication at Azure scale. The compute layer is optimized for throughput at the lowest latency.
- Recover faster from failures. When storage and compute are coupled, losing a host reduces both capacity and durability, and recovery means rebuilding a full repository copy. When they’re separate, losing a compute worker is closer to a cache miss: a replacement worker can start serving requests right away and fill its cache from durable storage as traffic arrives.
- Match capacity to demand. Compute workers can be added or removed as traffic changes instead of provisioning for peak load in advance. A repository going through a burst of activity, like a release or a new agent fleet coming online, can get extra capacity for the burst. Once it passes, that capacity goes away.
Together, these tenets allow us to support higher throughput and more concurrent work without abandoning the reliability and controls that our users need.
What comes next
We’re building an architecture designed to provide the highest throughput and reliability available. Reads and writes scale independently, and the system recovers gracefully from failures. In internal benchmarks, it has delivered up to 35 times higher write throughput, with read capacity that scales on its own to meet demand.
As automated development increases the frequency and concurrency of software change, GitHub will evolve its foundations without trading away the governance and control that teams rely on. We’re already putting that foundation in place. In the next post in this series, we’ll dive deeper into our future architecture and the journey that led us there.
The post Building Git infrastructure for agent-scale development appeared first on The GitHub Blog.
Improving SPIRE security and resiliency with AWS managed services
Post Syndicated from Brendan Paul original https://aws.amazon.com/blogs/security/improving-spire-security-and-resiliency-with-aws-managed-services/
In cloud-centered environments, establishing trust between workloads is fundamental to securing machine-to-machine communication. Traditional approaches such as API keys, shared secrets, and static service account credentials weren’t designed for the scale and ephemeral nature of cloud workloads. This drives an organizational need to shift from long-term, static credentials to short-lived cryptographic workload identities. Organizations are turning to SPIFFE (Secure Production Identity Framework for Everyone) as a core component of their workload identity strategy. SPIFFE is a set of open source standards for securely identifying software systems in dynamic and heterogeneous environments. Using SPIFFE can help address what’s known as the bottom turtle problem: the circular dependency where protecting one credential requires yet another credential.
SPIRE (the SPIFFE Runtime Environment) is an open source implementation of SPIFFE. Customers can use it to quickly experiment with the framework. However, there are operational and security considerations when deploying SPIRE in production, such as:
- Where cryptographic signing keys for workload identities are managed
- How to ensure availability and resiliency of the workload identity registry
- How the SPIRE Certificate Authority integrates into the existing enterprise PKI
- How workloads without direct connectivity to the SPIRE server validate identities
- How workload identities are delivered to serverless environments
- How fine-grained authorization based on SPIFFE IDs is enforced
In this post, we show you how to address each of these considerations by offloading SPIRE core functionality to AWS managed services. This post is accompanied by a Github Repository that guides you through the deployment of the reference architecture.
To learn the core concepts the SPIFFE framework, take a moment to familiarize yourself with the SPIFFE documentation.
SPIRE background
A SPIRE deployment is comprised of at least one SPIRE Server, at least one SPIRE Agent, and at least one workload.
Figure 1: SPIRE high-level architecture
The SPIRE Server is deployed on a central control plane instance (or instances) and manages identity issuance and stores workload identity registrations. The SPIRE server handles five key functions:
- RegistrationAPI: Create, update, and delete registration entries that define which workloads are entitled to specific SPIFFE IDs based on selectors (for example, Kubernetes namespace, Unix UID).
- DataStore: The SPIRE server uses a data store to keep track of the workload identity registration entries and the status of the SVIDs it has issued.
- KeyManager: Controls how the server manages private keys used to sign X.509-SVIDs and JWT-SVIDs.
- BundlePublisher: Publishes the local trust bundle to a store. The trust bundle is an object containing a trust domain’s cryptographic keys
- UpstreamAuthority: Dictates which root certificate authority is used to sign the SPIRE certificate authority (CA) certificate
The SPIRE Agent is distributed across your workloads, whether they run on Amazon Elastic Compute Cloud (Amazon EC2) instances, Amazon Elastic Container Service (Amazon ECS) containers, or Amazon Elastic Kubernetes Service (Amazon EKS) Pods. The SPIRE Agent handles:
- WorkloadAPI: Exposes the workload API to workloads, which enable workloads to retrieve their SVIDs and trust bundles from the SPIRE server.
- NodeAttestation: SPIRE requires that each agent attest and verify itself.
- WorkloadAttestation: The SPIRE agent collects workload metadata to determine its identity.
- SVIDStore: The SPIRE Agent also stores SVIDs in other destinations, such as AWS Secrets Manager or HashiCorp Vault making them available for workloads running in serverless environments.
After the server and agent are deployed, applications retrieve short-lived SPIFFE Verifiable Identity Documents (X.509 Certificates or JSON Web Tokens (JWTs)) using the workload API.
How to use AWS managed services for core SPIRE functions
In the following sections, we dive deeper into what each of these SPIRE functions does and the advantages of using each AWS service for the function.
Figure 2: SPIFFE architecture using AWS managed services
To deploy the infrastructure required to follow along with this post, follow the deployment process in the Github Repository.
Note: For each managed service, there’s a parameter in the source AWS CloudFormation templates that you can use to specify whether to provision the resource. For example, if you already have an AWS Private Certificate Authority (AWS Private CA) certificate authority deployed in your environment, you can use that for your SPIRE implementation and avoid provisioning a new certificate authority.
SPIRE Key Manager using AWS KMS
The SPIRE Key Manager controls the cryptographic keys that are used to sign SPIFFE Verifiable Identity Documents (SVIDs), which are either X.509 certificates or JWTs.
By using SPIRE, you can configure the cryptographic keys to be either stored in memory or on disk, or to use a plugin where external keys are used for signing SVIDs.
By using the AWS KMS plugin for SPIRE, you bring the security benefits of AWS KMS to your SPIRE implementation.
- HSMs: AWS KMS uses hardware security modules (HSMs) that have been validated under FIPS 140-3 Level 2
- Key protection: Your plaintext AWS KMS keys don’t leave the HSMs, aren’t written to disk, and are only used in the volatile memory of the HSMs for the time needed to perform your requested cryptographic operation
- Comprehensive auditing: Each request to use the keys for signing must be authenticated using Signature Version 4 (SigV4) and is audited and logged in AWS CloudTrail
- Fine-grained access control: Use AWS KMS Key Policies to restrict access to your signing keys
This configuration block indicates to SPIRE that AWS KMS keys are used to sign our SVIDs:
This code block exists in the SPIRE server configuration deployed in the Github repository.
SPIRE datastore using Amazon Aurora
The SPIRE server uses a data store to keep track of the workload identity registration entries as well as the status of the SVIDs it has issued. By default, the datastore lives as an SQLite database in memory. However, you can configure the SPIRE server to use Amazon Aurora. By doing this, you can decouple the database from the server and offload the operational overhead of managing the datastore to AWS.
Benefits:
- High availability: Automatic failover and multi-Availability Zone deployments
- Automated backups: Point-in-time recovery and automated backup retention
- Scalability: Scale compute and storage independently
- Security: Encryption at rest and in transit, with AWS Identity and Access Management (IAM)
- Monitoring: Built-in performance insights and enhanced monitoring
- Reduced operational burden: Automated patching, maintenance, and updates
To do this, the CloudFormation stack provisions:
- An Aurora PostgreSQL Cluster
- A database user with the capability to authenticate through AWS IAM
- Network connectivity between your SPIRE Server and the database
After the sample finishes deployment, the configuration in the SPIRE server is changed from the default:
To:
For more information, see Datastore SQL Plugin.
SPIRE certificate authority using AWS Private Certificate Authority
The SPIRE Server issues SVIDs to SPIRE agents and workloads. Given that these SVIDs are X.509 certificates, the SPIRE server needs to act as a certificate authority. In the default configuration of SPIRE, the server generates a self-signed certificate that it will use as the root CA certificate. In production scenarios, organizations often use their existing AWS Private CA certificate authority hierarchy.
Using the AWS Private CA plugin for SPIRE, you configure the SPIRE CA to act as an issuing certificate authority that chains to your existing root CA. This delivers three benefits for your SPIRE implementation:
- HSM-backed root of trust: The root of trust of your CA hierarchy is backed by fully managed, FIPS 140-2 Level 3 validated hardware security modules
- Non-exportable private keys: The private keys for the root CA can’t be exported
- Fully Managed Certificate Revocation Lists and OCSP endpoint: In case the issuing CA certificate needs to be revoked
Note: If the SPIRE CA certificate will be signed by an issuing CA, you need to ensure that the path length of that CA is at least 1. The sample code provisions a root CA, so this isn’t an issue if you’re following the sample code.
After the sample deployment is complete, the configuration of the SPIRE server configuration includes the UpstreamAuthority plugin to use your AWS Private CA certificate authority as the root of trust.
For more information, see the AWS Private CA upstream authority plugin documentation.
SPIRE bundle publisher using Amazon S3
In SPIFFE, trust bundles are sets of public keys that are used by destination workloads to verify the identity of source workloads. These bundles contain the root certificates and public keys used to sign JWT SVIDs.
The SPIRE Server supports exposing an endpoint containing this trust bundle or publishing the bundle to Amazon Simple Storage Service (Amazon S3).
By publishing the bundle to Amazon S3, you can:
- Decouple: Decouple bundle access from the SPIRE server availability
- Continuous validation: Allow destination workloads to continually validate source workload identities without requiring network access to the SPIRE server
- Global distribution: Distribute the trust bundle using Amazon CloudFront, improving performance through the CloudFront global edge network with over 750 Points of Presence
- Reduced load: Offload trust bundle retrieval traffic from the SPIRE server
- Version control: Maintain historical versions of trust bundles with S3 versioning
The following is a sample configuration for this plugin:
For more information, see S3 Bundle Publisher Plugin Reference.
SPIRE Agent SVID store using Secrets Manager
In a typical SPIRE implementation, SVIDs are retrieved from the workload API using the SPIRE agent by workloads. In serverless or container cases, it isn’t feasible or possible to run the SPIRE agent, especially in workloads running on AWS Lambda or an AI agent running in Amazon Bedrock AgentCore Runtime. In these scenarios, use the SVID store plugin to push the SVID to Secrets Manager.
In this scenario, you still need a SPIRE agent running on a node that performs node attestation with the SPIRE server and delivers the SVID to a secrets store. One of the available secrets store plugins is the AWS Secrets Manager SVIDStore Plugin.
Benefits:
- Encryption at rest: Secrets are encrypted by default using AWS managed or customer-managed AWS KMS keys
- Broad integration: Integrations with multiple compute types, including AWS Secrets Manager Lambda Extension, Amazon ECS, Amazon EMR, and 52 other AWS services
- Client-side caching: Secrets Manager provides client-side caching libraries to improve performance and reduce API calls
- Multi-region replication: Built-in multi-Region replication functionality for workloads with multi-region requirements
- Fine-grained access control: Use resource policies to control which workloads access which SVIDs
- Audit trail: All secret access is logged in CloudTrail
This plugin is configured in the agent, as opposed to the SPIRE server. A reference for the specific permissions required are in the plugin Github repository. In the sample code, see the configuration in the following agent.conf file:
When you need to specify a given workload’s SVID to be pushed to Secrets Manager, you add the storeSVID and selector parameters when you create the entry in the workload registry. If you’re managing your SPIRE server in Kubernetes, the appropriate command looks like:
Fine-grained access control using Verified Permissions
Up to this point, all the managed services we’ve talked about have been used for creating and distributing SVIDs and trust bundles to workloads. However, a key advantage of using a framework such as SPIFFE is that it allows resource owners to protect resources with fine-grained access control. One managed service that helps you do that in the context of SPIFFE and SPIRE implementations is Amazon Verified Permissions.
Verified Permissions is a scalable, fine-grained permissions management and authorization service that helps you build secure applications. You can use it to:
- Define declarative policies: Use Cedar policy language, which provides human-readable, declarative access control rules
- Centralize policy management: Separate authorization logic from application code
- Take advantage of the straightforward policy authoring, formal verification capabilities, and performance benefits of Cedar
- Make real-time decisions: Evaluate authorization requests in real-time based on multiple factors including user identity, workload identity (SPIFFE ID), resource attributes, and contextual information
- Integrate with identity providers: Support for OpenID Connect (OIDC) providers, enabling validation of JWT tokens including JWT-SVIDs
With Verified Permissions, you authenticate SPIRE issued SVIDs using the public keys associated with your trust domain, and deploy fine-grained authorization policies protecting your resources. To use Verified Permissions, you need a policy store, an identity source, and authorization policies.
The github sample creates a policy store on your behalf and sets up the necessary infrastructure to have SPIRE serve as an identity source for your policy store.
You will need policies to enforce fine grained authorization. The following sample policy forbids a specific workload from taking any action:
Or conversely allows a specific workload to take actions:
Experiment with these policies as you build on top of your SPIRE infrastructure.
Conclusion
This post demonstrates how organizations enhance their SPIRE deployments on AWS by replacing default implementations with AWS managed services. We showed you integrations with AWS KMS for SVID signing, Amazon RDS Aurora for managed datastore operations, AWS Private Certificate Authority for root of trust, Amazon S3 and CloudFront for global trust bundle distribution, AWS Secrets Manager for SVID storage, and Amazon Verified Permissions for fine-grained authorization. By adopting these integrations, organizations can improve security posture, reduce operational complexity, and enhance resiliency.
If you’re interested in learning SPIFFE on AWS, check out our SPIRE on AWS github sample or the SPIFFE and SPIRE on AWS Workshop for hands-on learning.
If you have feedback about this post, submit comments in the Comments section below.
Play Rabbithole—With a Little Celebrity Help
Post Syndicated from The Atlantic original https://www.youtube.com/shorts/1AIAE5WlqGo