Architecting conversational observability for cloud applications

Post Syndicated from Anton Aleksandrov original https://aws.amazon.com/blogs/architecture/architecting-conversational-observability-for-cloud-applications/

Modern cloud applications are commonly built as a collection of loosely coupled microservices running on services like Amazon Elastic Kubernetes Service (Amazon EKS), Amazon Elastic Container Service (Amazon ECS), or AWS Lambda. This architecture gives engineering teams flexibility and scalability, but its inherently distributed nature also makes troubleshooting more difficult. When something breaks, engineers often find themselves digging through logs, events, and metrics scattered across different observability layers. With Kubernetes, for example, without a deep understanding of the service, troubleshooting can turn into a time-consuming effort to manually correlate information from different sources.

In this post, we walk through building a generative AI–powered troubleshooting assistant for Kubernetes. The goal is to give engineers a faster, self-service way to diagnose and resolve cluster issues, cut down Mean Time to Recovery (MTTR), and reduce the cycles experts spend finding the root cause of issues in complex distributed systems.

Overview

One of the challenges of architecting a modern cloud application is keeping observability intact across many moving pieces. Anyone who has ever tailed logs in one terminal while running kubectl describe and curl commands in another, knows how tedious this can get. Distributed systems are powerful, but they’re also complex. Kubernetes, for example, offers strong orchestration capabilities, yet troubleshooting inside a cluster often means navigating multiple layers of abstractions such as pods, nodes, networking, logs, and events. On top of that, the system generates a large volume of telemetry, including kubelet logs, application logs, cluster events, and metrics. Making sense of these layers requires both expertise on the system and application knowledge.

This skill gap shows up in the numbers. According to the 2024 Observability Pulse Report, 48% of organizations say that lack of team knowledge is their biggest challenge to observability in cloud-native environments. MTTR has also been going up for three years straight, with most teams (82%) saying it can take more than an hour to resolve production issues.

When something goes wrong and your applications start to fail for an unknown reason, engineers often must start stitching together signals from multiple sources to find the root cause. That can be tedious for specialists, and it gets worse when the issue is intermittent or spans across services. Often, multiple teams need to get involved – application engineers may not know Kubernetes well, while platform teams may not have deep insight into the applications. This can result in longer troubleshooting cycles, potentially degraded user experience, and pulling engineers away from planned work that drives business goals.

Figure 1. A multitude of telemetry sources in Kubernetes clusters

This is where generative artificial intelligence (AI) can help. Users can build an AI assistant that combines large language model (LLM)-driven analysis and guidance with existing telemetry data. This assistant enables engineers to troubleshoot issues faster, in a self-service way, without requiring every team to become Kubernetes experts. In the following sections, we show how to build such an assistant for Amazon EKS, however, keep in mind that a similar approach can be extended to other compute services like Amazon ECS or AWS Lambda.

Solution architecture

Architecting this AI-powered troubleshooting assistant consists of three primary parts:

  • Deployment approach selection: The solution supports two architectures – a traditional Retrieval-Augmented Generation (RAG)-based chatbot and a modern Strands-based agentic system that uses the Strands Agents SDK with EKS MCP Server integration for direct EKS API access.
  • Telemetry collection and storage: Collecting telemetry from various sources and storing it as vector embeddings in Amazon OpenSearch (RAG approach) or as 1024-dimensional embeddings in Amazon S3 Vectors (Strands approach).
  • Interactive troubleshooting interface: Building either a web-based chatbot that retrieves relevant telemetry and injects it into LLM prompts, or a Slack integrated multi-agent system that uses MCP tools for real-time Kubernetes diagnostics.

For this architecture walkthrough, we focus on the RAG-based approach. The first step is setting up a pipeline that can reliably collect, process, and store telemetry data. This pipeline aggregates telemetry from the relevant data sources, such as application logs, kubelet logs, and Kubernetes events. In Kubernetes environments, this can be done with a telemetry processor and forwarder, such as Fluent Bit, which streams telemetry into Amazon Kinesis Data Streams. On the receiving end, we use a Lambda function to normalize collected data, Amazon Bedrock to generate vector embeddings, and OpenSearch Serverless to store this embedded representation for efficient retrieval. Because these services are serverless, we can avoid the overhead of managing infrastructure and can focus on the troubleshooting workflow itself.

Figure 2. Collecting telemetry from sources, generating embeddings, and saving in OpenSearch

Pro tip: for better performance and cost-efficiency, your Lambda functions should use batching when ingesting data from Kinesis, generating embeddings, and storing them in OpenSearch.

Once telemetry is collected, converted to embeddings, and stored in OpenSearch, the next step is building a chatbot that uses RAG. Using RAG means that when a user asks a question, the chatbot looks up semantically similar telemetry in OpenSearch, adds it to the prompt, and sends it to the LLM. Instead of generic answers, the model now has relevant telemetry and cluster-specific details it can use to generate useful next steps, such as precise kubectl commands for the troubleshooting assistant, as illustrated in the following diagram.

Figure 3. Chatbot is using user queries augmented with telemetry context to send kubectl commands to the troubleshooting assistant. 

One powerful aspect of this design is its iterative nature. The chatbot hands instructions to a troubleshooting assistant running in the cluster, which executes a set of allowlisted, read-only kubectl commands. The output comes back to the LLM, which can decide whether it needs to investigate further (by asking the troubleshooting assistant to run more kubectl commands), or present a clear resolution path to the engineer. This cycle gradually builds a richer picture of the issue by combining historical telemetry with real-time cluster state to speed up root cause analysis.

Figure 4. Iterative troubleshooting process.

Here’s the end-to-end troubleshooting flow illustrated in the preceding diagram:

  1. An engineer enters a query into the chatbot interface, for example “My pod is stuck in pending state. Investigate.”
  2. The chatbot sends the query to Bedrock, which converts it into vector embeddings.
  3. Using those embeddings, the chatbot retrieves semantically matching telemetry that was previously stored in OpenSearch.
  4. The chatbot generates an augmented prompt, which contains both the original query and semantically relevant telemetry, and passes it to the LLM. The LLM responds with a list of kubectl commands to run for further diagnostics.
  5. The chatbot forwards those commands to the troubleshooting assistant running in the EKS cluster. The agent executes them with a service account that has read-only permissions, following the principle of least privilege, and sends the output back.
  6. Based on the output, the chatbot asks LLM to decide whether to continue investigation (by asking the agent to run more commands), or whether it has enough context to produce an answer.
  7. Once enough information has been gathered (investigation concluded), the chatbot composes a final prompt, including the query, telemetry, and investigation results, and asks the LLM for a final resolution, which it then returns to the engineer.

Example implementation

Use the example repo to deploy the solution in your AWS account. Follow the instructions in README.md for provisioning and testing the sample project using Terraform. Resources provisioned by the example project incur costs in your AWS account. Make sure to clean up the project as described in the README.md to avoid unexpected costs.

The repository provides two deployment architectures controlled by the deployment_type Terraform variable:

  1. RAG-based deployment (default): See the ./terraform/modules directory for the “ingestion-pipeline” module that creates a Kinesis Data Stream and Lambda function to generate embeddings using "amazon.titan-embed-test-v2:0" and store them in OpenSearch. The “agentic-chatbot” module handles the Gradio web interface and kubectl command execution.
  2. Strands agentic deployment: this approach uses the Strands Agents SDK to create a multi-agent system with three specialized agents:
    1. Agent Orchestrator: Coordinates troubleshooting workflows
    2. Memory Agent: Manages conversation context and historical insights
    3. K8s Specialist: Handles Kubernetes diagnostics

The agentic system stores knowledge as 1024-dimensional embeddings in Amazon S3 Vectors, providing cost-optimized vector storage for AI agents. EKS MCP Server integration enabled direct EKS API access through standardized MCP tools located in ./apps/agentic-troubleshooting/src/tools/. Engineers interact via Slack bot integration, where the Strands agents can execute kubectl commands through the MCP protocol while maintaining Pod Identity security for AWS service access.

The following screenshot shows an example chatbot response to a query about a pod being stuck in pending state. The assistant generated and ran multiple kubectl commands to build the output and came up with recommendations for issue remediation.

Figure 5. EKS cluster troubleshooting, example output

See AWS re:Invent 2025 – Streamline Amazon EKS operations with Agentic AI and KubeCon – From Logs To Insights: Real-time Conversational Troubleshooting for Kubernetes with GenAI sessions for a deeper dive into solution implementation.

Security considerations

When implementing AI agents for Kubernetes environments, security must be a primary consideration throughout the architecture. The solution requires secure communication channels between the chatbot and EKS clusters, with interactions authenticated through AWS Identity and Access Management (AWS IAM) roles.

Permissions-wise, command execution security is critical. Implementing strict allowlists that only allow read-only kubectl operations to help prevent unauthorized cluster modifications while maintaining diagnostic capabilities. The troubleshooting assistant should also operate with minimal Kubernetes RBAC permissions, limited to viewing pods, services, events, and logs within specific namespaces.

Data protection measures must include sanitizing application logs before embedding generation to help prevent sensitive information exposure, encrypting the telemetry data in transit through Kinesis and at rest in OpenSearch using AWS Key Management Service (AWS KMS).

Follow the AWS Well-Architected Framework Security Pillar principles, deploy components within Amazon Virtual Private Cloud (Amazon VPC) using private subnets and VPC endpoints to minimize network exposure, implement comprehensive logging of troubleshooting activities for audit purposes, and validate user inputs to protect against prompt injection attacks that could manipulate the AI assistant’s behavior.

Conclusion

In this post, we walked through how to architect a generative AI-powered troubleshooting assistant that gives engineers a way to solve Kubernetes issues in a self-service way, without always needing service experts to step in. By combining telemetry analysis with AI-driven context, engineers can get to the root causes faster and keep MTTR low. Assistant’s ability to pull from multiple telemetry sources, run safe diagnostic commands, and provide actionable recommendations helps to make the troubleshooting process more efficient and less disruptive to ongoing work.

As distributed systems continue to grow in scale and complexity, solutions like the one described in this post become essential. Putting AI on top of your observability data helps to practically handle these challenges today, while also setting you up for more autonomous, resilient operations in the future.

Мила родино, ти си ергенски рай

Post Syndicated from Светла Енчева original https://www.toest.bg/mila-rodino-ti-si-ergenski-ray/

Мила родино, ти си ергенски рай

Първият сезон на „Ергенът: Любов в рая“ вече е минало. Човек може да гледа риалити шоута по различни причини. Например за да изпитва морално или естетическо превъзходство, да се подиграва, да се възмущава, да се вживява в екранни драми, да си намира повод да влиза в пререкания или просто да си убива времето. Може да гледа обаче и за да разбере – какви ценности и послания се внушават, как се „управляват“ постъпките на участниците и зрителските реакции, какво ни казва шоуто за обществото, в което живеем.

Що се отнася до авторката на настоящата статия – оправдавах своето guilty pleasure (гузно удоволствие) да си причинявам риалитито с обещание, което лекомислено дадох на редакцията на „Тоест“: че след първата си статия по темата съм съгласна да напиша и втора, когато сезонът свърши. Но колкото повече гледах, толкова повече „Ергенът: Любов в рая“ ми заприличваше на държавата, в която живеем.

Задкулисие, което вече не се крие

В България отдавна се говори за задкулисие – реална власт извън формалното управление, която „дърпа конците“ на държавата. През 2009 г. репликата на Ахмед Доган „Аз разпределям порциите на властта“ от изтекъл запис с негово участие на събитие на ДПС звучеше скандално. Днес #кой дърпа конците, отдавна не е тайна.

Задкулисието в ергенския рай е Продукцията – с главна буква, защото е самостоятелен герой със собствена роля. По принцип се очаква една продукция да е по-скоро невидима, освен в специални моменти, като церемонии, за да изглежда, че отношенията между участниците следват естествен ход. Нашата Продукция обаче се намесва, когато и както реши.

Под формата на глас зад кадър например Продукцията внушава на ергенка чувство за вина, защото върши нормалното за участниците в подобно риалити – колебае се в чувствата си към ерген, с когото се е запознала наскоро, и казва ту че не го харесва, ту че го харесва. В резултат ергенката цяла седмица плаче, тръшка се, че е виновна, и напомня „вината“ си дори в самия край на шоуто.

По аналогичен начин в България хора биват пречупвани и принудени да подпишат самопризнания под натиск, като бившия заместник на Коцев – Диан Иванов.

Властта на задкуслисието води до произволно прилагане на правилата.

Основно правило в ергенския рай е, че който не получи роза, си тръгва, освен ако водещата не му даде една от общо трите за цялото предаване спасителни златни рози. Един ерген обаче се колебае между две ергенки. Дава роза на едната, другата си тръгва и той се разкайва за избора си. Но не щеш ли, кахърният мъж кани на среща отхвърлената участничка, моли я да се върне и тя се съгласява. Как става тази работа, след като ергенката вече е напуснала предаването, без дългата ръка на Продукцията?

Но какво ли се учудваме – колко институции в България са оглавявани от хора, чийто мандат вече е изтекъл? Членове на регулатори, на Висшия съдебен съвет, директорът на БНТ… Най-очеизбождащият пример е с главния прокурор. Въпреки че съдът „не даде роза“ на Борислав Сарафов, постановявайки, че той е нелегитимен главен прокурор, засега нищо не може да го мръдне от поста.

Като стана дума за мандати, първоначалното впечатление за джендърната „мандатност“ в шоуто беше, че на една церемония жените дават рози на мъжете, а на следващата е обратното. В един период обаче имаше три поредни церемонии, в които мъжете даваха рози на жените. Защо? Ергенките все бяха повече. Кой обаче решава в „рая“ да влизат повече жени, отколкото мъже, ако не Продукцията?

В някои случаи произволните решения на Продукцията са обясними с гоненето на рейтинг. Например тя отне правото на избор на ергенка, която показваше колебание между двама ергени, като прати и тримата директно на финал – за да има интрига до самия край. По този начин удължи с няколко дни терзанията им, както и тормоза върху участничката. Но и това едва ли ни учудва – колко често в България законодателни решения се приемат след формални или предварително нагласени „обществени обсъждания“, без да се чуе гласът на хората, които те засягат?

С напредването на риалитито Продукцията все повече ставаше тема на разговор между ергените. Един от членовете на пратения на финал любовен триъгълник дори конкретно я визира, обосновавайки решението си да остане, вместо да си тръгне:

Просто напук и на инат, защото всички ме дразните. И искам и на вас да ви е гадно с мене. И на Продукцията.

Милиони, закони, кокошки, прошки и Ново начало

Понякога действията на Продукцията трудно могат да се оправдаят с рейтинга и сценария, а зад тях прозира неравно третиране на участниците.

В една игра ергените мъже трябва да се състезават по двойки в дърпане на въже. Няма предварително зададени правила как точно да се удържи въжето. Единият от двамата финалисти не разчита само на грубата си мъжка сила, а прилага тактика, с която побеждава. Това поражда възмущение сред участниците (макар той да не е единственият участник в играта, приложил подобна тактика). Продукцията предлага на ергенките да гласуват кой е победил, и те избират загубилия битката.

Ако отново потърсим паралели с обществения живот в България, няма да ни е трудно да открием такива. Ето, Конституционният съд и Върховният касационен съд сътвориха тълкувания на понятията за пол и традиции в Конституцията, каквито в Основния закон няма, само и само служебно да декласират групата на ЛГБТ хората.

Отвръщайки на канонадата от упреци, че не е спазил правилата (каквито е нямало), спечелилият с тактика и впоследствие служебно загубил ерген обръща внимание върху едно доста по-сериозно неспазване на правила. А именно, че

ергени са излезли да се почерпят извън района, в който се снима предаването.

Напускането на „рая“ е забранено и тъй като всичко се снима с камери, то не би могло да се осъществи без благословията на Продукцията.

„Бегълците“, които са сред най-яростните обвинители на декласирания ерген, не показват никакво притеснение, когато става ясно, че са нарушили едно от най-важните правила. Напротив, хвалят се колко хубаво са си изкарали и как са пили „шотчета“. На въпроса дали има привилегии, един от тях отговори:

Имам, да! В песента се пее: „За кокошка няма прошка, за милиони няма закони“! Кое не ти е ясно?

Със споменаването на известната поговорка ергенът нямаше предвид обичайния ѝ смисъл – изтъкване на социална несправедливост. Напротив, изразяваше гордост, тъй като самият той твърди, че е милионер. Един вид, полага му се да е над законите.

Над правилата е и ергенката, която беше кандидат-депутатка от ДПС – Ново начало на изборите на 27 октомври 2025 г.

В качеството си на половинка в „рая“ на ергена с милионите и тя е от пилите „шотчета“ извън района на снимките. Но това май не е единствената ѝ привилегия.

Едно от основните правила за участниците е, че те нямат право да използват телефоните си. Мобилните им устройства се вземат преди началото на снимките и им се връщат след финала. На кадър от предаването (на 25-тата секунда от това видео) обаче се вижда как ергенката от „Ново начало“ държи нещо, което има формата на телефон и свети като телефон. Ако следваме принципа, че най-простото обяснение често е вярното, това най-вероятно е мобилен телефон.

(Не)толерантност към насилието

Случаите на вербално насилие в „рая“, на тормоз, подигравки, обиди заради външния вид (т.нар. бодишейминг), унижаване и пр. са толкова много, че и няколко статии няма да стигнат за споменаването им. Физическото насилие обаче е забранено, научаваме от разговори на Продукцията с участници. Това правило беше моментално приложено за едва-що влязъл в шоуто ерген, който събори ергенка от стола ѝ. Сцената не беше излъчена изцяло в предаването, а само загатната, както е и редно, защото показването на насилието на екран го мултиплицира.

Аналогична би следвало да е реакцията и ако ергенка бутне друга в басейна. Но не е. Бутането в басейна е задача, която Продукцията дава в рамките на игра, която уж би трябвало да е забавна за всички. Задачата е възложена на ергенка, изпитваща ревност към друга, в момент, в който почти всички участници тормозят другата. Ревнивата ергенка блъска съперницата си с всички сили. След кратък полет момичето пада във водата. Кадърът с блъскането и полета се повтаря няколко пъти (няма да сложа линк, защото е унизително). Ергените се смеят. „Дано може да плува“, чува се развеселен глас. Момичето излиза от водата и сяда мокро до останалите, защото е длъжно да продължи участието си в играта.

За да направим връзка с обществения живот в България – има групи хора, насилието върху които остава без последствия и дори се толерира. Бежанци се задържат, ограбват и бият на границата; „ловци на бежанци“ се обявяват за герои; събарят се единствените жилища на роми и пр. – държавата внушава, че представителите на някои групи си го „заслужават“.

За ергенки и ергени, които налитат на бой, няма никакви последствия – щом не се е стигнало до реално нанасяне на удари, за Продукцията, изглежда, няма проблем. И някои „дребни“ форми на физическо насилие не се броят. Ерген постепенно засилва агресията си към друг – от вербален тормоз минава към блъскане на краката му с чекмедже. Продукцията го изгонва чак след като изплюва храна върху набелязаната жертва. За разлика от първия изгонен ерген, който ни се чу, ни се видя, този гордо дефилира по bTV, защитавайки гледната си точка.

Тук стигаме до темата за домашното насилие,

понеже вторият изгонен ерген си имаше половинка в „рая“. Забраняваше на избраницата си да говори с другите мъже в шоуто, както и на тях да общуват с нея. Изобщо, смяташе, че е най-добре тя да мълчи, а той да говори вместо нея. И тя преобладаващо си мълчеше, говореше основно в негово отсъствие, почти няма и синхрони с нея (реплики в разговор с Продукцията).

Когато Продукцията реши да изгони ергена, тя му даде възможност преди напускането си да говори и с участника, върху когото беше изплюл храна (уж да му се извини, но на практика не стана точно така), и с избраницата си. Нея той постави пред избор – да си тръгне с него или да остане. На практика обаче изборът не беше точно такъв – споменавайки как ще ѝ обясни причините за изгонването си в хотела, той вече беше предпоставил, че тя ще го последва.

Впоследствие стана ясно, че двамата продължават да са заедно и извън „рая“. Може би ергенката щеше да последва изгонения и ако имаше възможност да вземе решение самостоятелно, далеч от присъствието му. Няма как да знаем това. Продукцията обаче не ѝ помогна, а bTV допълнително легитимира тази основана на контрол и ограничения връзка.

Друг ерген пък забрани на половинката си да разговаря с ергенка, която я е обидила – с претенцията, че така всъщност я защитава. И дори ѝ постави условие: ако си позволи да общува с нея, той си тръгва. 

Но какво ли се учудваме – борбата с домашното насилие не е приоритет и на държавата. Важното е да се борим с „джендъра“ и да не допускаме хората в еднополови връзки до и без това не особено ефективно работещата система на защита.

В „рая“ не липсваше и сексуално насилие.

Някои ергени толкова искаха от половинките си да им „пуснат“, че понякога си позволяваха да ги понатиснат в леглото. Те се цупеха, че не постигат своето, и обвиняваха ергенките в липса на взаимност. Не вдяваха, че „не“-то, дори изречено през смях, си е „не“. Един участник беше особено настоятелен не само в леглото. Опита се да бръкне между краката на изгората си дори в двора. Не успя само защото тя е бивша спортистка и има здрави мускули.

На повечето ергенски мераци Продукцията не реагира, но след последния от описаните случаи попита ергенката дали да се намеси. Не го направи, защото участничката каза, че няма нужда. Така както институциите си измиват с ръцете с това, че много от пострадалите не си търсят правата.

Привидност и реалност

Както в политиката, така и в „рая“ не трябва да приемаме всичко за чиста монета. Не само защото нещата, които виждаме, са „за пред хората“ и поведението на участниците поне отчасти се дължи на факта, че играят роли, а и защото биваме сугестирани кое как да възприемаме.

Един от начините да се влияе на преценката ни е да се манипулира усещането ни за време. ГЕРБ, които са на власт с кратки прекъсвания от 2009 г., стоварват отговорността за всички беди върху периода между 2021 и 2024 г., когато не са участвали в правителството (макар че правителството на Николай Денков беше и тяхно).

В „рая“ времето е разтегливо дори повече, отколкото в главата на Бойко Борисов.

В началото на шоуто церемониите с розите се излъчваха веднъж седмично – в петък. Един период между две церемонии обаче се проточи цели три седмици. Така се създаде впечатлението, че участниците са прекарали много повече време заедно, отколкото реално са, и че някои драми са се разразявали в продължение на близо месец, а не по-малко от седмица.

Случваше се едно и също събитие да се излъчва в няколко поредни епизода, особено ако включва скандал. Други случки бяха показани мимоходом, а много така и не са стигнали до зрителите, понеже няма как да се излъчи всичко. Не по-маловажна обаче е

привидността в ергенските отношения,

въпреки че в някои случаи е трудно да се сложи рязка граница между привидност и реалност. Ергени и ергенки може да смятат, че обиждат, тормозят и унижават, за да става шоу, но обидите, тормозът и униженията са си истински. С любовта нещата стоят по-скоро по обратния начин. От екрана се леят огромни количества от нея, но такива са правилата на играта.

До финала стигнаха седем двойки и една тройка, като ергените (в случая с тройката – избраният от дамата ерген) трябваше да дадат на избраниците си или роза (еквивалент на „една студена вода“), или годежен пръстен. Рози получиха само две участнички, а пръстени – пет. Така че шоуто завърши с четири годежа и един полугодеж (тъй като пръстенът беше сложен на дясната ръка на ергенката от тройката със заръка да го премести на лявата, когато се почувства готова).

Годежът е обещание за съвместно бъдеще, но колко от двойките, дали си драматични любовни обещания, са заедно след края на предаването? Има данни само за една – двамата от бившата тройка, завършили с полугодеж. Именно техните отношения бяха поставяни под въпрос от другите участници по време на почти цялото шоу. Защо тя е избрала него, а не друг; защо не си намери по-подходящ и по-мъжествен; защо е „лека жена“, „волна пеперуда“, скача „от цвят на цвят“ и не може да избере? Защо той не е достатъчно „мъжествен“; гей ли е; защо ѝ дава свобода?

Злите езици говорят, че далеч от камерите се е оформила истинска тройка, а не за шоуто, точно между най-големите морализатори, упрекващи оцелялата двойка, но засега не разполагам с достатъчно данни, за да потвърдя това.

В заключение, ако човек има желание и нерви да си причинява риалити формати, те могат да ни промиват мозъците (което и основно правят), но пък и ние можем да ги използваме, за да се упражняваме да не вярваме на очите си и да развиваме критичното си мислене. Да се замисляме какво ни внушават да мразим и от кое да се разчувстваме, да имаме едно наум кой дава празни обещания и кой е превърнат в изкупителна жертва. За да не стане така, че отново да се доверим на поредните популисти.

Security updates for Thursday

Post Syndicated from jzb original https://lwn.net/Articles/1050117/

Security updates have been issued by Debian (ffmpeg, firefox-esr, libsndfile, and rear), Fedora (httpd, perl-CGI-Simple, and tinyproxy), Oracle (firefox, kernel, libsoup, mysql8.4, tigervnc, tomcat, tomcat9, and uek-kernel), SUSE (alloy, curl, dovecot24, fontforge, glib2, himmelblau, java-17-openjdk, java-21-openjdk, kernel, krb5, lasso, libvirt, mozjs128, mysql-connector-java, nvidia-open-driver-G07-signed-check, openssh, poppler, postgresql17, postgresql18, python-cbor2, python-Django, python310, python311-Django, runc, strongswan, tomcat11, and xwayland), and Ubuntu (binutils, libpng1.6, linux, linux-aws, linux-aws-5.4, linux-gcp, linux-gcp-5.4, linux-hwe-5.4,
linux-ibm, linux-ibm-5.4, linux-kvm, linux-oracle, linux-xilinx-zynqmp, linux, linux-aws, linux-aws-6.14, linux-gcp, linux-hwe-6.14, linux-raspi, linux, linux-aws, linux-gcp, linux-realtime, and qtbase-opensource-src).

New Research: Multifunction Printer (MFP) Security Concerns within the Enterprise Business Environment

Post Syndicated from Deral Heiland original https://www.rapid7.com/blog/post/new-research-multifunction-printer-mfp-security-concerns-within-the-enterprise-business-environment

Multifunction printers (MFPs) do far more than print. They scan, email, fax, store, and authenticate. That convenience comes with risk. Our latest report, Understanding Multifunction Printer (MFP) Security within the Enterprise Business Environment, from Rapid7’s Deral Heiland, Principal Security Researcher (IoT), and Sam Moses, Security Consultant, takes a clear look at where MFPs expand your attack surface and how to reduce that risk.

Why this research matters

MFPs are everywhere, often overlooked, and frequently underprotected. Many organizations deploy them without password changes, patch cycles, or network segmentation. Attackers notice. Because MFPs are attached to networks and can carry sensitive data, compromise can enable credential theft, data leakage, and lateral movement within the network.

The report tracks how long-standing and emerging weaknesses continue to affect MFP security. It highlights common risk areas such as weak authentication and limited patching practices, among others, that leave devices open to misuse or compromise. As these printers have grown more connected and feature-rich, the potential impact of a single vulnerable device has increased, especially when linked to core business systems or identity services.

The study also examines broader exposure trends across the enterprise landscape. Thousands of MFPs remain directly accessible from the internet, and vulnerability data shows that many models have faced serious flaws in recent years. Beyond technical issues, organizational processes like inconsistent patch management and poor decommissioning practices often allow sensitive data and credentials to linger on devices long after their use.

Penetration testing data collected by Rapid7 and Raxis confirms that these risks are not theoretical. Many organizations still deploy MFPs with default settings, leaving them open to credential theft and data access that can help attackers move deeper into the network.

The report introduces Praeda-II, a community tool designed for pentesters, auditors, and IT teams who need fast visibility into vulnerable printers, to identify risks in MFPs across modern models.

See the research

If your organization relies on networked printers, this research offers the insights you need. Read Understanding Multifunction Printer (MFP) Security within the Enterprise Business Environment to learn about key risks and practical steps to strengthen your printer security program.

Geopolitics and Cyber Risk: How Global Tensions Shape the Attack Surface

Post Syndicated from Jeremy Makowski original https://www.rapid7.com/blog/post/geopolitics-and-cyber-risk-how-global-tensions-shape-the-attack-surface

Geopolitics has become a significant risk factor for today’s organizations, transforming cybersecurity into a technical and strategic challenge heavily influenced by state behavior. International tensions and the strategic calculations of major cyber powers, including Russia, China, Iran, and North Korea, significantly shape the current threat landscape. Businesses can no longer operate as isolated entities; they now function as interconnected global ecosystems where employees, suppliers, cloud workloads, supply chains, and data flows intersect across multiple jurisdictions, each with its own unique set of political risks.

A region considered low-risk last month could become a high-risk zone overnight if a diplomatic dispute escalates. An overseas development team could suddenly become vulnerable if that region experiences sanctions, stricter regulations, or state pressure on the workforce.

Many organizations still underestimate this dynamic reality, relying on static risk models that assume relatively stable attack patterns. However, geopolitical decisions and internal vulnerabilities are often the drivers of the most sudden and consequential changes in exposure. For example, the announcement of sanctions can trigger retaliatory cyberattacks, a military buildup can unleash destructive campaigns, and a trade or intellectual property dispute can lead to large-scale espionage.

Cybersecurity leaders must therefore integrate geopolitical intelligence directly into their operational decision-making and risk assessment processes, recognizing that political forces, rather than technical errors, are often the primary trigger for increased vulnerability.

Geopolitics as a core driver of cyber risk

Geopolitics plays a decisive role in shaping the scale, direction, and sophistication of cybercriminal and state-sponsored activity, fundamentally altering the threat landscape for organizations worldwide. Geopolitical tensions and sanctions often create conditions in which state-aligned hackers operate with greater freedom, using cyber operations as tools for espionage, economic survival, political retaliation, or strategic influence. Isolated or sanctioned states often turn to cybercrime as an alternative source of revenue.

North Korea, for instance, intensifies financially motivated campaigns, including cryptocurrency theft and extortion, when economic pressure mounts. Iran, facing recurring sanctions and political isolation, tends to respond with retaliatory or disruptive cyber operations targeting sectors and institutions associated with adversarial nations.

China’s cyber activity often peaks during moments of heightened competition over technology and strategic resources, driving expansive espionage campaigns aimed at industries like aerospace, telecommunications, AI, and energy. Russia, meanwhile, escalates disruptive or destructive cyber actions during geopolitical confrontations or military conflicts, leveraging malware, industrial system interference, and coordinated information operations.

These patterns demonstrate how cyber risk extends far beyond technical vulnerabilities: organizations become targets because of their nationality, sector, technology assets, or global partnerships.

How geopolitical tensions influence threat actor behavior

Geopolitical tensions influence the behavior of threat actors by altering their objectives, aggression levels, and operational trade-offs in ways that directly impact global organizations. Russian groups, for example, will shift from covert intelligence collection to overt disruption, employing destructive malware, DDoS attacks, and infrastructure sabotage to exert pressure. Chinese actors are known to intensify long-term espionage and supply-chain infiltration, targeting IP, cloud providers, security firms, and development environments.

Iran responds to sanctions or regional tensions with opportunistic retaliation through data wiping, defacements, and financially motivated attacks. And when facing economic strain, North Korea expands cybercrime, including cryptocurrency theft, extortion, software supply-chain poisoning, and high-level financial fraud.

For organizations, these shifts manifest internally as newly observed attack patterns, such as targeted phishing aimed at political or strategic sectors, the exploitation of vulnerabilities relevant to conflicts, or supply-chain attacks aligned with espionage objectives. The unifying pattern is that geopolitical tensions cause attackers to reprioritize, whereby espionage becomes a means of destruction, revenue generation becomes a national strategy, and symbolic retaliation becomes an operational necessity. Security teams that do not account for these geopolitical triggers risk misjudging the scale, intent, and urgency of incoming threat campaigns.

Indicators that cyber escalation is coming

A cyber escalation is rarely an isolated phenomenon; it is usually accompanied by political and technical warning signs that can herald a wave of attacks. On the political front, organizations should monitor events such as sanctions announcements, diplomatic expulsions, military mobilizations, sudden breakdowns in negotiations, strategic military strikes, or public accusations of espionage. For example, tensions with Russia are often followed by cyber influence campaigns. Retaliatory cyberattacks are also common following the imposition of sanctions on the Islamic Republic of Iran. Increased cyber espionage campaigns coincide with periods of strategic competition with China, and financially motivated attacks intensify after economic pressure is exerted on North Korea.

On a technical level, the first warning signs manifest in one or more of the following ways:

  • An increase in sector-specific phishing attacks linked to political events
  • The reactivation of known command and control infrastructures
  • The formation of new politically-motivated hacktivist collectives
  • Access intermediaries launching campaigns to sell access points in sectors linked to ongoing conflicts

Internally, organizations may sometimes observe unusual activity from cybersecurity teams, such as unexpected code updates from maintenance managers located in politically sensitive regions, vendor outages correlated with geopolitical developments, or authentication anomalies linked to regions near ongoing crises. The most important pattern to recognize is convergence: when political escalation, external surveillance, and internal anomalies appear within the same time frame, organizations must assume that threat conditions have shifted from background noise to active risk and immediately adopt a strengthened defensive posture.

Adjusting defensive posture during geopolitical instability

Harden identity infrastructure against state-grade threats.

Identity has become a frontline asset in geopolitical conflict. In today’s environment, the boundaries between hacktivism, cybercrime, and state-sponsored activities are increasingly blurred, with governments at times guiding or amplifying these operations. Credential compromise is often the entry point that enables these broader campaigns. To mitigate this risk, organizations should enforce universal, phishing-resistant MFA, regularly review and tightly govern privileged roles, particularly in sensitive geographies, and adopt just-in-time access to minimize standing privileges. These measures materially reduce exposure and strengthen resilience against sophisticated, geopolitically motivated threat actors.

Conduct targeted threat hunts

  • Russia — Russian threat actors place a strong emphasis on disruption and destruction, particularly during periods of geopolitical conflict. They commonly deploy wiper malware that deletes or corrupts files and often pretend it’s ransomware. Threat hunters should watch for sudden mass file changes, system reboots, or the use of admin-level command-line tools immediately preceding damage. Russia also has advanced capabilities for ICS/OT manipulation, meaning unusual access to industrial controllers or configuration changes can be a strong indicator of potential compromise. Additionally, their operations often support information warfare, so defenders should look for compromised media or government accounts, unauthorized website changes, and targeted spear-phishing attacks tied to political events.
  • China — China focuses on long-term, stealthy access rather than quick disruption. They are known for supply-chain compromises, so unusual activity from vendor accounts or anomalies in software updates should be investigated. They frequently abuse cloud identity platforms, making it essential to monitor for impossible travel logins, token theft, MFA fatigue, or suspicious OAuth applications. Chinese groups also invest heavily in credential harvesting, often trying to quietly collect usernames, passwords, and tokens over long periods. Threat hunters should look for password spraying, attempts to dump credentials, or lateral movement linked to service or personal accounts that generally don’t access sensitive systems.
  • Iran — Iranian threat actors tend to be opportunistic and politically reactive, relying heavily on broad phishing campaigns. Organizations should monitor for spikes in failed logins, newly created email forwarding rules, and look-alike phishing domains. Iran also frequently conducts website defacements, so signs such as unexpected CMS admin logins, unauthorized web content changes, or DNS tampering are essential to hunt for. While generally less sophisticated than Russia or China, they can still deploy destructive malware, meaning defenders should watch for scripts or tools that mass-delete or encrypt files, suspicious scheduled tasks, and activity involving commodity RATs or .NET tools.
  • North Korea — North Korea’s cyber operations are primarily financially motivated, with a strong focus on cryptocurrency theft. Threat hunters should monitor for unauthorized access to wallet systems, unusual outbound connections to cryptocurrency platforms, or abnormal API calls associated with blockchain activity. They also excel at social engineering, especially targeting finance, HR, and engineering staff by posing as recruiters or job candidates. Indicators include suspicious attachments, communication from personal email accounts, or new “contractor” accounts accessing code or financial systems. Once inside a network, their activity is typically driven by exfiltration, so large or stealthy data transfers, especially to cloud storage or foreign VPNs, are significant warning signs.

Reprioritize assets exposed to geopolitical pressure.

Identify systems and identities that become high-value targets during periods of geopolitical tension, especially those associated with sensitive regions or government-linked operations. Immediately harden them with faster patching, tighter segmentation, stricter east–west controls, and increased telemetry to concentrate defenses where state-aligned actors are most likely to strike.

Reduce external exposure on high-value frontiers.

Reduce the attack surface by removing access paths favored by advanced adversaries. Disable legacy VPNs, retire unmonitored jump servers, tighten SSO/IdP trust paths, and eliminate unnecessary remote-admin or broad cloud access routes. Reducing weak entry points raises the cost of initial access for foreign intelligence units.

Harden response capabilities

Incident response teams must prepare for an increased likelihood of destructive or politically motivated attacks. Organizations should test their data destruction and destructive attack plans, validate their disaster recovery timelines, and ensure the restoration of offline or immutable backups. Management must be kept informed of evolving geopolitical risks, and cross-functional teams, including cybersecurity, legal, communications, and operations, must conduct crisis simulation exercises. Rapid response structures, such as crisis management teams, should be ready to be activated to facilitate fast decision-making under pressure. These measures are intended to help ensure that the organization can respond effectively even in the face of significant stress or disruption.

Building a geopolitical cyber attack surface map

Building a geopolitical map of the attack surface enables organizations to anticipate how political conditions may impact cyber risk. This involves understanding how people, technology, and third-party relationships are geographically distributed, and how those distributions intersect with jurisdictions that may impose legal, operational, or conflict-related risks. A robust map also integrates geopolitical assessments with business impact and criticality, enabling organizations to see where instability or state control could affect privileged access, essential services, or sensitive data.

The following steps describe how to perform an attack surface mapping based on geopolitical events. These steps are not derived from any single framework or source; they are a practical blend of best practices for mapping infrastructure, assessing geopolitical exposure, identifying weak points, and prioritizing remediation.

  • Map Internal Workforce: Create an authoritative inventory of the physical locations of all employees with technical or elevated privileges. Include full-time staff, contractors, and outsourced teams. Use HR, IAM, and staffing records to ensure accuracy and maintain updates as personnel relocate or roles change.
  • Map Infrastructure: Create a comprehensive list of regions that host your cloud services, data centers, disaster recovery sites, and replication routes. Document which workloads reside where, how traffic moves between regions, and what operational responsibilities each location carries. Capture both primary and failover arrangements.

  • Map Vendor & Subcontractor: This step requires suppliers to disclose the actual countries where engineering, customer support, managed services, and subcontracted tasks are performed. Validate this information through audits, questionnaires, or contractual obligations. Record each operational footprint, not just corporate registration locations.
  • Geopolitical Risk Scores: Apply a standardized scoring model to each region (e.g., Matteo Iacoviello Geopolitical Risk (GPR) index, BlackRock Geopolitical Risk Indicator (BGRI), or Bloomberg’s geopolitical risk scores). Inputs may include government stability indicators, international sanctions status, regulatory pressures, history of state intervention, and exposure to espionage or cyber operations. Use a consistent scoring range.
  • Overlay Business Criticality: Cross-reference each region’s risk score with the operational value of what that region supports. Identify where highly sensitive systems, privileged roles, or essential processes are located in areas with higher risk. Highlight areas where disruption would impact business continuity or security posture.
  • Identify Regional Strategic Points: Look for dependencies where a single region hosts an excessive number of critical people, systems, or vendors. This includes cloud regions serving multiple core workloads, a subcontractor with a heavily centralized team, or a country where several key staff reside. Flag these for targeted risk discussions.
  • Prioritize Remediation Measures: Develop a ranked set of actions based on the combined geopolitical and business impact. Potential responses include redistributing workloads across safer regions, shifting privileged roles, tightening access controls, enhancing monitoring for at-risk locations, or preparing contingency plans for rapid relocation or provider transition.

Conclusion

Geopolitics is now a key driver of cyber risk, redefining attacker profiles, motivations, and the organizations targeted and/or affected by collateral damage. Many vulnerabilities in modern businesses stem not from technical misconfigurations, but from the geopolitical interconnectedness of global supply chains, cloud architectures, distributed teams, and open-source ecosystems.

Traditional cybersecurity controls remain essential, but are insufficient on their own as they fail to account for laws, political incentives, national strategies, and human vulnerabilities influenced by the world’s most active cyber powers. To manage this reality, organizations must integrate geopolitical analysis into every layer of their security decision-making process, consider geography as a key security variable, and develop the agility to proactively adapt their posture to the evolving global context.

Celebrating the community: Irioluwa, Michelle, Jedidiah and Inioluwa

Post Syndicated from Sarah Lygoe original https://www.raspberrypi.org/blog/celebrating-the-community-irioluwa-michelle-jedidiah-and-inioluwa/

We love hearing from members of the community and sharing the stories of amazing young people, volunteers, and educators who are using their passion for technology to create positive change in the world around them.

In our latest story, we’re meeting a group of inspiring young innovators from Belfast — Michelle, 15, Inilouwa, 18, Jedidiah, 14 — and Irioluwa, 11, who are using their creativity and technical skills to tackle an issue that impacts millions of young women across the UK: period poverty.

Inspiring young innovators from Belfast — Michelle, 15, Inilouwa, 18, Jedidiah, 14 — and Irioluwa, 11

The power of community

The group first connected through Diverse Youth, a community space that gave them the chance to collaborate, learn, and grow.

At the youth centre, the girls were introduced to Code Club through mentor, and Jedidiah’s mum, Tiwa. When Tiwa offered them the opportunity to travel to Coolest Projects UK and showcase projects they had been working on, they knew they wanted to make something special…

Flow Body website screenshot

Tackling taboo topics

The idea for Flow Body began when Inilouwa was inspired by her sister’s experience at school.

“So, for me, it was my sister actually, that really inspired this project because she would tell me, her school at the time, they run out of period products very frequently, so not a lot of girls could access it.” 

Wanting to make a difference, the team created Flow Body, a website connecting young women to period charities in the UK and providing reliable information about periods, puberty, and related health issues.

Playing to their strengths

The technical side of the project was led by Inilouwa, with the rest of her teammates supporting her with research and content creation.

“So, alongside Michelle, I researched the diseases and what can come from period poverty and how different people get by with the lack of period products around,” Jedidiah said.

Their research highlighted just how widespread period-related health issues are.

The team discovered that one of the biggest challenges is access to accurate information, given the stigma around discussing menstruation, and using technology to solve this issue seemed like the perfect fit to them.

Inilouwa explained, “Having that access to information that’s tailored to young women, to young parents, to everyone all across the board would be really helpful, you know, in order to make periods less of a strange topic to discuss.”

Inilouwa, 18, speaking with presenter on stage at Coolest Projects UK

Celebrating teamwork

After taking home judges’ favourite in the web category, the girls could not recommend the experience enough. When asked what advice they would give to other young people thinking of taking the leap and entering Coolest Projects, the answer was simple…

“Do it. Don’t be afraid. Even if tech is not your thing, it’s not a lot of people’s things. But like you learn so much, you grow so much and it’s very fun. It is. It really is,” said Inilouwa.

If you are interested in getting involved in Coolest Projects, keep an eye on the website for exciting announcements for 2026 plans.

Want support on your coding journey? Find a Code Club near you to learn alongside a likeminded community.

The post Celebrating the community: Irioluwa, Michelle, Jedidiah and Inioluwa appeared first on Raspberry Pi Foundation.

[$] LWN.net Weekly Edition for December 11, 2025

Post Syndicated from corbet original https://lwn.net/Articles/1049161/

Inside this week’s LWN.net Weekly Edition:

  • Front: Rust in CPython; Python frozendict; Bazzite; IETF post-quantum disagreement; Distrobox; 6.19 merge window; Leaving the TAB.
  • Briefs: Let’s Encrypt retrospective; PKI infrastructure; Rust in kernel to stay; CNA series; Alpine 3.23.0; cmocka 2.0; Firefox 146; 2024 Free Software Awards; Quotes; …
  • Announcements: Newsletters, conferences, security updates, patches, and more.

Resolve and prevent operational incidents with AWS DevOps Agent and New Relic

Post Syndicated from Nava Ajay Kanth Kota original https://aws.amazon.com/blogs/devops/resolve-and-prevent-operational-incidents-with-aws-devops-agent-and-new-relic/

This post was co-written with Muthuvelan Swaminathan (Principal Partner Engineer) and Ruchika Bakolia (Software Engineer) from New Relic.

Modern distributed systems that generate massive volumes of metrics, traces, and logs are inherently complex. The process of correlating logs, comparing configurations and switching between tools during incident management makes manual root cause analysis a bottleneck that dramatically increases the mean time to detect and resolve. Instead of manually sifting through mountains of data, Site Reliability Engineers (SREs) and DevOps teams can leverage Agentic AI to automate and enhance the incident resolution process.

To address these challenges, New Relic partnered with AWS to integrate the New Relic Model Context Protocol (MCP) server with AWS DevOps Agent to access telemetry data providing automated root cause analysis and recommendations with cutting-edge artificial intelligence. AWS DevOps Agent is a frontier agent that resolves and proactively prevents incidents, continuously improving reliability and performance of applications in AWS, multi-cloud, and hybrid environments.

In this blog, we’ll explore the key features of both services, how to configure them and an example that shows how operation teams can correlate telemetry data, predict system anomalies and initiate remediation actions to significantly accelerate MTTR (Mean Time to Resolution).

New Relic AI MCP Server

The New Relic MCP Server is a standardized gateway that connects external AI agents such as AWS DevOps Agent to New Relic’s observability data and functions. It enables autonomous agents to query live data and execute actions without requiring custom API integrations.

As customers and partners build their own AI tools, there is no longer a need to maintain a bespoke API integration. MCP enables AI agents to seamlessly interact with their telemetry data on New Relic platform through an MCP client to leverage its capabilities and enhance their workflows.

AWS DevOps Agent

AWS DevOps Agent is a frontier agent that resolves and proactively prevents incidents, continuously improving reliability and performance. AWS DevOps Agent investigates incidents and identifies operational improvements as an experienced DevOps engineer would: by learning your resources and their relationships, working with your observability tools, runbooks, code repositories, and CI/CD pipelines, and correlating telemetry, code, and deployment data across all of them to understand the relationships between your application resources.

Key benefits for organizations 

The integration of in-depth observability with AWS DevOps Agent capabilities is designed to quickly resolve issues when they arise and prevent incidents for SRE and DevOps engineers. Here are few benefits:

  • Automated investigations: AWS DevOps Agent integrates with ticketing and alarming systems like ServiceNow to automatically launch investigations from incident tickets, accelerating incident response within your existing workflows to reduce meant time to resolution (MTTR).
  • Incident coordination: You can also initiate and guide investigations using interactive chat. AWS DevOps Agent acts as a member of your operations team, working directly within your collaboration tools like ServiceNow and Slack to share findings and coordinate responses. 
  • Root cause analysis: AWS DevOps Agent integrates with observability tools, code repositories, and CI/CD pipelines to correlate and analyze telemetry, code, and deployment data, sharing its explored hypotheses, observations, Through systematic investigations, AWS DevOps Agent identifies root cause of issues stemming from system changes, input anomalies, resource limits, component failures, and dependency issues across your entire environment.
  • Detailed mitigation plans: Once AWS DevOps Agent has identified the root cause, it provides detailed mitigations plans, which include actions to resolve the incident, validate success, and revert a change if needed. AWS DevOps Agent also provides agent-ready instructions that can be implemented by another frontier agent, for example, code improvements that can be implemented by Kiro autonomous agent.
  • Proactively future incidents: AWS DevOps Agent analyzes patterns across historical incidents to provide actionable recommendations that strengthen four key areas: observability, infrastructure optimization, deployment pipeline enhancement, and application resilience.

Onboarding

The onboarding process involves setting up an Agent Space and registering your existing New Relic servers. Onboarding does not require any new implementation.

Here are the high-level steps to create an AWS DevOps Agent Space and connect it to the New Relic MCP Server using an API-Key.

Setup Agent Space in AWS DevOps Agent

To create Agent Spaces, navigate to the AWS DevOps Agent page within the AWS Management Console. An Agent Space establishes the boundaries for the AWS DevOps Agent when accessing resources within a specific AWS account. To get started, click the create Agent Space button at the top right of the screen and enter the name, description and IAM roles.

Screen shot displaying orange Create Agent Space button in the AWS Console

AWS DevOps Agent creating agent space

Creating a New Relic association

 Navigate to the capabilities tab in the Agent Space

Screen shot with a red square around the tab for Capabilities in the AWS DevOps Agent -> AgentSpaces view in the AWS Console” width=”1430″ height=”849″></p>
<p style=Navigating to the capabilities tab in the Agent space

Go to the Telemetry section, select Add, then choose New Relic and click Next.

Screen shot with Add a new source radio button selected and Select source to add has radio button New Relic selected

Associating New Relic as the Telemetry provider in the Agent space

Upon successful registration of New Relic as a source, AWS DevOps Agent automatically generates a webhook URL. This URL is then used to receive alert notifications and trigger automated investigations.

Screen shot for Configure Webhook Connection displaying the Webhook URL and Webhook Secret, both are redacted by a black bar.

AWS DevOps Agent Webhook URL and Bearer secret key

The AWS DevOps Agent webhook requires a Bearer token to be included in the HTTP header for authentication purposes. This ensures that only authorized requests are processed. In New Relic, set up Amazon EventBridge as the alert destination. This configuration will trigger an AWS Lambda function that adds the Bearer token to the HTTP header and posts the alert payload to the AWS DevOps Agent webhook URL.

Use Case Walkthrough: Retail Chain – High Latency in shopping cart service resolution

This use case demonstrates how the integration of AWS DevOps Agent and New Relic MCP server empowers SRE and DevOps teams to access the untapped insights in your data to reduce MTTR and drive operational excellence.

Consider the following scenario: AWS DevOps Agent gets paged when the online boutique retail store application cart is experiencing P95 latency > 500ms for more than 2 minutes. This latency spike is critical and far exceeds the normal 5ms threshold, impacting the ability for customers to make purchases. In a typical scenario, the operations team would spend the first 15-30 minutes manually checking dependent services, alerts dashboard, and logs. This manual effort can be significantly reduced by configuring the New Relic observability platform with AWS DevOps Agent to automatically correlate telemetry data and surface the root cause faster.

To automatically remediate this issue, the online boutique application’s microservices are configured with New Relic’s APM agents that collect relevant metrics and send them to New Relic. When the latency exceeds a predefined threshold, an alert condition is triggered within New Relic. The triggered alert sends a notification to EventBridge, which in turn executes the Lambda function. The Lambda transforms the incoming payload into the required AWS DevOps Agent payload template. It then generates an HMAC signature to verify the message’s integrity and authenticity before dispatching it to the AWS DevOps Agent webhook endpoint.

Screen shot displaying new relic logo in the top left corner. The screen is divided into a navigation bar on the left, with Alerts selected. In the pane to the right, Alerts / Alerts Policies is displayed at the top, and below that a title Online Boutique High Latency appears. The notifications tab below that is selected.

Alert policy notifications in New Relic

The AWS DevOps Agent webhook triggers the agent to begin an automated investigation.

Screen shot displaying AWS DevOps Agent / GoldenPath_App in the title bar with Incident Response tab selected. Below that a heading is displayed for Online Boutique All Alerts followed by a timeline displaying User Request then Assistant Response

AWS DevOps Agent Incident response page

The New Relic MCP is first queried by the AWS DevOps Agent to retrieve telemetry data for the cart service GUID. Following this, the AWS DevOps Agent makes a second request to the New Relic MCP to formulate an investigation plan, which includes a list of related entities, their key metrics, and any associated change events for those dependencies.

Zoomed in screen shot of the previous screen with red boxes highlighting two areas in the timeline which say NewRelic MCP list related entities and NewRelic MCP list change events. Each shots the detail for the tool call.

AWS DevOps Agent and New Relic MCP interaction to list entities and related change events

Next, data gathering tasks are executed using New Relic MCP, following the investigation plan.

Time line screen similar to the previous screen shot with a red box around the timeline entry for Explore traces.

AWS DevOps Agent and New Relic MCP interaction to explore and analyze traces

Time line screen similar to the previous screen shot with a red box around the timeline entries for NewRelic MCP analyze entity logs and NewRelic MCP analyze golden metrics

AWS DevOps Agent and New Relic MCP interaction to explore and analyze logs and metrics

Continuing its analysis, the agent leverages New Relic’s MCP to examine entity logs, golden metrics, and traces, ultimately identifying the root cause for the latency spike.

Screen shot displaying AWS DevOps Agent / GoldenPath_App in the title bar with Incident Response tab selected. Below that a heading is displayed for Online Boutique All Alerts followed by a timeline displaying Update, Finding, and then Root cause. Root cause has a red box outlining it to draw attention.

AWS DevOps Agent Root Cause Analysis

You can review AWS DevOps Agent’s findings and the suggested root cause. The Site Reliability Engineer (SRE) can interact with the AWS DevOps Agent (side panel) in the chat panel to gain clarification on the steps of the ongoing investigation, enabling more effective monitoring and troubleshooting.

creen shot displaying AWS DevOps Agent / GoldenPath_App in the title bar with Incident Response tab selected. A timeline is visible on the lift and a chat window has been expanded on the right. The chat window contains a question and response.

AWS DevOps Agent Chat interface

You can review AWS DevOps Agent’s findings and the suggested root cause. If necessary, the SRE then executes the appropriate mitigation plan.

Conclusion

By integrating the New Relic MCP server with AWS DevOps Agent, organizations can quickly resolve issues when they arise and proactively prevent future incidents. This collaboration reduces Mean Time to Resolution (MTTR) and accelerates SREs and DevOps teams beyond manual, time-consuming investigations. It ensures rapid remediation of technical disruptions to minimize impact to the business. Ultimately, AWS DevOps Agent, the new frontier agent drives operational excellence, working in conjunction with the New Relic One Observability platform.

About New Relic
The New Relic Intelligent Observability Platform helps businesses eliminate interruptions in digital experiences. New Relic is an AI-strengthened platform that unifies and pairs telemetry data to provide clarity over your entire digital estate for proactive and predictive problem solving. That’s why businesses around the world run on New Relic to drive innovation, improve reliability, and deliver exceptional customer experiences to fuel growth.

Authors

Muthuvelan Swaminathan

Muthuvelan Swaminathan is a Principal Partner Architect at New Relic partnership organization building technical integrations with leading cloud providers and strategic partners. Through partner enablement, solution engineering and ecosystem alignment Muthuvelan helps drive product innovation at New Relic to ensure enterprises eliminate disruptions in their digital experiences for their customers.

Ruchika Bakolia

Ruchika Bakolia is a Software Engineer at New Relic. She is passionate about the intersection of AI and Cloud technologies, with extensive experience building and integrating solutions primarily on AWS. Ruchika enjoys traveling, reading, and exploring creative pursuits like pottery, always seeking out new experiences and challenges.

Nava Ajay Kanth Kota

Ajay Kota is a Senior Partner Solutions Architect at AWS, currently serving on the Amazon Partner Organization (APO) team collaborating closely with ISV Partners. With over 23 years of experience in enterprise computing infrastructure, Ajay brings deep expertise in cloud architecture, storage, backup, and cloud solutions. Before joining AWS, he led Storage, Backup, and Cloud teams, where he was responsible for developing Managed Services offerings across these domains.

10 Years of Let’s Encrypt Certificates

Post Syndicated from jzb original https://lwn.net/Articles/1049965/

Let’s Encrypt has published
a retrospective that covers the decade since it published its first
publicly trusted certificate in September 2015:

In March 2016, we issued our one millionth certificate. Just two years
later, in September 2018, we were issuing a million certificates every
day. In 2020 we reached a billion total certificates issued and as of
late 2025 we’re frequently issuing ten million certificates per
day. We’re now on track to reach a billion active sites, probably
sometime in the coming year.

Kroah-Hartman: Linux CVEs, more than you ever wanted to know

Post Syndicated from jzb original https://lwn.net/Articles/1049963/

Greg Kroah-Hartman is writing
a series of blog posts
about Linux becoming a Certificate
Numbering Authority (CNA):

It’s been almost 2 full years since Linux became a CNA (Certificate
Numbering Authority)
which meant that we (i.e. the kernel.org
community) are now responsible for issuing all CVEs for the Linux
kernel. During this time, we’ve become one of the largest creators of
CVEs by quantity, going from nothing to number 3 in 2024 to number 1
in 2025. Naturally, this has caused some questions about how we are
both doing all of this work, and how people can keep track of it.

So far, Kroah-Hartman has published the introductory post, as well
as a detailed
post about kernel version numbers
that is well worth reading.

How BASF’s Agriculture Solutions drives traceability and climate action by tokenizing cotton value chains using Amazon Managed Blockchain

Post Syndicated from Kevin S. Ridolfi original https://aws.amazon.com/blogs/architecture/how-basfs-agriculture-solutions-drives-traceability-and-climate-action-by-tokenizing-cotton-value-chains-using-amazon-managed-blockchain/

BASF Agricultural Solutions combines innovative products and digital tools with practical farmer knowledge. With over a century of experience, BASF offers a broad portfolio spanning seeds, crop protection, soil management, plant health, and digital agriculture solutions. Through collaboration with farmers, scientists, and partners, BASF strives to meet societal needs sustainably while creating a lasting agricultural legacy. Infosys is a global premier consulting and managed services partner of Amazon Web Services (AWS). Through this unique partnership, AWS helps customers integrate software, services, and processes to accelerate business transformation. This post explores the commitment of this partnership to driving positive change in the agricultural industry by using Amazon Managed Blockchain to tokenize food and cotton value chains for traceability, climate action, and circularity.

Global challenges and the agricultural industry

The world’s population is growing, with the UN projecting an estimated world population of 8.5 billion in 2030 and 10.4 billion by the end of the century. Along with this growth, as well as a global increase in standards of living, comes a rising demand for agricultural products such as fiber and food crops. At the same time, as society becomes more aware of the ecological impact of agriculture, both local communities and farmers are placing larger focus on a sustainable management of natural resources. The agricultural industry is uniquely positioned at the intersection of these two trends.

The agricultural industry faces numerous complex challenges that span both business and technical domains. From a business perspective, today’s agricultural supply chains have become incredibly complex, often involving multiple intermediaries across different countries. This complexity makes it difficult to ensure fair pricing and adequate compensation for farmers, who are often at the bottom of the value chain. Furthermore, verifying sustainable farming practices and organic certifications has become increasingly challenging, even as consumer demand for product authenticity and sustainability information continues to grow. Adding to these pressures, agricultural businesses must navigate increasing regulatory requirements for environmental regulation compliance and reporting, along with complex international trade regulations and documentation.

On the technical front, the industry struggles with limited digital infrastructure in rural farming areas, where internet connectivity and technology adoption remain significant hurdles. Data collection methods vary widely across different farms and regions, making it difficult to establish consistent metrics and reporting standards. Many agricultural businesses still operate with legacy systems that resist integration with modern tracking solutions, and the lack of standardization in agricultural data formats creates additional complications. Maintaining data integrity across multiple stakeholders has proven particularly challenging, as has the implementation of real-time tracking and tracing capabilities.

Cotton and fast fashion: Industry`s challenges

Cotton is the world’s most important natural fiber crop, with a yearly production of 126.5 million bales in the 2022–2023 season, enough to produce 25.3 billion pairs of jeans or 151.8 billion T-shirts. It also plays a major role in the fast fashion industry, where garments and clothing undergo a fast production and disposal cycle to quickly address customer attention and the latest fashion trends, with around 30% of clothing sold in the US being made with cotton. This accelerated production and disposal cycle comes at the expense of considerable environmental impact, with the fast fashion industry accounting for approximately 20% of the world’s water consumption and 10% of the world’s total CO2 emissions.

The cotton industry faces its own set of distinct challenges. Water usage stands as one of the most pressing concerns, with a single cotton T-shirt requiring approximately 2,700 liters of water to produce. Chemical usage tracking presents another significant challenge, as stakeholders must carefully monitor pesticide and fertilizer application throughout the growing process. Labor practices verification has become increasingly important, with brands and consumers demanding assurance of ethical working conditions throughout the supply chain.

Quality verification poses another crucial challenge, given that maintaining accurate documentation of cotton grade and characteristics is essential for pricing and processing. The industry’s global nature creates additional complexities in cross-border logistics, requiring careful management of international shipping and customs processes. Furthermore, the growing importance of sustainability certification has created new pressures to validate organic and sustainable farming practices with reliable, transparent documentation.

As consumer expectation of guaranteed fair practices, lower carbon emissions, and sustainable use of natural resources grows, so does the demand for traceability systems that can provide near real-time visibility into each step of the value chain by tracking sustainability information such as water consumption and CO2 emissions.

The potential for a blockchain-based solution

To address this demand, BASF identified blockchain as a foundational technology for a digital solution to deliver transparency along the value chain, targeting specific customer requirements for digital assets backed by information, validation, certificates, and know your business (KYB) policies for value chain partners.

Blockchain technology emerges as a particularly powerful solution to these challenges, offering unique capabilities that directly address many of the industry’s pain points. At its core, blockchain provides immutable record-keeping, creating permanent, tamper-proof records of transactions and events that ensure data integrity throughout the supply chain. This feature proves especially valuable in preventing fraudulent modification of sustainability certificates and maintaining the credibility of organic farming claims.

Smart contracts, a key feature of blockchain technology, enable the automation of compliance with agricultural standards and facilitate automatic payment execution based on predefined conditions. This automation significantly reduces administrative overhead in supply chain management and helps ensure fair compensation for farmers.

The technology’s traceability capabilities provide end-to-end visibility of cotton from seed to garment, enabling real-time tracking of sustainability metrics and creating transparent audit trails for certification purposes. This transparency helps brands and consumers verify the authenticity and sustainability of their cotton products while enabling farmers to demonstrate their commitment to sustainable practices.

Blockchain’s decentralized data management allows multiple stakeholders to maintain shared records without requiring a central authority, eliminating single points of failure in data storage and reducing dependency on central authorities. This decentralized approach proves particularly valuable in agricultural supply chains, where numerous parties need to access and verify information.

The implementation of token economics through blockchain creates new opportunities for incentivizing sustainable farming practices. Through tokenization, farmers can access new revenue streams, including carbon credits, while establishing more direct relationships with buyers. Additionally, blockchain’s digital identity capabilities provide secure authentication for supply chain participants, enabling granular access control to sensitive data and facilitating compliance with know your customer (KYC) and KYB requirements.

Solution overview

Using a permissioned blockchain based on open-source systems, BASF Agricultural Solutions has developed a novel way to promote data democratization and address the challenges of data recording, off-chain processes, and on-chain activities at scale. The solution enables value chain players to independently verify activities progressively, and an organizational structure within chain and off-chain monitors key performance indicators (KPIs) through a DAO (Distributed Autonomous Organization) interface.

To focus on building such a system rather than managing the underlying blockchain infrastructure, BASF selected Amazon Managed Blockchain alongside additional AWS services. Amazon Managed Blockchain simplifies BASF’s approach because it brings a suite of offerings that can be configured to build this solution without the need to add more layers and external or internal sources.

As a foundational system, Amazon Managed Blockchain augments the solution’s ability to generate smart certificates along with off-chain opportunities to further expand the offering as a platform, such as with AI and AWS Lambda. This fits into BASF’s vision to deliver best-in-class solutions for the farming community and deliver trusted information to communities that want to drive a positive impact for the planet.

The following are the key structural components of the solution:

  • Peers – These blockchain nodes run smart contracts (chain code) and maintain the ledger.
  • Ordering service – The ordering service makes sure a transaction meets the consensus requirements based on configured channel and endorsement policies for the installed chain code.
  • Fabric certificate authority (CA) – This component enrolls and generates blockchain identities needed to sign transactions.
  • AWS services – The solution uses various AWS services to perform operations on the blockchain efficiently. These services include:
    • Amazon Cognito – We use Amazon Cognito to onboard external users and clients to the platform.
    • AWS Fargate – A block listener is a custom service that listens to every block event from the blockchain and updates the off-chain storage accordingly. It’s hosted as a container on Fargate. Running a container using Fargate is more straightforward than other Kubernetes services because you don’t have to manage servers or clusters of Amazon Elastic Compute Cloud (Amazon EC2) instances. With Fargate, we no longer have to provision, configure, or scale clusters of virtual machines to run containers.
    • AWS Lambda – Middleware services are hosted as Lambda functions, which makes sure the services are automatically scalable by default and cost-efficient. This is important because we’re charged based on the number of requests for the function and the time it takes for the code to run.
    • Amazon OpenSearch Service – We use OpenSearch Service as an off-chain data store because the solution requires complex queries to aggregate the ledger data. The off-chain storage is kept in sync with the ledger and is restricted for direct updates. It can be updated only by an authorized application, based on ledger events.
    • AWS Secrets Manager – We use Secrets Manager to manage blockchain identities.
    • Amazon Simple Notification Service (Amazon SNS) – We use Amazon SNS to connect various services asynchronously.

The following diagram illustrates the solution architecture.

The solution architecture is extensible and scalable to meet the dynamic load requirements. It can seamlessly connect various data sources with appropriate connectors such as Salesforce, mobility platforms, third-party services, and more.

External users such as value chain players, retailers, and others who could benefit from tokens can access the platform through different methods. Generally, access of DAOs is done through business-to-customer (B2C) login, and API streams can be subscribed by end retailers for checkouts, point of sale (POS), and so on. Additionally, we provide internal access for admins and auditors to visualize the product flows.

Conclusion

Climate challenges are quite complex and require a joint approach between technology, the custodians of our planet (namely the farmers), and public chains that deliver the right protocols. BASF Agriculture Solutions represents the farming needs and the link to the right communities and crops on the ground, AWS brings in the right infrastructure and support of the cloud and scale, and Infosys brings in development support as a partner to both AWS and BASF.

BASF is connected to millions of farmers. BASF considers farming to be the biggest job on earth. Sustainable farming means bringing back lost biodiversity and increasing carbon capture within the soil. And sustainability overall requires additional effort by the farmers. Additionally, consumers like us make choices daily when it comes to our own purchase decisions, such as to buy sustainable products or take action that brings positive impact to the climate.

The solution outlined in this post creates a solution using blockchain as the base technology to enable a secure and reliable method for information sharing across all stakeholders. It’s the baseline to onboard use cases in the agriculture industry to enable end-to-end traceability with a 360-degree view. Smart contracts incentivize farmers and other stakeholders to follow the sustainable measures based on the information in the system, which is reviewed and authorized by validators. All the actions in the system are monitored and logged as immutable records, which enforce the information trust by default. This acts as a baseline for 100% traceability, tokenization for sustainable measures, and digital assets that can be exchanged and create a positive economy around sustainability. The design discussed in this post is flexible to onboard different use cases and can auto scale to meet dynamic data volumes.

We encourage you to join BASF, Infosys, and AWS in driving sustainability through trusted value chains that incentivize farmers, empower consumers, and create a positive economy around climate action. If you want to dive deep into topics surrounding sustainability and AWS architecture, we suggest visiting the AWS Architecture Blog.


About the Authors

[$] Mix and match Linux distributions with Distrobox

Post Syndicated from jzb original https://lwn.net/Articles/1049423/

Linux containers have made it reasonably easy to develop, distribute, and
deploy server applications along with all the distribution dependencies that they
need. For example, anyone can deploy and run a Debian-based PostgreSQL container
on a Fedora Linux host. Distrobox is a project that is designed to
bring the cross-distribution compatibility to the desktop and allow users to
mix-and-match Linux distributions without fussing with dual-booting, virtual
machines, or multiple computers. It is an ideal way to install
additional software on image-based systems, such as Fedora’s Atomic Desktops
or Bazzite, and also
provides a convenient way to move a development environment or
favorite applications to a new system.

The collective thoughts of the interwebz