Post Syndicated from The History Guy: History Deserves to Be Remembered original https://www.youtube.com/watch?v=9prYd6Oqd0M
On Moltbook
Post Syndicated from Bruce Schneier original https://www.schneier.com/blog/archives/2026/03/on-moltbook.html
The MIT Technology Review has a good article on Moltbook, the supposed AI-only social network:
Many people have pointed out that a lot of the viral comments were in fact posted by people posing as bots. But even the bot-written posts are ultimately the result of people pulling the strings, more puppetry than autonomy.
“Despite some of the hype, Moltbook is not the Facebook for AI agents, nor is it a place where humans are excluded,” says Cobus Greyling at Kore.ai, a firm developing agent-based systems for business customers. “Humans are involved at every step of the process. From setup to prompting to publishing, nothing happens without explicit human direction.”
Humans must create and verify their bots’ accounts and provide the prompts for how they want a bot to behave. The agents do not do anything that they haven’t been prompted to do.
I think this take has it mostly right:
What happened on Moltbook is a preview of what researcher Juergen Nittner II calls “The LOL WUT Theory.” The point where AI-generated content becomes so easy to produce and so hard to detect that the average person’s only rational response to anything online is bewildered disbelief.
We’re not there yet. But we’re close.
The theory is simple: First, AI gets accessible enough that anyone can use it. Second, AI gets good enough that you can’t reliably tell what’s fake. Third, and this is the crisis point, regular people realize there’s nothing online they can trust. At that moment, the internet stops being useful for anything except entertainment.
Do you have some rope? Then let’s teach about AI concepts
Post Syndicated from Jane Waite original https://www.raspberrypi.org/blog/do-you-have-some-rope-then-lets-teach-about-ai-concepts/
Teaching about AI concepts in schools is a tricky business as there are complicated ideas to be taught.
To teach complex concepts, in computer science, we often use an instructional approach called ‘unplugged’. We use the unplugged approach to teach computing concepts without a computer. Often unplugged activities include using an everyday analogy or a physical fun activity. For example, to teach about algorithms, students might learn how to make a jam sandwich where the recipe and following instructions accurately are similar to an algorithm and the steps within it used to write a program. The jam sandwich activity has now become a popular and key teaching experience for young students across the world, as it teaches a complex but fundamental idea in a simple and fun way.

At the January 2026 Raspberry Pi Foundation Research Seminar, Salomey Afua Addo, a researcher at the University of Cambridge, presented her work about how to teach about AI. She has specifically looked at this in the context of high school students in Ghana, where AI is now part of the mandatory curriculum. In Ghana, most schools do not have access to computers, therefore an unplugged approach to teach about AI is a good idea. Therefore, Salomey developed a set of unplugged activities to teach about a range of AI concepts.
Here, I focus on one of the activities that she presented — one that I think will become another ‘jam sandwich’ experience for students. So if you might teach about AI at some point, then read on.
Neural networks and rope: An unplugged activity
Salomey has designed an unplugged role-play activity to teach about neural networks and how they are trained to solve a problem. She focused on finding a familiar problem context for Ghanaian teachers and their students, and selected farming and crop disease. Students are asked to figure out what features about a farm are relevant for detecting diseases on cocoa trees. To solve the problem, students are given data about the farms (see Table 1). Giving students data, rather than preconceived rules about the context is key to the learning activity. Neural networks are data-driven — they provide a way to model given data so that we can make predictions. Here the features of farms, and importantly whether disease is or is not found in their cocoa trees, is the data that is used to train a model. The model is used to make predictions, which can then be used to improve farming by reducing crop disease.

Using farm data, students can learn how neural networks work, and they can do this through an unplugged role play — using ropes!
Here’s how Salomey’s classroom activity works. Sets of students act out the processes of training a neural network, including forward propagation, evaluation, and backpropagation. They take on the “roles’’ of some of the concepts of a neural network. One student acts as the supervisor, six students act as the input layer, two as the hidden layer, and one as the output layer.
Keeping it simple: Concepts and data
Key concepts are simplified for students:
- Forward propagation: The hidden layer players randomly select a set of farms (three of the six sets of input values), which reflects how weights are often set to random values at the start.
- Evaluation: The student acting as the output layer compares the prediction (whether crop disease is present or not) to the actual value for the farm to assess the error, similar to a loss function.
- Backpropagation: Inspired by MIT’s RAISE curriculum, this stage is modelled on establishing trust. Players in the hidden and output layers modify their trust in the previous layers (by adding or removing ropes) based on the accuracy of the prediction (if the farm has disease).
Simple numerical data about the features of the problem are given to the students, such as whether the “Temperature” is suitable (0=No, 1=Yes), if there are “Spots” on the plant (0=No, 1=Yes), if “Fertilizer’’ has been used, whether the “Leaf colour’’ is green or not (see Table 1). Importantly, each of the six features given are represented by the six “input layer” students. So each student can ‘process’ each feature as the data for a given farm is used to train the model. Cards are used to represent the data values passed between layers. And this is where the ropes come into play, as they are used to represent the connections between the nodes in the layers.

Instructions for each role
Written role-specific instructions are provided for the students to follow, for example, the Supervisor is given three steps to follow for the forward propagation stage, and the Input Layer students receive a different set of instructions and so on. The detail of the role play is shown in the instruction sheets (see Figure 1).

Why the ropes are important
Using ropes to connect the nodes becomes most important at the reverse propagation stage. The clever part of this is that we can show an increase or decrease in the strength of connection by adding or removing ropes. For me, this is the ‘jam sandwich’ effect. This, I think, is probably the most significant learning point. Here, the number of ropes that connect the nodes in the layers are changed based on the strength of evidence that a particular feature is indicated, by the data, to be relevant to the output. In this case, whether “Temperature”, for example, has an implied effect on cocoa disease or not — based on the data, not on any preconceived rule. Simply put, if a farm did have disease then a rope is added, if a farm did not then a rope is removed. Or at a more abstracted level, if a particular neuron contributes towards the correct prediction, a rope is added, otherwise a rope is removed. In a real neural network, backpropagation involves complex maths, such as calculus that would not be accessible to students of this age. Therefore, the rope is an analogy that replaces something that would be impossible for these students to grasp if it was taught using the real-world implementation.
Problem to be solved in the unplugged activity: Identify features that are relevant for detecting diseases
At the end of the activity, features (temperature, leaf color, family farm, etc.) with many rope connections are considered to be relevant for crop disease detection on the farm, whereas features with fewer rope connections are considered to be irrelevant for crop disease detection. The more ropes attached to a particular feature, e.g., temperature, represent its higher relevance in identifying crop disease on the farm.
Activity design, follow-on and evaluation
As part of the design of this activity, Salomey has simplified technical language so that throughout the role play students use everyday terms and she has chosen a context that is relatable for the students. For example, she uses the language of trust, and the new thickness of a rope connection, rather than using technical terms such as weight, loss function, and the error of the network.
Salomey also designed a follow-on activity that uses pen and paper. In this version of the activity, which she calls a board game, the students draw lines to connect the nodes in the layers. The thickness of the lines connecting the nodes represent the strength of the trust (see Figure 2).

Salomey also shared her evaluation of the resources. She conducted pre‑ and post‑intervention surveys with 39 teachers as part of the professional development on the AI teaching materials, and ten of those teachers implemented the unplugged activities in their classrooms. She reported that the teachers found the role-play activity was effective to demonstrate neural networks, that children worked independently to learn, and that some students who did not take part usually in class were engaged.
As well as sharing about her unplugged neural network activity, Salomey also talked about a set of AI stories that she has developed to teach about other aspects of AI applications. For example, the importance of fact-checking is demonstrated through a story about a young girl who fact-checked information she received from her friends about life in a city.
If you would like to find out more about Salomey’s work, you can find related materials on our seminar website.
Join our next seminar
Join us at our next seminar on Tuesday 17 March from 17:00 to 18:30 GMT to hear Rebecca Fiebrink (University of the Arts London) speak about teaching AI for creative practitioners. This will be the second seminar in our new series on how to teach about AI across disciplines. We hope to see you there!
To sign up and take part in our research seminars, click below:
You can also view the schedule of our upcoming seminars, and catch up on past seminars on our previous seminars page.
The post Do you have some rope? Then let’s teach about AI concepts appeared first on Raspberry Pi Foundation.
Избирателните списъци, демографията, регистрите и ефирната дъвка „мъртви души“
Post Syndicated from Боян Юруков original https://yurukov.net/blog/2026/spisaci-i-martvi-dushi/
Вчера публикувах текст разглеждаш последните избирателни списъци публикувани от ЦИК на база данните на ГРАО, както и промените в местата за гласуване и броя на секциите в страната. Има учудващо много промени.
Един интересен момент беше увеличаването на броя гласоподаватели спрямо изборите през октомври 2024-та. В предварителния списък публикуван на страницата на ЦИК има 6641768 души в избирателния списък. Това са данни подадени в ГРАО взети от базата ЕСГРАОН и включва всички български граждани навършили 18 години независимо къде се намират.
Виждаме увеличение в този предварителен списък спрямо октомври 2024-та – т.е. с 18 месеца разлика – с 1385 души. Имало е и увеличение през октомври 2024 спрямо вота през юни същата година с 934 (за 4 месеца). Спрямо юли 2024 и вота през април 2023-та обаче има намаление от 11835 души (19 месеца).
За това може да има няколко причини, които искам да разгледам тук. Осъзнавам, че темата с избирателните списъци винаги деградира в навикване за „мъртви души“ и дали българите зад граница следва да им се осигурява практическа възможност да упражняват конституционните си права. Знам също, че доста няма да дочетат статията ми, а ще извадя няколко числа и изменени цитата да подкрепят това, което вече са решили за себе си. Все пак, смятам, че е важно да се разисква темата, още повече, че нямам точен отговор защо има увеличение на избирателите в условията на демографска криза.

Предварителен спрямо окончателен списък
Първо трябва да се разбере, че тук разглеждам предварителния списък публикуван от ЦИК. По-късно ще бъде добавен окончателния, който ще бъде предоставен на изборните комисии в изборния ден. В него се махат всички, които нямат право да гласуват или по друга причина следва да бъдат махнати от списъците в България. Пример за първите са хората под запрещение. Ако подавате заявление за гласуване в чужбина, също бивате заличен от списъка по постоянен адрес и можете да гласувате в чужбина. Повече за това ще прочете тук.
В окончателните списъци виждаме 39121 по-малко души през октомври 2024-та спрямо предварителния, 46174 през юни и 56691 през 2023-та. От данните на Glasuvam.org знаем, че е имало 30684, 38269 и 38269 заявления за гласуване в чужбина. Това означава, че съответно 8437, 7905 и 18422 души са махнати съответно през изброените вотове по други причини. Като вземем окончателния списък и добавим искащите да гласуват в чужбина, виждаме намаление на избирателите с 402. За 4 месеца това е очаквано предвид данните, които ще разгледаме по-долу. Говори ни, че тези предварителни данни на ГРАО се чистят и въпреки увеличението, което виждаме сега, може все пак да има сравнимо и очакваното намаление в броя на избирателите.
Данните на ГРАО
Тъй като предварителните списъци са буквално взети от ГРАО, може да сравним с таблиците, които публикуват, за настоящ и постоянен адрес. За съжаление, там няма разбивка по възрасти. Там намираме обаче отговор на въпроса колко български граждани има всъщност. Той е в центъра на спира за тези списъци при всеки вот.
Независимо колко безсмислена е в последните 36 години, в България все още имаме концепцията за постоянен адрес. Това означава, че всеки, независимо къде се намира по света, трябва да има постоянен адрес в страната. Ако по някаква причина няма, се пише в общината. Придобилите българско гражданство чужденци, които не живеят в България, както и децата им родени след това гражданство, например, се пишат в Столична община. Имаме, разбира се, и настоящ адрес, който може да се смени с такъв в чужбина. При раждане в страната или чужбина и изваждане на български акт за раждане, за постоянен и настоящ адрес се взимат тези на майката или на бащата, ако майката не е български гражданин. При смърт където и да е се отписва човекът от адресите си.
Не на последно място, чужденци, които живеят в България и нямат гражданство, имат адресна регистрация, но не могат да участват на националните избори. Те обаче се включват в таблиците на ГРАО цитирани тук. От НСИ разбираме, че 72275 души от трети страни имат разрешение за пребиваване. Общия брой на хората с чуждо или без гражданство в България към края на 2024 са 147419 (показва се като 1-ви януари 2025 според методологията на Евростат). Това значи, че 75144 граждани на европейския съюз са пребивавали в България в края на 2024-та.
Според данните на ГРАО към 15-ти март 2026-та постоянен адрес в България имат 8226341 души. Настоящ имат 7405722. Това означава, че според ГРАО 820619 души са сменили настоящи си адрес в чужбина.
Имаме последни данни за населението от НСИ за 2024-та г. Те ни показват преценката им колко хора живеят предимно в България. Тази преценка е въз основа на икономическа и социална активност, миграция и прочие. Използват същите методи и дефиниции, както останалите Европейски държави. Все пак преброяванията показват отклонение, което налага корекция на населението. Същото виждаме и в други страни като Германия, където също имат оклонение от няколко процента.
С това наум и взимайки данните за постоянен адрес на ГРАО заедно с тези за чуждите граждани към декември 2024-та получаваме, че според НСИ към онзи момент е имало 1646569 български граждани живеещи в чужбина. Както обсъждах при анализа на диаспората ни в Германия, не всички са родени в България. Доста са деца на българи родени в чужбина, а други са получили българско гражданство и вече живеят зад граница. Има и доста, на които е било възстановено гражданството – предимно насилствено изселените български турци и децата им. Това число обаче не показва родените и напуснали България.
Данните на НСИ и президентството
Взимайки броят българи над 18 г. публикуван от ЦИК преди 18 месеца – 6640383 – може да се опитаме да изчислим колко следва да е към днешна дата. Ще вземем броя починали над 18 годишна възраст, броя родени през 2007 и 2008-ма и броя получили българско гражданство през последните 18 месеца. По-трудно е отколкото ще реши човек.
Нямаме все още данните за починалите през 2025-а и първите два месеца на 2026-та. Следва да вземем интервала между септември 2024 и март 2026-та, но може единствено да предположим. Последните данни са от 2024-та. Смъртните случаи на българи над 18 г. са около 100260 за 12 месеца. При липса на други данни и отправна точка може грубо да ги умножим по 150% и получаваме около 150390. Т.е. толкова трябва да махнем от населението. Разбира се, това са данните за съвсем друг период и може да има намаление на смъртните случаи в изминалата година, както и вариация в конкретните периоди. Отклонението може да е с 10-15 хиляди според кой месец включим и датите, на които е взета справката. Затова е важно да запомним, че при липсата на данни в реално време боравим с много условности.
Ражданията са също толкова сложни. По принцип, през 2007-та и 2008-ма е имало съответно 75335 и 77704 раждания. Тогава видяхме увеличение на ражданията, което, разбира се, медиите съобщиха като поредния „антирекорд“. Не намирам разбивка по месеци, но приблизително излиза 115370 живородени деца за 18 месеца. В това число обаче не влизат много от децата родени зад граница и получили българско гражданство. Взех една от старите ми справки получени от ГРАО и НСИ и приблизително може да очакваме, че 8375 деца са били родени в чужбина, регистрирани късно в България и са навършили 18 години от миналия вот насам. Така общо добавяме към избирателни списъци около 123745 души.
Към президентството има комисия, която предоставя или отнема българско гражданство. Покрай този процес имаше много скандали след като схемата за финансиране на партийни каси бяха осветени. Публикуват доклади, от които разбираме за броя такива решение. Имат доклад за 2025-та, но все още не за първите два месеца на 2026-та. Взимайки получилите или възстановили гражданство през 2025 и последните три месеца на 2024-та и извадим загубилите го, получаваме 18159. Исторически виждаме големи вариации през първото тримесечие от 2 до 7 хиляди. Ако вземем последните налични за 2025-та и направим предположение за януари и февруари 2026, както и септември 2024-та, получаваме приблизително 22297 лица, които са станали български граждани. За съжаление, нямаме разбивка по възрастови групи, но следейки процеса и материали по темата, мнозинството са пълнолетни.
Моята интерпретация
Ако приемем, че всички получили гражданство са пълнолетни и вземем оценката за родените и получили български акт за раждане и навършващи сега 18 г., разликата с оценката ми за смъртността остава около 4345 души намаление. Важното тук е, че условностите в числата горе означават огромна вариация в този резултат. Смъртните случаи може да са с 5000 по-малко предвид кой сезон взимаме за отправна точка, Тъй като данните ми са на няколко години, е възможно в последствие още деца родени в чужбина да са се регистрирали. Възможно е да има значително по-голям пик на получилите гражданство в последните два месеца.
4000 души са твърде малка разлика при толкова много условности и може да се направи обосновано предположение, че предварителните данни предоставени от ГРАО да отговарят на реалността. Още повече, че както стана ясно, в последствие ги изчистват на база изискванията на изборния кодекс, заявленията за гласуване и други сведения като общия брой по списъци в страната и в чужбина намалява спрямо предварителния списък между 8 и 18 хиляди при последните три вота.
При липса на окончателния списък и последните данни за населението и получилите гражданство и предвид разминаването в рамките на статистическата грешка предвид грубите сметки по-горе, не виждам сведения за манипулация или дописване на избирателни списъци. Разбира се, очаквам такива спекулации, както всеки път. Ще се радвам обаче да видя данните, на които почиват, защото често почиват на усещане и емоции. Последните са разбираеми и приемам притеснението, но имам нужда от нещо повече, ако ще говорим за измами и изборни манипулации.
Проблем има, дебат – не
Имам сериозни резерви към системата и процесите на ГРАО. Произтичат от разговори с хора в общините, НСИ и програмисти работили по системата. Отдавна трябваше да заличим концепцията за „постоянен адрес“, но е трудно и в политическия хаос и липсата на върховенство на закона това остава очаквано на заден план. Адресната регистрация е крайно пожелателна и къде заради липса на ресурс, къде заради нежелание общините не се занимават да следят темата. Тук може да добавим лисата на контрол над наемодателите от страна на НАП и масовото отдаване под наем без плащане на данък и като косвен ефект – отказ наемателя да се регистрира по настоящ адрес. Особено голям проблем е смяната на адреса при пренасяне в чужбина, който се задълбочава откакто имаме свободно пътуване.
Истината е, че подобни проблеми имат доста страни. Писал многократно как статистическата служба в Германия няма идея колко българи има в страната по аналогични причини. Важното обаче е ефектът им да се минимизира и най-малкото да знаем с какви условности боравим. Още по-важно е дали ефектът се променя през времето, т.е. дали е безпредметно да сравняваме година с година, защото данните са грешни, но в различна степен.
Всичко това се събира в частния случай на избирателни списъци. Доколкото горните неща са известни, ефектът им върху списъците е минимален защото задаваме прост въпрос – кой има ЕГН и е на 18 г. в деня на вота. Това е. Махат се някои хора според хипотези описани в Изборния кодекс, но като цяло базовите данни са прости. Всичко, което описах по-горе е опит за оценка с други източници на данните, за които ГРАО отговаря. Те не влизат в никоя стъпка от процесите на вота или изготвянето на списъците. Дори данните за българите зад граница не се гледат като правят секции.
Основно притеснение за т.н. „мъртви души“ е за починали хора, които обаче остават включени в списъците и се гласува от тяхно име. Чуват се анекдотни примери за това и не бих се учудил да има такива пропуски. Предполагам, че е и част от процеса на ГРАО в чистенето на избирателните списъци, още повече, че качеството на данните в ЕСГРАОН е единствено и само тяхно задължение. Предвид, че НСИ отчита над 36.7% смъртност над 95 г. (повече от 1 от 3-ма) и 17.2% смъртност между 80 и 94 г. всяка година, гледайки данните за населението на преклонна възраст, не мисля, че дори половината да са „фиктивни“, би имало някакъв ефект, за да има смисъл усилието да се подправят държавни бази данни. Единствено, може би, на местни избори и то ако ги концентрират в някоя малка община.
Далеч по-ефективен е сегашния подход да овладяване на изборни комисии, ЦИК и РИК чрез свързани лица, купуването на гласове при бездействие и чадър от местната полиция, гласуването под команда от местни тартори, овладяването на цели райони чрез нарочни проверки на регулатори и висящи обвинения от прокуратурата, за да се привлече местния бизнес, кмет и общински съветници да съучастват в контролирания вот и от там икономически и социално да заставят местното население да гласува за КОЙто трябва. Всичко това се случва днес и се е случвало с малки прекъсвания и изключения на места в последните десетилетия. Организацията е по-лесна и мащаба му е далеч по-голям, за да не си струва да се занимават с малкото т.н. „мъртви души“.
Отново, не значи, че няма грешки в системата. Писах за критиките си в тази посока. Не съм видял доказателства обаче, че някой се възползва от тях в мащаби, в които да има значение. Виждам доказателства за други изборни измами, за които прокуратурата и съда спят.
Избори 2026 – секциите в България и брой избиратели според ГРАО
Post Syndicated from Боян Юруков original https://yurukov.net/blog/2026/iz2026-izbirateli/
ЦИК са публикували адресите и броя гласоподавали за изборите през април. Очаквам да започнат спекулации по темата, затова ще опиша това, което виждам в таблиците на ГРАО предоставени на ЦИК. В следващата ми статия ще се заровя по-подробно в демографските данни.

По списъците има 6641768 гласоподавателя, което е увеличение от 1385 спрямо списъка от октомври 2024-та. Напомням, че в това число се включват всички българи в чужбина независимо кога са напуснали страната или дали са родени там. Те имат ЕГН и лични документи и затова са в списъците, но НСИ не ги брои към населението в страната, защото в последната година не са в България.
Най-голямо увеличение на гласоподавателите има в София – 19714 повече. Увеличение има и в областите Благоевград – 4156, Кърджали – 3143, Варна – 1775 и Бургас – 1451. Най-голямо намаление има в Перник – 3006, следван от Вража, Русе и Велико Търново с малко над 2000.
Има увеличение на секциите, в които ще се гласува в страната – общо 16. По области обаче се забелязва, че в София увеличават с 10 секции, в Пловдив – с 9, а в Бургас – с 2. В Софийска област намаляват с 4, а в Перник, Ловеч, Видин и Смолян – с 1 или 2 секции. Това показва обаче само броят секции. В София на едно място има средно по 5.2 секции, а в Пловдив, Варна и Сливен – по 2 до 2.3. Средното за всички останали е 1.37.
Местата също варират много. Общо 58 места или 0.81% от всички със секции в тях се променят спрямо миналия вот. Макар това да значи едно на всеки 123 места за гласуване да е променено, все пак е добра идея да проверите дали точно вас засяга, особено ако живеете в град. Може лесно да го сторите на страницата на ГРАО с име и ЕГН. (Страницата им ще бъде активна малко преди изборите)
Например, в Благоевградска област. Добавят секция в Благоевград в 2-ро ОУ Димитър Благоев, но махат секции в Петрич на три места – детските градини на ул. Солунска и Цар Симеон и сградата на социално подпомагане на ул. Славянска.
Макар в национален мащаб да има 16 повече секции, общо на 8 по-малко места ще има избори. Единствено в Хасково, Шумен, Сливен, Плевен, Кюстендил и Габрово изглежда няма да има промени. Най-много места Закриват във Варна и Перник – 9 и 8 съответно. Най-много откриват в София – 7 и в Търговище и Велико Търново – по 3.
Като гласоподаватели на секция, най-много има в София – 774, следван от Разград – 708, Пазарджик – 668 и Варна – 646. Най-малко има във Видин – 343. Средно, без да броим София, има 522 гласоподавателя на секция в страната. Ако вземем предвид активността по области през октомври 2024-та, грубо казано 192 души са гласували средно на секция. За сравнение в чужбина средният брой е 178.
Разбира се, всичко това е ужасно трудно да се проследи, тъй като за поредна година ЦИК и ГРАО не публикуват географски координати на секциите независимо, че ги имат. Налични са само адреси, които не са особено полезни, ако не познаваш добре мястото. За адреси на 580 места със секции е отбелязано само „кметството“. 228 – читалището. 140 – клуб на пенсионера. 290 – в някаква вариация на училище Христо Ботев.
Повече за секциите зад граница и проблемите с отварянето им може да прочетете в предишната ми статия. Тук може да следите всички новини за изборите в чужбина през 2026-та или да се абонирате за бюлетина ми в Glasvuvam.org.
How Cloudy translates complex security into human action
Post Syndicated from Ayush Kumar original https://blog.cloudflare.com/cloudy-upgrades-for-cloudflare-one/
Today’s security ecosystem generates a staggering amount of complex telemetry. For instance, processing a single email requires analyzing sender reputation, authentication results, link behavior, infrastructure metadata, and countless other attributes. Simultaneously, Cloud access security broker (CASB) engines continuously scan SaaS environments for signals that detect misconfigurations, risky access, and exposed data.
But while detections have become more sophisticated, explanations have not always kept pace.
Security and IT teams are often aware when something is flagged, but they do not always know, at a glance, why. End users are asked to make real-time decisions about emails that may impact the entire organization, yet they are rarely given clear, contextual guidance in the moment that matters.
Cloudy changes that.
Cloudy is our LLM-powered explanation layer, built directly into Cloudflare One. It translates complex machine learning outputs into precise, human-readable guidance for security teams and end users alike. Instead of exposing raw technical signals, Cloudy surfaces the reasoning behind a detection in a way that drives informed action.
For Cloudflare Email Security, this means helping users understand why a message was flagged before they escalate it to the security operations center, or SOC. For Cloudflare CASB, it means helping administrators quickly understand the risk and remediation path for SaaS findings without having to manually assess low-level signals.
This post outlines how we are extending Cloudy across Phishnet and API CASB to improve decision making, reduce unnecessary noise, and turn complex security signals into clear, actionable insight.
When an email is analyzed by Cloudflare Email Security, it is not evaluated by a single signal or model. Instead, a wide range of machine learning models analyze different parts of the message, from sender reputation and message structure to content, links, and behavioral patterns. This model set continues to grow as our machine learning team regularly trains and deploys new detections to keep pace with evolving threats.
Based on this analysis, messages are labeled with outcomes such as Malicious, Suspicious, Spam, Bulk, or Spoof. While these detections have been effective, we consistently heard feedback from customers that it was not always clear why a message was flagged. The decision was correct, they told us — but the reasoning behind it was often opaque to both end users and security teams.
To address this, we introduced the first version of Cloudy: LLM-powered summaries for detections. These summaries translate what our machine learning models are seeing into human readable explanations. Initially, these summaries were available in the Cloudflare dashboard to help SOC teams during investigations. Over the past few months, customer feedback has confirmed that these explanations significantly improve understanding in our detections.
As we continued speaking with customers, another challenge surfaced. Our Phishnet tool allows users to submit messages to the SOC when they believe an email may be suspicious. While this empowers employees to participate in security, many SOC teams told us their queues were being flooded with submissions that turned out to be clean messages.
The result was unnecessary backlog and slower response times for emails that actually required investigation.
At the same time, customers told us that traditional security awareness training was not always enough. Users still struggled to evaluate emails in the moment, when it mattered most. They wanted more contextual guidance directly within the workflow where decisions are made.
This upgrade is designed to address both of these problems. By bringing clearer explanations and contextual education directly into Phishnet, we aim to help users make better decisions while reducing noise for SOC teams, without sacrificing security.
As organizations and attack techniques have evolved, so has the role of the end user. Modern email threats increasingly rely on social engineering, subtle impersonation, and psychological pressure which places users directly in the decision path.
In response, users are being asked to act as an additional layer of defense. However, traditional security awareness tools often fall short. Training is typically delivered through periodic sessions or simulated phishing campaigns, disconnected from real messages and real decisions. When users encounter an unfamiliar email, they are left without enough context to confidently assess risk.
This gap commonly leads to one of two outcomes. Some users submit nearly every questionable message to the SOC, creating excessive noise and slowing down investigations. Others interact with messages they should not, simply because nothing in the moment signals clear risk.
By embedding Cloudy directly into Phishnet, we close this gap.
Users receive immediate, contextual explanations that help them understand what Cloudflare is seeing and why a message may be risky. This enables users to make informed decisions at the point of interaction, reduces unnecessary escalations to the SOC, and allows security teams to focus on the messages that truly require attention.
Over time, this approach shifts users from being a source of noise to becoming an effective part of the detection and response workflow. The result: stronger email security, without adding friction or burden to security teams.
In the next month, we will be upgrading our Phishnet reporting button to extend the Cloudy summaries.

The new Phishnet screens will show Cloudy summaries.
With this upgrade, end users receive a simplified, user-friendly version of Cloudy summaries at the moment they report a message. These summaries are generated in real time using Cloudflare Workers AI and run directly on Cloudflare’s global Workers platform when a user interacts with a message in Phishnet.
When a user clicks the Phishnet reporting button, the request triggers a Workers-based workflow that aggregates structured outputs from multiple detection models associated with that message. These model outputs include signals such as sender reputation, domain and infrastructure characteristics, authentication results, link and content analysis, and behavioral indicators collected during message processing.
The aggregated signals are then passed to Workers AI, where a series of purpose-built prompts generate a natural language explanation. Each prompt is designed to transform low-level detection outputs into a concise and human-readable summary. This process focuses on explanation rather than classification and does not alter the original disposition of the message.

How Cloudy transforms detections into clear explanations.
For this experience, we intentionally redesigned the summaries compared to those shown to administrators in the Cloudflare dashboard. During testing, we found that admin-focused summaries often relied on technical concepts that were difficult for non-technical users to interpret. Terms such as ASNs, IP reputation, or authentication failures required translation.
To ensure end users can understand the summaries, Phishnet emphasizes plain-language explanations while preserving the meaning of the underlying detections.
|
Signal |
What it means |
Cloudy translation for end users |
|
SPF Fail |
Sender explicitly not authorized by SPF |
This email failed a sender verification check. |
|
DKIM Fail |
Message signature does not validate |
The message integrity check failed, which can be a sign of tampering. |
|
DMARC Fail |
DMARC policy check failed |
The sender’s domain could not confirm this email is legitimate. |
|
Reply to Mismatch |
Reply To differs from From |
Replies may go to a different address than the sender shown. |
|
Domain Age |
Domain recently registered |
The sender domain is newly created, which is common in phishing. |
|
URL Low Reputation |
Destination URL has poor reputation |
The link destination has signals associated with risk. |
Because this workflow runs on the Cloudflare Workers platform, summaries are generated with low latency and at global scale — so users receive immediate feedback at the moment of interaction. This real-time context allows users to better understand why an email may be risky or why it appears safe before deciding whether to escalate it to the SOC.
We are currently beta testing this experience with Microsoft customers to ensure the summaries are accurate and reliable. Cloudy summaries are not trained on customer data. We are also applying additional validation to ensure the generated explanations do not hallucinate. Accuracy is critical at this stage as incorrect guidance could introduce real security risk.
Following the beta period, we plan to expand access to all Microsoft users. We will also bring similar upgrades to the Phishnet sidebar for Google Workspace users later in 2026.
But helping end users better understand what makes an email risky is only part of the story. We are also applying Cloudy to the administrative side of security operations, where clarity and speed matter just as much. Beyond Phishnet, Cloudy now translates complex CASB findings into structured explanations that help security and IT teams quickly understand risk, prioritize remediation, and take confident action across their SaaS environments.
Inside Cloudflare One, our SASE platform, CASB connects to the SaaS and cloud tools your teams already use. By talking to providers over API, CASB gives security and IT teams:
-
A consolidated view of misconfigurations, overshared files, and risky access patterns across apps like Microsoft 365, Google Workspace, Slack, Salesforce, Box, GitHub, Jira, and Confluence (CASB Integrations).
-
Continuous scanning for new issues as users collaborate, share, and adopt new tools.
-
Findings that are organized, searchable, and exportable for triage and reporting.

A typical CASB Findings page showing detections for a Microsoft 365 finding.
Until now, understanding what exactly triggered a CASB Finding — the detections that CASB makes across connected SaaS integrations — has been a black box. While the information was there to put together an explanation of why that file, that user, that configuration was triggering a CASB Finding Type, it wasn’t exactly obvious the reason why it was ultimately detected by our system.
With the introduction of Cloudy summaries in CASB, users receive a short description of the detection rationale with the specific details of the match listed out for easy comprehension.
Unlike a simple text summary, Cloudy for CASB provides a structured breakdown designed for immediate remediation. As seen in our beta testing across different providers, from Microsoft 365 to Dropbox, the model consistently parses findings into two distinct sections:
-
Risk: It identifies exactly why the finding matters. For instance, rather than just noting a ‘Suspended User,’ Cloudy clarifies that this ‘may indicate a compromised account or a user who should no longer have access to company data’.
-
Guidance: It offers immediate next steps. Instead of generic advice, it suggests specific actions, such as verifying if a suspension was intentional or reviewing an application’s legitimacy before revoking access.
This structure ensures that analysts can understand the gravity of a finding without needing deep expertise in the specific SaaS application involved.

An example Cloudy Summary in a CASB Posture Finding.
|
Finding Type |
Technical Signal |
Cloudy Translation (Risk & Guidance) |
|
Identity & Access |
Dropbox: Suspended User |
Risk: A suspended user account may indicate a compromised account or a user who should no longer have access to company data. Guidance: Verify that the suspension is intentional and that the user’s access has been properly revoked. |
|
Shadow IT |
Google Workspace: Installed 3rd-party app |
Risk: This installed application with Google Sign In access may pose a risk of unauthorized access to user data. Guidance: Review the application’s legitimacy and necessity, and consider revoking access if it is no longer needed. |
|
Email Security |
Microsoft 365: Domain DMARC record not present |
Risk: The absence of a DMARC record may leave the domain vulnerable to email spoofing and phishing attacks. Guidance: Configure a DMARC record for the domain to specify how to handle unauthenticated emails. |
|
Data Loss Prevention |
Microsoft 365: File publicly accessible + DLP Match |
Risk: This file being shared publicly with edit access may allow unauthorized modifications… especially given the potential sensitive content indicated by the DLP Profile match. Guidance: Review the file’s content… and consider restricting access if necessary. |
We know that when it comes to our customers getting to the bottom of identified security issues, time is of the essence. We believe that any amount of unnecessary uncertainty or lack of clarity around what’s going wrong just puts more time between an imperfect state and one that is more secure.
We built this feature on the same privacy-first foundations as all products at Cloudflare. Cloudy summaries in CASB are generated using Cloudflare Workers AI, ensuring that your data remains within our secure infrastructure during analysis. The models are not trained on your SaaS data, and the summaries are generated ephemerally to aid in triage. This allows your team to leverage the speed of AI without exposing sensitive internal documents or configurations to public models.
For Email Security, we will continue to expand how Cloudy supports both administrators and end users. Our focus is on delivering clearer explanations, better in context guidance, and deeper integration into daily workflows.
For CASB, we’re excited to look for opportunities where Cloudy can make it even easier for CASB administrators to understand what’s going on across their cloud and SaaS apps. Keep an eye out as we look to expand Cloudy coverage to allow administrators to query their findings using natural language, further reducing the time it takes to identify and remediate risks.
Looking ahead, this includes richer explanations for additional detection types, tighter feedback loops between user actions and detections, and continued improvements to how users and SOC teams collaborate through Phishnet. Our goal is to make Cloudy a core part of how organizations understand, trust, and act on email security decisions.
We provide all organizations (whether a Cloudflare customer or not) with free access to our Retro Scan tool, allowing them to use our predictive AI models to scan existing inbox messages in Microsoft 365.
Retro Scan will detect and highlight any threats found, enabling organizations to remediate them directly in their email accounts. With these insights, organizations can implement further controls, either using Cloudflare Email Security or their preferred solution, to prevent similar threats from reaching their inboxes in the future.
If you are interested in how Cloudflare can help secure your inboxes, sign up for a phishing risk assessment here.

From reactive to proactive: closing the phishing gap with LLMs
Post Syndicated from Sebastian Alovisi original https://blog.cloudflare.com/email-security-phishing-gap-llm/
Email security has always been defined by impermanence. It is a perpetual call-and-response arms race, where defenses are only as strong as the last bypass discovered and attackers iterate relentlessly for even marginal gains. Every control we deploy eventually becomes yesterday’s solution.
What makes this challenge especially difficult is that our biggest weaknesses are, by definition, invisible.
This problem is best illustrated by a classic example from World War II. Mathematician Abraham Wald was tasked with helping Allied engineers decide where to reinforce bomber aircraft. Engineers initially focused on the bullet holes visible on planes returning from missions. Wald pointed out the flaw: they were reinforcing the areas where planes could already take damage and survive. The true vulnerabilities were on the planes that never came back.

Email security faces an identical hurdle: our detection gaps are unseen. By integrating LLMs, we advance email phishing protection and move from reactive to proactive detection improvement.
The limits of reactive defense
Traditional email security systems improve primarily through user-reported misses. For example, if we marked a spam message as clean, customers can send us the original EML to our pipelines for our analysts to analyze and update our models. This feedback loop is necessary and valuable, but it is inherently reactive. It depends on someone noticing a failure after the fact and taking the time to report it.
That means detection improvements are often driven by what attackers already succeeded at, rather than by what they are about to exploit next.
To close this gap, we need a way to systematically observe the “planes that didn’t make it back.”
Large Language Models (LLMs) hit the mainstream market in late 2022 and early 2023, fundamentally changing how we process unstructured data. At their core, LLMs use deep learning and massive datasets to predict the next token in a sequence, allowing them to understand context and nuance. They are particularly well-suited for email security because they can read natural language and characterize complex concepts (like intent, urgency, and deception) across millions of messages.
Every day, Cloudflare processes millions of unwanted emails. Historically, it was not feasible to deeply characterize each message beyond coarse classifications. Manually mapping emails to nuanced threat vectors simply did not scale.
Now, Cloudflare has integrated LLMs into our email security tools to identify threats before they strike. By using the power of LLMs, as we’ll describe below, we can finally see a clear and comprehensive picture of the evolving threat landscape.

Our LLM-driven categorization shows clear spikes and persistent trends across several distinct categories, including “PrizeNotification” and “SalesOutreach”.
These LLM-generated tags provide Cloudflare analysts with high-fidelity signals in near real time. Tasks that previously required hours of manual investigation and complex querying can now be surfaced automatically, with relevant context attached. This directly increases the velocity at which we can build new targeted Machine Learning models or retrain existing ones to address emerging behaviors.
Because Cloudflare operates at global Internet scale, we can gather these insights earlier than ever before, often before a new technique becomes widely visible through customer-reported misses.
One of the clearest patterns we’ve identified using this new intelligence is the continued persistence of malicious messages structured to look like Sales Outreach-style phishing. These emails are designed to mimic legitimate B2B communication, often presenting opportunities to purchase or receive “special deals” on unique items or services, to lure targets into clicking malicious links or providing credentials.
Once LLM categorization surfaced Sales Outreach as a dominant vector, we moved from broad visibility to targeted data collection.
Using LLM-generated tags, we began systematically isolating messages that exhibited Sales Outreach characteristics across our global dataset. This produced a continuously growing, high-precision corpus of real-world examples, including confirmed malicious messages as well as borderline cases that traditional systems struggled to classify. From this corpus, we built a dedicated training pipeline.
First, we curated training data by grouping messages based on shared linguistic and structural traits identified by the LLMs. These traits included persuasive framing, manufactured urgency, transactional language, and subtle forms of social proof.
Next, we focused feature extraction on sentiment and intent rather than static indicators. The model learns how requests are phrased, how credibility is established, and how calls to action are embedded within otherwise normal business conversations.
Finally, we trained a purpose-built sentiment analysis model optimized specifically for Sales Outreach behavior. This avoided overloading a general phishing classifier and allowed us to tune precision and recall for this threat class.

The output of this model is a risk score that reflects how closely a message aligns with known Sales Outreach attack patterns. That score is evaluated alongside existing signals such as sender reputation, link behavior, and historical context to determine whether a message should be blocked, quarantined, or allowed.
This process is continuous. As attackers adapt their language, newly observed messages are fed back into the pipeline and used to refine the model without waiting for large volumes of user-reported misses. LLMs act as the discovery layer by surfacing new linguistic variants, while the specialized model performs fast and scalable enforcement.
This is what an all-out offensive looks like in practice. It is a feedback loop where large-scale language understanding drives focused, high-precision detection. The result is earlier intervention against a threat class that thrives on subtlety, and fewer malicious sales emails reaching the inbox.
The visibility unlocked by LLM-driven mapping fundamentally changed how we improve detections. Instead of waiting for attackers to succeed and relying on downstream user reports, we gained the ability to identify systemic gaps earlier and address them at the source. This shift from reactive remediation to proactive reinforcement translated directly into measurable customer impact.
The most immediate signal of success was a marked reduction in customer friction. Sales Outreach–related phishing has historically generated a high volume of user-reported misses, largely because these messages closely resemble legitimate business communication and often evade traditional rule-based or reputation-driven systems. As our targeted models came online and were continuously refined using LLM-derived insights, fewer of these messages reached end users in the first place.
The data reflects this change clearly. Average daily Sales Outreach submissions — messages that we labeled as clean but were in fact Sales Outreach phishing emails, flagged by end users — dropped from 965 in Q3 2025 to 769 in Q4 2025, representing a 20.4% reduction in reported misses in a single quarter.

This reduction is not just a metric improvement; it represents thousands fewer disruptive moments per day for security teams and end users alike. Each avoided submission is a phishing attempt that was stopped before it could erode trust, consume analyst time, or force a user to make a security judgment mid-workflow. We have seen this trend continue in Q1 of 2026 with average daily submissions decreasing by two-thirds.

In effect, LLMs allowed us to “see” the planes that never made it back. By illuminating previously invisible failure modes, we were able to reinforce defenses precisely where attackers were concentrating their efforts. The result is a system that improves not only detection rates, but also the day-to-day experience of the people relying on it.
Our work with LLMs is just beginning.
To stay ahead of the next evolution of attacks, we are moving toward a model of total environmental awareness by refining LLM specificity to extract forensic-level detail from every interaction. This granular mapping allows us to identify specific tactical signatures rather than relying on broad labels.
Simultaneously, we are deploying specialized machine learning models purpose-built to hunt for emerging, high-obfuscation vectors at the “fringes” that traditional defenses miss. By leveraging this real-time LLM data as a strategic compass, we can shift our human expertise away from known noise and toward the critical gaps where the next strike is likely to land.
By illuminating the “planes that didn’t make it back,” we are doing more than just reacting to missed email; we are systematically narrowing the battlefield. In the email arms race, the advantage belongs to the side that can see the invisible first.
We provide all organizations (whether a Cloudflare customer or not) with free access to our Retro Scan tool, allowing them to use our predictive AI models to scan existing inbox messages in Microsoft 365.
Retro Scan will detect and highlight any threats found, enabling organizations to remediate them directly in their email accounts. With these insights, organizations can implement further controls, either using Cloudflare Email Security or their preferred solution, to prevent similar threats from reaching their inboxes in the future.
If you are interested in how Cloudflare can help secure your inboxes, sign up for a phishing risk assessment here.
See risk, fix risk: introducing Remediation in Cloudflare CASB
Post Syndicated from Alex Dunbrack original https://blog.cloudflare.com/remediation-in-cloudflare-casb/
Starting today, Cloudflare CASB customers can do more than see risky file-sharing across their SaaS apps: they can fix it, directly from the Cloudflare One dashboard.
This launch marks a huge advancement for Cloudflare’s Cloud Access Security Broker (CASB). Since its release, Cloudflare’s API-based CASB has focused on providing robust, comprehensive visibility and detection. It also connects to the SaaS tools your business runs on, surfacing misconfigurations, and flagging overshared data before it becomes tomorrow’s incident.
With today’s release of Remediation – a new way to fix problems with just a click, right from the CASB Findings page – CASB begins its next chapter, and moves from telling you what’s wrong to helping you make it right.

An example of a Remediation Action (Remove Public File Sharing) in a CASB Finding.
Inside Cloudflare One, our SASE platform, CASB connects to the SaaS and cloud tools your teams already use. By talking to providers over API, CASB gives security and IT teams:
-
A consolidated view of misconfigurations, overshared files, and risky access patterns across apps like Microsoft 365, Google Workspace, Slack, Salesforce, Box, GitHub, Jira, and Confluence (CASB Integrations).
-
Continuous scanning for new issues as users collaborate, share, and adopt new tools.
-
Findings that are organized, searchable, and exportable for triage and reporting.
But until now, the actual fixing usually happened somewhere else, whether it’s inside each app’s admin UI, or through a ticket to the team that owns that tool. Remediation closes that loop.
The launch of CASB Remediation marks a major shift forward for the product and Cloudflare One, and we have a ton of big updates planned for the next year.
With today’s release, we focused on fixing file-share issues in Microsoft 365 and Google Workspace.
With Remediation, you can fix the highest-impact, most common file risks we see across customers, including:
-
Public links that let anyone on the Internet view or edit a file.
-
Files shared company-wide across your tenant or domain, even when just a handful of people should have access.
-
Files shared outside your organization to personal accounts and external domains.
-
All of the above, when they also match a DLP Profile. For example, a document full of customer records, credentials, or financial details.
When you trigger the ‘Remove sharing’ Remediation action on a supported finding, CASB immediately moves to remove the risky sharing configuration (for example, the public link or organization-wide access) from the file in question. And crucially, Remediation only removes risky sharing; it doesn’t delete files or change who owns them.

A new page to track the progress and success of Remediated CASB findings.
We chose to start with Microsoft 365 and Google Workspace because, for many organizations, that’s where the bulk of their business-critical documents live: internal financials, product roadmaps, customer contracts, HR notes, and more.
They’re also where “temporary” sharing tends to linger too long:
-
A spreadsheet shared “Anyone with the link can edit” for a quick review.
-
A doc made company-wide for an all-hands, then quietly forgotten.
-
A sheet of customer records shared to a contractor’s personal email.
For Microsoft 365, that means cleaning up risky shares in places like OneDrive and SharePoint. For Google Workspace, it means tightening sharing on Docs, Sheets, Slides, and other files stored in Drive.
Instead of exporting a CSV of risky files out of CASB, sending it to app owners, and hoping everyone gets around to fixing their share settings, you can drive the clean-up directly from CASB and know when those risks have actually been addressed.
And when you and your team use CASB Remediation, every action is logged in Cloudflare One’s Admin logs, so you can see who took action on which files and when, or export that activity to your security information and event management tool (SIEM).
When architecting the system that supports CASB Remediations, we knew it had to do three things really well:
-
Be fast, even at scale
-
Durable execution to handle surprises gracefully
-
Be easy for our customers to use
To meet these goals, we built a system using several Cloudflare products: Workers, Workflows, Queues, Workers KV, Secrets Store, and Hyperdrive.
When a remediation job is initiated, an API call is made to a Worker. That Worker writes the job to a Queue which is consumed by a second Worker to kick off a Workflow. Workers KV and Secrets Store are used to securely distribute credentials for use in the Workflow. The Workflow runs a series of steps to collect information and execute third-party API calls to complete the remediation. The final outcome of the action is recorded in a database via Hyperdrive.
At scale, we are guaranteed to encounter 429s from vendor APIs. Workflows’ native retries simplify handling this, and built-in step logging gives visibility into each retry. This means that there was no need for us to build a complex, single-purpose, state-tracking system or dozens of serverless functions for each action.

Performance results from load testing and early access customers have shown strong performance even under heavy load. The average (p50) end-to-end job completion time is 48 seconds, and the p90 is 72 seconds. Durable Execution (via Workflows) has made job management completely hands-off for our team, even when the Workflow encounters issues with third-party APIs. The simplicity of the final system has made troubleshooting issues fast and straightforward.
File-sharing Remediation for Microsoft 365 and Google Workspace is just the first step.
In the near term, we’re working on bringing our customers new Quarantine actions, which can move or isolate high-risk files to safer locations. We are also introducing Custom Webhook actions, hooks that let you trigger downstream workflows, like ticket creation, chat notifications, or your own automation.
And more broadly, we’re excited to explore ways to make CASB even more of an active control plane:
-
Autoremediation policies for carefully scoped, policy-driven fixes where you’re comfortable letting CASB take action automatically.
-
Custom CASB findings so you can define the exact patterns, data types, or access conditions that matter most to your organization.
-
Bulk Remediation that allows you to remediate many similar findings in a single operation.
-
Extending Remediation to additional SaaS integrations beyond Microsoft 365 and Google Workspace, so the same experience applies to tools like Box, Dropbox, Salesforce, GitHub, Slack, Atlassian, and more over time.
CASB Remediation requires a paid CASB license, but don’t let that stop you from trying CASB out today!
-
For existing Cloudflare One / CASB customers: Integrate your Microsoft 365 or Google Workspace tenant (or update your existing integration to Read-Write), and start remediating risky shares directly from the side panel within your file sharing-related finding types.
-
New to Cloudflare One? Sign up now for 50 free seats to begin using CASB immediately. For larger deployments, request a consultation with our experts.
From there, talk to our team about enabling CASB with Remediation for your Microsoft 365 and Google Workspace tenants so you can find and fix overshared files in one place.
We’re excited to see how you use Remediation to clean up long-lived file-sharing risks — and to help shape what CASB’s next generation of remediation capabilities looks like.
Optimizing Recommendation Systems with JDK’s Vector API
Post Syndicated from Netflix Technology Blog original https://netflixtechblog.com/optimizing-recommendation-systems-with-jdks-vector-api-30d2830401ec
By Harshad Sane
Ranker is one of the largest and most complex services at Netflix. Among many things, it powers the personalized rows you see on the Netflix homepage, and runs at an enormous scale. When we looked at CPU profiles for this service, one feature kept standing out: video serendipity scoring — the logic that answers a simple question:
“How different is this new title from what you’ve been watching so far?”
This single feature was consuming about 7.5% of total CPU on each node running the service. What started as a simple idea — “just batch the video scoring feature” — turned into a deeper optimization journey. Along the way we introduced batching, re-architected memory layout and tried various libraries to handle the scoring kernels.
Read on to learn how we achieved the same serendipity scores, but at a meaningfully lower CPU per request, resulting in a reduced cluster footprint.
Problem: The Hotspot in Ranker
At a high level, serendipity scoring works like this: A candidate title and each item in a member’s viewing history are represented as embeddings in a vector space. For each candidate, we compute its similarity against the history embeddings, find the maximum similarity, and convert that into a “novelty” score. That score becomes an input feature to the downstream recommendation logic.
The original implementation was straightforward but expensive. For each candidate we fetch its embedding, loop over the history to compute cosine similarity one pair at a time and track the maximum similarity score. Although it is easy to reason about, at Ranker’s scale, this results in significant sequential work, repeated embedding lookups, scattered memory access, and poor cache locality. Profiling confirmed this.

A flamegraph made it clear: One of the top hotspots in the service was Java dot products inside the serendipity encoder. Algorithmically, the hotspot was a nested loop structure of M candidates × N history items where each pair generates its own cosine similarity i.e. O(M×N) separate dot product operations.
Solution
The Original Implementation: Single video cosine loop
In simplified form the code looked like this:
for (Video candidate : candidates) {
Vector c = embedding(candidate); // D-dimensional
double maxSim = -1.0;
for (Video h : history) {
Vector v = embedding(h); // D-dimensional
double sim = cosine(c, v); // dot(c, v) / (||c|| * ||v||)
maxSim = Math.max(maxSim, sim);
}
double serendipity = 1.0 - maxSim;
emitFeature(candidate, serendipity);
}
The nested for loop with O(M×N) separate dot products brought upon its own overheads. One interesting detail we learned by instrumenting traffic shapes: most requests (about 98%) were single-video, but the remaining 2% were large batch requests. Because those batches were so large, the total volume of videos processed ended up being roughly 50:50 between single and batch jobs. This made batching worth pursuing even if it didn’t help the median request.
Step 1 : Batching, from Nested Loops to Matrix Multiply
The first idea was to stop thinking in terms of “many small dot products” and instead treat the work as a matrix operation. i.e. For batch candidates, implement a data layout to parallelize the math in a single operation i.e. matrix multiply. If D is the embedding dimension:
- Pack all candidate embeddings into a matrix A of shape M x D
- Pack all history embeddings into a matrix B of shape N x D
- Normalize all rows to unit length.
- Compute: cosine similarities as
[ C = A x B^T ]; where C is an M x N matrix of cosine similarities.
In pseudo‑code:
// Build matrices
double[][] A = new double[M][D]; // candidates
double[][] B = new double[N][D]; // history
for (int i = 0; i < M; i++) {
A[i] = embedding(candidates[i]).toArray();
}
for (int j = 0; j < N; j++) {
B[j] = embedding(history[j]).toArray();
}
// Normalize rows to unit vectors
normalizeRows(A);
normalizeRows(B);
// Compute C = A * B^T
double[][] C = matmul(A, B);
C[i][j] = cosine(candidates[i], history[j])
// Derive serendipity
for (int i = 0; i < M; i++) {
double maxSim = max(C[i][0..N-1]);
double serendipity = 1.0 - maxSim;
emitFeature(candidates[i], serendipity);
}
This turns M×N separate dot products into a single matrix multiply, which is exactly what CPUs and optimized kernels are built for. We integrated this into the existing framework by supporting both, encode()for single videos and batchEncode() for batches, while maintaining backward compatibility. At this point it seemed like we were “done”, but we weren't.
Step 2: When Batching Isn’t Enough
Once we had a batched implementation, we ran canaries and saw something surprising: about a 5% performance regression. The algorithm wasn’t the issue — turning M×N separate dot products into a matrix multiplication is mathematically sound. The problem was the overhead we introduced in the first implementation.
- Our initial version built double[][] matrices for candidates, history, and results on every batch. Those large, short-lived allocations created GC pressure, and the double[][] layout itself is non-contiguous in memory, which meant extra pointer chasing and worse cache behavior.
- On top of that, the first-cut Java matrix multiply was a straightforward scalar implementation, so it couldn’t take advantage of SIMD. In other words, we paid the cost of batching without getting the compute efficiency we were aiming for.
The lesson was immediate: algorithmic improvements don’t matter if the implementation details—memory layout, allocation strategy, and the compute kernel—work against you. That set up the next step for making the data layout cache-friendly and eliminating per-batch allocations before revisiting the matrix multiply kernel.
Step 3: Flat Buffers & ThreadLocal Reuse
We reworked the data layout to be cache-friendly and allocation-light. Instead of double[m][n], we moved to flat double[] buffers in row-major order. That gave us contiguous memory and predictable access patterns. Then we introduced a ThreadLocal<BufferHolder> that owns reusable buffers for candidates, history, and any other scratch space. Buffers grow as needed but never shrink, which avoids per-request allocation while keeping each thread isolated (no contention). A simplified sketch:
class BufferHolder {
double[] candidatesFlat = new double[0];
double[] historyFlat = new double[0];
double[] getCandidatesFlat(int required) {
if (candidatesFlat.length < required) {
candidatesFlat = new double[required];
}
return candidatesFlat;
}
double[] getHistoryFlat(int required) {
if (historyFlat.length < required) {
historyFlat = new double[required];
}
return historyFlat;
}
}
private static final ThreadLocal<BufferHolder> threadBuffers =
ThreadLocal.withInitial(BufferHolder::new);
This change alone made the batched path far more predictable: fewer allocations, less GC pressure, and better cache locality.
Now the remaining question was the one we originally thought we were answering: what’s the best way to do the matrix multiply?
Step 4: BLAS: Great in Tests, Not in Production
The obvious next step was BLAS (Basic Linear Algebra Subprograms). In isolation, microbenchmarks looked promising. But once integrated into the real batch scoring path, the gains didn’t materialize. A few things were working against us:
- The default netlib-java path was using F2J (Fortran-to-Java) BLAS rather than a truly native implementation.
- Even with native BLAS, we paid overhead for setup and JNI transitions.
- Java’s row-major layout doesn’t match the column-major expectations of many BLAS routines, which can introduce conversion and temporary buffers.
- Those extra allocations and copies mattered in the full pipeline, especially alongside TensorFlow embedding work.
BLAS was still a useful experiment — it clarified where time was being spent, but it wasn’t the drop-in win we wanted. What we needed was something that stayed pure Java, fit our flat-buffer architecture, and could still exploit SIMD.
Step 5: JDK Vector API to the rescue
A Short Note on the JDK Vector API: The JDK Vector API is an incubating feature that provides a portable way to express data-parallel operations in Java — think “SIMD without intrinsics”. You write in terms of vectors and lanes, and the JIT maps those operations to the best SIMD instructions available on the host CPU (SSE/AVX2/AVX-512), with a scalar fallback when needed. More crucially for us, it’s pure Java: no native dependencies, no JNI transitions, and a development model that looks like normal Java code rather than platform-specific assembly or intrinsics.
This was a particularly good match for our workload because we had already moved embeddings into flat, contiguous double[] buffers, and the hot loop was dominated by large numbers of dot products. The final step was to replace BLAS with a pure-Java SIMD implementation using the JDK Vector API. By this point we already had the right shape for high performance — batching, flat buffers, and ThreadLocal reuse. So the remaining work was to swap out the compute kernel without introducing JNI overhead or platform-specific code. We did that behind a small factory. At class load time, MatMulFactory selects the best available implementation:
- If jdk.incubator.vector is available, use a Vector API implementation.
- Otherwise, fall back to a scalar implementation with a highly optimized loop-unrolled dot product (implemented by my colleague Patrick Strawderman, inspired by patterns used in Lucene)
In the Vector API implementation, the inner loop computes a dot product by accumulating a * b into a vector accumulator using fma() (fused multiply-add). DoubleVector.SPECIES_PREFERRED lets the runtime pick an appropriate lane width for the machine. Here’s a simplified sketch of the inner loop:
// Vector API path (simplified)
for (int i = 0; i < M; i++) {
for (int j = 0; j < N; j++) {
DoubleVector acc = DoubleVector.zero(SPECIES);
int k = 0;
// SPECIES.length() (e.g. often 4 doubles on AVX2 and 8 doubles on AVX-512).
for (; k + SPECIES.length() <= D; k += SPECIES.length()) {
DoubleVector a = DoubleVector.fromArray(SPECIES, candidatesFlat, i*D + k);
DoubleVector b = DoubleVector.fromArray(SPECIES, historyFlat, j*D + k);
acc = a.fma(b, acc); // fused multiply-add
}
double dot = acc.reduceLanes(VectorOperators.ADD);
// handle tail k..D-1
similaritiesFlat[i*N + j] = dot;
}
}
Figure below shows how the Vector API utilizes SIMD hardware to process multiple doubles per instruction (e.g., 4 lanes on AVX2 and 8 lanes on AVX‑512). What used to be many scalar multiply-adds becomes a smaller number of vector fma() operations plus a reduction—same algorithm, much better use of the CPU’s vector units.

Fallbacks & Safety: When the Vector API Isn’t Available
Because the Vector API is still incubating, it requires a runtime flag: –add-modules=jdk.incubator.vector We didn’t want correctness or availability to depend on that flag. So we designed the fallback behavior explicitly: At startup, we detect Vector API support and use the SIMD batched matmul when available; otherwise we fall back to an optimized scalar path, with single-video requests continuing to use the per-item implementation.
That gives us a clean operational story: services can opt in to the Vector API for maximum performance, but the system remains safe and predictable without it.
Results in Production:
With the full design in place with batching, flat buffers, ThreadLocal reuse, and the Vector API, we ran canaries that run production traffic. We observed a ~7% drop in CPU utilization and ~12% drop in average latency. To normalize across any small throughput differences, we also tracked CPU/RPS (CPU consumed per request-per-second). That metric improved by roughly 10%, meaning we could handle the same traffic with about 10% less CPU, and we saw similar numbers hold after full production rollout.

At the function operator level, we saw the CPU drop from the initial 7.5% to a merely ~1% with the optimization in place. At the assembly level, the shift was clear: from loop-unrolled scalar dot products to a vectorized matrix multiply on AVX-512 hardware.

Closing Thoughts
This optimization ended up being less about finding the “fastest library” and more about getting the fundamentals right: choosing the right computation shape, keeping data layout cache-friendly, and avoiding overheads that can erase theoretical wins. Once those pieces were in place, the JDK Vector API was a great fit, as it let us express SIMD-style math in pure Java, without JNI, while still keeping a safe fallback path. Another bonus was the low developer overhead: compared to lower-level approaches, the Vector API let us replace a much larger, more complex implementation with a relatively small amount of readable Java code, which made it easier to review, maintain, and iterate on.
Have you tried the Vector API in a real service yet? I’d love to hear what workloads it helped (or didn’t), and what you learned about benchmarking and rollout in production.
Special thanks to Jason Koch, Patrick Strawderman, Daniel Huang, Fan Yang, and the Performance Engineering team at Netflix
Optimizing Recommendation Systems with JDK’s Vector API was originally published in Netflix TechBlog on Medium, where people are continuing the conversation by highlighting and responding to this story.
Trump’s Next Target
Post Syndicated from The Atlantic original https://www.youtube.com/shorts/xzPhuM3sO44
Radio Atlantic: Why the Iranian Opposition Might Fail to Seize on Khameini’s Fall
Post Syndicated from The Atlantic original https://www.youtube.com/shorts/rdJIdxBmUw0
[$] The ongoing quest for atomic buffered writes
Post Syndicated from corbet original https://lwn.net/Articles/1060063/
There are many applications that need to be able to write multi-block
chunks of data to disk with the assurance that the operation will either
complete successfully or fail altogether — that the write will not be
partially completed (or “torn”), in other words. For years, kernel
developers have worked on providing atomic writes as a way of satisfying
that need; see, for example, sessions from the Linux Storage, Filesystem,
Memory Management, and BPF (LSFMM+BPF) Summit from 2023, 2024,
and 2025 (twice). While atomic direct I/O is now supported by some filesystems, atomic
buffered I/O still is not. Filling
that gap seems certain to be a 2026 LSFMM+BPF topic but, thanks to an early
discussion, the shape of a solution might already be coming into focus.
Kash Patel #lastweektonight
Post Syndicated from LastWeekTonight original https://www.youtube.com/shorts/CVHWrUQLoIs
Høiland-Jørgensen: The inner workings of TCP zero-copy
Post Syndicated from corbet original https://lwn.net/Articles/1060953/
Toke Høiland-Jørgensen has posted an
overview of how zero-copy networking works in the Linux kernel.
Since the memory is being copied directly from userspace to the
network device, the userspace application has to keep it around
unmodified, until it has finished sending. The sendmsg()
syscall itself is asynchronous, and will return without waiting for
this. Instead, once the memory buffers are no longer needed by the
stack, the kernel will return a notification to userspace that the
buffers can be reused.
Standardize Amazon Redshift operations using Templates
Post Syndicated from Nidhi Nayak original https://aws.amazon.com/blogs/big-data/standardize-amazon-redshift-operations-using-templates/
Over the past year, Amazon Redshift has introduced capabilities that simplify operations and enhance productivity. Building on this momentum, we’re addressing another common operational challenge that data engineers face daily: managing repetitive data loading operations with similar parameters across multiple data sources. This intermediate-level post introduces AWS Redshift Templates, a new feature that you can use to create reusable command patterns for the COPY command, reducing redundancy and improving consistency across your data operations.
The challenge: Managing repetitive data operations at scale
Meet AnyCompany, a fictional data aggregation company that processes customer transaction data from over 50 retail clients. Each client sends daily delimited text files with similar structures:
While the data format is largely consistent across clients (pipe-delimited files with headers, UTF-8 encoding), the sheer volume of COPY commands required to load this data has become a development and maintenance overhead.
Their data engineering team faces several pain points:
- Repetitive parameter specification: Each COPY command requires specifying the same parameters for delimiter, encoding, error handling, and compression settings
- Inconsistency risks: With multiple team members writing COPY commands, slight variations in parameters lead to data ingestion failures
- Maintenance overhead: When they need to adjust error thresholds or encoding settings, they must update hundreds of individual COPY commands across their extract, transform, and load (ETL) pipelines
- Onboarding complexity: New team members struggle to remember all the required parameters and their optimal values
Additionally, a few clients send data in slightly different formats. Some use comma delimiters instead of pipes or have different header configurations. The team needs flexibility to handle these exceptions without completely rewriting their data loading logic.
Introducing Redshift Templates
You can address these challenges by using Redshift Templates to store commonly used parameters for COPY commands as reusable database objects. Think of templates as blueprints for your data operations where you can define your parameters once, then reference them across multiple COPY commands.
Template management best practices
Before exploring implementation scenarios, let’s establish best practices for template management to ensure your templates remain maintainable and secure.
- Use descriptive names that indicate purpose:
- Implement least privilege access:
- Query the system view to track template usage:
- Document each template, including:
- Purpose and use cases
- Parameter explanations
- Ownership and contact information
- Change history
Solution overview
Let’s explore how AnyCompany uses Redshift Templates to streamline their data loading operations.
Scenario 1: Standardizing client data ingestion
AnyCompany receives transaction files from multiple retail clients with consistent formatting. They create a template that encapsulates their standard loading parameters:
This template defines their standard approach:
DELIMITER '|'specifies pipe-delimited filesIGNOREHEADER 1skips the header rowENCODING UTF8facilitates proper character encodingMAXERROR 100allows up to 100 errors before failing, providing resilience for minor data quality issuesCOMPUPDATE OFFhelps prevent automatic compression analysis during loading for faster performanceSTATUPDATE ONkeeps table statistics current for query optimizationACCEPTINVCHARSreplaces invalid UTF-8 characters rather than failingTRUNCATECOLUMNStruncates data that exceeds column width rather than failing
Now, loading data from a standard client becomes remarkably straightforward:
Notice how clean and maintainable these commands are. Each COPY statement specifies only:
- The target table
- The Amazon Simple Storage Service (Amazon S3) source location
- The default AWS Identity and Access Management (IAM) role for authentication
- The template reference
The complex formatting and error handling parameters are neatly encapsulated in the template, facilitating consistency across the data loads.
Scenario 2: Handling client-specific variations with parameter overrides
AnyCompany has two clients (Client D, and E) who send comma-delimited files instead of pipe-delimited files. Rather than creating an entirely separate template, they can override specific parameters while still using the template’s other settings:
This demonstrates the Redshift Templates parameter hierarchy:
- Command-specific parameters (highest priority): Parameters explicitly specified in your COPY command take precedence
- Template parameters (medium priority): Parameters defined in the template are used when not overridden
- Amazon Redshift default parameters (lowest priority): Default values apply when neither command nor template specifies a value
This three-tier approach provides the perfect balance between standardization and flexibility. You maintain consistency where it matters while retaining the ability to handle exceptions gracefully.
Scenario 3: Simplified template maintenance
Six months after implementing templates, AnyCompany’s data quality team recommends increasing the error threshold from 100 to 500 to better handle occasional data quality issues from upstream systems. With templates, this change is trivial:
To remove a template when it’s no longer needed:
Scenario 4: Environment-specific templates for development and production
AnyCompany maintains separate templates for development and production environments, with different error tolerance levels:
This approach helps ensure that data quality issues are caught early in production while allowing flexibility during development and testing.
Key benefits
The key benefits of using templates include:
- Consistency and standardization: Templates help maintain consistency across different operations by making sure that the same set of parameters and configurations are used every time. This is particularly valuable in large organizations where multiple users work on the same data pipelines.
- Ease of use and timesaving: Instead of manually specifying the parameters for each command execution, users can reference a pre-defined template. This saves time and reduces the chances of errors caused by manual input.
- Flexibility with parameter overrides: While templates provide standardization, they don’t sacrifice flexibility. You can override a template parameter directly in your COPY command when handling exceptions or special cases.
- Simplified maintenance: When changes need to be made to parameters or configurations, updating the corresponding template propagates the changes across the instances where the template is used. This significantly reduces maintenance effort compared to manually updating each command individually.
- Collaboration and knowledge sharing: Templates serve as a knowledge base, capturing best practices and optimized configurations developed by experienced users. This facilitates knowledge sharing and onboarding of new team members, reducing the learning curve and facilitating consistent usage of proven configurations.
Additional use cases across industries
Templates can be used across industries.
Financial services: Standardizing regulatory data loads
A financial institution needs to load transaction data from multiple branches with consistent formatting requirements:
Healthcare: Loading patient data with strict standards
A healthcare analytics company standardizes their patient data ingestion across multiple hospital systems:
Retail: JSON data loading standardization
A retail company processes JSON-formatted product catalogs from various suppliers:
Conclusion
In this post, we introduced Redshift Templates and showed examples of how they can standardize and simplify your data loading operations across different scenarios. By encapsulating common COPY command parameters into reusable database objects, templates help remove repetitive parameter specifications, facilitate consistency across teams, and centralize maintenance. When requirements evolve, a single template update propagates quickly across the operations, reducing operational overhead while maintaining flexibility to override parameters for use cases.
Start using Redshift Templates to transform your data ingestion workflows. Create your first template for your most common data loading pattern, then gradually expand coverage across your pipelines. Your team will immediately benefit from cleaner code, faster onboarding, and simplified maintenance. To learn more about Redshift Templates and explore additional configuration options, see the Amazon Redshift documentation.
AWS Weekly Roundup: OpenAI partnership, AWS Elemental Inference, Strands Labs, and more (March 2, 2026)
Post Syndicated from Micah Walter original https://aws.amazon.com/blogs/aws/aws-weekly-roundup-openai-partnership-aws-elemental-inference-strands-labs-and-more-march-2-2026/
This past week, I’ve been deep in the trenches helping customers transform their businesses through AI-DLC (AI-Driven Lifecycle) workshops. Throughout 2026, I’ve had the privilege of facilitating these sessions for numerous customers, guiding them through a structured framework that helps organizations identify, prioritize, and implement AI use cases that deliver measurable business value.

AI-DLC is a methodology that takes companies from AI experimentation to production-ready solutions by aligning technical capabilities with business outcomes. If you’re interested in learning more, check out this blog post that dives deeper into the framework, or watch as Riya Dani teaches me all about AI-DLC on our recent GenAI Developer Hour livestream!
Now, let’s get into this week’s AWS news…
OpenAI and Amazon announced a multi-year strategic partnership to accelerate AI innovation for enterprises, startups, and end consumers around the world. Amazon will invest $50 billion in OpenAI, starting with an initial $15 billion investment and followed by another $35 billion in the coming months when certain conditions are met. AWS and OpenAI are co-creating a Stateful Runtime Environment powered by OpenAI models, available through Amazon Bedrock, which allows developers to keep context, remember prior work, work across software tools and data sources, and access compute.
AWS will serve as the exclusive third-party cloud distribution provider for OpenAI Frontier, enabling organizations to build, deploy, and manage teams of AI agents. OpenAI and AWS are expanding their existing $38 billion multi-year agreement by $100 billion over 8 years, with OpenAI committing to consume approximately 2 gigawatts of Trainium capacity, spanning both Trainium3 and next-generation Trainium4 chips.
Last week’s launches
Here are some launches and updates from this past week that caught my attention:
- AWS Security Hub Extended offers full-stack enterprise security with curated partner solutions — AWS launched Security Hub Extended, a plan that simplifies procurement, deployment, and integration of full-stack enterprise security solutions including 7AI, Britive, CrowdStrike, Cyera, Island, Noma, Okta, Oligo, Opti, Proofpoint, SailPoint, Splunk, Upwind, and Zscaler. With AWS as the seller of record, customers benefit from pre-negotiated pay-as-you-go pricing, a single bill, no long-term commitments, unified security operations within Security Hub, and unified Level 1 support for AWS Enterprise Support customers.
- Transform live video for mobile audiences with AWS Elemental Inference — AWS launched Elemental Inference, a fully managed AI service that automatically transforms live and on-demand video for mobile and social platforms in real time. The service uses AI-powered cropping to create vertical formats optimized for TikTok, Instagram Reels, and YouTube Shorts, and automatically extracts highlight clips with 6-10 second latency. Beta testing showed large media companies achieved 34% or more savings on AI-powered live video workflows. Deep dive into the Fox Sports implementation.
- MediaConvert introduces new video probe API — AWS Elemental MediaConvert introduced a free Probe API for quick metadata analysis of media files, reading header metadata to return codec specifications, pixel formats, and color space details without processing video content.
- OpenAI-compatible Projects API in Amazon Bedrock — Projects API provides application-level isolation for your generative AI workloads using OpenAI-compatible APIs in the Mantle inference engine in Amazon Bedrock. You can organize and manage your AI applications with improved access control, cost tracking, and observability across your organization.
- Amazon Location Service introduces LLM Context — Amazon Location launched curated AI Agent context as a Kiro power, Claude Code plugin, and agent skill in the open Agent Skills format, improving code accuracy and accelerating feature implementation for location-based capabilities.
- Amazon EKS Node Monitoring Agent is now open source — The Amazon EKS Node Monitoring Agent is now open source on GitHub, allowing visibility into implementation, customization, and community contributions.
- AWS AppConfig integrates with New Relic — AWS AppConfig launched integration with New Relic Workflow Automation for automated, intelligent rollbacks during feature flag deployments, reducing detection-to-remediation time from minutes to seconds.
For a full list of AWS announcements, be sure to keep an eye on the What’s New with AWS page.
Other AWS news
Here are some additional posts and resources that you might find interesting:
- Introducing Strands Labs — We created Strands Labs as a separate Git organization to support experimental agentic AI projects and push the frontier of agentic development. At launch, we’re making Strands Labs available with three projects. The first is Robots, the second is Robots Sim and the third is AI Functions.
- 6,000 AWS accounts, three people, one platform: Lessons learned — Architecture blog post on managing massive multi-account environments. Learn how ProGlove implemented a large-scale account-per-tenant model on AWS and how that model shifts complexity from service code to platform operations.
- Building intelligent event agents using Amazon Bedrock AgentCore and Amazon Bedrock Knowledge Bases — Practical guide to building event-driven agents. Check out how you can use Amazon Bedrock AgentCore components to rapidly productionize an event assistant—taking it from prototype to enterprise-ready deployment at scale.
From AWS community
Here are my personal favorite posts from AWS community:
- How to Run a Kiro AI Coding Workshop That Actually Works — Running a Kiro workshop at your company or user group? Here is the full step-by-step facilitator guide, resources, and references.
- RAG vs GraphRAG: When Agents Hallucinate Answers — This demo builds a travel booking agent with Strands Agents and compares RAG (FAISS) vs GraphRAG (Neo4j) to measure which approach reduces hallucinations when answering queries
- New output formats in AWS CLI v2 — You can now use two new features for the AWS Command Line Interface (AWS CLI) v2: structured error output and the “off” output format.
Upcoming AWS events
Check your calendar and sign up for upcoming AWS events:
- AWS at NVIDIA GTC 2026 — Join us at our AWS sessions, booths, demos, ancillary events in NVIDIA GTC 2026 on March 16 – 19, 2026 in San Jose. You can receive 20% off event passes through AWS and request a 1:1 meeting at GTC.
- AWS Summits — Join AWS Summits in 2026, free in-person events where you can explore emerging cloud and AI technologies, learn best practices, and network with industry peers and experts. Upcoming Summits include Paris (April 1), London (April 22), and Bengaluru (April 23–24).
- AWS Community Days — Community-led conferences where content is planned, sourced, and delivered by community leaders. Upcoming events include JAWS Days in Tokyo (March 7), Chennai (March 7), Slovakia (March 11), and Pune (March 21).
Browse here for upcoming AWS led in-person and virtual events, startup events, and developer-focused events.
That’s all for this week. Check back next Monday for another Weekly Roundup!
Texinfo 7.3 released
Post Syndicated from jzb original https://lwn.net/Articles/1060950/
Version 7.3 of Texinfo, the GNU documentation-formatting system, has been released.
It contains a number of new features, performance improvements, and enhancements.
How Twilio secured their multi-engine query platform with AWS Lake Formation
Post Syndicated from Aakash Pradeep, Venkatram Bondugula original https://aws.amazon.com/blogs/big-data/how-twilio-secured-their-multi-engine-query-platform-with-aws-lake-formation/
This is a guest post by Aakash Pradeep, Principal Software Engineer, and Venkatram Bondugula, Software Engineer at Twilio, in partnership with AWS.
Twilio is a cloud communications platform that provides programmable APIs and tools for developers to easily integrate voice, messaging, email, video, and other communication features into their applications and customer engagement workflows.
In this blog series we discuss how we built a multi-engine query platform at Twilio. The first part introduces the use case that led us to build a new platform and why we selected Amazon Athena alongside our open-source Presto implementation. This second part discusses how Twilio’s query infrastructure platform integrates with AWS Lake Formation to provide fine-grained access control to all their data.
At Twilio, we faced critical challenges in managing our multi-engine query platform across a complex data mesh architecture spanning multiple AWS accounts and Lines of Business. We needed a unified permissions model that could work consistently across different query engines like OSS Presto and Amazon Athena, eliminating the fragmented authentication experiences in our infrastructure. The growing demand for secure cross-account data sharing required moving beyond manual, multi-step provisioning processes that depended heavily on human intervention. Additionally, Twilio’s compliance and data stewardship requirements demanded fine-grained access controls at row, column, and cell levels, necessitating a scalable and flexible approach to permission management. By adopting the AWS Glue Data Catalog as our managed metastore and AWS Lake Formation for governance, we implemented Tag-Based Access Control (LF-TBAC) to simplify access management, enabled data sharing through automated workflows, and established a centralized governance framework that provided uniform permissions management across all AWS services.
Transitioning to a managed metastore and governance solutions
We discussed in part 1, how we were looking to move to managed services to alleviate us of the burden of managing the underlying infrastructure of a query platform. Along with our decision to adopt Amazon Athena, we also began to evaluate the adoption of Amazon EMR Serverless for our Spark workloads, which made us aware of the fact that we needed to migrate to a managed solution for our Apache Hive metastore.
We selected the AWS Glue Data Catalog as our managed metastore repository to support our enterprise-wide data mesh architecture. For managing permissions to the Data Catalog assets, we chose AWS Lake Formation, a service that enables data governance and security at scale using familiar database-like permissions. Lake Formation provides a unified permissions model as well as support for enabling data mesh architecture that we were seeking.
Lake Formation’s support for row, column, and cell-level access controls provides the fine-grained access control (FGAC) capabilities required by our compliance and data stewardship policies. Additionally, Lake Formation’s tag-based access control (LF-TBAC) feature allows us to define FGAC permissions based on tags attached to the Data Catalog resources, enabling flexible and scalable permission management.
Integrating Odin with AWS Lake Formation
Odin, our Presto-based gateway, serves as a central hub for query processing, managing authentication, routing, and the complete workflow throughout a query’s lifecycle. As the primary interface, Odin enables users to connect through JDBC or APIs from various BI tools, SQL IDEs, and other applications.
Beyond its core routing capabilities, Odin utilizes local caches implemented using Google’s Guava caching library to optimize performance across the platform. Guava delivers efficient in-memory caching for Java applications by storing data locally within the application instance, resulting in significantly faster retrieval times. Odin employs multiple Guava caching layers across various modules to ensure optimal response times for frequently accessed data and metadata.
Building on this performance foundation, Odin implements authentication and authorization layers to ensure secure and controlled access to data across multiple query engines. These security components work together to verify user identities and enforce data access policies, providing a unified security framework that abstracts away the complexities of individual engine implementations while maintaining strict governance standards.
The authentication layer
Different query engines like OSS Presto and Amazon Athena each implement their own authentication mechanisms. To create a consistent user experience, Odin provides a unified authentication layer that shields users from these underlying differences. Currently, Odin’s pluggable authentication system supports LDAP integration, with plans to expand this capability to include Okta authentication using IAM Identity center in the future.
The authorization layer
For data consumers using AWS Analytics services such as AWS Glue, Amazon EMR, and Athena through an IAM federated role-based access, AWS Lake Formation provided critical authorization capabilities for data governance through their existing integrations. However, we needed to extend its capabilities to integrate with OSS Presto. Additionally, our users for the query infrastructure platform were not mapped to an IAM user so would need to build a custom authorization layer in Odin to verify permissions and integrate with Lake Formation. Our challenge was creating a consistent way to control data access across all our query engines.
When a user runs a query, Odin’s authorization layer checks three key pieces of information:
- Table details: which database and table the query is accessing
- User permissions: what data tags the user has access to
- Resource tags: what security tags are attached to the requested table
We store user permissions in Amazon DynamoDB, which allows us to quickly look up what each user can access. By matching the user’s tags with the table’s Lake Formation tags, we can determine if the query should be allowed. To keep things fast, we cache this information temporarily, allowing us to expedite authorization for recent requests.
How the authorization works:
- Initial check: First, we see if this user recently ran a similar successful query (within the last 5 minutes).
- Gather information: We collect the table details, user permissions, and security tags—first checking our cache, then fetching from AWS Glue Data Catalog and Lake Formation if needed.
- Match permissions: We compare the user’s access tags stored in a DynamoDB table against the table’s security tags in Lake Formation.
- Make decision: If the user’s permissions match what’s required for their query action (like SELECT or INSERT), access is granted.
This approach allows us to make use of Lake Formation tag-based access control while keeping our authorization logic separate from the individual query engines. By using smart caching and efficient lookups, we can verify permissions in just milliseconds.
Building a data mesh
At Twilio, we have multiple line of business (LoBs) each managing their own data platform infrastructure. The individual platforms are spread across multiple AWS accounts, and primarily store data on Amazon S3 in variety of open table formats, such as Apache Hudi, Apache Iceberg, and Delta Lake. Each platform independently supports analytics and machine learning use cases, however, there was a growing need for secure sharing of data across LoBs. Additionally, we needed to enable self-service discovery and provisioning of access to the data with a centralized governance framework.
Data consumers bring their own AWS accounts and choice of tools, which include not only AWS services such as Amazon Athena, AWS Glue ETL jobs (Spark), and Amazon EMR, but also AWS partner solutions. To improve the process of access fulfillment, data auditability and lowering the operational overhead involved, we needed an automated framework in place that had minimal human intervention and oversight.
Implementing a data subscription workflow
Previously, consumers requiring access to specific data sets would need to go through multiple steps to secure access, which involved several dependencies and manual actions. To simplify this process and provide a self-service capability, we decided to build a custom integration solution between ServiceNow and AWS Lake Formation. At Twilio, ServiceNow is used extensively to automate workflows and build custom applications to connect disparate systems and improve operational efficiency.
We automated key parts of the data access process using Twilio’s standard tools: Git for version control, Terraform for infrastructure management, and custom scripts to execute the necessary AWS actions.
We automated three main use cases:
1. Sharing data between accounts
When one team needs to share data with another team or with our central governance account, the process starts with a Git pull request (PR). This triggers our custom Lake Formation automation tool, which:
- Connects to the source AWS account with admin permissions
- Sets up data sharing using the security tags (LF-Tags) specified in a YAML configuration file
- Completes the share using AWS Resource Access Manager (RAM)
- Creates resource links in the target account so the data appears in their catalog
- Updates ServiceNow with the newly shared database and table information
2. Granting permissions to user roles
When users request access to data, our automation tool grants tag-based permissions directly to their IAM roles in Lake Formation. This happens after approval of either a Git PR or ServiceNow ticket.
3. Granting access to individual users
For individual user access requests:
- Users submit a request in ServiceNow for specific tables
- After approval, ServiceNow calls our internal API that checks relevant Lake Formation tags
- The request is validated and sent to an Amazon Simple Queue Service (Amazon SQS) queue
- A consumer service processes the request, updates the user’s permissions in our DynamoDB table (which Odin uses for authorization checks), and includes retry logic for reliability
- Once complete, the service updates the ServiceNow ticket to notify the user
The overall subscription and authorization flow is as shown in the diagram below:

- Users submit a request in ServiceNow for access to a database, table, or LF-Tag
- The system retrieves the relevant LF-Tags from Lake Formation through our API integration
- Upon approval, the automation procedure adds the user to the User-To-Tag DynamoDB table, grants IAM role permissions in Lake Formation, and sets up cross-account sharing via RAM as needed
- Users submit SQL query to the Odin presto gateway
- Odin authorizes the user through LDAP
- Odin parsers the SQL query to identify the tables involved and the action being performed (SELECT, DDL, and more)
- Odin validates permissions using the User to LF-Tag mapping and Lake formation grants to authorize the SQL query based on granted permissions
- If authorized, Odin routes the query to Amazon Athena or Presto
Using standardized tools and processes to provide self-service capabilities to the users helped us scale the governance framework and support broader use cases. Important capabilities in Lake Formation, such as Tag-based access control (TBAC) and cross-account sharing of data, simplified developing automations and our overall approach to governance.
Lessons learned- Cache is king
“By adopting AWS Glue Data Catalog as our managed metastore and AWS Lake Formation for Tag-Based Access Control, we simplified access management and enabled data sharing by reducing auth overhead to just 6-10 milliseconds through caching and targeted scaling.”
As Odin began handling queries at scale, we encountered performance bottlenecks in our customized authorization process as we had to retrieve information from multiple services, particularly with complex queries spanning multiple tables. The authorization checks involved in the performance bottleneck frequently caused query timeouts which impacted overall system reliability. The root of the problem lay in our sequential authorization workflow: our system first had to parse each query to identify all tables requiring identity verification, then make separate API calls to the AWS Glue Data Catalog and Lake Formation for each table’s permissions. It became clear that we needed to optimize this authentication process to reduce response times and improve the overall query experience.
We also recognized there were different caching needs between our POST operations and GET/DELETE HTTP calls, so we decided to separate them into two different Application Load Balancer (ALB) target groups. For POST requests, which required Lake Formation authentication, we found that concentrating traffic through just 2-3 target instances distributed across multiple Availability Zones (AZ) was more efficient. This approach allowed authentication information to be effectively cached locally on these dedicated instances, dramatically reducing the volume of API calls to the Lake Formation service.
GET and DELETE requests follow a more simplified workflow. Since users have already completed initial authorization, there is no need to continue to perform authorization checks. Although they follow a simpler workflow, these requests have much higher volume with requests numbering into the 10s of millions per hour. Due to this scale, we opted to implement horizontal scaling to scale the target ALB to 10 Amazon EC2 instances to fetch the query history from the DynamoDB table. These EC2 instances make use of local LRU caching with a 5-minute expiration policy for authentication data.
By implementing authentication caching and adopting specialized approaches for different HTTP request types with targeted scaling groups, we successfully reduced Odin’s overall overhead to a maximum of 6-10 milliseconds for both authentication and authorization.
Conclusion and what’s next
In this post, we explored how we enhanced Odin, our unified multi-engine query platform, with authentication and authorization capabilities using AWS Lake Formation and a custom authorization workflow. By using AWS services including Lake Formation, AWS Glue Data Catalog, and Amazon DynamoDB alongside Twilio’s existing infrastructure, we created a scalable self-service governance framework that streamlines user access management, simplifies auditing, and enables seamless data sharing across our complex cloud environment. With this workflow automation, we eliminated operational overhead while building a secure, robust platform that serves as the foundation for Twilio’s data mesh architecture.
Going forward, we are focusing on strengthening our authentication and authorization framework by enabling trusted federation with an identity provider(IdP) through AWS IAM Identity Center, which integrates directly with Lake Formation. Using Trusted Identity Propagation capabilities supported by IAM IDC will allow us to establish a consistent governance flow based on a user identity and will allow us to unlock the full capabilities of AWS Lake Formation such as fine-grained access control with data filters.
To learn more and get started with building with AWS Lake Formation, see Getting started with Lake Formation, and How to build a data mesh architecture at scale using AWS Lake Formation tag-based access control.
About the authors
After Khamenei, What Now?
Post Syndicated from The Atlantic original https://www.youtube.com/watch?v=IbUuBh31c-Q
Welcome to our live show!
Post Syndicated from The Atlantic original https://www.youtube.com/watch?v=40RnL6t6r-Y