Post Syndicated from The Atlantic original https://www.youtube.com/shorts/qczqUFBa0Y8
Т.Е. от Е.Т. – епизод 33
Post Syndicated from Тоест original https://www.toest.bg/t-e-ot-e-t-epizod-33/

Знаете ли го този виц, дето един човек отишъл с кучето си за риба?
Отишъл човекът, отворил си такъмите, метнал въдицата и по едно време клъвнало. Засякъл хубаво с въдицата и издърпал на брега една крава, която най-любезно го попитала колко е часът. Той, разбира се, се опулил, нищо не успял да ѝ отговори и кравата си заминала.
След малко пак кълве, пак дърпа човекът, пак крава, пак го пита за часа. И след това пак и пак, и пак… Накрая човекът в ступор погледнал въпросително кучето си, а то му казало: „Копеле, и аз не знам к’во става.“ И си тръгнало.
И ние така – не знаем какво става. Изобщо. Където и да е. По какъвто и да е повод. Затова си имаме Е.Т. да ни разяснява.
Следете видеорубриката на Елена Телбис за „Тоест“ и във Facebook, Instagram и TikTok.
Скривалище
Post Syndicated from Тоест original https://www.toest.bg/skrivalishte/

Пространството
което
иска да те изяде
времето
което
те изпива
прегръщаш някого
потъваш в него
с плахата надежда
че ако сте двама
ще ви се размине
и за момент
се случва чудото –
пространството и времето
изчезват
отпуснати
в усмивката на Бог
приличате на думи
слети в рая
които ще сънуват
вечността
докато някой
проговори
Петър Чухов
Петър Чухов е автор на 17 стихосбирки, 3 книги с проза и книга за деца. Творбите му са преведени на 25 езика и публикувани в повече от 30 държави. Носител на различни награди, сред които Наградата на музея „Башо“ в Япония. Пише музика и текстове, свири в рок групи. Преподава творческо писане на поезия; работи в Столичната библиотека; ръководител е на Литературния клуб на библиотеката.
Тоест разговаряме – епизод 4
Post Syndicated from Владислав Севов original https://www.toest.bg/toest-razgovaryame-epizod-4/

В четвъртия епизод на „Тоест разговаряме“ посрещнахме Йовко Ламбрев – ИТ специалист и съосновател на нашата медия. Започнахме разговора от зараждането на идеята за „Тоест“ и причините да изберем този модел на финансиране от читателски дарения, а не от реклами.
Обсъдихме как изкуственият интелект вече променя обществото и пазара на труда, но най-големият риск идва от скоростта – хората, институциите и законите изостават. Разгледахме ИИ като инструмент, който може и да е в помощ на хората, но и да се използва за злоупотреби. Говорихме и за удара на ИИ върху младите кадри в компаниите – и как, ако не се обучават младши специалисти, утре няма да има кой да заеме мястото на старшите.
Коментирахме т.нар. chatcontrol и защо премахването на криптирането на комуникацията би засегнало най-вече обикновените хора, докато престъпниците ще заобиколят системата. Говорихме и за монопола на платформите върху вниманието и данните, за фалшивите новини и липсата на филтър в социалните мрежи.
Отбелязахме, че статии на Йовко отпреди 6–7 години в рубриката „Аз, киборгът“, като например „Що е то подкаст?“, „Електронна поща – правилната употреба“ и „Паролите са отживелица“, продължават да са сред най-четените – знак, че няма достатъчно текстове за повишаване на дигиталната грамотност. Насърчихме използването на платени онлайн услуги, в т.ч. и електронни пощи, и по-внимателното споделяне на лични данни (безплатното често се заплаща с данните, които предоставяме).
Завършихме с призив да поддържаме информационна хигиена и критично мислене – и да подкрепяме независимите медии.
Гледайте целия разговор в нашия YouTube канал:
Може да го чуете и като аудиозапис в SoundCloud:
На живо включихме и въпроси от публиката, но времето ни беше ограничено и не успяхме да обхванем всички. Помолих Йовко да отговори тук на още един въпрос на зрител:
Как ще изглежда спукването на ИИ балона? И как ще се отрази това на живота в България?
Спукването на който и да е балон нe е приятно събитие. Много инвеститори ще загубят пари. Със сигурност тежко ще пострадат много компании и особено тези, които са заложили всичко на ИИ. Компании като OpenAI, Anthropic и Nvidia ще са сред големите губещи. Гигантите ще оцелеят.
За Alphabet/Google например нямам никакви колебания. Или за Amazon/AWS. Компаниите, които предоставят инфраструктура, ще се възстановят по-бързо, защото нужда от такава винаги има. Не бих се тревожил много и за компаниите, фокусирани върху бизнес софтуера, като SAP.
По-сериозен крах ще повлече надолу цените на акциите на всички технологични компании, включително на най-големите. А това неминуемо ще доведе до свиване, спиране или замразяване на проекти и до съкращения, вероятно и в България.
В такава ситуация и клиентите на технологичния бранш забавят и изчакват с новите си проекти, докато отмине бурята и почувстват накъде ще задуха „вятърът след промяната“. А това влошава допълнително нещата. Ако технологичният бранш боледува, няма как да не пострада и световната икономика, особено на фона на днешния свят, който живее в непрестанна несигурност и в който всичко е с потенциала да „прелее чашата“.
Ако има глобален крах, той няма как да подмине България, но технологичният бранш тук е гъвкав, а в момента – и доста оптимизиран. Той така или иначе винаги следва пренесени отвън тенденции. У нас няма твърде рискови играчи и големи технологични лидери. Ще се пренастрои в някакъв момент, както винаги е правил. Поне по-виталната част от него. На живота в България всичко това би повлияло по-осезаемо, ако спукването предизвика световна финансова криза. Особено при ниската ефективност на труда тук.
Аз обаче по-скоро бих заложил на корекции и отрезвяване на инвеститорите и борсите по отношение на ИИ, а не на грандиозно спукване. Сигурно е едно – че дори и след огромен финансов катаклизъм отново ще има ИИ, най-вероятно концентриран в ръцете и възможностите на малцина.
Преди срещата ви помолихме да отговорите на кратката ни анкета. Ето и резултатите от нея:
Изкуственият интелект е изключителен помощник. Не съм допускал, че ще бъда свидетел на подобен инструмент в рамките на живота си.
Йовко Ламбрев е компютърен инженер с десетилетия опит в информационните технологии. Работил е дълги години за международни гиганти, като IBM, Siemens и SAP, а в момента е инфраструктурен архитект в българската ИТ компания „Новарто“. През 2003 г. основава OpenFest – и до днес най-голямата българска конференция, посветена на свободния софтуер и софтуера с отворен код. Съосновател е на две технологични компании и на Trakia.Tech – инкубатор и изследователски център, който подпомага технологични екипи и стартъпи чрез програми за иновации и дигитална трансформация. През 2018 г. заедно с Ан Фам, Лина Кривошиева и Владислав Севов създава онлайн медията „Тоест“. Оттогава е водещ на рубриката „Аз, киборгът“ в медията и член на настоятелството на Фондация „Тоест“.
Следващата среща на „Тоест разговаряме“ ще бъде със Светла Енчева, социоложка и правозащитничка, която е редовна авторка на „Тоест“ от почти самото начало. Разговорът ще се проведе на живо в YouTube Live на 13 декември, събота, от 16:00 ч.

В „Тоест разговаряме“ всеки месец ви срещаме с автори, които познавате добре от анализите или от рубриките им в „Тоест“, но този път ще ги видите и чуете в по-личен и непосредствен формат. Във видеоразговорите, предавани на живо, активно участие имате и вие, нашата публика – със своите въпроси, коментари и включване в тематичната анкета. Водещ на поредицата е Владислав Севов, дългогодишен телевизионен журналист и съосновател на „Тоест“.

„Тоест разговаряме“ е поредица, подкрепена от Институт „Отворено общество – София“ и съфинансирана от Европейския съюз в рамките на проекта Media Resilience. Изразените възгледи и мнения са само и изцяло на техните автори и не отразяват непременно възгледите и мненията на Европейския съюз, на Европейската изпълнителна агенция за образование и култура (EACEA) или на Институт „Отворено общество – София“ (ИООС). Нито Европейският съюз, нито EACEA, нито ИООС могат да бъдат държани отговорни за тях.
Arista Shows Scale-Across Network Switches at SC25
Post Syndicated from Cliff Robinson original https://www.servethehome.com/arista-shows-scale-across-network-switches-broadcom-at-sc25/
At SC25, Arista showed its Arista 7280R4 switches based on the Broadcom Jericho3 family for scale-across networking
The post Arista Shows Scale-Across Network Switches at SC25 appeared first on ServeTheHome.
Comic for 2025.11.28 – Tech Support
Post Syndicated from Explosm.net original https://explosm.net/comics/tech-support
New Cyanide and Happiness Comic
Bridge Clearance
Post Syndicated from xkcd.com original https://xkcd.com/3174/

Why Thanksgiving?
Post Syndicated from The History Guy: History Deserves to Be Remembered original https://www.youtube.com/shorts/OFzrS9r1BDw
The Smart Home Heartbeat You Didn’t Know You Needed ❤️📉
Post Syndicated from BeardedTinker original https://www.youtube.com/watch?v=C9gpu-00vNk
Научни новини: Динозаври, комети, ваксини и генни терапии
Post Syndicated from Михаил Ангелов original https://www.toest.bg/nauchni-novini-dinozavri-kometi-vaksini-genni-terapii/
Хищническа мистерия

Измежду могъщите същества, населявали планетата ни преди милиони години, може би най-пленяващ въображението е гигантският хищник тиранозавър (Tyrannosaurus rex). С впечатляващите си размери и запомнящи се роли в „Джурасик парк“ той се е превърнал в икона на динозавърското царство. В продължение на години палеонтолозите се опитват да разберат неговото поведение и развитие, като разполагат с по-малко от 40 сравнително пълни образци. Нови разкрития показват, че може би някои от представите ни за страшния звяр са грешни.
Историята започва през 40-те години на миналия век с откриването на череп в богатото на палеофлора и фауна находище Hell Creek. Тогава го определят като образец от нов вид горгозавър (Gorgosaurus), но по-късно започва активна дискусия за видовата му принадлежност, като с времето се оформят два лагера. Учените от единия смятат, че това е възрастен индивид от нов род – нанотиранус (Nanotyrannus), позовавайки се на морфологията на черепа и на това, че костите са сраснали. Опонентите им изказват мнението, че е по-вероятно находката да е от малък тиранозавър. Тъй като разполагат само с този единичен екземпляр, никоя от страните не може да даде надеждно доказателство.
Ситуацията остава патова до 2001 г., когато е открита нова находка – добре запазен образец, наречен Джейн. След няколко години подготовка Джейн става достъпна за изследване и предизвиква сериозно вълнение. Учените и от двете страни откриват доказателства за своята хипотеза. Да, зъбите са повече, челюстта има специфична дупка и някои кости са по-различни, но от друга страна, това може да се обясни с възрастови и индивидуални специфики. В крайна сметка Джейн успява да разубеди някои привърженици на идеята за отделен вид, накланяйки везните към мнението, че образците най-вероятно са млади тиранозаври.
Малко след като започват вълненията около Джейн, в същото находище е направено забележително откритие: почти цели скелети на два динозавъра – растителнояден трицератопс и млад тиранозавър, наречен Кървавата Мери (Bloody Mary), макар полът на динозавъра да е неясен. Привидно двамата са вплетени в битка, което дава името на фосила – „Воюващи динозаври“, но най-вероятно позата е случайна. Уви, за повече от десет години тази находка остава в частна колекция, скрита за учените, докато през 2020 г. скелетите са откупени и включени в музейна експозиция.

След като става достъпен за анализ, образецът предоставя много нова информация, като най-интригуващото е, че може да бъде определена възрастта на индивида. Поради специфики в растежа на костите в тях се образуват растежни пръстени, сходни с тези в дърветата. Така учените установяват, че динозавърът е бил поне на 14 години, когато е загинал, а в последните години растежът на костите е бил забавен. Това сочи, че Мери е почти възрастен индивид и е нямало да порасне много повече.
При тези новопредставени данни изглежда, че спорът за съществуването на Nanotyrannus като отделен род може да се смята за приключен. Учените дори дават предложение за два вида в този род – N. lancensis, чийто представител е Мери, и N. lethaeus, който е бил малко по-голям, представен от Джейн. Най-вероятно нанотиранусите са достигали малко над 2 метра и тегло около 700 кг (според кратка справка в интернет – колкото голяма крава или стар модел „Фолксваген“ костенурка) – около десет пъти по-малки от T. rex. Освен по размер те се различават и по стойка. Нанотиранусите имат много по-пропорционални крайници, което предполага, че са били по-пъргави от гигантските си сродници и са можели да тичат. Краката и „ръцете“ на Мери са с размери, сходни на тиранозавърските, въпреки че е по-дребна.
Освен че поставя вида в таксономичното дърво, публикацията повдига редица въпроси за представите ни за тиранозаврите до момента, част от които се основават на информация, получена всъщност от друг вид. Като вземем предвид колко богато е видовото разнообразие в наши дни, е много вероятно сходни грешки да са допуснати и при други видове. Така това изследване може да предизвика преразглеждане на много фосили.
Ваксини с двойно действие
Разработването на ваксини срещу ракови заболявания е едно от най-силно желаните постижения в медицинската наука. Макар и вече да има одобрени продукти, които предпазват от някои видове рак, те работят срещу вирусните причинители на заболяването. Към момента ваксините, които насочват имунната ни система към туморите още при възникването им, са основно обект на хипотези и медицински изследвания.
Но метаизследване на над 1000 пациенти представя интересно следствие от поставянето на ваксините срещу COVID-19. Както се оказва, освен че са с висока специфичност и ефективност за предпазване от вируса, те повишават общата активност на имунната система и могат да помогнат при терапията на някои туморни заболявания.
Зад това откритие стои по-стара разработка. В експерименти с мишки учените установяват, че за предизвикване на имунен отговор срещу тумори не е нужно да се целят в конкретен протеин в образуванието. Ваксините се базират на основната идея, че дават възможност на тялото да се запознае с патогена по безопасен начин, така че да подготви имунен отговор на него. Но освен това те предизвикват отделянето на сигнални протеини, наречени цитокини, включително интерферон. Оказва се, че той може да активира имунните клетки в туморните тъкани, които да обучат имунната система, така че тя да започне да ги атакува. Това обаче не е достатъчно, тъй като раковите клетки са разработили защита – отделят протеин, който дезактивира атакуващите ги Т-клетки.
За справяне с проблема учените прилагат двоен подход. Те разработват неспецифична иРНК ваксина за активиране на имунната система, и я комбинират с инхибиторни вещества, които се използват за терапия на някои туморни заболявания и потискат отделянето на защитния протеин. Така дават възможност на тялото да използва естествените си способности за справяне със злокачествените клетки.
Резултатите от тази експериментална разработка пораждат интересно предположение – щом не е нужна специфична ваксина, подобен ефект би трябвало да се наблюдава и при иРНК ваксините срещу COVID. Потвърждение на хипотезата идва от анализ на данни за продължителността на живота на над 1000 пациенти с рак на белия дроб и кожата. Тези, които са били ваксинирани в рамките на 100 дни от започване на имунотерапия с инхибитори, живеят почти два пъти по-дълго – 37 месеца срещу 21 месеца при неваксинираните. Подобно подобрение не се наблюдава при пациенти, получили ваксини срещу грип или пневмония, които не са базирани на иРНК.
Сходен ефект на повишаване на активността на имунната система е забелязан и при деца с атопичен дерматит. Поради връзката му с понижаването на имунитета той често е предвестник на други по-тежки заболявания. Също така децата, страдащи от него, са по-предразположени към инфекции, засягащи респираторната система. Метаизследване на почти 6000 пациенти под 17 години показва, че след ваксинация срещу COVID рискът от появата на ушни инфекции, пневмония, синузит и др. спада средно с 40%.
Данните подчертават важността на откритието на този нов клас ваксини. След като помогнаха за намаляване на жертвите от пандемията, изглежда, тепърва ще разгръщат потенциала си. Интересно е дали и как подобни разработки ще бъдат повлияни от политическия климат в САЩ, където някои щати подготвят законопроекти за забрана на иРНК ваксините. В комбинация с решението за орязване на бюджета за разработка на такива ваксини, което беше критикувано от редица международни и американски организации, има голяма вероятност следващите големи разработки в тази област да бъдат в портфолиото на европейски или азиатски компании.
Комета или извънземни?
Тази година се оказа изключително богата на комети, на които можем да се възхищаваме.
През януари C/2024 G3 (ATLAS) беше страхотна гледка в нощното небе на южното полукълбо. В началото на годината откриха и C/2025 A6 (Lemmon), която достигна най-близката си точка до Земята в края на октомври и все още е достатъчно ярка, за да може да се наблюдава с просто око.
C/2025 K1 (ATLAS) беше засечена през май и в момента е видима с по-силен бинокъл. Траекторията ѝ премина между Слънцето и Меркурий, което обикновено вещае неприятности за кометите, но когато се появи малко след завоя си покрай Слънцето в началото на октомври, изглеждаше, че е от изключенията, които успяват да останат цели. За съжаление, в средата на ноември C/2025 K1 (ATLAS) все пак се разчупи – със сигурност на две, а може би и на повече парчета.
В същия район, където може да се наблюдава K1, през септември се появи и C/2025 R2 (SWAN), която също е видима с по-силен бинокъл и премина през няколко фрагментации.

Но макар и вълнуващи, тези комети не заплениха вниманието на множество хора така, както го направи междузвездният пътник 3I/ATLAS. От откриването му през юли учените го следят с интерес, тъй като е едва третият засечен обект от междузвездното пространство, който пресича Слънчевата система. Наблюденията бяха трудни, понеже през немалка част от пътешествието си 3I/ATLAS беше закрит от Слънцето и нямаше как да се види пряко от Земята, затова го следяха космически обсерватории като „Хъбъл“ и „Джеймс Уеб“. Към него обаче бяха насочени и инструменти, които не са предвидени за подобни цели, като Mars Reconnaissance Orbiter, изучаващ повърхността на Марс.
Наред с публикациите за наблюденията на астрономите се появиха и множество конспиративни теории. Една от главните фигури зад тях е харвардският професор Ави Лоуб, който е известен с изказванията си, че е много вероятно да сме посетени от извънземни. В поредица от материали в Medium той посочи няколко „несъответствия“, които според него показват, че обектът всъщност е изкуствен. Например че орбитата му ще го преведе много близо до Земята, което е малко вероятно, ако се разчита на случайност. Или че опашката е нехарактерна за комета с такъв размер и това всъщност може да е следа от двигател. Бяха изказани и идеи, че промяната в орбитата му е невъзможна за естествен обект.
Ситуацията бе допълнително нажежена от НАСА, която дълго време не публикуваше снимки и информация за обекта. В началото това бе обяснено със спряното финансиране на правителствени агенции, но след възстановяването му имаше период на изчакване, който бе тълкуван от поддръжниците на хипотезата за изкуствения характер на обекта като потвърждение, че правителството се опитва да скрие нещо.
В крайна сметка Агенцията публикува множество нови кадри и всички участници в проведената пресконференция бяха категорични, че става въпрос за комета. Интересното е, че макар и по някои неща да прилича на кометите, които са постоянни обитатели на нашата Слънчева система, има и разлики – например 3I/ATLAS е много по-богат на никел. Учените спекулират, че най-вероятно скоро не е преминавал покрай друга звезда, поради което е възможно да е по-стар от Слънчевата система. Това, което знаем към момента, е, че идва от центъра на Млечния път и след като премине покрай нас, повече никога няма да се върне. Именно това подтиква астрономите да съберат възможно най-много информация, докато кометата е близо и можем да я наблюдаваме.
Изнесената информация не убеди Лоуб, който продължава да твърди, че данните са недостатъчни за изключване на хипотезата, че обектът е изкуствен. В средата на декември предстои най-близкото му преминаване покрай Земята (на безопасно разстояние от планетата ни). Тогава ще бъдат проведени множество наблюдения, които би трябвало да изяснят дали това всъщност не е преднамерено посещение.
Генни терапии за болестта на Хънтингтън
Болестта на Хънтингтън е изключително коварно невродегенеративно заболяване с летален изход, което има генетична основа. В началото симптомите се появяват по-скоро като психически смущения, но с времето, освен влошаване на когнитивните способности, при пациентите се наблюдава и загуба на възможността за координирани движения на тялото.
Заболяването е автозомно – засегнатият ген няма връзка с половите хромозоми.
Специфично е, че грешката не е „обикновена“ мутация, която води до бъг в даден ген, а натрупване на повтори от „букви“ в него. Така ако последователността CAG, кодираща аминокиселината глутамин, се намножи прекалено много в гена за протеина хънтингтин, полученият протеин става дефектен и води до разрушаването на някои видове неврони, като се наблюдава правопропорционална връзка между броя повтори и това колко рано започват симптомите.

Макар и с ясна генетична етиология, самият механизъм на намножаване на повторите все още не е напълно разгадан. Първоначалната хипотеза е, че това се случва при всяко наследяване на повредения ген – във всяко потомство бройката е по-висока. Но още през 2003 г. има данни за пациенти, при които се достигат 1000 повтори, което поставя хипотезата под въпрос, тъй като това предполага много дълга наследствена линия. В началото на годината добихме малко по-добра представа – някои видове клетки натрупват повтори вследствие на процес, наречен соматична експанзия. Затова и в началните стадии заболяването преминава без изявени симптоми, но повторите се множат в невронните клетки на пациентите. Акумулирането на около 80 CAG може да отнеме десетилетия, но след като достигнат този брой, повторите започват да се увеличават все по-бързо и в рамките на няколко години надвишават 150. Това се оказва границата, отвъд която невроните започват да загиват и да се проявяват симптомите на болестта.
Моделът обяснява наблюденията на лекарите през годините и внезапната поява на симптоми в по-късен етап от живота на пациентите. Добрата новина е, че бавната част от развитието на заболяването дава сравнително широк прозорец за намеса и прилагане на потенциална терапия. За съжаление, към момента няма лечение за болестта и изходът все още е летален, но по-доброто разбиране на процеса вече дава идеи за нови подходи, които могат да удължат безсимптомния период на пациентите или дори да спрат намножаването на повторите.
Един от методите е нарушаването на низовете от повтори в генома на пациентите. Използвайки т.нар. редактиране на бази, учените променят едната база в повтората (напр. C към A). Така процесът на соматична експанзия се обърква и натрупването на повтори спира, а в някои клетъчни линии се наблюдава дори намаляване на техния брой. Това е потвърдено и при мишки, върху които е приложена тази генна терапия.
Пред клиничните изпитвания на този тип редакции стоят редица пречки, най-важната от които е, че подобни повтори се срещат и на други места в човешкия геном, като създават риск за нарушаване на други процеси в организма. Засега данните показват, че редакциите извън целевия ген са само в участъци, които не кодират протеини, но това не означава, че те не са важни. Все пак идеята е интересна и ако бъде намерен начин редакцията да се насочи само в гена за болестта на Хънтингтън, терапевтичната ѝ стойност би била значителна.
Друг подход идва от компанията uniQure, която наскоро съобщи за 12 пациенти, при които влошаването на симптомите е забавено със 75% в рамките на три години. Това е постигнато не чрез директна редакция на гена за болестта на Хънтингтън, а чрез потискане на синтеза на дефектен протеин от него. За целта с помощта на аденовирус в генома на пациентите се вмъква малка ДНК последователност, която кодира не цял ген, а малка РНК молекула (микроРНК). Тя има способността да разпознае информационната РНК, произведена от дефектния ген, и да се прилепи към нея, правейки синтеза на протеин невъзможен.
За да се концентрира в целевите области на мозъка, векторът се инжектира в него през малки дупки в черепа, като процедурата е еднократна и трае около 12 часа. С времето вирусът се разпространява и към съседните региони на мозъка, потискайки образуването на дефектния протеин и в тях. Така терапията започва да действа веднага в най-засегнатите участъци, а впоследствие разширява зоната, която предпазва.
От предоставената към момента информация резултатите изглеждат обещаващи, но все пак трябва да се обърне внимание, че все още не са публикувани в рецензиран журнал, поради което не могат да бъдат подложени на критичен преглед от научната общност. Също така процедурата е изключително инвазивна, което може да постави пречки пред широкото ѝ прилагане. Въпреки това, ако компанията успее да предложи терапията комерсиално, най-вероятно ще има достатъчно пациенти, които биха се подложили на такава операция, като се има предвид как се развива заболяването.
Security updates for Thursday
Post Syndicated from jake original https://lwn.net/Articles/1048448/
Security updates have been issued by Debian (kdeconnect, libssh, and samba), Fedora (7zip, docker-buildkit, and docker-buildx), Oracle (bind, buildah, cups, delve and golang, expat, firefox, gimp, go-rpm-macros, haproxy, kernel, lasso, libsoup, libtiff, mingw-expat, openssl, podman, python-kdcproxy, qt5-qt3d, runc, squid, thunderbird, tigervnc, valkey, webkit2gtk3, xorg-x11-server, and xorg-x11-server-Xwayland), SUSE (buildah, cloudflared, containerd, expat, firefox, gnutls, helm, kernel, libxslt, mysql-connector-java, ongres-scram, openbao, openexr, openssh, podman, python311, python312, ruby2.5, rubygem-rack, runc, samba, sssd, tiff, unbound, and yelp), and Ubuntu (edk2, ffmpeg, h2o, python3.13, rust-openssl, and valkey).
The Xbox 360 Two Decades Later – An LGR Retrospective
Post Syndicated from LGR original https://www.youtube.com/watch?v=hlnbsRz6E2Y
Четвъртък, 27 Ноември 2025
Post Syndicated from georgi original http://georgi.unixsol.org/diary/archive.php/2025-11-27
Напомниха ми, че еврото идва след месец и малко, а програмката за обръщане на числа
словом работи само с левове. Затова запретнах ръкави и вече има Число словом: конвертор на числа към евро (код за сваляне).
Правилата са същите като за Число словом: конвертор на числа към български левове (код за сваляне),
Ползвайте с кеф и изпращайте в моя посока good vibes only (само благини).
How Alison Roman Does Thanksgiving
Post Syndicated from The Atlantic original https://www.youtube.com/watch?v=iMIwmfblWHo
The Best Camera under $1000?? Fuji X-T30 III Tested
Post Syndicated from Matt Granger original https://www.youtube.com/watch?v=mwTohKmN_Eg
Run Apache Spark and Iceberg 4.5x faster than open source Spark with Amazon EMR
Post Syndicated from Atul Payapilly original https://aws.amazon.com/blogs/big-data/run-apache-spark-and-iceberg-4-5x-faster-than-open-source-spark-with-amazon-emr/
This post shows how Amazon EMR 7.12 can make your Apache Spark and Iceberg workloads up to 4.5x faster performance.
The Amazon EMR runtime for Apache Spark provides a high-performance runtime environment with full API compatibility with open source Apache Spark and Apache Iceberg. Amazon EMR on EC2, Amazon EMR Serverless, Amazon EMR on Amazon EKS, Amazon EMR on AWS Outposts and AWS Glue use the optimized runtimes.
Our benchmarks show Amazon EMR 7.12 runs TPC-DS 3 TB workloads 4.5x faster than open source Spark 3.5.6 with Iceberg 1.10.0.
Performance improvements include optimizations for metadata caching, parallel I/O, adaptive query planning, data type handling, and fault tolerance. There were also some Iceberg specific regressions around data scans that we identified and fixed.
These optimizations let you match Parquet performance on Amazon EMR while keeping the key features of Iceberg key features: ACID transactions, time travel, and schema evolution.
Benchmark results compared to open source
To assess the performance of the Spark engine with the Iceberg table format, we performed benchmark tests using the 3 TB TPC-DS dataset, version 2.13, a popular industry standard benchmark. Benchmark tests for the Amazon EMR runtime for Apache Spark and Apache Iceberg were conducted on Amazon EMR 7.12 EC2 clusters compared to open source Apache Spark 3.5.6 and Apache Iceberg 1.10.0 on EC2 clusters.
Note: Our results derived from the TPC-DS dataset are not directly comparable to the official TPC-DS results due to setup differences.
The setup instructions and technical details are available in our GitHub repository. To minimize the influence of external catalogs like AWS Glue and Hive, we used the Hadoop catalog for the Iceberg tables. This uses the underlying file system, specifically Amazon S3, as the catalog. We can define this setup by configuring the property spark.sql.catalog.<catalog_name>.type. The fact tables used the default partitioning by the date column, which vary from 200–2,100 partitions. No precalculated statistics were used for these tables.
We ran a total of 104 SparkSQL queries in 3 sequential rounds, and the average runtime of each query across these rounds was taken for comparison. The average runtime for the 3 rounds on Amazon EMR 7.12 with Iceberg enabled was 0.37 hours, demonstrating a 4.5x speed increase compared to open source Spark 3.5.6 and Iceberg 1.10.0. The following figure presents the total runtimes in seconds.

The following table summarizes the metrics.
| Metric | Amazon EMR 7.12 on EC2 | Amazon EMR 7.5 on EC2 | Open source Apache Spark 3.5.6 and Apache Iceberg 1.10.0 |
| Average runtime in seconds | 1349.62 | 1535.62 | 6113.92 |
| Geometric mean over queries in seconds | 7.45910 | 8.30046 | 22.31854 |
| Cost* | $4.81 | $5.47 | $17.65 |
*Detailed cost estimates are discussed later in this post.
The following chart demonstrates the per-query performance improvement of Amazon EMR 7.12 relative to open source Spark 3.5.6 and Iceberg 1.10.0. The extent of the speedup varies from one query to another, with the fastest up to 13.6x faster for q23b, with Amazon EMR outperforming open source Spark with Iceberg tables. The horizontal axis arranges the TPC-DS 3TB benchmark queries in descending order based on the performance improvement seen with Amazon EMR, and the vertical axis depicts the magnitude of this speedup as a ratio.

Cost comparison breakdown
Our benchmark provides the total runtime and geometric mean data to assess the performance of Spark and Iceberg in a complex, real-world decision support scenario. For additional insights, we also examine the cost aspect. We calculate cost estimates using formulas that account for EC2 On-Demand instances, Amazon Elastic Block Store (Amazon EBS), and Amazon EMR expenses.
- Amazon EC2 cost (includes SSD cost) = number of instances * r5d.4xlarge hourly rate * job runtime in hours
- 4xlarge hourly rate = $1.152 per hour
- Root Amazon EBS cost = number of instances * Amazon EBS per GB-hourly rate * root EBS volume size * job runtime in hours
- Amazon EMR cost = number of instances * r5d.4xlarge Amazon EMR cost * job runtime in hours
- 4xlarge Amazon EMR cost = $0.27 per hour
- Total cost = Amazon EC2 cost + root Amazon EBS cost + Amazon EMR cost
The calculations reveal that the Amazon EMR 7.12 benchmark yields a 3.6x cost efficiency improvement over open source Spark 3.5.6 and Iceberg 1.10.0 in running the benchmark job.
| Metric | Amazon EMR 7.12 | Amazon EMR 7.5 | Open source Apache Spark 3.5.6 and Apache Iceberg 1.10.0 |
| Runtime in seconds | 1349.62 | 1535.62 | 6113.92 |
|
Number of EC2 instances (Includes primary node) |
9 | 9 | 9 |
| Amazon EBS Size | 20gb | 20gb | 20gb |
|
Amazon EC2 (Total runtime cost) |
$3.89 | $4.42 | $17.61 |
| Amazon EBS cost | $0.01 | $0.01 | $0.04 |
| Amazon EMR cost | $0.91 | $1.04 | $0 |
| Total cost | $4.81 | $5.47 | $17.65 |
| Cost savings | Amazon EMR 7.12 is 3.6x better | Amazon EMR 7.5 is 3.2x better | Baseline |
In addition to the time-based metrics discussed so far, data from Spark event logs show that Amazon EMR scanned approximately 4.3x less data from Amazon S3 and 5.3x fewer records than the open source version in the TPC-DS 3 TB benchmark. This reduction in Amazon S3 data scanning contributes directly to cost savings for Amazon EMR workloads.
Run open source Apache Spark benchmarks on Apache Iceberg tables
We used separate EC2 clusters, each equipped with 9 r5d.4xlarge instances, for testing both open source Spark 3.5.6 and Amazon EMR 7.12 for Iceberg workload. The primary node was equipped with 16 vCPU and 128 GB of memory, and the 8 worker nodes together had 128 vCPU and 1024 GB of memory. We conducted tests using the Amazon EMR default settings to showcase the typical user experience and minimally adjusted the settings of Spark and Iceberg to maintain a balanced comparison.
The following table summarizes the Amazon EC2 configurations for the primary node and 8 worker nodes of type r5d.4xlarge.
| EC2 Instance | vCPU | Memory (GiB) | Instance storage (GB) | EBS root volume (GB) |
| r5d.4xlarge | 16 | 128 | 2 x 300 NVMe SSD | 20 GB |
Prerequisites
The following prerequisites are required to run the benchmarking:
- Using the instructions in the emr-spark-benchmark GitHub repository, set up the TPC-DS source data in your S3 bucket and on your local computer.
- Build the benchmark application following the steps provided in Steps to build spark-benchmark-assembly application and copy the benchmark application to your S3 bucket. Alternatively, copy spark-benchmark-assembly-3.5.6.jar to your S3 bucket.
- Create Iceberg tables from the TPC-DS source data. Follow the instructions on GitHub to create Iceberg tables using the Hadoop catalog. For example, the following code uses an Amazon EMR 7.12 cluster with Iceberg enabled to create the tables:
Note: The Hadoop catalog warehouse location and database name from the preceding step. We use the same Iceberg tables to run benchmarks with Amazon EMR 7.12 and open source Spark.
This benchmark application is built from the branch tpcds-v2.13_iceberg. If you’re building a new benchmark application, switch to the correct branch after downloading the source code from the GitHub repository.
Create and configure a YARN cluster on Amazon EC2
To compare Iceberg performance between Amazon EMR on Amazon EC2 and open source Spark on Amazon EC2, follow the instructions in the emr-spark-benchmark GitHub repository to create an open source Spark cluster on Amazon EC2 using Flintrock with 8 worker nodes.
Based on the cluster selection for this test, the following configurations are used:
Make sure to replace the placeholder <private ip of primary node>, in the yarn-site.xml file, with the primary node’s IP address of your Flintrock cluster.
Run the TPC-DS benchmark with Apache Spark 3.5.6 and Apache Iceberg 1.10.0
Complete the following steps to run the TPC-DS benchmark:
- Log in to the open source cluster primary node using
flintrock login $CLUSTER_NAME. - Submit your Spark job:
- Choose the correct Iceberg catalog warehouse location and database that has the created Iceberg tables.
- The results are created in
s3://<YOUR_S3_BUCKET>/benchmark_run. - You can track progress in
/media/ephemeral0/spark_run.log.
Summarize the results
After the Spark job finishes, retrieve the test result file from the output S3 bucket at s3://<YOUR_S3_BUCKET>/benchmark_run/timestamp=xxxx/summary.csv/xxx.csv. This can be done either through the Amazon S3 console by navigating to the specified bucket location or by using the Amazon Command Line Interface (AWS CLI). The Spark benchmark application organizes the data by creating a timestamp folder and placing a summary file within a folder labeled summary.csv. The output CSV files contain 4 columns without headers:
- Query name
- Median time
- Minimum time
- Maximum time
With the data from 3 separate test runs with 1 iteration each time, we can calculate the average and geometric mean of the benchmark runtimes.
Run the TPC-DS benchmark with Amazon EMR runtime for Apache Spark
Most of the instructions are similar to Steps to run Spark Benchmarking with a few Iceberg-specific details.
Prerequisites
Complete the following prerequisite steps:
- Run
aws configureto configure the AWS CLI shell to point to the benchmarking AWS account. Refer to Configure the AWS CLI for instructions. - Upload the benchmark application JAR file to Amazon S3.
Deploy Amazon EMR cluster and run the benchmark job
Complete the following steps to run the benchmark job:
- Use the AWS CLI command as shown in Deploy EMR on EC2 Cluster and run benchmark job to deploy an Amazon EMR on EC2 cluster. Make sure to enable Iceberg. See Create an Iceberg cluster for more details. Choose the correct Amazon EMR version, root volume size, and same resource configuration as the open source Flintrock setup. Refer to create-cluster for a detailed description of the AWS CLI options.
- Store the cluster ID from the response. We need this for the next step.
- Submit the benchmark job in Amazon EMR using
add-stepsfrom the AWS CLI:- Replace
<cluster ID>with the cluster ID from Step 2. - The benchmark application is at
s3://<your-bucket>/spark-benchmark-assembly-3.5.6.jar. - Choose the correct Iceberg catalog warehouse location and database that has the created Iceberg tables. This should be the same as the one used for the open source TPC-DS benchmark run.
- The results will be in
s3://<your-bucket>/benchmark_run.
- Replace
Summarize the results
After the step is complete, you can see the summarized benchmark result at s3://<YOUR_S3_BUCKET>/benchmark_run/timestamp=xxxx/summary.csv/xxx.csv in the same way as the previous run and compute the average and geometric mean of the query runtimes.
Clean up
To help prevent future charges, delete the resources you created by following the instructions provided in the Cleanup section of the GitHub repository.
Summary
Amazon EMR optimizes the runtime for Spark when used with Iceberg tables, achieving 4.5x faster performance than open source Apache Spark 3.5.6 and Apache Iceberg 1.10.0 with Amazon EMR 7.12 on TPC-DS 3 TB, v2.13. This represents a significant advancement from Amazon EMR 7.5, which delivered 3.6x faster performance and closes the gap to parquet performance on Amazon EMR so customers can use the benefits of Iceberg without a performance penalty.
We encourage you to keep up to date with the latest Amazon EMR releases to fully benefit from ongoing performance improvements.
To stay informed, subscribe to the RSS feed for the AWS Big Data Blog, where you can find updates on the Amazon EMR runtime for Spark and Iceberg, as well as tips on configuration best practices and tuning recommendations.
About the authors
Atul Felix Payapilly is a software development engineer for Amazon EMR at Amazon Web Services.
Akshaya KP is a software development engineer for Amazon EMR at Amazon Web Services.
Hari Kishore Chaparala is a software development engineer for Amazon EMR at Amazon Web Services.
Giovanni Matteo is the Senior Manager for the Amazon EMR Spark and Iceberg group.
Apache Spark encryption performance improvement with Amazon EMR 7.9
Post Syndicated from Sonu Kumar Singh original https://aws.amazon.com/blogs/big-data/apache-spark-encryption-performance-improvement-with-amazon-emr-7-9/
The Amazon EMR runtime for Apache Spark is a performance-optimized runtime for Apache Spark that is 100% API compatible with open source Apache Spark. With Amazon EMR release 7.9.0, the EMR runtime for Apache Spark introduces significant performance improvements for encrypted workloads, supporting Spark version 3.5.5.
For compliance and security requirements, many customers need to enable Apache Spark’s local storage encryption (spark.io.encryption.enabled = true) in addition to Amazon Simple Storage Service (Amazon S3) encryption (such as server-side encryption (SSE) or AWS Key Management Service (AWS KMS)). This feature encrypts shuffle files, cached data, and other intermediate data written to local disk during Spark operations, protecting sensitive data at rest on Amazon EMR cluster instances.
Industries subject to regulations such as the Health Insurance Portability and Accountability Act (HIPAA) for healthcare, Payment Card Industry Data Security Standard (PCI-DSS) for financial services, General Data Protection Regulation (GDPR) for personal data, and Federal Risk and Authorization Management Program (FedRAMP) for government often require encryption of all data at rest, including temporary files on local storage. While Amazon S3 encryption protects data in object storage, Spark’s I/O encryption secures the intermediate shuffle and spill data that Spark writes to local disk during distributed processing—data that never reaches Amazon S3 but might contain sensitive information extracted from source datasets. Generally, encrypted operations require additional computational overhead that can impact overall job performance.
With the built-in encryption optimizations of Amazon EMR 7.9.0, customers might see significant performance improvements in their Apache Spark applications without requiring any application changes. In our performance benchmark tests, derived from TPC-DS performance tests at 3 TB scale, we observed up to 20% faster performance with the EMR 7.9 optimized Spark runtime compared to Spark without these optimizations. Individual results may vary depending on specific workloads and configurations.
In this post, we analyze the results from our benchmark tests comparing the Amazon EMR 7.9 optimized Spark runtime against Spark 3.5.5 without encryption optimizations. We walk through a detailed cost analysis and provide step-by-step instructions to reproduce the benchmark.
Results observed
To evaluate the performance improvements, we used an open source Spark performance test utility derived from the TPC-DS performance test toolkit. We ran the tests on two nine-node (eight core nodes and one primary node) r5d.4xlarge Amazon EMR 7.9.0 clusters, comparing two configurations:
- Baseline: EMR 7.9.0 cluster with a bootstrap action installing Spark 3.5.5 without encryption optimizations
- Optimized: EMR 7.9.0 cluster using the EMR Spark 3.5.5 runtime with encryption optimizations
Both tests used data stored in Amazon Simple Storage Service (Amazon S3). All data processing was configured identically except for the Spark runtime version.
To maintain benchmarking consistency and ensure a consistent, equivalent comparison, we disabled Dynamic Resource Allocation (DRA) in both test configurations. This approach eliminates variability from dynamic scaling and so we can measure pure computational performance improvements.
The following table shows the total job runtime for all queries (in seconds) in the 3 TB query dataset between the baseline and Amazon EMR 7.9 optimized configurations:
| Configuration | Total runtime (seconds) | Geometric mean (seconds) | Performance improvement |
| Baseline (Spark 3.5.5 without optimization) | 1,485 | 10.24 | |
| EMR 7.9 (with encryption optimization) | 1,176 | 8.15 | 20% faster |
We observed that our TPC-DS tests with the Amazon EMR 7.9 optimized Spark runtime completed about 20% faster based on total runtime and 20% faster based on geometric mean compared to the baseline configuration.
The encryption optimizations in Amazon EMR 7.9 deliver performance benefits through:
- Improved shuffle and decryption operations reducing overhead during data exchange without compromising security
- Better memory management for intermediate results
Cost analysis
The performance improvements of the Amazon EMR 7.9 optimized Spark runtime directly translate to lower costs. We realized an approximately 20% cost savings running the benchmark application with encryption optimizations compared to the baseline configuration, because of reduced hours of EMR, Amazon Elastic Compute Cloud (Amazon EC2) and Amazon Elastic Block Store (Amazon EBS) using General Purpose SSD (gp2).
The following table summarizes the cost comparison in the us-east-1 AWS Region:
| Configuration | Runtime (hours) | Estimated cost | Total EC2 instances | Total vCPU | Total memory (GiB) | Root device (EBS) |
| Baseline: Spark 3.5.5 without optimization, 1 primary and 8 core nodes | 0.41 | $5.28 | 9 | 144 | 1152 | 64 GiB gp2 |
| Amazon EMR 7.9 with optimization, 1 primary and 8 core nodes | 0.33 | $4.25 | 9 | 144 | 1152 | 64 GiB gp2 |
Cost breakdown
Formulas used:
- Amazon EMR cost – Number of instances × EMR hourly rate × Runtime hours
- Amazon EC2 cost – Number of instances × EC2 hourly rate × Runtime hour)
- Amazon EBS cost – (EBS cost per GB per month ÷ hours in a month) × EBS volume size × number of instances × runtime hours
Note: EBS is priced monthly ($0.1 per GB per month), so we divide by 730 hours to convert to an hourly rate. EMR and EC2 are already priced hourly, so no conversion is needed.
Baseline configuration (0.41 hours):
- Amazon EMR cost – 9 × $0.27 × 0.41 = $1.00
- Amazon EC2 cost – 9 × $1.152 × 0.41 = $4.25
- Amazon EBS cost – ($0.1/730 × 64 × 9 × 0.41) = $0.032
- Total cost – $5.28
EMR 7.9 optimized configuration (0.33 hours):
- Amazon EMR cost – (9 × $0.27 × 0.33) = $0.80
- Amazon EC2 cost – (9 × $1.152 × 0.33) = $3.42
- Amazon EBS cost – ($0.1/730 × 64 × 9 × 0.33) = $0.025
- Total cost: $4.25
Total cost savings: 20% per benchmark run, which scales linearly with your production workload frequency.
Set up EMR benchmarking
For detailed instructions and scripts, see the companion GitHub repository.
Prerequisites
To set up Amazon EMR benchmarking, start by completing the following prerequisite steps:
- Configure your AWS Command Line Interface (AWS CLI) by running
aws configureto point to your benchmarking account, - Create an S3 bucket for test data and results.
- Copy the TPC-DS 3TB source data from a publicly available dataset to your S3 bucket using the following command:
Replace
<YOUR-BUCKET-NAME>with the name of the S3 bucket you created in step 2. - Build or download the benchmark application JAR file (spark-benchmark-assembly-3.3.0.jar)
- Ensure you have appropriate AWS Identity Access Management (IAM) roles for EMR cluster creation and Amazon S3 access
Deploy the baseline EMR cluster (without optimization)
Step 1: Launch EMR 7.9.0 cluster with bootstrap action
The baseline configuration uses a bootstrap action to install Spark 3.5.5 without encryption optimizations. We have made the bootstrap script publicly available in an S3 bucket for your convenience.
Create the default Amazon EMR roles:
Now create the cluster:
Note: The bootstrap script is available in a public S3 bucket at s3://spark-ba/install-spark-3-5-5-no-encryption.sh. This script installs Apache Spark 3.5.5 without the encryption optimizations present in the Amazon EMR runtime.
Step 2: Submit the benchmark job to the baseline cluster
Next submit the Spark job using the following commands:
Deploy the optimized EMR cluster (with encryption optimization)
Step 1: Launch EMR 7.9.0 cluster with Spark runtime
The optimized configuration uses the EMR 7.9.0 Spark runtime without any bootstrap actions:
Example:
Step 2: Submit the benchmark job to optimized cluster
ext submit the Spark job using the following commands:
Benchmark command parameters explained
The Amazon EMR Spark step uses the following parameters:
- EMR step configuration:
- Type=Spark: Specifies this is a Spark application step
- Name=”EMR-7.9-Baseline-Spark-3.5.5″: Human-readable name for the step
- ActionOnFailure=CONTINUE: Continue with other steps if this one fails
- Spark submit arguments:
- –deploy-mode client: Run the driver on the master node (not cluster mode)
- –class com.amazonaws.eks.tpcds.BenchmarkSQL: Main class for the TPC-DS benchmark
- Application parameters:
- JAR file:
s3://<YOUR-BUCKET-NAME>/jar/spark-benchmark-assembly-3.3.0.jar - Input data
: s3://<YOUR-BUCKET-NAME>/blog/BLOG_TPCDS-TEST-3T-partitioned(3 TB TPC-DS dataset) - Output location:
s3://<YOUR-BUCKET-NAME>/blog/BASELINE_TPCDS-TEST-3T-RESULT(S3 path for results) - TPC-DS tools path:
/opt/tpcds-kit/tools(local path on EMR nodes) - Format:
parquet(output format) - Scale factor:
3000(3 TB dataset size) - Iterations:
3(run each query 3 times for averaging) - Collect results: false (don’t collect results to driver)
- Query list:
"q1-v2.4,q10-v2.4,...,ss_max-v2.4"(all 104 TPC-DS queries) - Final parameter:
true(enable detailed logging and metrics)
- JAR file:
- Query coverage:
- All 104 standard TPC-DS benchmark queries (
q1-v2.4throughq99-v2.4) - Plus the
ss_max-v2.4query for additional testing - Each query runs 3 times to calculate average performance
- All 104 standard TPC-DS benchmark queries (
Summarize the results
- Download the test result files from both output S3 locations:
- The CSV files contain four columns (without headers):
- Query name
- Median time (seconds)
- Minimum time (seconds)
- Maximum time (seconds)
- Calculate performance metrics for comparison:
- Average time per query:
AVERAGE(median, min, max)for each query - Total runtime: Sum of all median times
- Geometric mean:
GEOMEAN(average times)across all queries - Speedup: Calculate the ratio between baseline and optimized for each query
- Average time per query:
- Create comparison analysis:
Speedup = (Baseline Time - Optimized Time) / Baseline Time * 100%
Testing configuration details
The following table summarizes the test environment used for this post:
| Parameter | Value |
| EMR release | emr-7.9.0 (both configurations) |
| Baseline Spark version | 3.5.5 (installed through bootstrap action) |
| Baseline bootstrap script | s3://spark-ba/install-spark-3-5-5-no-encryption.sh (public) |
| Optimized spark version | Amazon EMR Spark runtime |
| Cluster size | 9 nodes (1 primary and 8 core) |
| Instance type | r5d.4xlarge |
| vCPUs per node | 16 |
| Memory per node | 128 GB |
| Instance storage | 600 GB SSD |
| EBS volume | 64 GB gp2 (2 volumes per instance) |
| Total vCPUs | 144 (9 × 16) |
| Total memory | 1152 GB (9 × 128) |
| Dataset | TPC-DS 3TB (Parquet format) |
| Queries | 104 queries (TPC-DS v2.4) |
| Iterations | 3 runs per query |
| DRA | Disabled for consistent benchmarking |
Clean up
To avoid incurring future charges, delete the resources you created:
- Terminate both EMR clusters:
- Delete S3 test results if no longer needed:
- Remove IAM roles if created specifically for testing
Key findings
- Up to 20% performance improvement using the Amazon EMR 7.9’s Spark runtime with no code changes required
- 20% cost savings because of reduced runtime
- Significant gains for shuffle-heavy, join-intensive workloads
- 100% API compatibility with open source Apache Spark
- Simple migration from custom Spark builds to EMR runtime
- Easy benchmarking using publicly available bootstrap scripts
Conclusion
You can run your Apache Spark workloads up to 20% faster and at lower cost without making any changes to your applications by using the Amazon EMR 7.9.0 optimized Spark runtime. This improvement is achieved through numerous optimizations in the EMR Spark runtime, including enhanced encryption handling, improved data serialization, and optimized shuffle operations.
To learn more about Amazon EMR 7.9 and best practices, see the EMR documentation. For configuration guidance and tuning advice, subscribe to the AWS Big Data Blog.
Related resources:
If you’re running Spark workloads on Amazon EMR today, we encourage you to test the EMR 7.9 Spark runtime with your production workloads and measure the improvements specific to your use case.
About the authors
Turkey?
Post Syndicated from The History Guy: History Deserves to Be Remembered original https://www.youtube.com/shorts/tjL5SWQhQyI
Run Apache Spark and Apache Iceberg write jobs 2x faster with Amazon EMR
Post Syndicated from Atul Payapilly original https://aws.amazon.com/blogs/big-data/run-apache-spark-and-apache-iceberg-write-jobs-2x-faster-with-amazon-emr/
Amazon EMR runtime for Apache Spark offers a high-performance runtime environment while maintaining API compatibility with open source Apache Spark and Apache Iceberg table format. Amazon EMR on EC2, Amazon EMR Serverless, Amazon EMR on Amazon EKS, Amazon EMR on AWS Outposts and AWS Glue use the optimized runtimes.
In this post, we demonstrate the write performance benefits of using the Amazon EMR 7.12 runtime for Spark and Iceberg compares to open source Spark 3.5.6 with Iceberg 1.10.0 tables on a 3TB merge workload.
Write Benchmark Methodology
Our benchmarks demonstrate that Amazon EMR 7.12 can run 3TB merge workloads over 2 times faster than open source Spark 3.5.6 with Iceberg 1.10.0, delivering significant improvements for data ingestion and ETL pipelines while providing the advanced features of Iceberg including ACID transactions, time travel, and schema evolution.
Benchmark workload
To evaluate the write performance improvements in Amazon EMR 7.12, we chose a merge workload that reflects common data ingestion and ETL patterns. The benchmark consists of 37 basic merge operations on TPC-DS 3TB tables, testing the performance of INSERT, UPDATE, and DELETE operations. The workload is inspired by established benchmarking approaches from the open source community, including Delta Lake’s merge benchmark methodology and the LST-Bench framework. We combined and adapted these approaches to create a comprehensive test of Iceberg write performance on AWS. We also started with an initial focus on copy-on-write performance only.
Workload characteristics
The benchmark executes 37 basic sequential merge queries that modify TPC-DS fact tables. The 37 queries are organized into three categories:
- Inserts (queries m1-m6): Adding new records to tables with varying data volumes. These queries use source tables with 5-100% new records and zero matches, testing pure insert performance at different scales.
- Upserts (queries m8-m16): Modifying existing records while inserting new ones. These upsert operations combine different ratios of matched and non-matched records—for example, 1% matches with 10% inserts, or 99% matches with 1% inserts—representing typical scenarios where data is both updated and augmented.
- Deletes (queries m7, m17-m37): Removing records with varying selectivity. These range from small, targeted deletes affecting 5% of files and rows to large-scale deletions, including partition-level deletes that can be optimized to metadata-only operations.
The queries operate on the table state created by previous operations, simulating real ETL pipelines where subsequent steps depend on earlier transformations. For example, the first six queries insert between 607,000 and 11.9 million records into the web_returns table. Later queries then update and delete from this modified table, testing read-after-write performance. Source tables were generated by sampling the TPC-DS web_returns table with controlled match/non-match ratios for consistent test conditions across the benchmark runs.
The merge operations vary in scale and complexity:
- Small operations affecting 607,000 records
- Large operations modifying over 12 million records
- Selective deletes requiring file rewrites
- Partition-level deletes optimized to metadata operations
Benchmark configuration
We ran the benchmark on identical hardware for both Amazon EMR 7.12 and open source Spark 3.5.6 with Iceberg 1.10.0:
- Cluster: 9 r5d.4xlarge instances (1 primary, 8 workers)
- Compute: 144 total vCPUs, 1,152 GB memory
- Storage: 2 x 300 GB NVMe SSD per instance
- Catalog: Hadoop Catalog
- Data format: Parquet files on Amazon S3
- Table format: Apache Iceberg (default: copy-on-write mode)
Benchmark results
We compared benchmark results for Amazon EMR 7.12 to open source Spark 3.5.6 and Iceberg 1.10.0. We ran the 37 merge queries in three sequential iterations, and the average runtime across these iterations was taken for comparison. The following table shows the results averaged across three iterations:
| Amazon EMR 7.12 (seconds) | Open Source Spark 3.5.6 + Iceberg 1.10.0 (seconds) | Speedup |
| 443.58 | 926.63 | 2.08x |
The average runtime for the three iterations on Amazon EMR 7.12 with Iceberg enabled was 443.58 seconds, demonstrating a 2.08x speed increase compared to open source Spark 3.5.6 and Iceberg 1.10.0. The following figure presents the total runtimes in seconds.

The following table summarizes the metrics.
| Metric | Amazon EMR 7.12 on EC2 | Open source Spark 3.5.6 and Iceberg 1.10.0 |
| Average runtime in seconds | 443.58 | 926.63 |
| Geometric mean over queries in seconds | 6.40746 | 18.50945 |
| Cost* | $1.58 | $2.68 |
*Detailed cost estimates are discussed later in this post.
The following chart demonstrates the per-query performance improvement of Amazon EMR 7.12 relative to open source Spark 3.5.6 and Iceberg 1.10.0. The extent of the speedup varies from one query to another, with the fastest up to 13.3 times faster for query m31, with Amazon EMR outperforming open source Spark with Iceberg tables. The horizontal axis arranges the TPC-DS 3TB benchmark queries in descending order based on the performance improvement seen with Amazon EMR, and the vertical axis depicts the magnitude of this speedup as a ratio.

Performance optimizations in Amazon EMR
Amazon EMR 7.12 achieves over 2x faster write performance through systematic optimizations across the write execution pipeline. These improvements span multiple areas:
- Metadata-only delete operations: When deleting entire partitions, EMR can now optimize these operations to metadata-only changes, eliminating the need to rewrite data files. This significantly reduces the time and cost for partition-level delete operations.
- Bloom filter joins for merge operations: Enhanced join strategies using bloom filters reduce the amount of data that needs to be read and processed during merge operations, particularly benefiting queries with selective predicates.
- Parallel file write out: Optimized parallelism during the write phase of merge operations improves throughput when writing filtered results back to Amazon S3, reducing overall merge operation time. We balanced the parallelism with read performance for overall optimized performance on the entire workload.
These optimizations work together to deliver consistent performance improvements across diverse write patterns. The result is significantly faster data ingestion and ETL pipeline execution while maintaining Iceberg’s ACID assurances and data consistency of Iceberg.
Cost comparison
Our benchmark provides the total runtime and geometric mean data to assess the performance of Spark and Iceberg in a complex, real-world decision support scenario. For additional insights, we also examine the cost aspect. We calculate cost estimates using formulas that account for EC2 On-Demand instances, Amazon Elastic Block Store (Amazon EBS), and Amazon EMR expenses.
- Amazon EC2 cost (includes SSD cost) = number of instances * r5d.4xlarge hourly rate * job runtime in hours
- 4xlarge hourly rate = $1.152 per hour
- Root Amazon EBS cost = number of instances * Amazon EBS per GB-hourly rate * root EBS volume size * job runtime in hours
- Amazon EMR cost = number of instances * r5d.4xlarge Amazon EMR cost * job runtime in hours
- 4xlarge Amazon EMR cost = $0.27 per hour
- Total cost = Amazon EC2 cost + root Amazon EBS cost + Amazon EMR cost
The calculations reveal that the Amazon EMR 7.12 benchmark yields a 1.7x cost efficiency improvement over open source Spark 3.5.6 and Iceberg 1.10.0 in running the benchmark job.
| Metric | Amazon EMR 7.12 | Open source Spark 3.5.6 and Iceberg 1.10.0 |
| Runtime in seconds | 443.58 | 926.63 |
| Number of EC2 instances(Includes primary node) | 9 | 9 |
| Amazon EBS Size | 20gb | 20gb |
| Amazon EC2(Total runtime cost) | $1.28 | $2.67 |
| Amazon EBS cost | $0.00 | $0.01 |
| Amazon EMR cost | $0.30 | $0 |
| Total cost | $1.58 | $2.68 |
| Cost savings | Amazon EMR 7.12 is 1.7 times better | Baseline |
Run open source Spark benchmarks on Iceberg tables
We used separate EC2 clusters, each equipped with nine r5d.4xlarge instances, for testing both open source Spark 3.5.6 and Amazon EMR 7.12 for Iceberg workload. The primary node was equipped with 16 vCPU and 128 GB of memory, and the eight worker nodes together had 128 vCPU and 1024 GB of memory. We conducted tests using the Amazon EMR default settings to showcase the typical user experience and minimally adjusted the settings of Spark and Iceberg to maintain a balanced comparison.
The following table summarizes the Amazon EC2 configurations for the primary node and eight worker nodes of type r5d.4xlarge.
| EC2 Instance | vCPU | Memory (GiB) | Instance storage (GB) | EBS root volume (GB) |
| r5d.4xlarge | 16 | 128 | 2 x 300 NVMe SSD | 20 GB |
Benchmarking instructions
Follow the steps below to run the benchmark:
- For the open source run, create a Spark cluster on Amazon EC2 using Flintrock with the configuration described previously.
- Setup the TPC-DS source data with Iceberg in your S3 bucket.
- Build the benchmark application jar from the source to run the benchmarking and get the results.
Detailed instructions are provided in the emr-spark-benchmark GitHub repository.
Summarize the results
After the Spark job finishes, retrieve the test result file from the output S3 bucket at s3://<YOUR_S3_BUCKET>/benchmark_run/timestamp=xxxx/summary.csv/xxx.csv. This can be done either through the Amazon S3 console by navigating to the specified bucket location or by using the Amazon Command Line Interface (AWS CLI). The Spark benchmark application organizes the data by creating a timestamp folder and placing a summary file within a folder labeled summary.csv. The output CSV files contain four columns without headers:
- Query name
- Median time
- Minimum time
- Maximum time
With the data from three separate test runs with one iteration each time, we can calculate the average and geometric mean of the benchmark runtimes.
Clean up
To help prevent future charges, delete the resources you created by following the instructions provided in the Cleanup section of the GitHub repository.
Summary
Amazon EMR is consistently enhancing the EMR runtime for Spark when used with Iceberg tables, achieving write performance that is over 2 times faster than open source Spark 3.5.6 and Iceberg 1.10.0 with EMR 7.12 on 3TB merge workloads. This represents a significant improvement for data ingestion and ETL pipelines, helping to deliver 1.7x cost reduction while maintaining the ACID assurances of Iceberg. We encourage you to keep up to date with the latest Amazon EMR releases to fully benefit from ongoing performance improvements.
To stay informed, subscribe to the RSS feed for the AWS Big Data Blog, where you can find updates on the EMR runtime for Spark and Iceberg, as well as tips on configuration best practices and tuning recommendations.
About the authors
Atul Felix Payapilly is a software development engineer for Amazon EMR at Amazon Web Services.
Akshaya KP is a software development engineer for Amazon EMR at Amazon Web Services.
Hari Kishore Chaparala is a software development engineer for Amazon EMR at Amazon Web Services.
Giovanni Matteo is the Senior Manager for the Amazon EMR Spark and Iceberg group.
Medidata’s journey to a modern lakehouse architecture on AWS
Post Syndicated from Mike Araujo original https://aws.amazon.com/blogs/big-data/medidatas-journey-to-a-modern-lakehouse-architecture-on-aws/
This post was co-authored by Mike Araujo Principal Engineer at Medidata Solutions.
The life sciences industry is transitioning from fragmented, standalone tools towards integrated, platform-based solutions. Medidata, a Dassault Systèmes company, is building a next-generation data platform that addresses the complex challenges of modern clinical research. In this post, we show you how Medidata created a unified, scalable, real-time data platform that serves thousands of clinical trials worldwide with AWS services, Apache Iceberg, and a modern lakehouse architecture.
Challenges with legacy architecture
As the Medidata clinical data repository expanded, the team recognized the shortcomings of the legacy data solution to provide quality data products to their customers across their growing portfolio of data offerings. Several data tenants began to erode. The following diagram shows Medidata’s legacy extract, transform, and load (ETL) architecture.

Built upon a series of scheduled batch jobs, the legacy system proved ill-equipped to provide a unified view of the data across the entire ecosystem. Batch jobs ran at different intervals, often requiring a sufficient degree of scheduling buffer to make sure upstream jobs completed within the expected window. As the data volume expanded, the jobs and their schedules continued to inflate, introducing a latency window between ingestion and processing for dependent consumers. Different consumers operating from various underlying data services further magnified the problem as pipelines had to be continuously built across a variety of data delivery stacks.
The expanding portfolio of pipelines began to overwhelm existing maintenance operations. With more operations, the opportunity for failure expanded and recovery efforts further complicated. Existing observability systems were inundated with operational data, and identifying the root cause of data quality issues became a multi-day endeavor. Increases in the data volume required scaling considerations across the entire data estate.
Additionally, the proliferation of data pipelines and copies of the data in different technologies and storage systems necessitated expanding access controls with enhanced security features to make sure only the correct users had access to the subset of data to which they were permitted. Making sure access control changes were correctly propagated across all systems added a further layer of complexity to consumers and producers.
Solution overview
With the advent of Clinical Data Studio (Medidata’s unified data management and analytics solution for clinical trials) and Data Connect (Medidata’s data solution for acquiring, transforming, and exchanging electronic health record (EHR) data across healthcare organizations), Medidata introduced a new world of data discovery, analysis, and integration to the life sciences industry powered by open source technologies and hosted on AWS. The following diagram illustrates the solution architecture.

Fragmented batch ETL jobs were replaced by real-time Apache Flink streaming pipelines, an open source, distributed engine for stateful processing, and powered by Amazon Elastic Kubernetes Service (Amazon EKS), a fully managed Kubernetes service. The Flink jobs write to Apache Kafka running in Amazon Managed Apache Kafka (Amazon MSK), a streaming data service that manages Kafka infrastructure and operations, before landing in Iceberg tables backed by the AWS Glue Data Catalog, a centralized metadata repository for data assets. From this collection of Iceberg tables, a central, single source of data is now accessible from a variety of consumers without additional downstream processing, alleviating the need for custom pipelines to satisfy the requirements of downstream consumers. Through these fundamental architectural changes, the team at Medidata solved the issues presented by the legacy solution.
Data availability and consistency
With the introduction of the Flink jobs and Iceberg tables, the team was able to deliver a consistent view of their data across the Medidata data experience. Pipeline latency was reduced from days to minutes, helping Medidata customers realize a 99% performance gain from the data ingestion to the data analytics layers. Due to Iceberg’s interoperability, Medidata users saw the same view of the data regardless of where they viewed that data, minimizing the need for consumer-driven custom pipelines because Iceberg could plug into existing consumers.
Maintenance and durability
Iceberg’s interoperability provided a single copy of the data to satisfy their use cases, so the Medidata team could focus its observation and maintenance efforts on a five-times smaller subset of operations than previously required. Observability was enhanced by tapping into the various metadata components and metrics exposed by Iceberg and the Data Catalog. Quality management transformed from cross-system traces and queries to a single analysis of unified pipelines, with an added benefit of point in time data queries thanks to the Iceberg snapshot feature. Data volume increases are handled with out-of-box scaling supported by the entire infrastructure stack and AWS Glue Iceberg optimization features that include compaction, snapshot retention, and orphan file deletion, which provide a set-and-forget experience for solving a number of common Iceberg frustrations, such as the small file problem, orphan file retention, and query performance.
Security
With Iceberg at the center of its solution architecture, the Medidata team no longer had to spend the time building custom access control layers with enhanced security features at each data integration point. Iceberg on AWS centralizes the authorization layer using familiar systems such as AWS Identity and Access Management (IAM), providing a single and durable control for data access. The data also stays entirely within the Medidata virtual private cloud (VPC), further reducing the opportunity for unintended disclosures.
Conclusion
In this post, we demonstrated how legacy universe of consumer-driven custom ETL pipelines can be replaced with a scalable, high-performant streaming lakehouses. By putting Iceberg on AWS at the center of data operations, you can have a single source of data for your consumers.
To learn more about Iceberg on AWS, refer to Optimizing Iceberg tables and Using Apache Iceberg on AWS.