On the Security of Password Managers

Post Syndicated from Bruce Schneier original https://www.schneier.com/blog/archives/2026/02/on-the-security-of-password-managers.html

Good article on password managers that secretly have a backdoor.

New research shows that these claims aren’t true in all cases, particularly when account recovery is in place or password managers are set to share vaults or organize users into groups. The researchers reverse-engineered or closely analyzed Bitwarden, Dashlane, and LastPass and identified ways that someone with control over the server­—either administrative or the result of a compromise­—can, in fact, steal data and, in some cases, entire vaults. The researchers also devised other attacks that can weaken the encryption to the point that ciphertext can be converted to plaintext.

This is where I plug my own Password Safe. It isn’t as full-featured as the others and it doesn’t use the cloud at all, but it’s actual encryption with no recovery features.

Cloudflare One is the first SASE offering modern post-quantum encryption across the full platform

Post Syndicated from Sharon Goldberg original https://blog.cloudflare.com/post-quantum-sase/

During Security Week 2025, we launched the industry’s first cloud-native post-quantum Secure Web Gateway (SWG) and Zero Trust solution, a major step towards securing enterprise network traffic sent from end user devices to public and private networks.

But this is only part of the equation. To truly secure the future of enterprise networking, you need a complete Secure Access Service Edge (SASE). 

Today, we complete the equation: Cloudflare One is the first SASE platform to support modern standards-compliant post-quantum (PQ) encryption in our Secure Web Gateway, and across Zero Trust and Wide Area Network (WAN) use cases.  More specifically, Cloudflare One now offers post-quantum hybrid ML-KEM (Module-Lattice-based Key-Encapsulation Mechanism) across all major on-ramps and off-ramps.

To complete the equation, we added support for post-quantum encryption to our Cloudflare IPsec (our cloud-native WAN-as-a-Service) and Cloudflare One Appliance (our physical or virtual WAN appliance that establish Cloudflare IPsec connections). Cloudflare IPsec uses the IPsec protocol to establish encrypted tunnels from a customer’s network to Cloudflare’s global network, while IP Anycast is used to automatically route that tunnel to the nearest Cloudflare data center. Cloudflare IPsec simplifies configuration and provides high availability; if a specific data center becomes unavailable, traffic is automatically rerouted to the closest healthy data center. Cloudflare IPsec runs at the scale of our global network, and supports site-to-site across a WAN as well as outbound connections to the Internet.

The Cloudflare One Appliance upgrade is generally available as of appliance version 2026.2.0. The Cloudflare IPsec upgrade is in closed beta, and you can get on the list by reaching out to your account team [email protected].

Post-quantum cryptography matters now

Quantum threats are not a “next decade” problem. Here is why our customers are prioritizing post-quantum cryptography (PQC) today:

The deadline is approaching. At the end of 2024, the National Institute of Standards and Technology (NIST) sent a clear signal (that has been echoed by other agencies): the era of classical public-key cryptography is coming to an end. NIST set a 2030 deadline for depreciating RSA and Elliptic Curve Cryptography (ECC) and transitioning to PQC that cannot be broken by powerful quantum computers. Organizations that haven’t begun their migration risk being out of compliance and vulnerable as the deadline nears.

Upgrades have historically been tricky. While 2030 might seem far away, upgrading cryptographic algorithms is notoriously difficult. History has shown us that depreciating cryptography can take decades: we found examples of MD5 causing problems 20 years after it was deprecated. This lack of crypto agility — the ability to easily swap out cryptographic algorithms — is a major bottleneck. By integrating PQ encryption directly into Cloudflare One, our SASE platform, we provide built-in crypto agility, simplifying how organizations offer remote access and site-to-site connectivity.

Data may already be at risk. Finally, “Harvest Now, Decrypt Later” is a present and persistent threat, where attackers harvest sensitive network traffic today and then store it until quantum computers become powerful enough to decrypt it. If your data has a shelf life of more than a few years (e.g. financial information, health data, state secrets) it is already at risk unless it is protected by PQ encryption.

The two migrations on the road to quantum safety: key agreement and digital signatures

Transitioning network traffic to post-quantum cryptography (PQC) requires an overhaul of two cryptographic primitives: key agreement and digital signatures.  

Migration 1: Key establishment. Key agreement allows two parties to establish a shared secret over an insecure channel; the shared secret is then used to encrypt network traffic, resulting in post-quantum encryption. The industry has largely converged on ML-KEM (Module-Lattice-based Key-Encapsulation Mechanism) as the standard PQ key agreement protocol. 

ML-KEM has been widely adopted for use in TLS, usually deployed alongside classical Elliptic Curve Diffie Hellman (ECDHE), where the key used to encrypt network traffic is derived by mixing the outputs of the ML-KEM and ECDHE key agreements. (This is also known as “hybrid ML-KEM”). Well over 60% of human-generated TLS traffic to Cloudflare’s network is currently protected with hybrid ML-KEM. The transition to hybrid ML-KEM has been successful because it:

Because ML-KEM runs in parallel with classical ECDHE, there is no reduction in security and compliance as compared to the classical ECDHE approach.  

Migration 2: Digital signatures. Meanwhile, digital signatures and certificates protect authenticity, stopping active adversaries from impersonating the server to the client. Unfortunately, PQ signatures are currently larger in size than classical ECC algorithms, which has slowed their adoption. Fortunately, the migration to PQ signatures is less urgent, because PQ signatures are designed to stop active adversaries armed with powerful quantum computers, which are not known to exist yet. Thus, while Cloudflare is actively contributing to the standardization and rollout of PQ digital signatures, the current Cloudflare IPsec upgrade focuses on upgrading key establishment to hybrid ML-KEM.  

The U.S. Cybersecurity & Infrastructure Security Agency (CISA) recognized the nature of these two migrations in its January 2026 publication, “Product Categories for Technologies That Use Post-Quantum Cryptography Standards.”

Breaking new ground with IPsec 

To achieve a SASE fully protected with post-quantum encryption, we’ve upgraded our Cloudflare IPsec products to support hybrid ML-KEM in the IPsec protocol.

The IPsec community’s journey toward post-quantum cryptography has been very different from that of TLS. TLS is the de facto standard for encrypting public Internet traffic at Layer 4  — e.g. from a browser to a content delivery network (CDN) — so security and vendor interoperability are at the forefront of its design. Meanwhile, IPsec is a Layer 3 protocol that commonly connects devices built by the same vendor (e.g. two routers), so interoperability has historically been less of a concern. With this in mind, let’s take a look at IPsec’s journey into the quantum future. 

Pre-Shared Keys? Quantum key distribution?

RFC 8784, published in May 2020, was intended to be the post-quantum update to IPsec Internet Key Exchange v2 (IKEv2), which is used to establish the symmetric keys used to encrypt IPsec network traffic. RFC 8784 implies the use of either long-lived pre-shared keys (PSK) or quantum key distribution (QKD). Neither of these approaches are very palatable.

RFC 8784 proposes mixing a PSK with a key derived from Diffie Hellman Exchange (DHE), essentially running PSK in hybrid with DHE. This approach protects against harvest-now-decrypt-later attackers, but does not offer forward secrecy against quantum adversaries. 

Forward secrecy is a standard desideratum of key agreement protocols. It ensures that a system is secure even if the long-lived key is leaked. The PSK approach in RFC 8784 is vulnerable to an harvest-now-decrypt-later adversary that also obtains a copy of a long-lived PSK, and can then decrypt traffic in the future (by breaking the DHE key agreement) once powerful quantum computers become available.

To solve this forward secrecy issue, RFC 8784 can instead be used to mix the key from the classical DHE with a freshly generated key derived from a QKD protocol.

QKD uses quantum mechanics to establish a shared, secret cryptographic key between two parties. Importantly, for QKD to work, the parties must have specialized hardware or be connected by a dedicated physical connection. This is a significant limitation, rendering QKD useless for common Internet use cases like connecting a laptop to a distant server over Wi-Fi. These limitations are also why we never invested in deploying QKD for Cloudflare IPsec. The U.S. National Security Agency (NSA), Germany’s BSI and the UK National Cyber Security Centre have also warned against relying solely on QKD.

But what about interoperability? 

RFC 9370 landed in May 2023, specifying the use of hybrid key agreement rather than PSK or QKD. But unlike TLS, which only supports using post-quantum ML-KEM in parallel with classical DHE, this IPsec standard allows using up to seven different key agreements to run at the same time in parallel with classical Diffie Helman. Moreover, it doesn’t specify details about what these key agreements should be, leaving it up to the vendors to choose their algorithms and implementations. Palo Alto Networks, for example, took this seriously and built support for over seven different PQC ciphersuites into its next generation firewall (NGFW), most of which do not interoperate with other vendors and some of which have not yet been standardized by NIST.

Over the years, TLS has gone in the opposite direction, reducing the number of registered ciphersuites from hundreds in TLS 1.2, down to around five in TLS 1.3. This philosophy of reducing “ciphersuite bloat” is also in line with NIST’s SP 800 52 from 2019.  The rationale for reducing “ciphersuite bloat” includes: 

  • Improved interoperability across vendors and regions

  • Lower risk of attacks that exploit downgrades to weak ciphersuites 

  • Lower risk of security problems due to misconfiguration

  • Lower risk of implementation flaws by reducing the size of the codebase

This is why we didn’t initially build support for RFC 9370. 

Standards that are finally on the right track

It’s also why we were excited when the IPsec community put forth draft-ietf-ipsecme-ikev2-mlkem. This Internet-Draft standardizes PQ exchange for IPsec in the same way PQ key exchange has been widely deployed for TLS: hybrid ML-KEM. The new draft fills in the gaps in RFC 9370, by specifying how to run the ML-KEM as the additional key exchange in parallel with classical Diffie Hellman in IKEv2. 

Now that this specification is available, we’ve moved forward with supporting post-quantum IPsec in our Cloudflare IPsec products. 

Cloudflare IPsec goes post-quantum

Cloudflare IPsec is a WAN Network-as-a-Service solution that replaces legacy private network architectures by connecting data centers, branch offices, and cloud VPCs to Cloudflare’s global IP Anycast network. 

With Cloudflare IPsec, Cloudflare’s network acts as the IKEv2 Responder, awaiting connection requests from an IPsec initiator, which is a branch connector device in the customer’s network. Cloudflare IPsec supports IPsec sessions initiated by branch connectors that include our own Cloudflare One Appliance, along with branch connectors from a diverse set of vendors, including Cisco, Juniper, Palo Alto Networks, Fortinet, Aruba and others.

We’ve implemented production hybrid ML-KEM support in the Cloudflare IPsec IKEv2 Responder, as specified in draft-ietf-ipsecme-ikev2-mlkem. The draft requires a first key exchange to run using a classical Diffie Helman key exchange. The derived key is used to encrypt a second key exchange that is run using ML-KEM. Finally, the keys derived by the two exchanges are mixed and the result is used to secure the data plane traffic in IPsec ESP (Encapsulating Security Payload) mode. ESP mode uses symmetric cryptography and is thus already quantum safe without any additional upgrades.  We’ve tested our implementation against the IPsec Initiator in the strongswan reference implementation.

You can see the ciphersuite used in the IKEv2 negotiation by viewing the Cloudflare IPsec logs.

We chose to implement hybrid ML-KEM rather than “pure” ML-KEM, i.e. only ML-KEM without DHE running in parallel, for two reasons. First, we’ve used hybrid ML-KEM across all of our other Cloudflare products, since this is the approach adopted across the TLS community. And second, it provides a “belt-and-suspenders” security: ML-KEM provides protection against quantum harvest-now-decrypt-later attacks, while DHE provides a tried-and-true algorithm against non-quantum adversaries.

An invitation for interoperability

The full value of this implementation can be realized only via interoperability. For this reason, we are inviting other vendors that are building out support for IPsec Initiators in their branch connectors per draft-ietf-ipsecme-ikev2-mlkem to test against our Cloudflare IPsec implementation. Cloudflare customers looking to test out interoperability with third-party branch connectors while we are in closed beta can get in touch with us by reaching out to your account team at [email protected]. We plan to GA and build out interoperability with other vendors as more begin to come online with support for draft-ietf-ipsecme-ikev2-mlkem.

Quantum-safe hardware: the Cloudflare One Appliance

Many of our customers purchase their branch connector (hardware or virtualized) from Cloudflare, rather than a third-party vendor. That’s why the Cloudflare One Appliance — our plug-and-play appliance that connects your local network to Cloudflare One — has also been upgraded with post-quantum encryption.

Cloudflare One Appliance does not use IKEv2 for key agreement or session establishment, opting instead to rely on TLS. The appliance periodically initiates a TLS handshake with the Cloudflare edge, shares a symmetric secret over the resulting TLS connection, then injects that symmetric secret into the ESP layer of IPsec, which then encrypts and authenticates the IPsec data plane traffic. This design allowed us to avoid building out IKEv2 Initiator logic, and makes the Connector easier to maintain using our existing TLS libraries. 

Thus, upgrading Cloudflare One Appliance to PQ encryption was just a matter of upgrading TLS 1.2 to TLS 1.3 with hybrid ML-KEM — something we’ve done many times on different products at Cloudflare. 

How do I turn this on? And what does it cost?

As always, this upgrade to Cloudflare IPsec comes at no extra cost to our customers. Because we believe that a secure and private Internet should be accessible to all, we’re on a mission to include PQC in all our products, without specialized hardware, at no extra cost to our customers and end users.

Customers using the Cloudflare One Appliance obtained this upgrade to PQC in version 2026.2.0 (released 2026-02-11). The upgrade is pushed automatically (with no customer action required) according to each appliance’s configured interrupt window.

For customers using Cloudflare IPsec with another vendor’s branch connector appliance, we will be interoperating with these once more support for draft-ietf-ipsecme-ikev2-mlkem comes online. You can also contact us directly to get access to closed beta and request that we interoperate with a specific vendor’s branch connector by reaching out to your account team at [email protected]. 

The full picture: post-quantum SASE

The value proposition for a post-quantum SASE is clear: organizations can obtain immediate end-to-end protection for their private network traffic by sending it over tunnels protected by hybrid ML-KEM. This protects traffic from  harvest-now-decrypt-later attacks, even if the individual applications in the corporate network are not yet upgraded to PQC.


The diagram above shows how post-quantum hybrid ML-KEM is offered in various Cloudflare One network configurations.  It includes the following on-ramps:

and the following off-ramps:

The diagram below highlights a sample network configuration that uses the Cloudflare One Client on-ramp to connect a device to a server behind a Cloudflare One Appliance offramp. The end user’s device connects to the Cloudflare network (link 1) using MASQUE with hybrid ML-KEM. The traffic then travels across Cloudflare’s global network over TLS 1.3 with hybrid ML-KEM (link 2). Traffic then leaves the Cloudflare network over a post-quantum Cloudflare IPsec link (link 3) that is terminated at a Cloudflare One Appliance appliance. Finally it connects to a server inside the customer’s environment. Traffic is protected by post-quantum cryptography as it travels over the public Internet, even if the server itself does not support post-quantum cryptography.


Finally, we note that traffic that on-ramps to Cloudflare One and then egresses to the public Internet can also be protected by our post-quantum Cloudflare Gateway, our Secure Web Gateway (SWG).  Here’s a diagram showing how the SWG works:


 As discussed in an earlier blog post, our SWG can already support hybrid ML-KEM on traffic from SWG to the origin server (as long as the origin supports hybrid ML-KEM), and on traffic from the client to the SWG (if the client supports hybrid ML-KEM, which is the case for most modern browsers). Importantly, any traffic that onramps to the SWG via a device that has Cloudflare One Client installed is still protected with hybrid ML-KEM — even if the web browser itself does not yet support post-quantum cryptography. This is due to the post-quantum MASQUE tunnel that the Cloudflare One Client establishes to Cloudflare’s global network.  The same is true of traffic that onramps to the SWG via a post-quantum Cloudflare IPsec tunnel.

Putting it all together, Cloudflare One now offers post-quantum encryption on our TLS, MASQUE and IPsec on-ramp and off-ramps, and for private network traffic, and to traffic that egresses to the public Internet via our SWG. 

The future is quantum-safe

By completing the post-quantum SASE equation with Cloudflare IPsec and the Cloudflare One Appliance, we have extended post-quantum encryption across all our major on-ramps and off-ramps. We have intentionally chosen the path of interoperability and simplicity — the hybrid ML-KEM approach that the IETF and NIST have championed, rather than locking our customers into proprietary implementations, “ciphersuite bloat,” or unnecessary hardware upgrades. 

This is the promise of Cloudflare One: a SASE platform that is not only faster and more reliable than the legacy architectures it replaces, but one that provides post-quantum encryption. Whether you are securing a remote worker’s browser or a multi-gigabit data center link, you can now do so with the confidence that your data is protected from harvest-now-decrypt-later attacks and other future-looking threats.  

You can sign up here to get a full demo of our post-quantum capabilities across the Cloudflare One SASE platform. We are proud to lead the industry into this new era of cryptography, and we invite you to join us in building a scalable, standards-compliant, and post-quantum Internet.

Седмицата (16–21 февруари)

Post Syndicated from Надежда Радулова original https://www.toest.bg/sedmitsata-16-21-fevruari/

Седмицата (16–21 февруари)

„Скръбта е твар перната“ (по великолепния роман на Макс Портър), точно както „Любовта е немирна птица“ (по „Кармен“ на Бизе). И скръбта, и любовта ни съпътстват през целия ни живот – от мига, в който се отделим от майчиното тяло, до мига, в който телата ни се превърнат в прах. И скръбта, и любовта обаче са дарове крехки, които лесно могат да бъдат пропилени, прокудени наистина като птици, а с тях и възможността, дадена ни да се самолекуваме, да преживяваме раната, през която влизаме в света и в езика.

В прочутото си есе от 1917 г. „Траур и меланхолия“ Зигмунд Фройд описва човешката реакция при загуба на любим човек или идеал. И проследява сложния и нелек път, по който именно траурът ще върне скърбящия субект на страната на живота, ще го разграничи от изгубения обект, ще му даде възможност да продължи напред. Това е процес, осеян с подводни камъни, който невинаги върви по план и би могъл да доведе до отключване на меланхолия, депресия и други психични страдания. Да ни превърне във вечни пленници на загубата. И на скръбта. Траурът всъщност е грижа за душата и възпрепятстването му крие огромни рискове.

През последните две седмици станахме свидетели и участници в една всеобща мрачна демонстрация на това колко е трудно да скърбим. И колко непосилно е за нас да уважаваме начините, по които другите скърбят. В голяма степен го дължим на редица държавни институции, които с безотговорното си (бих допълнила и безпросветно) поведение пред медиите ни нанесоха колективна травма и откриха фронт за идентификации и деидентификации, за поляризиране в обществото. Всичко това – удобно и навреме. Точно преди изборите.

Нормализирането на ситуацията в момента изглежда трудно постижимо и изисква усилия не само от институциите, отговорни за тази ситуация, но и от всички нас. Защото освен публичната, видимата „сцена“, на която се играе тази древногръцка по мащаба си трагедия, всеки от нас има своята вътрешна, смълчана „сцена“, на която се извършват душевните ни кръвопролития. Без да прогледнем за тях, катарзисът е невъзможен.

За да постигнем зрелост в отношението ни като общество към случая „Петрохан“, а и към всеки случай, от който (без)отговорни представители на институции и кликбейт медии извайват сензационни сюжети, е необходимо все по-умно и предпазливо да се ориентираме в морето от трудно проверима информация. Затова колко уязвими сме пред заливащата ни медийна пяна, сред десетките конкуриращи се за място в кошмарите ни конспиративни теории, пише Александър Драганов в есето си „Има ли компас в информационния океан?“.

Има ли компас в информационния океан?
С колкото повече информация разполагаме, толкова повече се давим в нея. Александър Драганов дава някои ценни съвети как да отсяваме зърното от плявата и да не ставаме жертва на собствените си или чужди конспирации.
Седмицата (16–21 февруари)

Текстът на Александър се явява естествено продължение на други два скорошни анализа от предишните ни броеве, които има смисъл да препрочетем и тази седмица:

„Няма места.“ Журналистиката след журналистите
Ако се чудите какво стана с медиите в България, този текст ви е напълно достатъчен, за да си дадете отговори на много въпроси. А иначе, вие въпроси може и да си задавате, но в много медии у нас вече няма кой да ги задава – гледайте какво нещо… Защо стана така – от Дарина Сарелска.
Седмицата (16–21 февруари)

Манипулацията – този стар и все тъй полезен занаят
Свикнали сме да свързваме Държавна сигурност със задкулисието. Бившият служител на ДС Полковник А. обаче иска да бъде полезен на обществото. В дебютната си статия за „Тоест“ той ни дава насоки как да разберем кога ни манипулират. И има още много неща да ни каже. Дали ще го чуем, зависи от нас.
Седмицата (16–21 февруари)

Въпросът за паралелните вселени (или балони, ако повече ви харесва), в които съществуваме, не опира само до начина, по който възприемаме и обработваме информацията за света около нас. В образователната сфера например също има здрачни, неосветени от държавата зони, които – както напоследък показаха и петроханският случай, и самоубийството на Билгин – стават видими само покрай поредната трагедия, свързана с деца в училищна възраст. Повече за алтернативните форми на начално, основно и средно образование, както и за отсъствието на ангажимент и грижа от страна на държавните образователни политики към тях може да научите от анализа на Светла Енчева „Паралелните образователни реалности в България“.

Паралелните образователни реалности в България
Независимо дали детето учи в класна стая или вкъщи, гаранция за неговата сигурност няма. Така излиза, защото системата реагира избирателно, а родителите все по-често търсят алтернативи. Накрая държавата се оказва изненадана от всичко, което сама е допуснала.
Седмицата (16–21 февруари)

А според Светла далечните последствия от този престъпен тип незаинтересованост могат да бъдат наистина зловещи:

В каквато и форма да се обучава детето, няма гаранция, че държавата ще се загрижи за най-добрия му интерес. На базата на практиката можем да предположим, че вероятността да не се загрижи е по-голяма. Това е зле не само за децата, а и за самата държава – в нея съществуват алтернативни вселени, за които тя си няма и понятие. Тази държава не разбира откъде се пръкват „локалите“, защо деца се самоубиват и защо наказанията не ги правят по-добри. Тя в един момент може и да се сдобие с фундаменталисти и терористи (все едно дали обучавани в училище, или вкъщи) и да се чуди откъде ѝ е дошло.

Продължаваме с друга важна тема, за която обикновено си даваме сметка отново когато вече е късно – системата на пътната безопасност и колко много елементи от различни редове включва тя, за да е в действителност работеща. За управление на риска, а не на последствията, които твърде често са изгубени човешки животи, настоява в експертната си статия „Пътната безопасност отвъд знаците“ Симеон Иванов от „Екипът на София“.

Пътната безопасност отвъд знаците
Замисляли ли сте се защо понякога пътните знаци и маркировки са „врата в полето“, а в други случаи може дори да няма такива, но правилата да се спазват? Симеон Иванов от „Екипът на София“ ни разказва колко много неща накуп трябва да се имат предвид, за да е работеща пътната безопасност.
Седмицата (16–21 февруари)

Подобна сложна система от условия и правила за безопасност обикновено се изгражда и в социалната минивселена на платформите за запознанства. Тази седмица наш водач из онлайн стъргалото на мюсюлманските сайтове е Атанас Шиников, който ни разказва за практическата функционалност, но и за моралните гаранции за съществуването на въпросните форуми, стимулиращи срещи, раздели, но и трайни партньорства. Повече за непознатите у нас Muzz, Muslima, Salams, Baklava и пр. четете в „Тиндър/Миндър, халал, сайтове и приложения за запознанства“.

Тиндър/Миндър, халал, сайтове и приложения за запознанства
14 февруари отмина, но както се пее в песента, love is in the air. А спомняйки си за един обилно награден филм, може да допълним: everywhere, all at once. Така е и в арабските сайтове и приложения за запознанства според Атанас Шиников. Да ги разгледаме!
Седмицата (16–21 февруари)

Съвсем в тон с криминалната лента, която денонощно се прожектира в главите ни през последните седмици, в рубриката „На второ четене“ Антония Апостолова ни представя роман за убийство: „Нощта на професор Андершен“ от норвежкия писател Даг Сулста. Но както се оказва, това не е скандинавска кримка, тъй като не престъплението и неговото разкриване са в основата на сюжета, а вътрешните колебания и сложни умишления на единствения свидетел.

На второ четене: „Нощта на професор Андершен“
На Бъдни вечер професор Андершен става случаен свидетел на убийство. Въпросът е дали да се обади в полицията… Ето ви вътъка на класически криминален роман. Не такъв обаче е случаят. Авторът Даг Сулста ни изненадва, а Антония Апостолова разказва защо и как.
Седмицата (16–21 февруари)

Ако не сте били свидетели на живо на видеопоредицата ни, то сега може да поправите тази грешка, като се върнете към седми епизод на „Тоест разговаряме“ с Владислав Севов, в който водещият на редовната ни рубрика „Научни новини“ Михаил Ангелов коментира границите и възможностите на съвременната наука – от генетичното редактиране с CRISPR, през космическите изследвания и ваксините, до бъдещето на храните и добива им. Освен да гледате целия епизод в нашия YouTube канал, вече може да го чуете и като аудиозапис в SoundCloud.

Тоест разговаряме – епизод 7
Как се разказва за наука разбираемо и отговорно? И защо критичното мислене е ключово в епохата на псевдонаучни твърдения и конспиративни интерпретации? Отговори на тези въпроси ни дава биологът и автор на „Научни новини“ Михаил Ангелов в седмия епизод на „Тоест разговаряме“.
Седмицата (16–21 февруари)

За феновете на рубриката, които са проследили епизода на живо, има специален бонус – Михаил отговаря на допълнителен зрителски въпрос, свързан с промените в начините, по които ще отглеждаме храната си в следващите десет години.

Много по-кратък времеви хоризонт, само два-три месеца, има новото служебно правителство на България начело с Андрей Гюров, което положи клетва в четвъртък и от което се очаква да осигури гладкото провеждане на честни парламентарни избори. Според Емилия Милчева правителството на Гюров е

скроено по изпитаната сватбена традиция – нещо старо, нещо ново, нещо назаем и нещо синьо; засега е трудно да се прецени кое е в повече – старото, новото или синьото. 

При всички положения обаче, твърди Емилия, волята на декемврийските протести срещу Борисов и Пеевски е чута и оттук нататък остава да видим дали това правителство ще се окаже „мост към нови коалиции“. Междувременно експресната оставка на вицепремиера Стоил Цицелков показа, че този кабинет няма да се размине с натиск от определени политически формации, както и от някои медии. А министърът на правосъдието Андрей Янкулов побърза да свика Висшия съдебен съвет на 26 февруари, за да се избере нов временен главен прокурор и да се реши веднъж завинаги проблемът с нелегитимно заемащия поста Сарафов. Повече по динамично развиващата се тема с новия служебен кабинет четете в „Боят настана“.

Боят настана
Въпросът е кой с кого и срещу кого ще се бие, докато всички ние стоим отстрани с агиткаджийския си манталитет. Едно е ясно – кал ще лети във всички посоки, което е удобно за замазване на очите на хората. От Емилия Милчева.
Седмицата (16–21 февруари)

Над угнетяващите купчини от сериозни и ултрасериозни теми, които ни притискат и задушават от всички страни, и в новия 41-ви епизод Е.Т. прелита като пакостливия Пък, призовава към бойна готовност и за пореден път показва, че има много възможни интонации, с които да изречем гнева, болката и омерзението си.

И за да се върнем там, откъдето тръгнахме, а именно при трудното изкуство да скърбим и да уважаваме чуждата скръб, от сърце ви препоръчвам филма на Клинт Бентли „Сънища с влакове“ (по едноименния роман на Денис Джонсън). Пожелавам ви спокойна и светла събота с едноименната песен от филма на Ник Кейв и Брайс Деснър:

Благодарим ви за подкрепата и грижата! Без тях това, което правим, не би имало смисъл.

Cloudflare outage on February 20, 2026

Post Syndicated from David Tuber original https://blog.cloudflare.com/cloudflare-outage-february-20-2026/

On February 20, 2026, at 17:48 UTC, Cloudflare experienced a service outage when a subset of customers who use Cloudflare’s Bring Your Own IP (BYOIP) service saw their routes to the Internet withdrawn via Border Gateway Protocol (BGP).

The issue was not caused, directly or indirectly, by a cyberattack or malicious activity of any kind. This issue was caused by a change that Cloudflare made to how our network manages IP addresses onboarded through the BYOIP pipeline. This change caused Cloudflare to unintentionally withdraw customer prefixes.

For some BYOIP customers, this resulted in their services and applications being unreachable from the Internet, causing timeouts and failures to connect across their Cloudflare deployments that used BYOIP. A subset of 1.1.1.1, specifically our destination one.one.one.one, was also impacted. The total duration of the incident was 6 hours and 7 minutes with most of that time spent restoring prefix configurations to their state prior to the change.

Cloudflare engineers reverted the change and prefixes stopped being withdrawn when we began to observe failures. However, before engineers were able to revert the change, ~1,100 BYOIP prefixes were withdrawn from the Cloudflare network. Some customers were able to restore their own service by using the Cloudflare dashboard to re-advertise their IP addresses. We resolved the incident when we restored all prefix configurations.

We are sorry for the impact to our customers. We let you down today. This post is an in-depth recounting of exactly what happened and which systems and processes failed. We will also outline the steps we are taking to prevent outages like this from happening again.

How did the outage impact customers?

This graph shows the amount of prefixes advertised by Cloudflare during the incident to a BGP neighbor, which correlates to impact as prefixes that weren’t advertised were unreachable on the Internet:


Out of the total 6,500 prefixes advertised to this peer, 4,306 of those were BYOIP prefixes. These BYOIP prefixes are advertised to every peer and represent all the BYOIP prefixes we advertise globally. 

During the incident, 1,100 prefixes out of the total 6,500 were withdrawn from 17:56 to 18:46 UTC. Out of the 4,306 total BYOIP prefixes, 25% of BYOIP prefixes were unintentionally withdrawn. We were able to detect impact on one.one.one.one and revert the impacting change before more prefixes were impacted. At 19:19 UTC, we published guidance to customers that they would be able to self-remediate this incident by going to the Cloudflare dashboard and re-advertising their prefixes.

Cloudflare was able to revert many of the advertisement changes around 20:20 UTC, which caused 800 prefixes to be restored. There were still ~300 prefixes that were unable to be remediated through the dashboard because the service configurations for those prefixes were removed from the edge due to a software bug. These prefixes were manually restored by Cloudflare engineers at 23:03 UTC. 

This incident did not impact all BYOIP customers because the configuration change was applied iteratively and not instantaneously across all BYOIP customers. Once the configuration change was revealed to be causing impact, the change was reverted before all customers were affected. 

The impacted BYOIP customers first experienced a behavior called BGP Path Hunting. In this state, end user connections traverse networks trying to find a route to the destination IP. This behavior will persist until the connection that was opened times out and fails. Until the prefix is advertised somewhere, customers will continue to see this failure mode. This loop-until-failure scenario affected any product that uses BYOIP for advertisement to the Internet. One.one.one.one, which is a subset of 1.1.1.1, is a prefix onboarded as a BYOIP prefix, and was impacted in this manner. This prefix, which was Cloudflare-maintained but using our own products, allowed us to detect this issue quickly. A full breakdown of the services impacted is below.

Service/Product Impact Description
Core CDN and Security Services Traffic was not attracted to Cloudflare, and users connecting to websites advertised on those ranges would have seen failures to connect
Spectrum Spectrum apps on BYOIP failed to proxy traffic due to traffic not being attracted to Cloudflare
Dedicated Egress Customers who used Gateway Dedicated Egress leveraging BYOIP or Dedicated IPs for CDN Egress leveraging BYOIP would not have been able to send traffic out to their destinations
Magic Transit End users connecting to applications protected by Magic Transit would not have been advertised on the Internet, and would have seen connection timeouts and failures

There was also a set of customers who were unable to restore service by toggling the prefixes on the Cloudflare dashboard. As engineers began reannouncing prefixes to restore service for these customers, these customers may have seen increased latency and failures despite their IP addresses being advertised. This was because the addressing settings for some users were removed from edge servers due an issue in our own software, and the state had to be propagated back to the edge. 

We’re going to get into what exactly broke in our addressing system, but to do that we need to cover a quick primer on the Addressing API, which is the underlying source of truth for customer IP addresses at Cloudflare.

Cloudflare’s Addressing API

The Addressing API is an authoritative dataset of the addresses present on the Cloudflare network. Any change to that dataset is immediately reflected in Cloudflare’s global network. While we are in the process of improving how these systems roll out changes as a part of Code Orange: Fail Small, today customers can configure their IP addresses by interacting with public-facing APIs which configure a set of databases that trigger operational workflows propagating the changes to Cloudflare’s edge. This means that changes to the Addressing API are immediately propagated to the Cloudflare edge.

Advertising and configuring IP addresses on Cloudflare involves several steps:

  • Customers signal to Cloudflare about advertisement/withdrawal of IP addresses via the Addressing API or BGP Control

  • The Addressing API instructs the machines to change the prefix advertisements

  • BGP will be updated on the routers once enough machines have received the notification to update the prefix

  • Finally, customers can configure Cloudflare products to use BYOIP addresses via service bindings which will assign products to these ranges

The Addressing API allows us to automate most of the processes surrounding how we advertise or withdraw addresses, but some processes still require manual actions. These manual processes are risky because of their close proximity to Production. As a part of Code Orange: Fail Small, one of the goals of remediation was to remove manual actions taken in the Addressing API and replace them with safe workflows.

How did the incident occur?

The specific piece of configuration that broke was a modification attempting to automate the customer action of removing prefixes from Cloudflare’s BYOIP service, a regular customer request that is done manually today. Removing this manual process was part of our Code Orange: Fail Small work to push all change towards safe, automated, health-mediated deployment. Since the list of related objects of BYOIP prefixes can be large, this was implemented as part of a regularly running sub-task that checks for BYOIP prefixes that should be removed, and then removes them. Unfortunately, this regular cleanup sub-task queried the API with a bug.

Here is the API query from the cleanup sub-task:

 resp, err := d.doRequest(ctx, http.MethodGet, `/v1/prefixes?pending_delete`, nil)

And here is the relevant part of the API implementation:

	if v := req.URL.Query().Get("pending_delete"); v != "" {
		// ignore other behavior and fetch pending objects from the ip_prefixes_deleted table
		prefixes, err := c.RO().IPPrefixes().FetchPrefixesPendingDeletion(ctx)
		if err != nil {
			api.RenderError(ctx, w, ErrInternalError)
			return
		}

		api.Render(ctx, w, http.StatusOK, renderIPPrefixAPIResponse(prefixes, nil))
		return
	}

Because the client is passing pending_delete with no value, the result of Query().Get(“pending_delete”) here will be an empty string (“”), so the API server interprets this as a request for all BYOIP prefixes instead of just those prefixes that were supposed to be removed. The system interpreted this as all returned prefixes being queued for deletion. The new sub-task then began systematically deleting all BYOIP prefixes and all of their related dependent objects including service bindings, until the impact was noticed, and an engineer identified the sub-task and shut it down.

Why did Cloudflare not catch the bug in our staging environment or testing?

Our staging environment contains data that matches Production as closely as possible, but was not sufficient in this case and the mock data we relied on to simulate what would occur was insufficient.

In addition, while we have tests for this functionality, coverage for this scenario in our testing process and environment was incomplete. Initial testing and code review focused on the BYOIP self-service API journey and were completed successfully. While our engineers successfully tested the exact process a customer would have followed, testing did not cover a scenario where the task-runner service would independently execute changes to user data without explicit input.

Why was recovery not immediate?

Affected BYOIP prefixes were not all impacted in the same way, necessitating more intensive data recovery steps. As a part of Code Orange: Fail Small, we are building a system where operational state snapshots can be safely rolled out through health-mediated deployments. In the event something does roll out that causes unexpected behavior, it can be very quickly rolled back to a known-good state. However, that system is not in Production today.

BYOIP prefixes were in different states of impact during this incident, and each of these different states required different actions:

  • Most impacted customers only had their prefixes withdrawn. Customers in this configuration could go into the dashboard and toggle their advertisements, which would restore service. 

  • Some customers had their prefixes withdrawn and some bindings removed. These customers were in a partial state of recovery where they could toggle some prefixes but not others.

  • Some customers had their prefixes withdrawn and all service bindings removed. They could not toggle their prefixes in the dashboard because there was no service (Magic Transit, Spectrum, CDN) bound to them. These customers took the longest to mitigate, as a global configuration update had to be initiated to reapply the service bindings for all these customers to every single machine on Cloudflare’s edge.

How does this incident relate to Code Orange: Fail Small?

The change we were making when this incident occurred is part of the Code Orange: Fail Small initiative, which is aimed at improving the resiliency of code and configuration at Cloudflare. As a brief primer of the Code Orange: Fail Small initiatives, the work can be divided into three buckets:

  • Require controlled rollouts for any configuration change that is propagated to the network, just like we do today for software binary releases.

  • Change our internal “break glass” procedures and remove any circular dependencies so that we, and our customers, can act fast and access all systems without issue during an incident.

  • Review, improve, and test failure modes of all systems handling network traffic to ensure they exhibit well-defined behavior under all conditions, including unexpected error states.

The change that we attempted to deploy falls under the first bucket. By moving risky, manual changes to safe, automated configuration updates that are deployed in a health-mediated manner, we aim to improve the reliability of the service.

Critical work was already ongoing to enhance the Addressing API’s configuration change support through staged test mediation and better correctness checks. This work was ongoing in parallel with the deployed change. Although preventative measures weren’t fully deployed before the outage, teams were actively working on these systems when the incident occurred. Following our Code Orange: Fail Small promise to require controlled rollouts of any change into Production, our engineering teams have been reaching deep into all layers of our stack to identify and fix all problematic findings. While this outage wasn’t itself global, the blast radius and impact were unacceptably large, further reinforcing Code Orange: Fail Small as a priority until we have re-established confidence in all changes to our network being as gradual as possible. Now let’s talk more specifically about improvements to these systems.

Remediation and follow-up steps

API schema standardization

One of the issues in this incident is that the pending_delete flag was interpreted as a string, making it difficult for both client and server to rationalize the value of the flag. We will improve the API schema to ensure better standardization, which will make it much easier for testing and systems to validate whether an API call is properly formed or not. This work is part of the third Code Orange workstream, which aims to create well-defined behavior under all conditions.

Better separation between operational and configured state

Today, customers make changes to the addressing schema that are persisted in an authoritative database, and that database is the same one used for operational actions. This makes manual rollback processes more challenging because engineers need to utilize database snapshots instead of rationalizing between desired and actual states. We will redesign the rollback mechanism and database configuration to ensure that we have an easy way to roll back changes quickly and also to introduce layers between customer configuration and Production.

We will snap shot the data that we read from the database and are applying to Production, and apply those snapshots in the same way that we deploy all our other Production changes, mediated by health metrics that can automatically stop the deployment if things are going wrong. This means that the next time we have a problem where the database gets changed into a bad state, we can near-instantly revert individual customers (or all customers) to a version that was working.

While this will temporarily block our customers from being able to make direct updates via our API in the event of an outage, it will mean that we can continue serving their traffic while we work to fix the database, instead of being down for that time. This work aligns with the first and second Code Orange workstreams, which involves fast rollback and also safe, health-mediated deployment of configuration.

Better arbitrate large withdrawal actions

We will improve our monitoring to detect when changes are happening too fast or too broadly, such as withdrawing or deleting BGP prefixes quickly, and disable the deployment of snapshots when this happens. This will form a type of circuit breaker to stop any out-of-control process that is manipulating the database from having a large blast radius, like we saw in this incident.

We also have some ongoing work to directly monitor that the services run by our customers are behaving correctly, and those signals can also be used to trip the circuit breaker and stop potentially dangerous changes from being applied until we have had time to investigate. This work aligns with the first Code Orange workstream, which involves safe deployment of changes.

Below is the timeline of events inclusive of deployment of the change and remediation steps:

Time (UTC) Status Description
2026-02-05 21:53 Code merged into system Broken sub-process merged into code base
2026-02-20 17:46 Code deployed into system Address API release with broken sub-process completes
2026-02-20 17:56 Impact Start Broken sub-process begins executing. Prefix advertisement updates begin propagating and prefixes begin to be withdrawn – IMPACT STARTS –
2026-02-20 18:13 Cloudflare engaged Cloudflare engaged for failures on one.one.one.one
2026-02-20 18:18 Internal incident declared Cloudflare engineers continue investigating impact
2026-02-20 18:21 Addressing API team paged Engineering team responsible for Addressing API engaged and debugging begins
2026-02-20 18:46 Issue identified Broken sub-process terminated by an engineer and regular execution disabled; remediation begins
2026-02-20 19:11 Mitigation begins Cloudflare Engineers begin to restore serviceability for prefixes that were withdrawn while others focused on prefixes that were removed
2026-02-20 19:19 Some prefixes mitigated Customers begin to re-advertise their prefixes via the dashboard to restore service. – IMPACT DOWNGRADE –
2026-02-20 19:44 Additional mitigation continues Engineers begin database recovery methods for removed prefixes
2026-02-20 20:30 Final mitigation process begins Engineers complete release to restore withdrawn prefixes that still have existing service bindings. Others are still working on removed prefixes – IMPACT DOWNGRADE –
2026-02-20 21:08 Configuration update deploys Engineering begins global machine configuration rollout to restore prefixes that were not self-mitigated or mitigated via previous efforts – IMPACT DOWNGRADE –
2026-02-20 23:03 Configuration update completed Global machine configuration deployment to restore remaining prefixes is completed. – IMPACT ENDS –

We deeply apologize for this incident today and how it affected the service we provide our customers, and also the Internet at large. We aim to provide a network that is resilient to change, and we did not deliver on our promise to you. We are actively making these improvements to ensure improved stability moving forward and to prevent this problem from happening again.

The collective thoughts of the interwebz