Metasploit Wrap-Up 03/20/2026

Post Syndicated from Brendan Watters original https://www.rapid7.com/blog/post/pt-metasploit-wrap-up-03-20-2026

♫ I Just Called ♫ To Say ♫ 7f45 4c46 0201 0100 0000 0000 0000 0000 0300 3e00 0100♫

This release contains 2 new exploit modules, 2 enhancements, and 7 bug fixes. Community contributor Chocapikk submitted both exploit modules this release: one targeting AVideo-Encoder’s getImage.php file and another targeting FreePBX. Leading the enhancements is a granularization for LDAP queries allowing the omission of SACL data on security descriptors, as without the proper permissions the entire query of the security descriptor will fail if the SACL data is even just a part of the query.

New module content (2)

AVideo Encoder getImage.php Unauthenticated Command Injection

Authors: Valentin Lobstein [email protected] and arkmarta

Type: Exploit

Pull request: #21076 contributed by Chocapikk

Path: linux/http/avideo_encoder_getimage_cmd_injection

AttackerKB reference: CVE-2026-29058

Description: Adds an exploit module for CVE-2026-29058, an unauthenticated OS command injection in AVideo Encoder’s getImage.php endpoint.

FreePBX filestore authenticated command injection

Authors: Cory Billington and Valentin Lobstein [email protected]

Type: Exploit

Pull request: #20719 contributed by Chocapikk

Path: unix/http/freepbx_filestore_cmd_injection

AttackerKB reference: CVE-2025-64328

Description: Adds a new Metasploit exploit module for FreePBX filestore authenticated command injection (CVE-2025-64328) with automatic vulnerable-version detection and full documentation, and renames the XorcomCompletePbx HTTP mixin to CompletePBX updating affected modules accordingly.

Enhancements and features (2)

  • #20730 from zeroSteiner – This update modifies the ldap_query module to skip querying the SACL (System Access Control List) on security descriptors by default. This behavior is now controlled by a new option, LDAP::QuerySacl. This change is necessary when using a non-privileged user to query security descriptors via LDAP; otherwise, querying the SACL will cause the entire query to be blocked, resulting in no security descriptors being returned.
  • #20997 from Nayeraneru – This adds a new OptTimedelta datastore option type. It enables module authors to specify a time duration and users to set it with a human-friendly syntax.

Bugs fixed (7)

  • #20960 from g0tmi1k – This adds a DHCPINTERFACE option to the DHCP server mixin, allowing modules that start that server to specify a particular interface to bind to.
  • #21020 from g0tmi1k – This makes a small change to the docs by removing two lines that were previously duplicated.
  • #21024 from Aaditya1273 – Fixes a bug in the JSON-RPC msfrpcd functionality that incorrectly required SSL certificates to be present even when disabled with msfrpcd -S.
  • #21025 from Hemang360 – Fixes a crash when calling the HTTP cookie jar with non-string values.
  • #21028 from SilentSobs – Fixes a crash when using the reload_all command no module is present.
  • #21081 from Hemang360 – Fixes a crash when using the windows/exec with non-ascii characters.
  • #21139 from jheysel-r7 – This fixes a bug in the ldap_esc_vulnerable_cert_finder module that was preventing authentication from working when making a WinRM connection.

Documentation added (1)

  • #21074 from jeanmtr – Adds documentation for the pop3_login module.

You can always find more documentation on our docsite at docs.metasploit.com.

Get it

As always, you can update to the latest Metasploit Framework with msfupdate and you can get more details on the changes since the last blog post from GitHub:

If you are a git user, you can clone the Metasploit Framework repo (master branch) for the latest. To install fresh without using git, you can use the open-source-only Nightly Installers or the commercial edition Metasploit Pro

Agama 19 released

Post Syndicated from jzb original https://lwn.net/Articles/1064067/

Version
19
of the Agama installer for openSUSE and SUSE has been
released. This release includes major changes in Agama’s architectural
design, organization of the web interface, and more.

We always wanted Agama to follow the schema […] in which the core of
the installer could be controlled through a consistent and simple
programming interface (an API, in developers jargon). In that schema,
the web-based user interface, the command-line tools and the
unattended installation are built on top of that generic API.

But previous versions of Agama were full of quirks that didn’t
allow us to define an API that would match our quality standards as a
solid foundation to build a simple but comprehensive installer. Agama
19 represents a quite significant architectural overhaul, needed to
leave all those quirks behind and to define mechanisms that can be the
cornerstone for any future development.

LWN last looked at
Agama
in September 2025.

[$] A truce in the Manjaro governance struggle

Post Syndicated from jzb original https://lwn.net/Articles/1063717/

Members of the Manjaro Linux distribution’s community have published
a “Manjaro 2.0 Manifesto”
that contains a list of complaints and a demand to restructure the project to provide
a clear separation between the community and Manjaro as a company. The manifesto
asserts that the project’s leadership is not acting in the best interests of the
community, which has caused developers to leave and innovation to stagnate. It
also demands a handover of the Manjaro trademark and other assets to a
to-be-formed nonprofit association. The responses on the Manjaro forum showed widespread support
for the manifesto; Philip Müller, project lead and CEO of the Manjaro
company, largely stayed out of the discussion. However, he surfaced
on March 19 to say he was “open to serious discussions“, but only
after a nonprofit had actually been set up.

NVIDIA DGX Station Systems Available At Last GB300 and GB200 Workstations For Your Desktop

Post Syndicated from Ryan Smith original https://www.servethehome.com/nvidia-dgx-station-systems-available-at-last-gb300-gb200-workstations-for-your-desktop/

While the bulk of NVIDIA’s focus at their most recent GTC trade show has been on their forthcoming Vera Rubin platform for obvious reasons, right now in the present the company is still in the middle of delivering their Grace Blackwell family of offerings. This includes the GB300 Blackwell Ultra accelerator itself, as well as […]

The post NVIDIA DGX Station Systems Available At Last GB300 and GB200 Workstations For Your Desktop appeared first on ServeTheHome.

Negotiating with the Board: Translating Active Risk into Financial Exposure

Post Syndicated from Trevor Christiansen original https://www.rapid7.com/blog/post/pt-translating-active-into-risk-financial-exposure-board-negotiating-vm

Security leaders rarely struggle to produce data. The challenge is turning that data into something the board can use to make decisions.

Walk into a board meeting with a slide showing 1,200 critical vulnerabilities and 44 internet-facing assets, and you will likely see polite acknowledgment rather than meaningful discussion. The question that follows tends to cut through quickly: what does this mean for the business?

Boards allocate capital based on financial exposure, not vulnerability counts. A list of findings describes workload, but directors are responsible for revenue protection, liability, and risk to the balance sheet. When security reporting remains technical, it sits outside the way investment decisions are made elsewhere in the organization. The issue is less about communication and more about framing the problem in terms the business already understands.

From severity to risk

CVSS measures theoretical severity, but it does not measure business risk. A high score indicates that a flaw could be dangerous, yet it does not tell you whether the vulnerability is reachable in your environment, whether exploit code exists, or whether it is likely to affect revenue in the near term. It answers a useful engineering question, but it does not answer the question the board is asking.

That question is about likelihood and impact. Most enterprise risk frameworks define risk in those terms, and that is how financial decisions are made. The gap becomes clear when two vulnerabilities appear similar on a dashboard but carry very different consequences. A high-CVSS issue on a segmented lab system may present little business risk, while a moderately severe vulnerability on an internet-facing production system with active exploit activity can expose regulated data and revenue streams.

What is often missing in that comparison is threat context. Understanding how attackers behave, which vulnerabilities they are exploiting, and where access paths actually exist changes how risk is interpreted. Active Risk in InsightVM brings those elements together by combining exploit telemetry, attacker behavior, and asset context to estimate the likelihood that a vulnerability will be used. When that likelihood is paired with business impact, the conversation shifts toward exposure rather than severity.

From CVSS scores to financial exposure

Prioritization alone does not translate into board-level decisions. Knowing what is most likely to be exploited is necessary, but it is not sufficient when the goal is to justify investment.

FAIR provides a way to bridge that gap. The model defines risk as a combination of how often a loss event is likely to occur and how much that event would cost. In practical terms:

Annualized Loss Exposure (ALE) = Loss Event Frequency × Probable Loss Magnitude

Active Risk informs the likelihood side of that equation by grounding it in observed attacker behavior and exploit activity. FAIR converts that likelihood into financial terms, allowing security teams to describe exposure in a way that aligns with how capital is allocated.

Instead of reporting that a set of vulnerabilities is “high risk,” the discussion becomes more concrete. A team might say that a group of issues represents several million dollars in annualized exposure across systems tied to revenue. That is a number that can be evaluated alongside other business risks, rather than interpreted as a technical signal.

A practical example

Consider two vulnerabilities identified during a scan. The first is a CVSS 9.8 issue on a segmented guest Wi-Fi router. It is severe from a technical standpoint, but it has no access to sensitive data, no path into production systems, and no evidence of active exploitation.

The second is a vulnerability with a moderate CVSS score on an internet-facing customer database. Public exploit code exists, and the system stores regulated data tied directly to revenue and compliance obligations.

On a scanner dashboard, the first may appear more urgent. When viewed through a financial lens, the second carries greater risk.

Assume an annual probability of exploitation of 20 percent for the database scenario. If the potential impact includes $750,000 in incident response, $1.2 million from several days of business interruption, $600,000 in legal and regulatory costs, and $1 million in customer churn and reputational damage, the total loss for a single event is $3.55 million.

Applying the FAIR model results in approximately $710,000 in annualized exposure. That figure reflects the risk carried by that single vulnerability on a production system.

By contrast, even if the Wi-Fi router vulnerability had a 5 percent probability of exploitation and a $50,000 impact, the resulting exposure would be around $2,500. Both findings may appear critical in a technical report, but only one represents a material financial concern.

This is where Active Risk and FAIR work together. One identifies where attackers are likely to act, and the other expresses the consequence in financial terms. The combination changes how vulnerabilities are evaluated and how priorities are set.

Visualizing exposure across your environment

Once risk is expressed in financial terms, the next step is to understand how that exposure is distributed. Boards tend to think in terms of portfolios rather than individual issues, and the same principle applies to cybersecurity.

In most environments, exposure is not evenly spread. A relatively small number of systems and vulnerabilities account for a large portion of potential loss. Internet-facing services, systems tied to revenue, and assets with known exploit activity often sit at the higher end of that distribution.

This creates a practical way to focus effort. Rather than attempting to address every vulnerability equally, teams can identify where exposure is concentrated and reduce risk in those areas first. In many cases, addressing a small number of issues can significantly reduce overall exposure, particularly when those issues sit on systems that are both reachable and business-critical.

A before-and-after view helps make this visible. If an organization reduces modeled exposure from several million dollars to a substantially lower figure through targeted remediation, the result can be explained in terms of reduced downside risk rather than increased patching activity. Over time, tracking that change shows whether investments are producing measurable outcomes.

Making risk actionable

By the time exposure is expressed in financial terms, the discussion in the boardroom has already shifted. The focus moves away from counts and severity toward risk, trade-offs, and acceptable levels of exposure.

One of the first issues that arises in that context is the assumption that risk should be driven to zero. In practice, eliminating all exposure is neither achievable nor economically sensible. Reducing risk always involves trade-offs, and those trade-offs become clearer when expressed in financial terms.

If an organization has already reduced exposure significantly, but further reduction requires a disproportionate increase in cost, the decision becomes one of balance. The question is no longer why risk still exists, but whether the remaining exposure aligns with the organization’s tolerance.

The same logic applies when discussing budget. Requests framed in operational terms, such as additional headcount or tooling, are difficult to evaluate in isolation. When those requests are tied to measurable reductions in exposure, the relationship between cost and benefit becomes clearer.

For example, if additional resources reduce several million dollars of modeled exposure at a fraction of that cost, the investment can be assessed alongside other initiatives using the same financial lens. At that point, the discussion is no longer about capacity. It is about risk reduction.

Putting security in business terms

Reducing exposure also affects how the organization is perceived externally. Cyber insurance underwriting, for example, increasingly considers factors such as attack surface, exploit availability, and remediation speed. Demonstrating that exposure is measured and reduced over time can influence how risk is priced.

The same applies during customer due diligence. Being able to explain where risk exists, how it is prioritized, and how it has been reduced provides evidence of maturity. It shows that security is being managed deliberately rather than reactively.

Aligning to risk tolerance

Productive board discussions tend to end with agreement on acceptable levels of exposure. Without a financial view, every issue can appear urgent. With it, prioritization becomes more grounded.

Leadership can evaluate whether the level of risk being carried is consistent with business objectives, and whether further investment is warranted. That shifts vulnerability management from a process focused on volume to one focused on where exposure is concentrated and how it can be reduced most effectively.

Clear exposure, clearer decisions

Vulnerability management has often been treated as an operational activity centered on patching and scanning. When combined with threat context and financial modeling, it becomes part of enterprise risk management.

Instead of reporting how many vulnerabilities exist, security leaders can describe how much exposure the organization carries. Instead of focusing on activity, they can show how targeted actions reduce risk over time. That framing aligns cybersecurity with the same decision-making process used across the rest of the business.

When exposure is clear, decisions become clearer. Leadership can determine where to accept risk, where to transfer it, and where to invest in reduction. The conversation with the board moves away from technical detail and toward measurable impact, which is where security becomes part of strategy rather than an isolated function.

Security updates for Friday

Post Syndicated from jzb original https://lwn.net/Articles/1063990/

Security updates have been issued by AlmaLinux (capstone, glibc, grub2, kernel, libarchive, libpng, mysql, and python3.11), Debian (evolution-data-server, imagemagick, and snapd), Fedora (bpfman, chromium, cpp-httplib, dotnet10.0, openssh, polkit, and vim), Mageia (graphicsmagick, imagemagick, openssh, and perl-YAML-Syck), Oracle (capstone, grub2, kernel, mysql, and python-pyasn1), Red Hat (container-tools:rhel8, rhc, yggdrasil, and yggdrasil-worker-package-manager), SUSE (cargo1.92, cargo1.93, chromedriver, coturn, curl, freerdp, jq, kernel, libssh, php-composer2, python311-uv, python312, qemu, tomcat, util-linux, vim, and virtiofsd), and Ubuntu (exiv2, freerdp3, glance, linux, linux-aws, linux-aws-hwe, linux-gcp, linux-gcp-4.15, linux-hwe, linux-kvm, linux-oracle, and linux-aws-fips, linux-fips, linux-gcp-fips).

CVE-2026-31381, CVE-2026-31382: Gainsight Assist Information Disclosure and Cross-Site Scripting (FIXED)

Post Syndicated from Christopher O’Boyle original https://www.rapid7.com/blog/post/ve-cve-2026-31381-cve-2026-31382-gainsight-assist-information-disclosure-xss-fixed

Overview

Rapid7 Labs recently identified a chain of security vulnerabilities in the Gainsight Assist plugin and its interactions with the associated domain app.gainsight.com. These vulnerabilities include an Information Disclosure flaw (CVE-2026-31381) and a Reflected Cross-Site Scripting (XSS) vulnerability (CVE-2026-31382). By chaining these vulnerabilities, an attacker can move from passive information gathering to active client-side exploitation.

The XSS vulnerability was remediated by Gainsight via a server side code-level fix on March 6, 2026. A patched update to the Chrome and Outlook plugins to remediate the Information Disclosure were released on March 9, 2026.

Product description

Gainsight Assist is a plugin that allows users to access Gainsight email templates and easily sync inbound and outbound emails to the Timeline within the Gainsight Customer Success (CS) product directly from their email platform.

Credit

These vulnerabilities were discovered and reported to the Gainsight team by Christopher O’Boyle, Cybersecurity Advisor at Rapid7. The vulnerabilities are being disclosed in accordance with Rapid7’s vulnerability disclosure policy. Rapid7 is grateful to the Gainsight team for their assistance and collaboration.

Vulnerability details

CVE

Description

CVSS

CVE-2026-31381

Information Disclosure: An attacker can extract user email addresses (PII) exposed in base64 encoding via the state parameter in the OAuth callback URL.

5.3 (Medium)

CVE-2026-31382

Reflected XSS / HTML Injection: The error_description parameter is vulnerable to Reflected XSS. An attacker can bypass the domain’s WAF using a Safari-specific onpagereveal payload.

6.1 (Medium)

The testing target was the Gainsight Assist plugin and its interactions with the app.gainsight.com domain, used as a callback mechanism that processes authentication data and error descriptions following user login attempts.

CVE-2026-31381: Information disclosure

During testing involving Salesforce and Okta authentication channels, an OAuth callback flow failure was observed. The resulting error message exposed the user’s email address (PII) within a Base64 encoded state parameter in the URL. Because Base64 is merely obfuscation and not encryption, these email addresses can be easily harvested from server logs, proxies, or browser history by third parties.

CVE-2026-31382: Reflected XSS and HTML injection

The Gainsight callback URL contained an error_description parameter that was found to be vulnerable to content spoofing and HTML Injection. While Gainsight employs a Web Application Firewall (WAF) that successfully blocks most standard JavaScript execution, Rapid7 researchers bypassed this protection using a browser-specific payload targeting Safari’s onpagereveal event.

When the victim opens the malicious URL in Safari, the onpagereveal payload executes automatically without further user interaction. By injecting HTML content and spoofing the error page, an attacker can create a legitimate-looking prompt instructing the user to switch to a Safari browser to ensure the payload fires.

<body onpagereveal=open("https://www.rapid7.com")>
We have detected a browser compatibility issue for 
this step, this can only be completed on Safari <br><br>
Please copy the URL from the address bar above and 
paste it in a Safari browser...

Figure 1: Example of the injected HTML payload instructing the user to utilize Safari.

Chaining for Impact

When combined, these vulnerabilities create a high-impact attack path:

  1. Target identification: The login error page includes the user’s attempted login email address in a Base64-encoded state parameter in the URL. Anyone with visibility into that URL (e.g., via the browser address bar, existing access to internal logs, or XSS on that page) can decode the state value to recover the email address. The vulnerability pertains to the data included in the URL rather than granting access to logs or history.

  2. Luring the victim: Using HTML injection on the trusted app.gainsight.com domain, the attacker crafts a highly convincing phishing link to send to the targeted user.

  3. XSS execution: Once the victim opens the link in Safari, the onpagereveal payload executes. Because the payload can recursively call the exact same URL, it can cause an infinite loop leading to client-side resource exhaustion, log flooding, or the delivery of malware.

Vendor statement

“Gainsight values the work of the security research community and appreciates Rapid7’s collaboration. We have fully remediated the identified vulnerabilities through a platform-wide update that strengthens our input validation and WAF configurations. Our forensic investigation found no evidence of exploitation or impact to customer data. We continue to prioritize transparency and supporting our customers to build a more resilient and secure community together. “

Mitigation guidance

As of March 6, 2026, Gainsight has implemented a code-level fix to remediate these findings. Customers should ensure they are utilizing the latest version of the Gainsight Assist plugin.

Disclosure timeline

  • January 30, 2026: Rapid7 makes initial outreach to Gainsight.

  • February 1, 2026: Gainsight confirms outreach and requests details. Rapid7 provides vulnerability details.

  • February 11, 2026: Gainsight confirms receipt, states that the vulnerability has been reproduced, and acknowledges that triage has begun.

  • March 5, 2026: Gainsight and Rapid7 meet to discuss agreed impact, remediation, and next steps.

  • March 6, 2026: Gainsight implements a server-side, code-level fix to remediate the XSS issue.

  • March 9, 2026: Gainsight implements an update to the Chrome and Outlook plugins for the information disclosure vulnerability.

  • March 12, 2026: Gainsight requests disclosure date of March 20, 2026.

  • March 13, 2026: Rapid7 accepts the disclosure date of March 20, 2026.

  • March 20, 2026: This disclosure.

Computing and AI for all: From classrooms to national dialogue in India

Post Syndicated from Mamta Manaktala original https://www.raspberrypi.org/blog/computing-and-ai-for-all-from-classrooms-to-national-dialogue-in-india/

On 7 February 2026, we witnessed something special in Bhubaneswar: a day where classroom experience took centre stage in conversations about computing and AI education in India at our Computing and AI for All conference.

At this event, we brought together educators, researchers, and system leaders for shared learning, reflection, and professional dialogue where classroom practice was placed at the heart of the discussion.

An event to connect conversations

Across computing and AI education in India, conversations about policy, pedagogy, research, and future readiness often take place in separate spaces. At Computing and AI for All, we wanted to connect these strands in a national forum where classroom practitioners, researchers, and system leaders could speak with each other about computing and AI education. 

The theme ‘Computing and AI for all: Classroom to policy’ connected:

  • Ecosystem and policy
  • Pedagogy and innovation
  • AI and future readiness
  • Research, evidence and impact

Bringing these conversations together helped shape a richer dialogue than we’ve seen in many traditional education forums.

Grounding in classroom experience

The response from educators to the opportunity to attend was both energising and encouraging. 165 educators and 22 invited guests joined us on the day. The event also featured 35 paper and poster presentations, reflecting the breadth of classroom practice, research, and innovation that teachers were keen to share.

An aerial view of the attendees of the Computing and AI for All event.

The conference also brought together diverse contributors. Alongside teachers, there was representation from organisations such as Computer Science Teachers’ Associations, the University of Southampton, Quest Alliance, and Learning Links Foundation. Also part of the event were attendees from government bodies including TTWREIS – Telangana Tribal Welfare Residential Educational Institutes; PSSS – Panchasakha Shikhya Setu Sangathan (a Govt. of Odisha initiative); and IIITMK Kerala. The presence of an international delegate from Nepal further broadened the exchange of perspectives.

First and foremost, the event was rooted in classroom practice. 88% of participants were ICT teachers, with strong representation from Odisha and Telangana, where we have been working on partnerships to support computing educators. A significant proportion of the educators present were attending a computing conference in person for the first time, so it was a milestone in their professional journey.

Teachers present a poster at the Computing and AI for All event.

Teachers stepping onto a national stage to present their classroom experiments, research, and reflections signaled an important shift: computing and AI education in India is no longer something designed solely for teachers. It is increasingly being shaped by teachers and informed by their lived classroom realities and professional expertise.

Ideas, evidence, and inspiration

The conference featured a full-day academic programme beginning with a formal inauguration and an introduction to our work in India.

The group performing the formal inauguration of the Computing and AI for All event.

The keynote by the Foundation’s Director of Research and Impact, Shuchi Grover — ‘Democratizing the future: Why computing and AI literacy matters for every child, everywhere”’ — set an equity-driven, future-focused tone. It reminded us that AI literacy is not optional; it’s foundational for learning and opportunity.

In the panel discussion on ‘Embedding ethical AI in everyday teaching: From principles to classroom practice’, teachers, researchers and leaders exchanged ideas about how to integrate responsible AI without losing sight of learning goals.

Across sessions, teachers and researchers shared a wide range of classroom-centred insights, including:

  • Using generative AI for algebraic discovery
  • Designing AI-supported personalised learning models
  • Inclusive coding practices with Scratch
  • Using AI to reduce teacher workload
  • Low-cost computing innovations in K–12 classrooms
  • Hands-on technical explorations, such as Kali Linux on Raspberry Pi

These sessions reflected both pedagogical creativity and deep engagement with everyday classroom challenges.

The structured poster gallery created space for peer feedback and deeper discussion around critical themes such as:

  • AI for first-generation learners
  • Classroom research evidence
  • AI tools to support ICT engagement

What educators told us

In the feedback we captured, educators described the conference as:

  • Highly educative
  • A space for cross-state collaboration
  • A confidence-building platform for presenting work

One participant shared: “I’m honoured to be part of [the conference]. It was inspiring to listen to educational experts sharing valuable insights on AI tools and education.”

Many participants asked us to expand the event, suggesting a two-day format, extended hands-on AI workshops, and practical demonstrations they could take back to their classrooms.

Looking ahead

Our Computing and AI for All event formed a national, educator-led platform linking classroom innovation, AI literacy, research evidence, and policy dialogue.

What stood out most was the spirit of participation:

  • Teachers sharing real classroom experiments
  • Researchers listening
  • Policy voices engaging
  • First-time presenters stepping into a national forum

If AI is shaping the future of work and society, then teachers need to play a part in shaping the future of AI education. Computing and AI for All was one move in this direction that we were pleased to facilitate — and it is only the beginning.

The post Computing and AI for all: From classrooms to national dialogue in India appeared first on Raspberry Pi Foundation.

Коремно възлизане през април

Post Syndicated from Емилия Милчева original https://www.toest.bg/koremno-vuzlizane-prez-april/

Коремно възлизане през април

Номинираният за победител на изборите Румен Радев започна да спада още преди да е победил на 19 април. 

Поне това показват данните на социологическата агенция „Маркет Линкс“ тази седмица, които дават преднина в рамките на статистическата грешка от 3% на Радев и „Прогресивна България“ пред ГЕРБ–СДС. Резултатът изглежда така: 21,1% срещу 18,6%.

Преди месец същата агенция му даде 25,6%, а след това „Алфа Рисърч“ го изстреля на 32,6%, докато ГЕРБ отстоеше на разстояние почти една партия със своите 19,7%.

Отечество любезно, аз ще те спася!
Точно преди 40 години Тина Търнър изпя We don’t need another hero. Колко продължения на реалност а ла „Лудия Макс“ са ни необходими, за да спрем да повтаряме същата грешка? Емилия Милчева за новия спасител, задаващ се на хоризонта, и за останалите месии, които играят като за последно десет.
Коремно възлизане през април

Еуфорията от появата на новия спасител изпарява ли се, заземява ли се, или има нужда от презареждане? Румен Радев все така прилича на президент, прекъснал предсрочно мандата си, но не и на политик, решен да докаже на българските граждани, че си струва да се гласува за него. На „Дондуков“ 2 той нямаше конкуренция. Ролята му не изискваше да се включва по злободневни теми и слоганите от типа „Битка с олигархията!“ минаваха за позиция, без да се налага да бъдат политика. 

Но вън от Президентството въпросите вече са: кои водят битката, как я водят, с кого точно. А Радев още не може да смени костюма на президент, както Симеон Сакскобургготски докрая си остана Царя – нито със, нито без корона.

Има време за висок резултат. Също и за по-нисък

До изборите има месец и Радев би могъл да повиши резултата си. Зависи колко ще е смел в програмата и в публичните си изяви. Пряк дебат с Борисов ще е особено интересен, но залозите ще са в полза на обиграния и речовит лидер на ГЕРБ, който има вроден талант за политическо шоу. Радев след две минути слушане доскучава.

Програмата на „Прогресивна България“, от която на 19 март бяха представени два приоритета –

демонтаж на олигархичния модел и ускорено икономическо развитие на страната, –

е микс от всички добри намерения, заявени и написани в някоя институционална стратегия, плюс няколко нови идеи, като разработването на изкуствен интелект срещу корупцията в обществените поръчки. На тази иновация се разчита да блокира съмнителни търгове, да информира медиите и да сезира компетентните органи за тях.

Останалото е дежавю. И подобно на други предишни програми, „Прогресивна България“ остава на ниво общи цели, без да е ясен инструментариумът за постигането им. Радикално нов модел не се предлага. Но Радев поне недвусмислено посочи към какво се стреми: квалифицирано мнозинство от 160 депутати около „Прогресивна България“, което да смени състава на Висшия съдебен съвет и да бъде избран нов главен прокурор. 

Знаете обаче, че това може и да не се случи. Зависи от математиката в Народното събрание, зависи от силите, които движат присъстващите в него. Няма да останем заложник на евентуален политически блокаж, а ще продължим да работим паралелно – за реформиране на службите и МВР, за разкриване и пресичане на корупционни схеми, за събиране на годни за съда доказателства, защото все някога ще им дойде времето. За още по-интензивна работа с нашите партньорски служби и правителства, така че да установяваме активи, банкови сметки, имоти на български граждани зад граница с незаконен произход.

От ключово значение е математиката, защото тя ще определи колко коалиционни партньори ще са му необходими и кои. Дали ще са двама, или трима? „Прогресивна България“, ПП–ДБ и БСП – Обединена левица (ако прескочи 4-те процента) или „Прогресивна България“ и ГЕРБ–СДС (и Пеевски през задния вход) – никой не затваря вратата за Радев и Радев не затваря вратата за никого. 

Мълчанието на Радев струва злато
Наблюдаваме феномена „Румен Радев на Шрьодингер“. Докато мълчи, той едновременно е срещу корупцията и мафията в правосъдието, за диалог с Русия, против еврото, за ЕС… Това ще свърши с обявяването на партийните му листи. И после? Коментар от Емилия Милчева.
Коремно възлизане през април

Още кандидати за протестен вот

Докато Радев разгръща настъплението си, а дръжките на томахавките, заровени от ПП и ДБ покрай реденето на листите им, още стърчат, на терена се нареждат и други кандидати за протестния вот. 

Сред тях е Антикорупционният блок, обединяващ четири формации. Едната е „Единение“, чийто лидер Иван Христанов беше заместник-министър на земеделието в правителството на Кирил Петков (ПП), а сега е служебен министър на земеделието в кабинета на Андрей Гюров. Там е и „Ние идваме“ на някогашното острие на БСП Мая Манолова, бивша омбудсманка. Другата партия е Зелено движение, което беше в „Демократична България“ до 2024 г. В Антикорупционния блок влиза и „Средна европейска класа“ (СЕК) на бургаския предприемач Константин Бачийски. Кирил Петков и Асен Василев използваха регистрацията на СЕК, както и на партия ВОЛТ, за да се явят като коалиция на изборите през 2021 г., защото още не бяха учредили партията си „Продължаваме промяната“.

Появи се и „Сияние“ – гражданско движение, регистрирано от Николай Попов, бащата на убитата на пътя 12-годишна Сияна. То има подкрепата на организацията „Ангели на пътя“, която обединява близки и роднини на жертви на войната по пътищата, а бързото събиране на подписи показа и по-широка подкрепа. Само за няколко часа бяха събрани над 6500 подписа.

Чашата преля и повече няма да се молим на никого. Ще направим така, че ние да променяме, а не да чакаме промяна, която не идва никога. Не за власт, а за живот. За справедливост. Последвайте ме.

Николай Попов

За да влезе в парламента, Попов е в съдружие със същата тази партия ВОЛТ на Настимир Ананиев, мандатоносител на „Продължаваме промяната“. ПП и ВОЛТ се разделиха преди изборите през октомври 2024 г. Ананиев, който за пореден път прелита към нова формация, е известен и като човека, патентовал марката „Реформаторски блок“ през 2014 г.

За следващите избори пък се формира партия „Ускорение“, също от политици, преминали през други партии. Нейни създатели са служебният министър на енергетиката и кмет на столичния район „Средец“ Трайчо Трайков и бившият депутат и бивш лидер на Зелено движение, който днес е кандидат за депутат от ПП–ДБ Владислав Панев.

Със или без утре – ние решаваме
Ще има ли достатъчно гориво в гражданската енергия, заляла площадите, за да просветли бъдещето и то да бъде „отворено“ за всички, а не да се реализира като поредното празно обещание? От Емилия Милчева.
Коремно възлизане през април

Листи с медалисти

Тази седмица „Прогресивна България“ прикова вниманието върху себе си с листите си, в които има известни спортисти, като волейболната звезда Владимир Николов, олимпийската шампионка по карате Ивет Горанова, плувеца Петър Стойчев, гимнастика Йордан Йовчев. Какво изниква в умовете ни, когато чуем име на успял спортист, донесъл медали и слава на България? Националното знаме и химнът, който пее просълзен на стълбицата. Няма по-силен патриотичен символ, далеч по-добър от кресливите възгласи на „Възраждане“ за лева.

На фона на листите на системните партии, пълни предимно с партократи на избираеми места, Радев изненадва. Макар че и при него има стари муцуни – бивши министри от служебните му кабинети, символ на периода, в който осигуряваха политическа стабилност, също и хора, преминали през други партии, бивши военни и представители на местната власт. 

Успелите спортисти са национално известни лица и не е необходимо време за изграждане на политически образ. Затова се използват като електорален магнит, особено от формации, които искат бързо да мобилизират подкрепа.

Освен това нямат изградена политическа мрежа и зависят от този, който ги е издигнал. 

Хората ги свързват с успех, медали, гордост. За разлика от повечето професионални политици, те биват възприемани като личности, успели с честен труд, не с връзки или политически комбинации, а и се предполага, че са дисциплинирани. За общество, в което доверието в политиците е ниско, това е голямо предимство. 

С първите си телевизионни участия като политици обаче някои от тях предизвикаха присмех. Владо Николов, чиито „лични убеждения са много десни“, повлече крак по БНТ.

Не трябва да делим на Изток и на Запад. Ние сме Европа. Искам България да има мнение накъде да поеме Европа, а не да правим каквото каже Меркел [канцлерка на Германия до 2021 г. – б.а.]. 

Друг, като Петър Стойчев, водач на листата в Смолян, мотивира по bTV избора на Радев така: 

Всеки един от нас е победител и е минал през много трудности, за да защити името си и честта на родината.

Листите на останалите формации са израз на стремежа да се осигурят избираеми места на партийното ръководство, тъй като с появата на Радев всички губят депутати, а някои ще останат и пред вратите на 52-рия парламент. Корекциите не са особено значими, но две промени произведоха шум – санкционираният за корупция по „Магнитски“ Владислав Горанов поведе листата на ГЕРБ–СДС във Варна, а настоящият депутат от ПП–ДБ Явор Божанков изпадна от листите на коалицията. Очевидно Бойко Борисов залага на познати лица и управленски опит вместо на нови фигури. За публичния образ и антикорупционната легитимност не се тревожи особено, нито за санкциите по „Магнитски“, които, изглежда, че безпокоят единствено санкционираните – пречат им на бизнеса. 

Отпадането на Божанков, също и на Даниел Лорер, беше поставено като условие от лидера на ПП Асен Василев през януари. Божанков е от гражданската квота и му беше поискана оставката като депутат, а Лорер беше изключен от ПП, след като двамата отказаха да гласуват заедно с „Възраждане“ кандидатурата на Силви Кирилов от „Има такъв народ“ за шеф на парламента. 

Предизборната кампания ще мине бързо и се очертава да е ожесточена предвид битката между първия и втория („Прогресивна България“ и ГЕРБ–СДС), както и между третия и четвъртия (ПП–ДБ и ДПС). Големият въпрос сега е каква коалиция очаква България и как ще се нарича следващият главен прокурор. 

Внимание, Румен Радев възлиза на лоста. 

Proton Mail Shared User Information with the Police

Post Syndicated from Bruce Schneier original https://www.schneier.com/blog/archives/2026/03/proton-mail-shared-user-information-with-the-police.html

404 Media has a story about Proton Mail giving subscriber data to the Swiss government, who passed the information to the FBI.

It’s metadata—payment information related to a particular account—but still important knowledge. This sort of thing happens, even to privacy-centric companies like Proton Mail.

Равносметка за 51-ото Народно събрание

Post Syndicated from Bozho original https://blog.bozho.net/blog/4572

Вчера беше последният работен ден на това Народно събрание.

Малко числа: вносител (и съавтор) съм на 77 законопроекта. И на 40 предложения между четенията. Задал съм 172 въпроса на министри. Имам близо 400 изказвания в зала и в комисии.

Но числата не са най-важните за равносметката – важно е съдържанието и политическият прочит.

От съдържанието, като най-голямо усилие, мога да откроя закона за изкуствения интелект – технология, с която България изостава на всички фронтове, и е нужно насърчителни мерки на законово ниво, за да „не изпуснем влака“. Такъв позитивен дневен ред е нужен в политическия живот, защото непрестанното плюене не носи просперитет.

А относно политическия прочит – този парламент можеше да сложи санитарен кордон около пълзящото завладяване на държавата. И мнозинството в него отказа да го направи. Затова санитарния кордон го сложиха гражданите, които с многохилядни протести свалиха правителството, съучастващо в завладяването на институциите, така че те все повече да работят за частни интересни и все по-малко – за обществения интерес.

Едно от много малкото позитивни неща, което се случиха, беше еврозоната – успех, дължащ се на работата на много правителства, а и на нашия натиск върху разколебаното правителство в началото на мандата.

Извън това, дневният ред беше доминиран почти изцяло от желанията на Пеевски да я преяде с власт. И това свърши както винаги свършва – с протести, които го свалиха, и нелепи обяснения от трибуната как протестите били на Сорос.

Много от цитираните по-горе законопроекти бяха или с антикорупционна цел, или с цел дигитализация и модернизация, или с цел намаляване на административната тежест и провеждане на десни политики. Почти нито един от тях не беше приет.

На тези избори търсим подкрепа, така че всички нереализирани правилни, десни, антикорупционни политики да намерят мнозинство и да се случат. За да живеят българските граждани по-добре.

#7 #105

Материалът Равносметка за 51-ото Народно събрание е публикуван за пръв път на БЛОГодаря.

На второ четене: „Родословно дърво“

Post Syndicated from Стефан Иванов original https://www.toest.bg/na-vtoro-chetene-rodoslovno-durvo/

„Родословно дърво“ от Мария Роса Лохо

На второ четене: „Родословно дърво“

превод от испански Елица Колева, София: изд. „ЖАР-Жанет Аргирова“, 2020

Това на пръв поглед е роман, който изглежда като семейна хроника, но постепенно се оказва нещо далеч по-сложно. Това е литературен опит да се мисли за паметта, идентичността и историята чрез фигурата на рода. В основата му стои простият, почти антропологичен въпрос доколко човек принадлежи на себе си и доколко е резултат от една верига от предходни животи, истории и митове. Още в началото разказвачката поставя този въпрос като своеобразна програма на текста. Родът не е единственото дърво, а цяла вселена от съдби, които се преплитат, отразяват и повтарят.

В този смисъл романът се движи по една граница, позната от големите семейни епоси на световната литература. Неговият генетичен код може да бъде проследен едновременно в европейската традиция на родовата хроника и в латиноамериканския магически реализъм. От европейската страна стои моделът на романи като „Буденброкови“ на Томас Ман, мащабна история на семейство, в която времето постепенно разрушава социалните и моралните структури на рода. При Ман родът е икономическа и културна институция, която се разпада под натиска на модерността. При Лохо обаче родът не се разпада толкова, колкото се разклонява и митологизира. Той не е буржоазна институция, а почти природен феномен, нещо като биологична джунгла от характери, странности и съдби.

Тук се появява и другият важен литературен паралел, традицията на латиноамериканската фамилна сага, чийто най-емблематичен пример е „Сто години самота“ на Габриел Гарсия Маркес. В романа на Маркес родът Буендия живее в пространство, където историята, митът и магията са почти неразличими. Подобна атмосфера се усеща и при Лохо, макар и в по-тиха и по-иронична форма. Историите за предците, за урочасани жени, мистични свещеници и своенравни матриарси звучат едновременно като семейни анекдоти и като фолклорни легенди. В тях реалността постоянно се плъзга към фантастичното. Например разказът за Маруха може да бъде прочетен и като история за психическо страдание, и като легенда за магическо проклятие.

Тази двусмисленост е една от най-интересните стратегии на романа. Той не се опитва да обясни света рационално, но и не настоява на магическото. Вместо това оставя читателя в едно пространство на колебание, където суеверието, религията и психологията съществуват едновременно. В този смисъл романът се родее и с творчеството на Уилям Фокнър, където историята на рода също се разказва чрез фрагменти, слухове и полузабравени истории. И при Фокнър, и при Лохо

родовата памет се оказва ненадежден архив, нещо средно между документ и мит.

Особено силна линия в романа е образът на жените, които се оказват истинските носители на родовата енергия. Докато официалната история обикновено поставя мъжете в центъра и те са войници, свещеници, политически фигури, то семейната история на Лохо постепенно разкрива, че именно жените са архитектите на рода. Прапрабабата Мария Антония е почти легендарна фигура, дребна, но властна жена, която управлява имота си с твърда ръка, стреля по бирници и раздава жито на бедните.

Този образ може да бъде поставен в интересен диалог с други литературни матриарси. Той напомня например на Урсула от „Сто години самота“, но и на героините на Федерико Гарсия Лорка, особено в „Къщата на Бернарда Алба“. При Лорка женската власт е едновременно трагична и жестока, тя съществува в рамките на патриархалния ред, но често се превръща в негово най-строго продължение. Подобно напрежение се усеща и в романа на Лохо. Жените са подчинени на социалните правила, но същевременно са тези, които ги налагат и поддържат.

На второ четене: „Родословно дърво“

Повествователната структура на книгата също е показателна за нейния постмодерен характер. Разказвачката постоянно напомня, че много от историите са непълни, откъслечни или предадени чрез слухове. Родът е архив на устната традиция, която се променя с всяко ново поколение. Тук романът влиза в диалог с писатели като Хорхе Луис Борхес и В. Г. Зебалд, които също превръщат паметта в литературен лабиринт. При Борхес миналото често се оказва мрежа от въображаеми текстове, а при Зебалд е архив от отломки, снимки и полуизчезнали истории. „Родословно дърво“ използва сходна техника и семейната памет се превръща в своеобразна библиотека от призрачни разкази.

В исторически план романът проследява и миграцията, един от големите процеси на модерността. Галисийците, които напускат Испания и се отправят към Южна Америка, са част от огромна вълна от европейски емигранти в края на XIX и началото на XX век. За тях Аржентина е обещание за нов живот, но и пространство, в което старата идентичност постепенно се размива. В този аспект книгата може да бъде сравнена с литературата на изгнанието, например с творчеството на Владимир Набоков или Милан Кундера, макар контекстът да е различен. И при тях миграцията създава чувство за раздвоена принадлежност, човек никога не е напълно у дома си нито в старата, нито в новата страна.

Тази тема има особено силен отзвук, ако се опитаме да прочетем романа през перспективата на съвременността. В началото на XXI век идеята за родословие преживява своеобразен ренесанс благодарение на генетичните тестове, дигиталните архиви и социалните мрежи. Хиляди хора по света се опитват да реконструират семейната си история чрез ДНК анализи и генеалогични бази данни. Романът на Лохо обаче подсказва нещо, което тези технологии често забравят.

Произходът е разказ. Гените могат да покажат откъде идват телата ни, но не могат да възстановят историите, които са изградили нашата идентичност.

Тук се появява и един любопитен паралел с българската действителност. Галисийският свят, описан в романа – суров, селски, белязан от бедност и силни семейни структури – изненадващо напомня на българското село от началото на XX век. И в двата случая виждаме общество, в което религиозните вярвания, суеверията и семейната солидарност играят решаваща роля. Болестите се обясняват чрез магия или проклятие, а социалният ред се поддържа чрез авторитета на рода.

Разликата е, че за галисийците голямото бягство е към Америка, докато за българите през последните десетилетия посоката е към Западна Европа. Но логиката е сходна. Икономическата периферия изтласква хората навън, а семейните истории се разкъсват между различни континенти. В този смисъл романът на Лохо може да бъде прочетен и като своеобразно огледало на съвременната демографска съдба на България – страна, в която милиони семейни истории също са разделени между различни държави.

Най-ироничният и може би най-философски момент в книгата е начинът, по който тя се отнася към самата идея за родословие. От една страна, родът е представен като мощна сила, нещо, което определя характера, съдбата и дори странностите на героите. От друга страна, самият разказ показва колко хаотична и случайна е тази система. Родословното дърво се оказва борхесов лабиринт от истории, в който всяко поколение се опитва да открие някакъв смисъл.

Така книгата постепенно се превръща в роман не толкова за миналото, колкото за нашия начин да го разказваме. Родът е литературна конструкция, история, която постоянно се пренаписва. И ако книгата оставя някакъв философски извод, той също е леко ироничен. Човек може да търси корените си безкрайно, но накрая неизбежно открива, че

всяко родословно дърво е съставено от паднали клони, пропуснати истории и няколко упорити легенди, които семейството отказва да забрави.


Никой от нас не чете единствено най-новите книги. Тогава защо само за тях се пише? „На второ четене“ е рубрика, в която отваряме списъците с книги, публикувани преди поне година, четем ги и препоръчваме любимите си от тях. За нея медията „Тоест“ е отличена с Националната награда „Христо Г. Данов“ (2025) за принос в представянето на българската книга.

Рубриката е част от партньорската програма Читателски клуб „Тоест“, благодарение на която активните дарители на „Тоест“ получават 20% отстъпка от коричната цена на всички книги на включените издателства. Изборът на заглавия обаче е единствено на авторите Стефан Иванов, Севда Семер и Антония Апостолова, които биха ви препоръчали тези книги и ако имаше как да се разходите с тях в книжарницата. 

Powering the agents: Workers AI now runs large models, starting with Kimi K2.5

Post Syndicated from Michelle Chen original https://blog.cloudflare.com/workers-ai-large-models/

We’re making Cloudflare the best place for building and deploying agents. But reliable agents aren’t built on prompts alone; they require a robust, coordinated infrastructure of underlying primitives.

At Cloudflare, we have been building these primitives for years: Durable Objects for state persistence, Workflows for long running tasks, and Dynamic Workers or Sandbox containers for secure execution. Powerful abstractions like the Agents SDK are designed to help you build agents on top of Cloudflare’s Developer Platform.

But these primitives only provided the execution environment. The agent still needed a model capable of powering it. 

Starting today, Workers AI is officially in the big models game. We now offer frontier open-source models on our AI inference platform. We’re starting by releasing Moonshot AI’s Kimi K2.5 model on Workers AI. With a full 256k context window and support for multi-turn tool calling, vision inputs, and structured outputs, the Kimi K2.5 model is excellent for all kinds of agentic tasks. By bringing a frontier-scale model directly into the Cloudflare Developer Platform, we’re making it possible to run the entire agent lifecycle on a single, unified platform.

The heart of an agent is the AI model that powers it, and that model needs to be smart, with high reasoning capabilities and a large context window. Workers AI now runs those models.

The price-performance sweet spot

We spent the last few weeks testing Kimi K2.5 as the engine for our internal development tools. Within our OpenCode environment, Cloudflare engineers use Kimi as a daily driver for agentic coding tasks. We have also integrated the model into our automated code review pipeline; you can see this in action via our public code review agent, Bonk, on Cloudflare GitHub repos. In production, the model has proven to be a fast, efficient alternative to larger proprietary models without sacrificing quality.

Serving Kimi K2.5 began as an experiment, but it quickly became critical after reviewing how the model performs and how cost-efficient it is. As an illustrative example: we have an agent that does security reviews of Cloudflare’s codebases. This agent processes over 7B tokens per day, and using Kimi, it has caught more than 15 confirmed issues in a single codebase. Doing some rough math, if we had run this agent on a mid-tier proprietary model, we would have spent $2.4M a year for this single use case, on a single codebase. Running this agent with Kimi K2.5 cost just a fraction of that: we cut costs by 77% simply by making the switch to Workers AI.

As AI adoption increases, we are seeing a fundamental shift not only in how engineering teams are operating, but how individuals are operating. It is becoming increasingly common for people to have a personal agent like OpenClaw running 24/7. The volume of inference is skyrocketing.

This new rise in personal and coding agents means that cost is no longer a secondary concern; it is the primary blocker to scaling. When every employee has multiple agents processing hundreds of thousands of tokens per hour, the math for proprietary models stops working. Enterprises will look to transition to open-source models that offer frontier-level reasoning without the proprietary price tag. Workers AI is here to facilitate this shift, providing everything from serverless endpoints for a personal agent to dedicated instances powering autonomous agents across an entire organization.

The large model inference stack

Workers AI has served models, including LLMs, since its launch two years ago, but we’ve historically prioritized smaller models. Part of the reason was that for some time, open-source LLMs fell far behind the models from frontier model labs. This changed with models like Kimi K2.5, but to serve this type of very large LLM, we had to make changes to our inference stack. We wanted to share with you some of what goes on behind the scenes to support a model like Kimi.

We’ve been working on custom kernels for Kimi K2.5 to optimize how we serve the model, which is built on top of our proprietary Infire inference engine. Custom kernels improve the model’s performance and GPU utilization, unlocking gains that would otherwise go unclaimed if you were just running the model out of the box. There are also multiple techniques and hardware configurations that can be leveraged to serve a large model. Developers typically use a combination of data, tensor, and expert parallelization techniques to optimize model performance. Strategies like disaggregated prefill are also important, in which you separate the prefill and generation stages onto different machines in order to get better throughput or higher GPU utilization. Implementing these techniques and incorporating them into the inference stack takes a lot of dedicated experience to get right. 

Workers AI has already done the experimentation with serving techniques to yield excellent throughput on Kimi K2.5. A lot of this does not come out of the box when you self-host an open-source model. The benefit of using a platform like Workers AI is that you don’t need to be a Machine Learning Engineer, a DevOps expert, or a Site Reliability Engineer to do the optimizations required to host it: we’ve already done the hard part, you just need to call an API.

Beyond the model — platform improvements for agentic workloads

In concert with this launch, we’ve also improved our platform and are releasing several new features to help you build better agents.

Prefix caching and surfacing cached tokens

When you work with agents, you are likely sending a large number of input tokens as part of the context: this could be detailed system prompts, tool definitions, MCP server tools, or entire codebases. Inputs can be as large as the model context window, so in theory, you could be sending requests with almost 256k input tokens. That’s a lot of tokens.

When an LLM processes a request, the request is broken down into two stages: the prefill stage processes input tokens and the output stage generates output tokens. These stages are usually sequential, where input tokens have to be fully processed before you can generate output tokens. This means that sometimes the GPU is not fully utilized while the model is doing prefill.

With multi-turn conversations, when you send a new prompt, the client sends all the previous prompts, tools, and context from the session to the model as well. The delta between consecutive requests is usually just a few new lines of input; all the other context has already gone through the prefill stage during a previous request. This is where prefix caching helps. Instead of doing prefill on the entire request, we can cache the input tensors from a previous request, and only do prefill on the new input tokens. This saves a lot of time and compute from the prefill stage, which means a faster Time to First Token (TTFT) and a higher Tokens Per Second (TPS) throughput as you’re not blocked on prefill.

Workers AI has always done prefix caching, but we are now surfacing cached tokens as a usage metric and offering a discount on cached tokens compared to input tokens. (Pricing can be found on the model page.) We also have new techniques for you to leverage in order to get a higher prefix cache hit rate, reducing your costs.

New session affinity header for higher cache hit rates

In order to route to the same model instance and take advantage of prefix caching, we use a new x-session-affinity header. When you send this header, you’ll improve your cache hit ratio, leading to more cached tokens and subsequently, faster TTFT, TPS, and lower inference costs.

You can pass the new header like below, with a unique string per session or per agent. Some clients like OpenCode implement this automatically out of the box. Our Agents SDK starter has already set up the wiring to do this for you, too.

curl -X POST \
"https://api.cloudflare.com/client/v4/accounts/{ACCOUNT_ID}/ai/run/@cf/moonshotai/kimi-k2.5" \
  -H "Authorization: Bearer {API_TOKEN}" \
  -H "Content-Type: application/json" \
  -H "x-session-affinity: ses_12345678" \
  -d '{
    "messages": [
      {
        "role": "system",
        "content": "You are a helpful assistant."
      },
      {
        "role": "user",
        "content": "What is prefix caching and why does it matter?"
      }
    ],
    "max_tokens": 2400,
    "stream": true
  }'

Redesigned async APIs

Serverless inference is really hard. With a pay-per-token business model, it’s cheaper on a single request basis because you don’t need to pay for entire GPUs to service your requests. But there’s a trade-off: you have to contend with other people’s traffic and capacity constraints, and there’s no strict guarantee that your request will be processed. This is not unique to Workers AI — it’s evidently the case across serverless model providers, given the frequent news reports of overloaded providers and service disruptions. While we always strive to serve your request and have built-in autoscaling and rebalancing, there are hard limitations (like hardware) that make this a challenge.

For volumes of requests that would exceed synchronous rate limits, you can submit batches of inferences to be completed asynchronously. We’re introducing a revamped Asynchronous API, which means that for asynchronous use cases, you won’t run into Out of Capacity errors and inference will execute durably at some point. Our async API looks more like flex processing than a batch API, where we process requests in the async queue as long as we have headroom in our model instances. With internal testing, our async requests usually execute within 5 minutes, but this will depend on what live traffic looks like. As we bring Kimi to the public, we will tune our scaling accordingly, but the async API is the best way to make sure you don’t run into capacity errors in durable workflows. This is perfect for use cases that are not real-time, such as code scanning agents or research agents.

Workers AI previously had an asynchronous API, but we’ve recently revamped the systems under the hood. We now rely on a pull-based system versus the historical push-based system, allowing us to pull in queued requests as soon as we have capacity. We’ve also added better controls to tune the throughput of async requests, monitoring GPU utilization in real-time and pulling in async requests when utilization is low, so that critical synchronous requests get priority while still processing asynchronous requests efficiently.

To use the asynchronous API, you would send your requests as seen below. We also have a way to set up event notifications so that you can know when the inference is complete instead of polling for the request. 

// (1.) Push a request in queue
// pass queueRequest: true
let res = await env.AI.run("@cf/moonshotai/kimi-k2.5", {
  "requests": [{
    "messages": [{
      "role": "user",
      "content": "Tell me a joke"
    }]
  }, {
    "messages": [{
      "role": "user",
      "content": "Explain the Pythagoras theorem"
    }]
  }, ...{<add more requests in a batch>} ];
}, {
  queueRequest: true,
});


// (2.) grab the request id
let request_id;
if(res && res.request_id){
  request_id = res.request_id;
}
// (3.) poll the status
let res = await env.AI.run("@cf/moonshotai/kimi-k2.5", {
  request_id: request_id
});

if(res && res.status === "queued" || res.status === "running") {
 // retry by polling again
 ...
}
else 
 return Response.json(res); // This will contain the final completed response 

Try it out today

Get started with Kimi K2.5 on Workers AI today. You can read our developer docs to find out model information and pricing, and how to take advantage of prompt caching via session affinity headers and asynchronous API. The Agents SDK starter also now uses Kimi K2.5 as its default model. You can also connect to Kimi K2.5 on Workers AI via Opencode. For a live demo, try it in our playground.

And if this set of problems around serverless inference, ML optimizations, and GPU infrastructure sound  interesting to you — we’re hiring!


The collective thoughts of the interwebz