Post Syndicated from xkcd.com original https://xkcd.com/3114/

Post Syndicated from xkcd.com original https://xkcd.com/3114/

Post Syndicated from Akira Mikami original https://aws.amazon.com/blogs/big-data/realizing-ocean-data-democratization-furuno-electrics-initiatives-using-amazon-datazone/
This is a guest post authored by Akira Mikami, a technical expert at Furuno Electric. The content and opinions in this post are those of the third-party author and AWS is not responsible for the content or accuracy of this post.
Since successfully commercializing the world’s first fish finder in 1948, Furuno Electric has been developing unique ultrasonic and electronic technologies in the marine electronics field. Under the company motto of “making the invisible visible”, they’ve have expanded their business centered on marine sensing technology and are now extending into subscription-based data businesses using Internet of Things (IoT) data. They’re are actively promoting the planning and development of data businesses to realize their new management vision outlined in FURUNO GLOBAL VISION NAVI NEXT 2030.
Like many manufacturing companies, Furuno Electric faced significant changes in revenue structure and technical architecture as they transitioned from traditional business to data-driven business. To succeed in this transformation, it was essential to build a foundation that promotes data utilization across the entire organization.
This post demonstrates how Furuno Electric built their system using Amazon DataZone and other Amazon Web Services (AWS) services to address technical infrastructure fragmentation, establish proper security governance, and develop an effective data business promotion system as part of their journey transitioning from a traditional manufacturing company to a data-driven business.
Furuno Electric faced three specific challenges in promoting their data business: technical infrastructure fragmentation and duplication, lack of security governance, and underdeveloped data business promotion system.
Project managers in the data business were independently designing and building data infrastructure, resulting in duplication of components for data collection, processing, and storage. This situation created wasteful development investments, hindered effective use of common data, and caused inefficient states that took time to launch businesses. Marine data services including fishing vessel data collection and sharing system, FWC, and Furuno Open Platform (FOP) had similar functions implemented separately for each project along the functional axes of data collection, processing, visualization, and analysis, resulting in unnecessary workload across the organization.
Security measures were considered and implemented separately by each department, and although checklists existed, they weren’t applied uniformly. This resulted in a lack of consistency in security measures, duplicate consideration costs for each department, and uncertainty in the comprehensiveness of measures. Integrated risk management was also difficult.
The organizational structure wasn’t prepared for the iterative development processes and long-term revenue models specific to data businesses, and there was a lack of mechanisms for cross-departmental data utilization and joint development. The distributed operational structure across departments made it difficult to rapidly deploy and continuously improve data businesses. In the process of creating data businesses, it became necessary to build entirely different customer relationships compared to traditional product sales businesses. In terms of organizational management and talent strategy, there was a need to transition from a top-down, risk-averse, specialized skill-focused structure to a bottom-up, challenge-oriented structure that emphasizes communication skills and diversity.
Furuno Electric built a data management foundation centered on Amazon DataZone, Amazon Simple Storage Service (Amazon S3), AWS Glue, and AWS Control Tower, a comprehensive solution designed to address each of the three challenges mentioned in the preceding section.
To address technical infrastructure fragmentation and duplication, they built Junction Architecture of Business Raw Data (JuBuRaw), a platform that consolidates common components for data collection, storage, management, and authentication. Using AWS Cloud Development Kit (AWS CDK) to code the infrastructure, they achieved standardization and automation of environment construction. This provides consistency and reproducibility, making it easier to add new systems and migrate existing systems to the common platform. Merely by executing CDK, a standard data pipeline (using AWS IoT Core, Amazon S3, AWS Glue, Amazon Kinesis, Amazon API Gateway, and AWS Lambda) for a specific system is automatically built. This eliminates duplicate design and development within the organization, reducing business launch time and improving fixed cost management. By standardizing common functions, they reduced the management and operation costs of existing systems and enabled the launch of new systems in half the time compared to before.
The following diagram is the overall JuBuRaw architecture.

To address the lack of security governance, they implemented a comprehensive security framework centered on AWS Control Tower to apply consistent security policies across multiple accounts. With automated monitoring systems using AWS Security Hub, AWS Config, and AWS CloudTrail and an integrated authentication system using AWS IAM Identity Center, they provide security consistency while reducing operational costs and management burden.
With the organization’s management account at the top, they placed AWS Control Tower, AWS Organizations, and AWS IAM Identity Center to achieve hierarchical security management. By adopting a multi-layered defense structure consisting of account baselines with AWS CloudTrail and AWS Config enabled, log archive environments, and audit and security operation environments, consistent security policies are applied to all accounts, enabling early detection and response to security incidents. This integrated approach has reduced the workload for security responses. This configuration is shown in the following diagram.

To address the underdeveloped data business promotion system, they introduced Amazon DataZone to streamline data discovery, sharing, and governance across the organization. They clarified the role division between the infrastructure management team and the data management team, centralizing data security policies, quality management, and metadata standardization. With a project-based collaboration environment, they promoted cross-departmental data utilization, establishing a foundation to support the creation and continuous monetization of data businesses.
In parallel with the introduction of technical solutions, they implemented organizational reforms to support medium- to long-term data utilization. The new organizational structure consists of three main roles: the infrastructure management team, the data management team, and the chief data officer. The following chart shows this organizational structure.

The infrastructure management team is responsible for maintaining and developing the technical foundation of the platform, managing multiple accounts using AWS Control Tower, applying and monitoring security baselines, and tracking infrastructure version management and changing history. By specializing in common technologies, they can provide a stable platform.
The data management team is responsible for data quality management and continual improvement using AWS Glue Data Quality, standardization and maintenance of metadata, definition and application of data security policies, management of Amazon DataZone data Catalog, and providing data governance using Amazon DataZone. To maximize the value of data, they focus on deeply understanding business requirements and data characteristics and performing appropriate data management.
The chief data officer is responsible for formulating data business strategies and determining direction, promoting coordination and collaboration between teams, making decisions regarding the evolution of the data management foundation, and fostering a data utilization culture throughout the organization. From a strategic perspective, they oversee the whole and bridge business goals and technology.
This clear division of roles has established an operational structure for effective data utilization, accelerating the data business creation process. Additionally, clarifying data ownership has improved data quality and reliability, promoting data utilization across the organization. This structure is sustainable and can flexibly respond to technological changes and changes in the business environment.
As a concrete application example of the integrated data platform JuBuRaw and organizational structure explained in the previous section, we introduce the migration project of the existing service SHIPS. This use case is a comprehensive migration case that uses all three solution elements of data collection, management, and utilization mentioned earlier.
Furuno Electric provides a system called SHIPS that plots ship position information and monitors the status of equipment installed on ships. By migrating this existing service to the JuBuRaw foundation, the several functional enhancements are expected.
In terms of data integration enhancement, by using the data catalog function of Amazon DataZone, it becomes easier to integrate not only ship position information but also various data sources such as internal systems, IoT devices, other company systems, automatic identification system (AIS) data, and weather and sea condition data. This enables swift data analysis and comprehensive ship management, which means operators can detect potential issues and implement preventive measures before they develop into serious problems. Particularly important is that by storing this data in a common data lake and retaining them as master data, they create an environment where the data can be easily used by other applications.
For security enhancement, organizations can use Amazon DataZone federated governance with publish-subscribe (pubsub) workflow mechanism and fine-grained access control capabilities. This means they can implement detailed permissions management specifically for data assets, rows, and columns while maintaining unified access control and data governance across multiple AWS accounts and organizational boundaries.
In this case, by using the new integrated data management foundation, it becomes possible to integrate individually designed and built data foundations, improving both efficiency and functionality. A consistent data flow from data sources to the data platform and then to individual applications is realized, enabling flexible data utilization centered on the data lake. Linkage with each application can also be easily realized from the data lake, providing expandability for future data utilization.
This SHIPS migration case is a comprehensive approach using the solution elements of the JuBuRaw foundation and is expected to serve as a reference model for future system migrations. It’s expected to achieve both service quality improvement and operational cost reduction.
Based on the data management foundation they’ve built, Furuno Electric aims to further expand and deepen data utilization. As part of their plan to continue and expand digital transformation, they’re currently starting with the migration of SHIPS, but plan to gradually migrate other IoT-related services (such as FOP, FWC, and Ichidake) to the new data management foundation in the future. This is expected to further strengthen the foundation for company-wide data utilization and enhance synergies between services.
Continuous enhancement of secure data sharing and access control is also essential. With the increase in data and expansion of utilization scope, the importance of security and access control will further increase. They’ll optimize the balance between data protection and utilization while incorporating practices accumulated through operations.
Additionally, Furuno Electric is exploring the expansion of their data management capabilities to Amazon SageMaker, specifically using Amazon SageMaker Catalog integrated with Amazon DataZone. This integration will enable them to seamlessly extend their existing data analytics governance workflows into artificial intelligence and machine learning (AI/ML) workloads. By applying the same data discovery, data sharing, and access control foundation across both data analytics and AI model development, they can accelerate the development of new AI-powered services. The unified governance framework will also provide secure and efficient AI adoption throughout the organization.
Through these initiatives, Furuno Electric is realizing their company motto of “making the invisible visible” in the field of data business as well. The integrated data platform JuBuRaw isn’t just an integration of technical foundations but serves as a foundation to support organizational culture transformation and the creation of new business models. As seen in the SHIPS migration case, using this foundation not only enhances existing services but also expands possibilities for new data utilization.
Through building a data foundation that can flexibly respond to business growth and changes while using a cloud-based environment, Furuno Electric has successfully led their digital transformation. They’ll continue to provide new value to customers through the democratization of marine data and accelerate the transition to data-driven business.
This case serves as a reference for many manufacturing companies promoting data utilization, showing that approaches from both technical and organizational perspectives are key to success. As Furuno Electric’s initiatives demonstrate, data democratization and effective utilization play an important role in the digital transformation of manufacturing.
Akira Mikami is a technical expert who played a central role in the FURUNO Data Platform (JuBuRaw) Construction Project at Furuno Electric Co., Ltd. Specializing in data platform construction and architecture, he led the implementation of cloud solutions utilizing AWS. He contributed to achieving efficient data management and strengthening team collaboration, leading the project to success.
Junpei Ozono is a Sr. Go-to-market (GTM) Data & AI solutions architect at Amazon Web Services (AWS) in Japan. He drives technical market creation for data and AI solutions while collaborating with global teams to develop scalable GTM motions. He guides organizations in designing and implementing innovative data-driven architectures powered by AWS services, helping customers accelerate their cloud transformation journey through modern data and AI solutions. His expertise spans across modern data architectures including data mesh, data lakehouse, and generative AI, so customers can build scalable and innovative solutions on Amazon Web Services (AWS).
Mitsuhiko Nishida is an Enterprise Solutions Architecture Automotive & Manufacturing Group Solutions Architect at Amazon Web Services (AWS) in Japan. He serves as a field Solutions Architect for manufacturing customers, helping them solve their business challenges. With expertise in generative AI and manufacturing IT, he guides the design and implementation of innovative solutions leveraging cutting-edge technologies. He supports manufacturing customers in building efficient architecture powered by AWS services to accelerate their cloud transformation journey and contribute to their digital transformation initiatives.
Post Syndicated from Jeremy Spell original https://aws.amazon.com/blogs/big-data/geospatial-data-lakes-with-amazon-redshift/
Data lake architectures help organizations offload data from premium storage systems without losing the ability to query and analyze the data. This architecture can be useful for geospatial data, where builders might have terabytes of infrequently accessed data in their databases that they want to cost-effectively maintain. However, this requires for their data lake query engine to support geographic information systems (GIS) data types and functions.
Amazon Redshift supports querying spatial data, including the GEOMETRY and GEOGRAPHY data types and functions that are used in querying GIS systems. Additionally, Amazon Redshift lets you query geospatial data both in your data lakes on Amazon S3 and your Redshift data warehouse, giving you the choice of how you can access your data. Additionally, AWS Lake Formation and support for AWS Identity and Access Management (IAM) in Esri’s ArcGIS Pro gives you a way to securely bridge data between your geospatial data lakes and map visualization tools. You can set up, manage, and secure geospatial data lakes in the cloud with a few clicks.
In this post, we walk through how to set up a geospatial data lake using Lake Formation and query the data with ArcGIS Pro using Amazon Redshift Serverless.
In our example, a county public health department has used Lake Formation to secure their data lake that contains public health information (PHI) data. Epidemiologists within the county want to create a map for the clinics providing vaccination for their communities. The county’s GIS analysts need access to the data lake to create the required maps without being able to access the PHI data.
This solution uses Lake Formation tags to allow column-level access in the database to the public information that includes the clinic names, addresses, zip codes, and longitude/latitude coordinates without allowing access to the PHI data within the same tables. We use Redshift Serverless and Amazon Redshift Spectrum to access this data from ArcGIS Pro, a GIS mapping software from Esri, an AWS Partner.
The following diagram shows the architecture for this solution.

The following is a sample schema for this post.
Description |
Column Name |
Geoproperty Tag |
Patient ID |
patient_id |
No |
Clinic ID |
clinic_id |
Yes |
Address of Clinic |
clinic_address |
Yes |
Clinic Zip Code |
clinic_zip |
Yes |
Clinic City |
clinic_city |
Yes |
First Name Patient |
first_name |
No |
Last Name Patient |
last_name |
No |
Patient Address |
patient_address |
No |
Patient Zip Code |
patient_zip |
No |
Vaccination Type |
vaccination_type |
No |
Latitude of Clinic |
clinic_lat |
Yes |
Longitude of Clinic |
clinic_long |
Yes |
In the following sections, we walk through the steps to set up the solution:
You should have the following prerequisites:
To create the environment for the demo, complete the following steps:
The CloudFormation template creates the following components:
samp-clinic-db-{ACCOUNT_ID}samp-clinical-glue-dbsamp-glue-crawlersamp-clinical-rs-wgsamp-clinical-rs-nsdemo-RedshiftIAMRole-{UNIQUE_ID}samp-clinical-glue-rolegeopropertyThe next step is to create a data lake in our demo environment and then use an AWS Glue crawler to populate the AWS Glue database and update the schema and metadata in the AWS Glue Data Catalog.
The CloudFormation stack created the S3 bucket we will use as well as the AWS Glue database and crawler. We have provided a fictious test dataset that will represent the patient and clinical information. Download the file and complete the following steps:
The crawler run should only take a minute to complete, and will populate a table named clinic-sample-s3_ACCOUNT_ID with a fictious dataset.
You will see that the dataset contains fields that contain PHI and personally identifiable information (PII).

We now have a database set up and the Data Catalog populated with the schema and metadata we will use for the rest of the demo.
In this next set of steps, we demonstrate how to secure PHI data to maintain compliance and empower GIS analysts to work effectively. To secure the data lake, we use AWS Lake Formation. In order to properly set up Lake Formation permissions, we need to gather details on how access to the data lake is established.
The Data Catalog provides metadata and schema information that enables services to access data within the data lake. To access the data lake from ArcGIS Pro, we use the ArcGIS Pro Redshift connector, which allows a connection from ArcGIS Pro to Amazon Redshift. Amazon Redshift can access the Data Catalog and provide connectivity to the data lake. The CloudFormation template created a Redshift Serverless instance and namespace and an IAM role that we will use to configure this connection. We still need to set up Lake Formation permissions so that GIS analysts can only access publicly available fields and not those containing PHI or PII. We will assign a Lake Formation tag on the columns containing the publicly available information and assign permissions to the GIS analysts to allow access to columns with this tag.
By default, the Lake Formation configuration allows Super access to IAMAllowedPrinciples; this is to maintain backward compatibility as detailed in Changing the default settings for your data lake. To demonstrate a more secure configuration, we will remove this default configuration.

IAMAllowedPrincipals and choose Revoke.clinic-sample-s3_ACCOUNT_ID and choose Edit schema.geoproperty. Assign geoproperty as the key and true for the value on all the clinic_ fields, then choose Save.Next, we need to grant the Amazon Redshift IAM role permission to access fields tagged with geoproperty = true.
demo-RedshiftIAMRole-UNIQUE_ID.geoproperty for the key and true for the value.Next, we need to perform the initial configuration of Amazon Redshift required for database operations. We use an AWS Secrets Manager secret created by the template to make sure password access is managed securely in accordance with AWS best practices.
For more information about these options, refer to Configuring your AWS account.

The query editor will require credentials to connect to the serverless instance; these have been created by the template and stored in Secrets Manager.
Redshift-admin-credentials).
An external schema in Amazon Redshift is a feature used to reference schemas that exist in external data sources. For information on creating external schemas, see External schemas in Amazon Redshift Spectrum. We use an external schema to provide access to the data lake in Amazon Redshift. From ArcGIS Pro, we will connect to Amazon Redshift to access the geospatial data.
The IAM role used in the creation of the external schema needs to be associated with the Redshift namespace. This has already been set up by the CloudFormation template, but it’s a good practice to verify that the role is set up correctly before proceeding.
sample-rs-namespace).
On the Security and encryption tab, you should see the IAM role created by CloudFormation. If this role or the namespace isn’t present, verify the stack in AWS CloudFormation before proceeding.


sample-glue-database:Because the associated role has been granted access to columns tagged with geoproperty = true, only those fields will be returned, as shown in the following screenshot (the data in this example is fictionalized).

For the data to be viewable from ArcGIS Pro, we will need to create a view. Now that the schemas have been established, we can create the view that can be accessed from ArcGIS Pro.
Amazon Redshift provides many geospatial functions that can be used to create views with fields used by ArcGIS Pro to add points onto a map. We will use one of these functions because the dataset contains latitude and longitude.
Use the following SQL code in the Amazon Redshift Query Editor to create a new view named clinic_location_view. Replace {ACCOUNT_ID} with your own account ID.
The new view that is created under your local schema will have a column named geom containing map-based points that can be used by ArcGIS Pro to add points during map creation. The points in this example are for the clinics providing vaccines. In a real-world scenario, as new clinics are built and their data is added to the data lake, their locations would be added to the map created using this data.
For this demo, we use a database user and group to provide access for ArcGIS Pro clients. Enter the following SQL code into the Amazon Redshift Query Editor to create a database user and group:
After the commands are complete, use the following code to grant permissions to the group:
In order to add the database connection to ArcGIS Pro, you need the endpoint for the Redshift Serverless workgroup. You can access the endpoint information on the sample-rs-wg workgroup details page on the Redshift Serverless console. The Redshift namespaces and workgroups are listed by default, as shown in the following screenshot.

You can copy the endpoint in the General information section. This endpoint will need to modified; the :5439/dev will need to be removed when configuring the connector in ArcGIS Pro.

| Make sure the Amazon Redshift ODBC connector has already been installed; this is required in order to make the connection. |
.com from the endpoint).
If your ArcGIS Pro client doesn’t have access to the endpoint, you will receive an error during this step. A network path must exist between the ArcGIS Pro client and the Redshift Serverless endpoint. You can set up the network path with Direct Connect, AWS Site-to-Site VPN, or AWS Client VPN. Although it’s not recommended for security reasons, you can also configure Amazon Redshift with a publicly available endpoint. Be sure you consult your security and network teams for best practices and policy guidance before allowing public access to your Redshift Serverless instance.
If a network path exists and you’re having issues connecting, verify the security group rules allow communication inbound from your ArcGIS Pro subnet over the port your Redshift Serverless instance is running on. The default port is 5439, but you can configure a range of ports depending on your environment; see Connecting to Amazon Redshift Serverless for more information.
If connectivity is successful, ArcGIS Pro will add the Amazon Redshift connection under Connection File Name.
clinic_location_view).ArcGIS Pro will add the points from the view onto the map. The final map displayed has the symbology edited to use red crosses to represent the clinics instead of dots.

After you have finished the demo, complete the following steps to clean up your resources:
data-with-geocode.csv file.In this post, we reviewed how to set up Redshift Serverless to use geospatial data contained within a data lake to enhance maps in ArcGIS Pro. This technique helps builders and GIS analysts use available datasets in data lakes and transform it in Amazon Redshift to further enrich the data before presenting it on a map. We also showed how to secure a data lake using Lake Formation, crawl a geospatial dataset with AWS Glue, and visualize the data in ArcGIS Pro.
For additional best practices for storing geospatial data in Amazon S3 and querying it with Amazon Redshift, see How to partition your geospatial data lake for analysis with Amazon Redshift. We invite you to leave feedback in the comments section.
Jeremy Spell is a Cloud Infrastructure Architect working with Amazon Web Services (AWS) Professional Services. He enjoys architecting and building solutions for customers. In his free time Jeremy makes Texas style BBQ, and spends time with his family and church community.
Jeff Demuth is a solutions architect who joined Amazon Web Services (AWS) in 2016. He focuses on the geospatial community and is passionate about geographic information systems (GIS) and technology. Outside of work, Jeff enjoys traveling, building Internet of Things (IoT) applications, and tinkering with the latest gadgets.
Post Syndicated from Netflix Technology Blog original https://netflixtechblog.com/netflix-tudum-architecture-from-cqrs-with-kafka-to-cqrs-with-raw-hollow-86d141b72e52
By Eugene Yemelyanau, Jake Grice

Tudum.com is Netflix’s official fan destination, enabling fans to dive deeper into their favorite Netflix shows and movies. Tudum offers exclusive first-looks, behind-the-scenes content, talent interviews, live events, guides, and interactive experiences. “Tudum” is named after the sonic ID you hear when pressing play on a Netflix show or movie. Attracting over 20 million members each month, Tudum is designed to enrich the viewing experience by offering additional context and insights into the content available on Netflix.
At the end of 2021, when we envisioned Tudum’s implementation, we considered architectural patterns that would be maintainable, extensible, and well-understood by engineers. With the goal of building a flexible, configuration-driven system, we looked to server-driven UI (SDUI) as an appealing solution. SDUI is a design approach where the server dictates the structure and content of the UI, allowing for dynamic updates and customization without requiring changes to the client application. Client applications like web, mobile, and TV devices, act as rendering engines for SDUI data. After our teams weighed and vetted all the details, the dust settled and we landed on an approach similar to Command Query Responsibility Segregation (CQRS). At Tudum, we have two main use cases that CQRS is perfectly capable of solving:

The high-level diagram above focuses on storage & distribution, illustrating how we leveraged Kafka to separate the write and read databases. The write database would store internal page content and metadata from our CMS. The read database would store read-optimized page content, for example: CDN image URLs rather than internal asset IDs, and movie titles, synopses, and actor names instead of placeholders. This content ingestion pipeline allowed us to regenerate all consumer-facing content on demand, applying new structure and data, such as global navigation or branding changes. The Tudum Ingestion Service converted internal CMS data into a read-optimized format by applying page templates, running validations, performing data transformations, and producing the individual content elements into a Kafka topic. The Data Service Consumer, received the content elements from Kafka, stored them in a high-availability database (Cassandra), and acted as an API layer for the Page Construction service and other internal Tudum services to retrieve content.
A key advantage of decoupling read and write paths is the ability to scale them independently. It is a well-known architectural approach to connect both write and read databases using an event driven architecture. As a result, content edits would eventually appear on tudum.com.
Did you notice the emphasis on “eventually?” A major downside of this architecture was the delay between making an edit and observing that edit reflected on the website. For instance, when the team publishes an update, the following steps must occur:
By introducing a highly-scalable eventually-consistent architecture we were missing the ability to quickly render changes after writing them — an important capability for internal previews.
In our performance profiling, we found the source of delay was our Page Data Service which acted as a facade for an underlying Key Value Data Abstraction database. Page Data Service utilized a near cache to accelerate page building and reduce read latencies from the database.
This cache was implemented to optimize the N+1 key lookups necessary for page construction by having a complete data set in memory. When engineers hear “slow reads,” the immediate answer is often “cache,” which is exactly what our team adopted. The KVDAL near cache can refresh in the background on every app node. Regardless of which system modifies the data, the cache is updated with each refresh cycle. If you have 60 keys and a refresh interval of 60 seconds, the near cache will update one key per second. This was problematic for previewing recent modifications, as these changes were only reflected with each cache refresh. As Tudum’s content grew, cache refresh times increased, further extending the delay.
As this pain point grew, a new technology was being developed that would act as our silver bullet. RAW Hollow is an innovative in-memory, co-located, compressed object database developed by Netflix, designed to handle small to medium datasets with support for strong read-after-write consistency. It addresses the challenges of achieving consistent performance with low latency and high availability in applications that deal with less frequently changing datasets. Unlike traditional SQL databases or fully in-memory solutions, RAW Hollow offers a unique approach where the entire dataset is distributed across the application cluster and resides in the memory of each application process.
This design leverages compression techniques to scale datasets up to 100 million records per entity, ensuring extremely low latencies and high availability. RAW Hollow provides eventual consistency by default, with the option for strong consistency at the individual request level, allowing users to balance between high availability and data consistency. It simplifies the development of highly available and scalable stateful applications by eliminating the complexities of cache synchronization and external dependencies. This makes RAW Hollow a robust solution for efficiently managing datasets in environments like Netflix’s streaming services, where high performance and reliability are paramount.
Tudum was a perfect fit to battle-test RAW Hollow while it was pre-GA internally. Hollow’s high-density near cache significantly reduces I/O. Having our primary dataset in memory enables Tudum’s various microservices (page construction, search, personalization) to access data synchronously in O(1) time, simplifying architecture, reducing code complexity, and increasing fault tolerance.

In our simplified architecture, we eliminated the Page Data Service, Key Value store, and Kafka infrastructure, in favor of RAW Hollow. By embedding the in-memory client directly into our read-path services, we avoid per-request I/O and reduce roundtrip time.
The updated architecture yielded a monumental reduction in data propagation times, and the reduced I/O led to faster request times as an added bonus. Hollow’s compression alleviated our concerns about our data being “too big” to fit in memory. Storing three years’ of unhydrated data requires only a 130MB memory footprint — 25% of its uncompressed size in an Iceberg table!
Writers and editors can preview changes in seconds instead of minutes, while still maintaining high-availability and in-memory caching for Tudum visitors — the best of both worlds.
But what about the faster request times? The diagram below illustrates the before & after timing to fulfil a request for Tudum’s home page. All of Tudum’s read-path services leverage Hollow in-memory state, leading to a significant increase in page construction speed and personalization algorithms. Controlling for factors like TLS, authentication, request logging, and WAF filtering, homepage construction time decreased from ~1.4 seconds to ~0.4 seconds!

An attentive reader might notice that we have now tightly-coupled our Page Construction Service with the Hollow In-Memory State. This tight-coupling is used only in Tudum-specific applications. However, caution is needed if sharing the Hollow In-Memory Client with other engineering teams, as it could limit your ability to make schema changes or deprecations.
In the next episode, we’ll share how Tudum.com leverages Server Driven UI to rapidly build and deploy new experiences for Netflix fans. Stay tuned!
Thanks to Drew Koszewnik, Govind Venkatraman Krishnan, Nick Mooney
Netflix Tudum Architecture: from CQRS with Kafka to CQRS with RAW Hollow was originally published in Netflix TechBlog on Medium, where people are continuing the conversation by highlighting and responding to this story.
Post Syndicated from BeardedTinker original https://www.youtube.com/shorts/OZRpBh6hewc
Post Syndicated from jzb original https://lwn.net/Articles/1028558/
Few, if any, web sites or web-based services have gone unscathed by
the locust-like hordes of AI crawlers looking to consume (and then
re-consume) all of the world’s content. The Anubis project is designed to
provide a first line of defense that blocks mindless bots—while
granting real users access to sites without too much hassle. Anubis is
a young project, not even a year old. However, its development is
moving quickly, and the project seems to be enjoying rapid
adoption. The most recent release of Anubis, version
1.20.0, includes a feature that many users have been interested in
since the project launched: support for challenging clients without
requiring users to have JavaScript turned on.
Post Syndicated from Зорница Христова original https://www.toest.bg/po-bukvite-radkova-mindova-dench/

Пловдив: изд. „Жанет 45“, 2024
„Лаш насам, лаш натам, сега има, сега няма, тука памет, тука забрава – коментира Стефан Русинов в разказа си за гостуването на Ю Хуа в България, – политическата обстановка е като природна стихия, която завихря безпомощните хорица и тласка умовете им в какви ли не посоки.“
Темата на литературните срещи с Ю Хуа, Мария Степанова и прочее… беше паметта. И тази тема продължава да ме човърка, искам да видя какво се случва с паметта тук, у нас.
Паметта е свързана с идентичността (страдащият от амнезия не знае кой е), с доброто и злото (без памет няма отговорност), дори с желанието за живот.

Затова ме вълнуват книгите, които насочват окото на читателя към слепите петна на паметта, опитват се да поправят изличеното от забравата, да го възстановят. Това обаче не е лесна задача – изисква пренареждане на цялостния наратив на историята, тоест иска от читателя не просто да научи нещо ново, а да се отучи да разказва нещата, както е свикнал.
Да видим.
„И леглото ни е зеленина“ от Виолета Радкова говори за две малко засягани теми в нашата литература – обикновените българи в земите, останали извън Княжество България, че и Източна Румелия, и българската емиграция в началото на XX век. Първата е предмет донякъде и на „Рана“ на Захари Карабашлиев, но там детството на героя в Беломорска Тракия е еднозначно идилично, като в народна песен – по дваж на година жнеме и вършеме, по триж на година гюл-трендафил цъфти. В „И леглото ни е зеленина“ този гюл си има бодли. На Виолета Радкова ѝ стиска да напише, че децата в някогашните големи семейства са заменими – едно умряло, не го мисли толкова, карай нататък, гледай живите. Стиска ѝ и несигурността на жените да не е етнически обусловена, ей ги на – лошите ни застрашават, добрите ни бранят честта. И изобщо да не задълбава в травмата на етническия конфликт, в който българите са жертви, да ги види не само като носители на трагична съдба, а като хора.
Второто, което ми е много интересно, е връзката бежанци–мигранти. Трябва да призная, че аз самата съм пристрастна, изследвам тази връзка от няколко години покрай собствена родова история, която искам да придобие книжна плът; впрочем в „Кой е Сава Попов?“ Свобода Цекова и Антон Стайков също отбелязват как тази траектория не е била никак рядка. Огромният бежански проблем след Балканската война е карал много младежи от Тракия и Македония да търсят късмета си зад океана.
В литературен план това е доста рязка смяна на контекста – в нашата литературна традиция разказът за селския бит и разказът за Америка, пък и изобщо за далечния свят почти не се допират, и толкоз. Любопитно е как подхожда Радкова към това: с бавното отваряне на героя си към обитавания от него град, с ритуала за четене на вестникарски заглавия, през който скицира и контекста, задава важна разлика – жените тук, жените там – и добавя щипка хумор с вметката, че жените там са свободни да се самоубиват от любов на воля.
Впрочем такива „емигрантски романи“, които преминават от селското към космополитно-урбанистичното, не са рядкост в съвременната световна литература – ей на, „Мидълсекс“ на Юдженидис тръгва от съседна Гърция, и то пак в контекста на етническите трагедии. На тях си им отива и щипка магически реализъм, в този смисъл историята с ръцете на героя, а и раздвояването на неговата траектория биха се приели добре от англоезичната публика. Би било хубаво да предизвика дискусия и тук, в половината ни родови истории има едно такова „ами ако“ – „ами ако прадядо не се беше върнал“, или „ако чичо си беше дошъл“. Не знам дали за нашата публика това усложняване на сюжета ще работи (в комбинация с тематичната новост); надявам се, а и авторката е сторила възможното да направи фабулата увлекателна и четивна, със
Чух оплакване, че на моменти е прекалено умел – че бързото отгръщане на страници те кара да го довършиш, преди да си успял да се стъписаш и замислиш – но това да ни е проблемът. Чудесна книга.
София: изд. „Изток-Запад“, 2025
Книгата на Людмила Миндова за Прехода и за нейната майка Свобода Стефанова поема друг курс към паметта. Не търси линейния разказ, сюжета, не се опитва да намести всичко в единна обяснителна фабула, да го навърже на една нишка… Защото това значи да изпуснеш някакви важни неща, а други да посмачкаш, за да се вместят в картинката. В тази книга е осезаемо важно това да не се случва, осезаемо важна е верността към събитията, верността и пълнотата на разказа.
Залозите са големи: думите трябва да уловят и запазят на този свят същината на любим човек; и заедно с това трябва да уловят и осмислят същината на времето, което е изтекло. Припряността може да е подмяна, дори на историята с литературата, а
Да добави сетивността. Да извади до факта нещата, с които той е в емоционална връзка, за да може читателят да види назад в емоционалната си история какво се е случило, какво е преживял, какво го е разтърсило, натъжило, изпълнило с надежда, защо е реагирал така. И какво е станало с онези, които е срещнал по пътя си.

Спомняте ли си Кашпировски например? Телевизионния хипнотизатор, дето подканяше зрителите да сложат чаша вода пред телевизора, та той дистанционно да ѝ придаде лечебни свойства? Какво стана с него – но и какво стана с нас, когато времето, дето трябваше да ни лекува от травмите на комунизма, всъщност ни даде да пием студена вода, простете елементарната аналогия. Как е свързано едното с другото – доносничеството с унищожението на паметта, поезията с въздуха, отровеният въздух с отровеното куче, вчера с днес?
Всякак. Поетическият усет на Людмила Миндова ѝ е подсказал отлична метафора: мемъри картите, с които играят децата – търсенето на съответствия насред един пасианс, който се реди на сляпо, с гръб към теб. В рамките на една глава-карта виждаме събрани Свети Стефан, Сан Стефано и Свобода Стефанова; в друга – дядо Торбалан и Луис Корвалан (Марин също разказваше как в детството си ги е бъркал тези двамата); в трета – мемъри пяната и „Човешкият отпечатък“ на Цветан Тодоров, хазарта и the hazard, опасността; в четвърта – четвъртъкът на атентата в „Света Неделя“ и четвъртъкът на Народния съд, повтарящите се имена в нашите лични истории. И така нататък.
Това са римите на реалността. Възможните, неочевидни връзки, които могат да звучат случайно, повърхностно – тази стъпва върху алюзията, тази върху фонетиката, какво следва, миризми? Текстура, допир? Но именно те удържат паметта, когато централният разказ е строшен като предно стъкло на автомобил. Потъването в книгата на Людмила Миндова е като потъване във времето – но има какво да те задържи. И покрай цялата тъга и мрак има и памет за музика, за поезия, за дух. За свобода. За Свобода.
превод Иглика Василева, София: изд. „Локус“, 2025
Свобода, страст, смях, усет за лекотата и дълбочините, езикова находчивост, характер – какъв разкош са Шекспировите героини, а? Ако имате дъщери, дайте им да четат колкото може повече Шекспир, обградете ги с тези великолепни жени, които бърборят, отстояват своето, преобличат се, пътуват, шегуват се, влюбват се, не остават никому длъжни и само в краен случай полудяват или се обявяват за умрели, за да избегнат обществения скандал!
Джуди Денч е била всяка от тях и нейните интервюта за Шекспировите ѝ роли са като сбирка на женската компания – не на благонравен чай (въпреки че „Чай с дамите“ за Джуди Денч и нейните колежки от „Чай с Мусолини“ Маги Смит и Джоун Плоурайт беше чудесен). По-скоро в кръчмата „Дърти Дък“ или зад кулисите на „Олд Вик“.

Спомените за театъра са нещо друго, защото те са спомен за различните хора, които даден актьор е бил, в случая – за лейди Макбет, за Титания, за Беатриче, Виола, Мария, Гертруда, Офелия, Регана, Хермиона, Пердита, Изабела, Жулиета…
Спомен „отвътре“ – какво движи лейди Макбет? Властолюбие или както мисли актрисата, силната връзка със съпруга ѝ, желанието ѝ да му помогне да постигне неговата си мечта, ако и да е твърде мекушав за нея? Защо Беатриче и Бенедикт се нападат така, предишна връзка ли е това и как е свършила? Знае ли Гертруда за престъплението на Клавдий, на колко години следва да е Титания… Трябва ли всички тези отговори да се изиграят? Не. Но според Джуди Денч все пак трябва да знаеш – за себе си.
За читателя на интервютата усещането е като сам да изпробваш роля подир роля. Бих ли могла да съм, примерно, мистрис Скокли? Ще ми писне ли от Бертрам на мястото на Елена? И т.н. Интересен тип история – този, който те приканва да си припомниш възможните свои алтернативи, възможните пътища. Но и реална история на едно славно време, на Олд Вик и Кралския Шекспиров театър, на Джон Гилгуд и Иън Маккелън, и Ванеса Редгрейв, и т.н. И другото славно време, разбира се, елизабетинското. И цялата утопична идея, че сме способни на приемственост, че сме способни да запазим смисъла на миналото и да го направим свой.
Чудесен превод на Иглика Василева, отлично оформление на Емил Марков. Да бяха направили още една стъпка от Locus и да бяха наели коректор, който да оправи очевидните пропуски при набирането, разменените или пропуснатите букви и пр. Същото при „Изток-Запад“, макар че там освен липсата на коректор буди недоумение и корицата. Защо?! При такива хубави книги.
В емблематичната си колонка, започната още през 2008 г. във в-к „Култура“, Марин Бодаков ни представяше нови литературни заглавия и питаше с какво точно тези книги ни променят. Вярваме, че е важно тази рубрика да продължи. От човек до човек, с нова книга в ръка.
Post Syndicated from jake original https://lwn.net/Articles/1029418/
Security updates have been issued by Debian (sslh), Oracle (container-tools:rhel8, gnome-remote-desktop, golang, javapackages-tools:201801, jq, libvpx, libxml2, mpfr, and perl-File-Find-Rule-Perl), Red Hat (glib2, libblockdev, and sudo), Slackware (git), SUSE (avif-tools, containerd, djvulibre, gpg2, helm, kernel, libpoppler-cpp2, libxml2, libxml2-2, openssl-3, perl-YAML-LibYAML, python-cryptography, python-setuptools, python311-pycares, tomcat10, and wireshark), and Ubuntu (djvulibre, git, libyaml-libyaml-perl, and protobuf).
Post Syndicated from Colm MacCarthaigh original https://aws.amazon.com/blogs/security/establishing-a-european-trust-service-provider-for-the-aws-european-sovereign-cloud/
Last month, we announced new sovereign controls and governance structure for the AWS European Sovereign Cloud. The AWS European Sovereign Cloud is a new, independent cloud for Europe, designed to help customers meet their evolving sovereignty needs, including stringent data residency, operational autonomy, and resiliency requirements. Launching by the end of 2025, the AWS European Sovereign Cloud will be entirely located within the European Union (EU) and operate as an independent cloud for Europe. Last month, we announced plans to launch a dedicated European certificate authority (CA), or trust service provider, to support autonomous trust service operations within the AWS European Sovereign Cloud.
We are actively building out the first AWS Region of the AWS European Sovereign Cloud in the state of Brandenburg, Germany. We are on track for launch and AWS services are being deployed, configured, and tested for autonomous operations in the AWS European Sovereign Cloud. The AWS European Sovereign Cloud infrastructure will be physically and logically separate from other Regions. We designed the AWS European Sovereign Cloud to have no critical dependencies on non-EU infrastructure. Everything needed to operate the AWS European Sovereign Cloud is in the EU: the talent, the technology, the infrastructure, and the leadership. In addition to independent infrastructure, there will be zero operational control outside of EU borders. Only AWS employees, residing in the EU, will control day-to-day operations, including access to data centers, technical support, and customer service for the AWS European Sovereign Cloud.
For the first time, we will provide a dedicated sovereign European trust service provider (EU-TSP). This EU-TSP will autonomously operate its own CA key materials and perform certificate issuance functions within the AWS European Sovereign Cloud. A trust service provider is an entity that manages the policies and operations for a set of root and subordinate certificate authorities. A root CA is a cryptographic building block and root of trust upon which end entity certificates can be issued. It represents a private key for signing (issuing) certificates and a root certificate that identifies the root CA and binds the private key to the name of the CA. In short, the EU-TSP is an autonomous trust service provider in Europe, for Europe.
The EU-TSP will be the public root of trust for the AWS European Sovereign Cloud, helping to maintain the confidentiality and integrity of network communications. The EU-TSP will provide the default CA used by AWS service endpoints, AWS Certificate Manager (ACM), and ACM integrated services. For AWS European Sovereign Cloud customers, this means that even in the event of a material loss of connectivity outside of the EU, the EU-TSP will continue to provide trust services autonomously.
We recently completed the cryptographic key signing ceremony for our EU-TSP at a secure EU location, witnessed by external, third-party auditors. The resulting root CAs have been submitted for inclusion to popular web browsers used by AWS customers. This EU-TSP will be operated in accordance with the requirements of the Certificate Authority/Browser Forum. All the key material for the EU-TSP is located within EU borders, and only EU residents have the ability to operate, control, or reconfigure the EU-TSP.
To maintain verifiable trust, we will engage independent EU-based auditors to assure the EU-TSP controls are designed appropriately, operate effectively, and can help customers satisfy their compliance obligations. We will make the audit reports publicly available.
The EU-TSP will be active and providing autonomous trust services when the AWS European Sovereign Cloud launches at the end of 2025. To learn more, visit AWS European Sovereign Cloud.
If you have feedback about this post, submit comments in the Comments section below. If you have questions about this post, contact AWS Support.
Post Syndicated from Anton Dort-Golts original https://blog.cloudflare.com/quicksilver-v2-evolution-of-a-globally-distributed-key-value-store-part-1/
Quicksilver is a key-value store developed internally by Cloudflare to enable fast global replication and low-latency access on a planet scale. It was initially designed to be a global distribution system for configurations, but over time it gained popularity and became the foundational storage system for many products in Cloudflare.
A previous post described how we moved Quicksilver to production and started replicating on all machines across our global network. That is what we called Quicksilver v1: each server has a full copy of the data and updates it through asynchronous replication. The design served us well for some time. However, as our business grew with an ever-expanding data center footprint and a growing dataset, it became more and more expensive to store everything everywhere.
We realized that storing the full dataset on every server is inefficient. Due to the uniform design, data accessed in one region or data center is replicated globally, even if it’s never accessed elsewhere. This leads to wasted disk space. We decided to introduce a more efficient system with two new server roles: replica, which stores the full dataset and proxy, which acts as a persistent cache, evicting unused key-value pairs to free up some disk space. We call this design Quicksilver v1.5 – an interim step towards a more sophisticated and scalable system.
To understand how those two roles helped us reduce disk space usage, we first need to share some background on our setup and introduce some terminology. Cloudflare is architected in a way where we have a few hyperscale core data centers that form our control plane, and many smaller data centers distributed across the globe where resources are more constrained. Quicksilver has dozens of servers in the core data centers with terabytes of storage called root nodes. In the smaller data centers, though, things are different. A typical data center has two types of nodes: intermediate nodes and leaf nodes. Intermediate servers replicate data either from the other intermediate nodes or directly from the root nodes. Leaf nodes serve end user traffic, and receive updates from intermediate servers, effectively being leaves of a replication tree. Disk capacity varies significantly between node types. While root nodes aren’t facing an imminent disk space bottleneck, it’s a definite concern for leaf nodes.
Every server – whether it’s a root, intermediate, or leaf – hosts 10 Quicksilver instances. These are independent databases, each used by specific Cloudflare services or products such as the DNS, CDN or WAF.

Figure 1. Global Quicksilver
Let’s consider the role distribution. Instead of hosting ten full datasets on every machine within a data center, what if we deploy only a few replicas in each? The remaining servers would be proxies, maintaining a persistent cache of hot keys and querying replicas for any cache misses.

Figure 2. Role allocation for different Quicksilver instances
Data centers across our network are very different in size, ranging from hundreds of servers to a single rack with just a few servers. To ensure every data center has at least one replica, the simplest initial step is an even split: on each server, place five replicas of some instances and five proxies for others. The change immediately frees up disk space, as the cached hot dataset on a proxy should be smaller than a full replica. While it doesn’t remove the bottleneck entirely, it could, in theory, lead to an up to 50% reduction in disk space usage. More importantly, it lays the foundation for a new distributed design of Quicksilver, where queries can be served by multiple machines in a data center, paving the way for further horizontal scaling. Additionally, an iterative approach helps to battle-proof the code changes earlier.
Before committing to building Quicksilver v1.5, we wanted to be sure that the proxy/replica design would actually work for our workload. If proxies needed to cache the entire dataset for good performance, then it would be a dead end, offering no potential disk space benefits. To assess this, we built a data pipeline which pushes accessed keys from all across our network to ClickHouse. This allowed us to estimate typical sizes of working sets. Our analysis revealed that:
in large data centers approximately, 20% of the keyspace was in use
in small data centers this number dropped to just about 1%
These findings gave us confidence that the caching approach should work, though it wouldn’t be without its challenges.
When talking about caches, the first thing that comes to mind is an in-memory cache. However, this cannot work for Quicksilver for two main reasons: memory usage and the “cold cache” problem.
Indeed, with billions of stored keys, even a fraction of them would lead to an unmanageable increase in memory usage. System restarts should not affect performance, which means that cache data must be preserved somewhere anyway. So we decided to make the cache persistent and store it in the same way as full datasets: in our embedded RocksDB. Thus, cached keys normally sit on disk and can be retrieved on-demand with low memory footprint.
When a key cannot be found in the proxy’s cache, we request it from a replica using our internal distributed key-value protocol, and put it into a local cache after processing.
Evictions are based on RocksDB compaction filters. Compaction filters allow defining custom logic executed in background RocksDB threads responsible for compacting files on disk. Each key-value pair is processed with a filter on a regular basis, evicting least recently used data from the disk when available disk space drops below a certain threshold called a soft limit. To track keys accessed on disk, we have an LRU-like in-memory data structure, which is passed to the compaction filter to set last access date in key metadata and inform potential evictions.
However, with some specific workloads there is still a chance that evictions will not keep up with disk space growth, and for this scenario we have a hard limit: when available disk space drops below a critical threshold, we temporarily stop adding new keys to the cache. This hurts performance, but it acts as a safeguard, ensuring our proxies remain stable and don’t overflow under a massive surge of requests.
Quicksilver has, from the start, provided sequential consistency to clients: if key A was written before B, it’s not possible to read B and not A. We are committed to maintaining this guarantee in the new design. We have experienced Hyrum’s Law first hand, with Quicksilver being so widely adopted across the company that every property we introduced in earlier versions is now relied upon by other teams. This means that changing behaviour would inevitably break existing functionality and introduce bugs.
However, there is one thing standing in our way: asynchronous replication. Quicksilver replication is asynchronous mainly because machines in different parts of the world replicate at different speeds, and we don’t want a single server to slow down the entire tree. But it turns out in a proxy-replica design, independent replication progress can result in non-monotonic reads!
Consider the following scenario: a client sequentially writes keys A, B, C, .. K one after another to the Quicksilver root node. These keys are asynchronously replicated through data centers across our network with varying latency. Imagine we have a proxy on index 5, which has observed keys from A to E, and two replicas:
replica_1 is at index 2 (slightly behind the proxy), having only received A and B
replica_2 at index 9, which is slightly ahead due to a faster replication path and has received all keys from A to I

Figure 3. Asynchronous replication in QSv1.5
Now, a client performs two successive requests on a proxy, each time reading the keys E, F, G, H and I. For simplicity, we assume these keys are not cacheable (for example, due to low disk space). The proxy’s first remote request is routed to replica_2, which already has all keys and responds back with values. To prevent hot spots in a data center, we load balance requests from proxies, and the next one lands on replica_1, which hasn’t received any of the requested keys yet, and responds with a “not found” error.
So, which result is correct?
The correct behavior here is that of Quicksilver v1, which we aim to preserve. If the server on replication index 5 were a replica instead of a proxy, it would have seen updates for keys A through E inclusive, resulting in E being the only key in both replies, while all other keys cannot be found yet. Which means responses from both replica_1 and replica_2 are wrong!
Therefore, to maintain previous guarantees and API backwards compatibility, Quicksilver v1.5 must address two crucial consistency problems: cases where the replica is ahead of the proxy, and conversely, where it lags behind. For now let’s focus on the case where a proxy lags behind a replica.
In our example, replica_2 responds to a request from a proxy “from the past”. We cannot use any locks for synchronizing two servers, as it would introduce undesirable delays to the replication tree, defeating the purpose of asynchronous replication. The only option is for replicas to maintain a history of recent updates. This naturally leads us to implementing multiversion concurrency control (MVCC), a popular database mechanism for tracking changes in a non-blocking fashion, where for any key we can keep multiple versions of its values for different points in time.
With MVCC, we no longer overwrite the latest value of a key in the default column family for every update. Instead, we introduced a new MVCC column family in RocksDB, where all updates are stored with a corresponding replication index. Lookup for a key at some index in the past goes as follows:
First we search in the default column family. If a key is found and the write timestamp is not greater than the index of a requesting proxy, we can use it straight away.
Otherwise, we begin scanning the MVCC column family, where keys have unique suffixes based on latest timestamps for which they are still valid.
In the example above, replica_2 has MVCC enabled and has keys A@1 .. K@11 in its default column family. The MVCC is initially empty, because no keys have been overwritten yet. When it receives a request for, say, key H with target index 5, it first makes a lookup in a default column family and finds the given key, but its timestamp is 8, which means this version should not be visible to the proxy yet. It then scans the MVCC, finds no matching previous versions and responds with “not found” to the proxy. Should key H be updated twice at indexes 4 and 8, we would have placed the initial version into MVCC before overwriting it in the default column family, and the proxy would receive the first version in response.
If a key E is requested at index 5, replica_2 can find it quickly in the default column family and return it back to the proxy. There is no need to read from MVCC, as the timestamp of the latest version (5) satisfies the request.
Another corner case to consider is deletions. When a key is deleted and then re-written, we need to explicitly mark the period of removal in MVCC. For that we’ve implemented tombstones – a special value format for absent keys.
Finally, we need to make sure that key history is not growing uncontrollably, using up all of the disk space available. Luckily we don’t actually need to record history for a long period of time, it just needs to cover the maximum replication index difference between any two machines. And in practice, a two-hour interval turned out to be way more than enough, while adding only about 500 MB of extra disk space usage. All records in the MVCC column family older than two hours are garbage collected, and for that again we use custom RocksDB compaction filters.
Now we know how to deal with proxies lagging behind replicas. But what about the opposite case, when a proxy is ahead of replicas?
The simplest solution is for replicas to not serve requests with a target index higher than its own. After all, it cannot know about keys from the future, whether they will be added, updated, or removed. In fact, our first implementation just returned an error when the proxy was ahead, as we expected it to happen quite infrequently. But after rolling out gradually to a few data centers, our metrics made it clear that the approach was not going to work.
This led us to analyze which keys are affected by this kind of replication asymmetry. It’s definitely not keys added or updated a long time ago, because replicas would already have the changes replicated. The only problematic keys are those updated very recently, which the proxy already knows about, but the replica does not.
With this insight, we realized that the issue should be solved on the proxies rather than on the replica side. By preserving all recent updates locally, the proxy can avoid querying the replica. This became known as the sliding window approach.
The sliding window retains all recent updates written in a short, rolling timeframe. Unlike cached keys, items in the window cannot be evicted until they move outside of the window. Internally, the sliding window is defined by lower and upper boundary pointers. These are kept in memory, and can easily be restored after a reload from the current database index and the pre-configured window size.

Figure 4. The sliding window shifts when replication updates arrive
When a new update event arrives from the replication layer we add it to the sliding window by moving both the upper and lower boundary one position higher. Thereby, we maintain the fixed size of the window. Keys written before the lower bound can be evicted by the compaction filter, which is aware of current sliding window boundaries.
Another problem arising with our distributed replica-proxy design is negative lookups – requests for keys which don’t exist in the database. Interestingly, in our workloads we see about ten times more negative lookups than positive ones!
But why is it a problem? Unfortunately, each negative lookup will be a cache miss on a proxy, requiring a request to a replica. Given the volume of requests and proportion of such lookups, it would be a disaster for performance, with overloaded replicas, overused data center networks, and massive latency degradation. We needed a fast and efficient approach to identifying non-existing keys directly at the proxy level.
In v1, negative lookups are the quickest type of requests. We rely on a special probabilistic data structure – Bloom filters – used in RocksDB to determine if the requested key might belong to a certain data file containing a range of sorted keys (called Sorted Sequence Table or SST) or definitely not. 99% of the time, negative lookups are served using only this in-memory data structure, avoiding the need for disk I/O.
One approach we considered for proxies was to cache negative lookups. Two problems immediately arise:
How big is the keyspace of negative lookups? In theory, it’s infinite, but the real size was unclear. We can store it in our cache only if it is small enough.
Cached negative lookups would no longer be served by the fast Bloom filters. We have row and block caches in RocksDB, but the hit rate is nowhere near the filters for SSTs, which means negative lookups would end up going to disk more often.
These turned out to be dealbreakers: not only was the negative keyspace vast, greatly exceeding the actual keyspace (by a thousand times for some instances!), but clients also need lookups to be really fast, ideally served from memory.
In pursuit of probabilistic data structures which could give us a dynamic compact representation of a full keyspace on proxies, we spent some time exploring Cuckoo filters. Unfortunately, with 5 billion keys it takes about 18 GB to have a false positive rate similar to Bloom filters (which only require 6 GB). And this is not only about wasted disk space — to be fast we have to keep it all in memory too!
Clearly some other solution was needed.
Finally, we decided to implement key and value separation, storing all keys on proxies, but persisting values only for cached keys. Evicting a key from the cache actually results in the removal of its value.
But wait, don’t the keys, even stripped of values, take a lot of space? Well, yes and no.
The total size of pure keys in Quicksilver is approximately 11 times smaller than the full dataset. Of course, it’s larger than any representation by probabilistic data structure, but there are some very desirable properties to such a solution. Firstly, we continue to enjoy fast Bloom filter lookups in RocksDB. Another benefit is that it unlocks some cool optimizations for range queries in a distributed context.
We may revisit it one day, but so far it has worked great for us.
Having solved all of the above challenges, one bit remained to be sorted out to make distributed query execution work: how can proxies discover replicas?
Within the local data center it is fairly easy. Each one runs its own consul cluster, where machines are registered as services. Consul is well integrated with our internal DNS resolvers, and with a single DNS request, we can get the names of all replicas running in a data center, which proxies can directly connect to.
However, data centers vary in size, servers are constantly added and removed, and having only local discovery would not be enough for the system to work reliably. Proxies also need to find replicas in other nearby data centers.
We had previously encountered a similar problem with our replication layer. Initially, the replication topology was statically defined in a configuration and distributed to all servers, such that they know from which sources they should replicate. While simple, this approach was quite fragile and tedious to operate. It led to a rigid replication tree with suboptimal overall performance, unable to adapt to network changes.
Our solution to this problem was the Network Oracle – a special overlay network based on a gossip protocol and consisting of intermediate nodes in our data centers. Each member of this overlay constantly exchanges status and metainformation with other nodes, which helps us see active members in near-real time. Each member runs network probes measuring round-trip time to its peers, making it easy to find closest (in terms of RTT) active intermediate nodes to form a low-latency replication tree. Introducing the Network Oracle was a major improvement: we no longer needed to reconfigure the topology, watch intermediate nodes or entire data centers go down, or investigate frequent replication issues. Replication is now a completely self-organized and self-healing dynamic system.
Naturally, we decided to reuse the Network Oracle for our discovery mechanism. It consists of two subproblems: data center discovery and specific service lookup. We use the Network Oracle to find the closest data centers. Adding all machines running Quicksilver to the same overlay would be inefficient because of significant increase of network traffic and message delivery times. Instead, we use intermediate nodes as sources of network proximity information for the leaf nodes. Knowing which data centers are nearby, we can directly send DNS queries there to resolve specific services – Quicksilver replicas in this case.
Proxies maintain a pool of connections to active replicas and distribute requests among them to smooth out the load and avoid hotspots in a data center. Proxies also have a health-tracking mechanism, monitoring the state of connections and errors coming from replicas, and temporarily deprioritizing or isolating potentially faulty ones.

Figure 5. Internal replica request errors
To demonstrate its efficiency, we graphed errors coming from replica requests, which showed that such errors almost disappeared after introducing the new discovery system.
Our objective with Quicksilver v1.5 was simple: gain some disk space without losing request latency, because clients rely heavily on us being fast. While the replica-proxy design delivered significant space savings, what about latencies?
Proxy

Replica

Figure 6. Proxy-replica latency comparison
Above, we have the 99.9% percentile of request latency on both a replica and proxy during a 24-hour window. One can hardly find a difference between the two. Surprisingly, proxies can even be slightly faster than replicas sometimes, likely because of smaller datasets on disk!
Quicksilver v1.5 is released but our journey to a highly scalable and efficient solution is not over. In the next post we’ll share what challenges we faced with the following iteration. Stay tuned!
This project was a big team effort, so we’d like to thank everyone on the Quicksilver team – it would not have come true without you all.
Aleksandr Matveev
Aleksei Surikov
Alex Dzyoba
Alexandra (Modi) Stana-Palade
Francois Stiennon
Geoffrey Plouviez
Ilya Polyakovskiy
Manzur Mukhitdinov
Volodymyr Dorokhov
Post Syndicated from Sean Sayers original https://www.raspberrypi.org/blog/new-hello-world-podcast-series-bringing-computer-science-into-every-classroom/
The Hello World podcast is back, accompanying the latest issue of Hello World magazine. This new three-part miniseries explores some of the topics from issue 27 of Hello World, which focuses on integrating computing education across the curriculum.
Whether you’re a seasoned educator or just starting your journey with computing education, this podcast audio and video series is full of insights, inspiration, and practical tips from educators and experts around the world.
If you’re already subscribed, the new episodes will appear automatically in your favourite podcast app every Tuesday.
In our first episode, Raspberry Pi Foundation CEO Philip Colligan, CBE, sits down with teacher Janine Kirk to discuss why, in the age of AI, it’s more important than ever for young people to learn to code. Their conversation draws on ideas from our downloadable position paper, which is also featured in issue 27 of Hello World magazine.

Next week, we’ll bring you the buzz from the Computer Science Teachers Association’s annual conference in Cleveland, USA. We’re speaking with educators at the conference to hear how they’re integrating computer science across subjects, and you’ll hear their top classroom tips for teaching CS in context.

The miniseries wraps up with an in-depth discussion about AI education around the world. Hosted by Ben Garside, Senior Learning Manager for Experience AI, the conversation features Leonida Soi, Learning Manager in Kenya; Monika Katkute-Gelzine, CEO of Vedliai in Lithuania; and Aimy Lee, COO of Penang Science Cluster in Malaysia. Monika and Aimy work with us in our global Experience AI partner programme.
Each of these three podcast episode builds on the themes in the latest Hello World issue, where you’ll find inspiration and practical tips from educators who are integrating CS across a variety of subjects and for all school ages.
Subscribe to the Hello World podcast wherever you get your podcasts to never miss an episode, and to help us reach more teachers. If you’re subscribe to Hello World magazine (it’s free), we’ll also let you know when new podcast episodes are available.
And, don’t forget to share this new podcast series with your fellow educators.
The post New Hello World podcast series: Bringing computer science into every classroom appeared first on Raspberry Pi Foundation.
Post Syndicated from Bruce Schneier original https://www.schneier.com/blog/archives/2025/07/using-signal-groups-for-activism.html
Good tutorial by Micah Lee. It includes some nonobvious use cases.
Post Syndicated from Matt Granger original https://www.youtube.com/watch?v=i8jSNqHfbyM
Post Syndicated from The Atlantic original https://www.youtube.com/watch?v=i4CJu31LMwg
Post Syndicated from Светла Енчева original https://www.toest.bg/tragediyata-v-solingen-i-dvoiniyat-standart-za-maltsinstvata/

Кога България е загрижена за малцинствата, техните права, дискриминацията и насилието срещу представителите на малцинствени групи? Когато става дума за българи. И то не по гражданство, а по етнос. Пример за това е отношението към българите в Северна Македония. Българските евродепутати внасят стотина предложения за промени в тази посока в доклада за напредъка на югозападната ни съседка по пътя ѝ към ЕС. Нашите медии следят под лупа всеки случай на насилие върху някой идентифициращ се като българин в Северна Македония.
Същевременно отношението към собствените ни малцинства е, меко казано, проблематично. България не само не признава наличието им на конституционно равнище, а системно третира представителите им като граждани втора ръка. И не се трогва особено, когато техните права се нарушават или когато върху тях се упражнява насилие.
Помните ли имената на Катя, Кънчо, Галия и Емили Жилови? Това е четиричленното българско семейство от турски етнически произход, загинало при пожар в германския град Золинген (наричан на български и Солинген) на 25 март 2024 г. – майка, баща и двете им деца. Родителите така и не навършват 30 години, Галия остава завинаги на три, а Емили – едва на няколко месеца. Ранените при пожара са десетки. А някои от оцелелите, като Айше и Нихат К., скочили от прозореца с малкия си син, получават сериозни и дълготрайни здравословни проблеми. Нихат е в кома седмици наред след трагедията, Айше и до днес трябва да се подлага на операции. Детето също не остава без поражения.
Първоначално в България се наблюдава известен интерес към случая. Има дори репортажи от мястото на събитието, например на bTV. В онзи период водещата версия е, че заподозреният за палежа германец Даниел Чала е нямал расистки мотив, а е извършил престъплението за отмъщение, защото е изгонен от сградата поради неплащане на наем. Имигранти в Золинген обаче излизат на протест с настояване причините за трагедията да се разследват по-сериозно. На него се веят български, турски и германски знамена. Случаят събужда травматичния спомен за палежа в същия град през 1993 г., извършен от четирима десни екстремисти, чийто жертви са пет турски жени и момичета.
През 2025 г. започва делото срещу Даниел Чала и постепенно излизат все повече данни, които са пренебрегнати от разследващите в началото и сочат към престъпление от омраза. В дома на обвиняемия е намерена доста нацистка литература, включително „Моята борба“ на Адолф Хитлер. Към Чала сочи и дълго неразследван пожар в близкия до Золинген град Вупертал. Той е възникнал скоро след като приятелката на Чала се е изнесла от същата сграда във Вупертал.
Въпреки смущаващите разкрития обаче, в България новините по случая не стават водещи. Отразяват го рехаво, преобладаващо протоколно и резюмират информацията от германската преса. С едно изключение.
На 3 юли в „Свободна Европа“ излиза статията на съдебно-криминалния журналист Борис Митов – „Кошмарът в Солинген. Как полицията в Германия пропусна расистки мотив за палеж, убил българско семейство“. За изготвянето ѝ не просто се обобщават известните по случая данни, а и се вземат интервюта от адвокати, ангажирани по делото: представляващия роднините на загиналата Катя Жилова – Радослав Радославов, както и Фатих Зингал и Седа Башай-Йълдъз. Башай-Йълдъз коментира:
Моят дългогодишен професионален опит, за съжаление, се потвърди – обвиняемият е действал от ксенофобски подбуди. Всеки път, когато е влизал в конфликт с чужденци, си е отмъщавал чрез палежи и е предизвиквал смъртта на невинни хора.
Какъв медиен и институционален отзвук предизвиква в България статията на Митов? Засега – никакъв, макар в нея да става дума за български граждани, вероятни жертви на престъпление от омраза.
В социалните мрежи статията също не се радва на особен интерес. Към редакционното приключване на настоящата публикация материалът на Митов е споделен във Facebook едва седем пъти. Затова пък някои от коментарите под него са пример за реч на омразата. Те внушават, че жертвите не са българи, а турци и/или роми и затова не заслужават нито институционално, нито медийно внимание. Камо ли справедливост. Важното е не дали са писани от тролове, или не (някои от авторите им имат съмнително малко контакти), а че не срещат особен отпор. Ето няколко примера:

В Instagram се намират повече хуманни коментари върху статията на Борис Митов, но има и обвинения към жертвите, че всъщност са роми, които се правят на турци:

Традегията в Золинген е поредният пример, че Германия не е имунизирана срещу избирателно отношение към престъпленията от омраза. Освен общо трите пожара (в Золинген от 2024 г. и 1993 г.) и Вупертал (от 2022 г.) си струва да припомним и историята с Беате Чепе и партньорите ѝ в любовта и престъпленията Уве Мундлос и Уве Бьонхарт, които между 2000 и 2007 г. убиват десет души. Жертвите им са осем турци, една гъркиня и германска полицайка. Деянията им години наред остават неразкрити, защото разследващите не правят връзка между членовете на групичката и не допускат, че може да става дума за престъпления от омраза.
След разкриването на групата на Чепе Германия дава вид, че си е взела поука. Делото срещу Чепе е обект на силен медиен и обществен интерес, а животът на адвокатите ѝ се обръща с главата надолу. От една страна, клиентката им ги разиграва, както си поиска, от друга, те са възприемани като злодеи, защото я защитават. На едната адвокатка дори се налага да се пресели в друг град, за да избегне част от тормоза и да има работа.
Въпреки това обаче немските разследващи продължават да допускат същите грешки, което позволява на радикализирани германци като Даниел Чала да убиват нееднократно. Когато едно масово убийство се извършва от чужденец, особено мюсюлманин, обикновено първата хипотеза е тероризъм, дори да става дума за психически проблем. А ако то е дело на етнически германец, се търси психически или битов проблем.
Представителите на етническите малцинства се възприемат – и третират – не само като не-българи, а и като непритежаващи основни човешки права, включително правото на живот, и незаслужаващи справедливост. Това не се променя, колкото и пъти Европейският съд за правата на човека да осъжда България за етническа дискриминация.
Затова хора като кмета на столичния район „Илинден“ Емил Бранчевски ще нарушават международното законодателство и ще продължават да трупат имидж с цената на оставянето на десетки ромски семейства без дом. Затова след края на ерата на Ахмед Доган българските турци ще продължават да гласуват под натиск и от страх от нов Възродителен процес.
Но ако питате нашите патриоти, най-важно е Северна Македония да спазва човешките права на представителите на малцинствата и да не ги дискриминира. Нашите си малцинства са си наша работа и никой не трябва да ни се меси в отношението към тях.
Post Syndicated from Matt Granger original https://www.youtube.com/shorts/kf01eMxKPDQ
Post Syndicated from corbet original https://lwn.net/Articles/1028368/
Inside this week’s LWN.net Weekly Edition:
Post Syndicated from The History Guy: History Deserves to Be Remembered original https://www.youtube.com/shorts/zVq0g5vcvYI
Post Syndicated from Explosm.net original https://explosm.net/comics/moment
New Cyanide and Happiness Comic