Saving another 100TB of RAM with math (and Rust)
Recorded: Sept. 18, 2026, 8 p.m.
| Original | Summarized |
Saving another 100TB of RAM with math (and Rust) | Cloudflare Blog Skip to contentAll CategoriesAIDevelopersRadarProduct NewsSecurityPolicy & LegalZero TrustSpeed & ReliabilityLife at CloudflarePartnersSwitch Site LanguageEnglishDeutschEspañolEspañol (Latinoamérica)FrançaisItaliano日本語한국어繁體中文简体中文PortuguêsРусскийBahasa IndonesiaภาษาไทยTiếng ViệtPolskiالعربيةעבריתSvenskaNederlandsTürkçeLoginDashboardContact SalesBlogDeep DiveEngineeringOpen Source+4Show 4 more tags7 TagsShow 7 tagsPost TagsDeep DiveEngineeringOpen SourceOptimizationPerformancePingoraRustAll tagsMatching tagsNo tags found1.1.1.12FAAbuseAccessAccess Control Lists (ACLs)AccessibilityAccount TakeoverAcquisitionsAddressingAdvanced Certificate ManagerAdvanced DDoSAdvertisingAegisAEOAfricaAfroflareAgent CloudAgent Development LifecycleAgent ReadinessAgentsAgents WeekAgents Week 2026AIAI BotsAI GatewayAI SearchAI WAFAI WeekAI-SPMAlertmanagerAlways OnlineAMDAMPAnalyticsAnonymousAnti MalwareAnycastAPIAPI GatewayAPI SecurityAPI ShieldAPJCAppleApplication SecurityApplication ServicesArea 1 SecurityArgo Smart RoutingASCIIAsiaAthenian ProjectAtlassianAttacksAudit LogsAustinAustraliaAuthenticationAuthyAuto RagAutomatic HTTPSAutomatic Platform OptimizationAutomationAutoMinifyAwardsAWSBaiduBandwidth AllianceBandwidth CostsBest PracticesBetaBetter InternetBGPBillingBirthday WeekBlack FridayBlackbirdBot Fight ModeBot ManagementBotnetBotsBPFBrandBrand ProtectionBrazilBrowser InsightsBrowser RenderingBrowser RunBug BountyBugsBYOIPCacheCache PurgeCache ReserveCache RulesCaliforniaCanadaCap'n ProtoCAPTCHACareersCASBCategoriesCDNCDNJSCertificate AuthorityCertificate PinningCertificate TransparencyCertificationCFSSLChallenge PageChatGPTChinaChina NetworkChristmasChromeCIO WeekCISAClaireCLIClickHouseClientlessClientless Web IsolationCloud ConnectorCloud Email SecurityCloudflare AccessCloudflare AppsCloudflare Area 1Cloudflare CallsCloudflare Email ServiceCloudflare for CampaignsCloudflare for SaaSCloudflare for StartupsCloudflare GatewayCloudflare HistoryCloudflare ImagesCloudflare Media PlatformCloudflare MeetupsCloudflare NetworkCloudflare OneCloudflare One ClientCloudflare One User Risk ScoreCloudflare One WeekCloudflare OSCloudflare PagesCloudflare PolishCloudflare QueuesCloudflare RealtimeCloudflare StreamCloudflare TunnelCloudflare TVCloudflare WorkersCloudflare Workers (PT)Cloudflare Workers KVCloudflare Workers KV (ES)Cloudflare Zero TrustCloudforce OneCloudyCode OrangeCoinbaseColombiaCommunityComplianceCompressionConfig RulesConfiguration ManagementCongestion ControlConnectivityConnectivity CloudConsumer ServicesContainersContent Independence DayContent ScanningContextCoreCOVID-19Crawler HintsCrowdStrikeCrypto WeekCryptographyCSAM ReportingCustomer SuccessCustomer ZeroCustomersCVECVE-2023-50387Cyber ReadinessCybersecurityD1DashboardDataData CatalogData CenterData LocalizationData Localization SuiteData LossData Loss PreventionData PlatformData Privacy DayData ProtectionData SovereigntyData Transfer BucketDatabaseDDoSDDoS AlertsDDoS ReportsDebuggingDeep DiveDescalerDesignDeskopeDeveloper DocumentationDeveloper PlatformDeveloper SpotlightDeveloper WeekDevelopersDevelopers StorageDevice SecurityDevOpsDEXDigital Experience MonitoringDigital ForensicsDisruptDistributedDistributed SystemsDistributed WebDiversityDLPDMARCDNSDNS FilteringDNS FloodDNS SecurityDNSSECDogfoodingDoHDomain RankingsDomain Scoped RolesdosdDrupalDue ProcessDurable ExecutionDurable ObjectsEarly HintsEarth DayeBPFEC2eCommerceEdgeEdge ComputingEdge DatabaseEdge RulesEducationEgressElasticElection SecurityElectionsElliptic CurvesEmailEmail RoutingEmail SecurityEmail WorkersEmDashEmissionsEmployee Resource GroupsEncrypted SNIEncryptionEngineeringEnterpriseEntropyEPYCEthereumEuropeEuropean UnionEventsExploitFacebookFancy BearFast FontsFCCFeature FlagsFedRAMPFedRAMP HighFedRAMP ModerateFirefoxFirewallFirmwareFloridaFootballFormal MethodsForresterFortranFoundation DNSFounders' LetterFranceFraudFreeFreedom of SpeechFront EndFull StackFull Stack WeekFunGA WeekGartnerGatebotGDPRGen XGeneral AvailabilityGenerative AIGeo Key ManagerGermanyGitHubGoGoogleGoogle AnalyticsGoogle CloudGoogle WorkspaceGovernment InnovationGrace HopperGrafanaGraphQLGreenGrinchGrowthgRPCGuest PostHackathonHalloweenHardwareHashiCorpHertzbleedHeuristicsHistoryHolidaysHolocaustHong KongHosting ConHostnamesHTTP2HTTP3HTTPSHuman RightsHurricaneHybrid CloudHyperdriveI'm Under Attack ModeIBMICANNiCloud Private RelayIdentityIETFIETF StandardsIL4Image OptimizationImage RecognitionImage ResizingImage StorageImpactImpact WeekIncident ReportIncident ResponseIndiaIndicators of CompromiseIndonesianInferenceInfrastructureInfrastructure as CodeInsightsIntelInterconnectionInternal DNSInternet PerformanceInternet QualityInternet RegulationInternet ShutdownInternet SummitInternet TrafficInternet TrendsInternship ExperienceIntrusion DetectionInvestorsIoCsiOSIoTIPFSIPsecIPv4IPv6IRAPIsraelItalyIWDJAMstackJapanJavaScriptJengoJengo PolicyJoomlaJudeoflareKafkaKernelKey ValueKeyless SSLKeyTrapKillnetKoreaKubernetesLangChainLatencyLatin AmericaLatinflareLavaRandLazarus groupLeaked Credential ChecksLegalLegal Patents SableLGBTQIA+Life at CloudflareLinuxLisbonLive StreamingLlamaLLMLoad BalancingLocalizationLog PushLog4JLog4ShellLoggingLogsLUAMachine LearningMagecartMagic FirewallMagic Network MonitoringMagic TransitMagic WANMagic WAN ConnectorMalicious JavaScriptMalwareManaged ComponentsManaged RulesMarch of CloudflareMASQUEMCPMeerkatMeetUpMerisMessage ProtocolMexicoMicro-frontendsMicrosoftMicrosoft 365Microsoft AzureMiddle EastMigration HubMilestonesMiniflareMirageMiraiMitelMitigationMixed Content ErrorsMLopsMobileMobile SDKModel Context ProtocolMoldovaMonitoringMulti-CloudMulti-UserMySQLNaaSNet NeutralityNetworkNetwork InterconnectNetwork Performance UpdateNetwork ProtectionNetwork ServicesNetworkingNew YearNGINXNinjasNISTNode.jsNorth AmericaNotebooksNotificationsNSEC3OAuthObservabilityOceaniaOCSPOfficesOktaOlympicsOnboardingOpen APIOpen SourceOpenAIOpenBMCOpenDNSOpenSSLOpenTelemetry OptimizationOrigin RulesOutageOxyPacific NorthwestPage RulesPage ShieldParallelsPartnersPartnershipPassword-reusePasswordsPasswords (PT)PatentsPay Per CrawlPAYGOPaymentsPCI CertifiedPeeringPerformancePerformance OptimizationPhishingphpPhythonPingoraPipelinesPlanetScalePlansPlatform EngineeringPlatform WeekPleskPolicy & LegalPoliticsPortugalPost MortemPost-QuantumPostgresPrecursorPrepared StatementsPrismaPrivacyPrivacy PassPrivacy WeekPrivate IPPrivate NetworkProduct DesignProduct NewsProgrammingProgramming (PT)Project Fair ShotProject GalileoProject Honey PotProject PangeaProject SafekeepingProject TurpentinePrometheusProtocolsProudflareProxyingPublic SectorPythonPython WorkersQuantizationQueuesQUICQUICHEQuicksilverR2R2 Super SlurperRadarRadar AlertsRadar APIRadar MapsRailgunRandomnessRansom AttacksRapid ResetRaspberry PiRate LimitingRC4RDDoSReactReading ListReal-timeRecruitingRegional ServicesRegistrarReliabilityRemote Browser IsolationRemote Desktop Protocol Remote WorkReplicationResearchResolverRestreamingRetreatReverse EngineeringREvilRisk ManagementRoad to Zero TrustRocket LoaderRocksDBRoutingRouting SecurityRPCRPKIRRDNSRSARussiaRustRust WorkersSaaSSAAS SecuritySableSaltSamplingSandboxSASESave The WebSDKSearch EngineSecrets StoreSecure Web GatewaySecuritySecurity AnalyticsSecurity CenterSecurity PostureSecurity Posture ManagementSecurity Service EdgeSecurity Weeksecurity.txtSEOServer PushServerlessServerless (PT)Serverless AIServerless WeekServersSIEMSigned Exchanges (SXG)SIMSingaporeSingle Sign On (SSO)Smart PlacementSmart ShieldSnippetsSOC as a ServiceSouth AfricaSouth AmericaSpainspdySpectrumSpeedSpeed & ReliabilitySpeed BrainSpeed WeekSpoofingSportsSQLSRESSESSHSSLStandardsStartup Enterprise PlanStatisticsStopTheHackerStorageSumo LogicSuper BowlSupercloudSupply Chain AttacksSupportSustainabilitySWAGSWGSwiftSwitzerlandSXSWSYNSYN FloodSyriaTCPTeamTeams DashboardTech TalksTechCrunchTechnical WritingTerraformTestimonialsTestingTexasThanksgivingThe Serverlist NewsletterThreat DataThreat FeedsThreat IntelligenceThreat OperationsThreat ReportThreatsTiered CacheTikTokTLSTLS 1.3ToolsTorTracingTrafficTransform RulesTransparencyTrendsTrust & SafetyTTFBTTLTURNTURN ServerTurnstileTypeScriptUDPUkraineUnited KingdomUniversal SSLURL ScannerUSAUser ResearchVDIVectorizeVetflareVideoVisibilityViteVoIPVPCVPNVulnerabilitiesWAFWAF Attack ScoreWAF RulesWaiting RoomWARPWARP ConnectorWASMWeb Application FirewallWeb Asset DiscoveryWeb3WebAssemblyWebinarsWebMCPWebPWebRTCWebSocketsWildebeestWomenflareWordPressWorkers AIWorkers LaunchpadWorkers LogsWorkers ObservabilityWorkers SitesWorkers UnboundWorkers VPCWorkflowsWorld IPv6 DayWranglerx402Year in ReviewZ3ZarazZero Day ThreatsZero TrustZero Trust WeekZone VersioningArtificial IntelligenceWorkersClient-Side SecurityOptimizationPerformancePingoraRustDeep DiveEngineeringOpen SourceOptimizationPerformancePingoraRustSeptember 18, 2026Saving another 100TB of RAM with math (and Rust)Kevin Guthrie, Mariia Iurchenko, Zaidoon Abd Al Hadi, and Ivan Babrou13 minute readCOPY URLCloudflare operates at a scale so big that even after working here for years, it doesn’t seem real. We have thousands of servers all over the world with petabytes of RAM and millions of CPU cores, and all of it is pushed to the max. As vast as those resources feel, they are still finite, and when you need every service to run on every node, it doesn’t leave room for wasted space.At this scale, small improvements are greatly magnified, so even 1%-at-a-time improvements are worth celebrating. And some tweaks add up to a lot more: in this post, we’ll look at how small changes to a single algorithm reduced the memory footprint of one of our Pingora-based services significantly. That allowed us to reclaim more than 100TB of RAM globally, on top of the 100TB of memory the DNS team was able to shed last month.Waste notMaintaining equitable resource sharing between teams is not easy, especially in large organizations. One of the ways Cloudflare ensures the balance is kept is through the tireless efforts of the wonderful Performance team. This story starts with a ticket filed by Ivan who found: Excessive memory usage from pingora-ketama in Pingora Backend Router. The finding was that our internal load-balancing service, Pingora Backend Router (yes, PBR), was using significantly more memory than expected — specifically in structures associated with pingora-ketama, which is our open-source library for handling consistent hashing.In order to talk about how we addressed this seeming overuse of memory, we need to talk about what consistent hashing even is, why we are using it in PBR, and how it became so memory hungry. Along the way, we’ll learn some Rust and even a little math.Consistent hashingConsistent hashing is a widely used method for distributing tasks across multiple servers in a way that does not require large changes when servers are added or removed. Internally we use it to route cacheable requests to servers by URL. This allows us to keep only one copy of a file stored per data center and gives a stable way to find the location of each file. We have mentioned this system before, but let’s take the time to walk through how and why this algorithm is used and how it works.The key concept of consistent hashing is that while hash functions can accept any kind of input, their output is limited to a single unsigned integer (32, 64, or 128-bit integers depending on which hash function). This allows us to relate tasks and servers to each other in a consistent way. Most discussions of consistent hashing have you think of that output space as a continuous, circular ring that wraps around from its max value to zero. This depiction makes for some nice visualizations, but it can also make the simple concept of integer ranges seem more complicated than it needs to be. For our discussion, we’ll represent the 32-bit output of our hash function as a number line.Now, let’s say we have a set of servers, A, B, & C, and a set of tasks t-z. We can map each onto the number line based on the hash of their representative values, so something like IP addresses for servers and cache keys for tasks.Assigning tasks to servers is now just a matter of finding the first server to the left of each task. We can represent this visually by coloring in the region of hashes that will be associated with each server. Notice that the range covered by server C wraps around to the beginning, hence the idea that hashes exist in a ring.And that’s it. At a base level, consistent hashing is this simple — but it doesn’t take long to see that there is room for improvement. Notice that the range covered by server A in our example is significantly larger than that of either B or C. This is a problem because the fraction of the requests a server handles is going to be proportional to the size of its range on the number line. Ideally we would like to guarantee each server will have an equal size, but because hashes are essentially random numbers, we have to talk about the size of the regions in terms of statistics. 😨Math and consequencesFirst: don’t panic. I promise I'm not about to lie to you and that we will stay safely within the bounds of a day-one probability lesson. When we talk about statistical distributions, there are two big factors that help us quantify uncertainty in helpful ways: expected value and standard deviation. In (over-)simplified terms, expected value gives us a point where measurements based on a distribution will be centered, and standard deviation tells how close to that central point most measurements are likely to be.For consistent hashing, we can calculate these factors for the fractional size of the range associated with one of N servers. (Details on where this formula comes from later).$$m \begin{align*} m$$In terms of concrete numbers, let’s say we have 100 servers. The formulas above give:$$m \text{Exp}=1/100 = 1\% \\ \text{SD}= \frac{1}{100}\sqrt{\frac{100-1}{100+1}} \approx 0.99\% m$$That tells us that we can expect that the range each server handles will be centered around 0.99% of the total and most of the lengths to fall within 1% of what's expected. This sounds good until we realize that that’s 0.99% of the total length. We need to scale the standard deviation by the expected value to see how big the error is as a fraction of the target size. This value is called the coefficient of variation. $$m \text{CV} = \frac{\text{SD}}{\text{Exp}} = \sqrt{\frac{N-1}{N+1}} m$$At $m N=100, \text{CV} \approx 99\% m$ — meaning some servers will likely be working 99% harder than they should be (handling twice as many requests) while others could be doing practically nothing! Now that we have a way to predict how evenly loaded servers will be using consistent hashing, we can start working on improvements.What if we add hashes?The simplicity of consistent hashing is a double-edged sword. It’s easy to understand and implement because everything is turned into easily-relatable hashes on the same numberline, but any improvements to the system will also need to be relatable to that numberline. That means the solution to any consistent hashing problem can only be more hashes. It’s less like a golden hammer (a tool with which all problems look like nails) and more like a golden nail in that it turns all tools into hammers.To solve the problem of imbalanced workloads, we can add multiple hashes to represent each server instead of just one. We’ll get to the math behind this momentarily, but it should make some intuitive sense that while each individual range has a large standard deviation, adding a bunch together should make their total size even out. If we take our three-server example from the above diagrams and add two more hashes at random for each server, we see that it helps even out each server’s workload. This is an admittedly contrived example. The random nature of the system means there’s no guarantee how much improvement you will get from adding 2 additional hashes per server, but it should make some intuitive sense that combining more of these hash segments together produces a more even distribution. Each segment in the sum has a chance of balancing another. Maybe one is too short; maybe one is too long. This is essentially what the law of large numbers tells us should happen… The obvious problem is it only works for large numbers. In NGINX, the baseline number of hashes per server is hardcoded to 160, and Pingora uses the same value as the default. I’ll spare you the math for now, but if we go back to our 100-server example, if we use 160 points per server instead of just one, the coefficient of variation (which we can think of like an error margin) drops from about 99% to about 8%, a significant improvement.What if we add more hashes?We saw above that increasing the number of hashes per server by a constant amount allows us to improve how evenly workloads are distributed per server, but what if we don’t want to distribute the work evenly? In Cloudflare’s case, we have some servers that have more storage space than others, so it would be better to have the number of requests allotted to a server be proportional to its disk space. One way to accomplish this is with the ketama algorithm. The naming is a little funny because the algorithm is named after the library where it was first implemented, and the library was named … well you can google it 😶🌫️.The whole algorithm boils down to: For any two servers, $m S_1m$ & $mS_2m$, if we want the requests served by $mS_1m$ to be $mw\timesm$ more than those served by $mS_2m$, the number of hashes associated with $mS_1m$ needs to be $mH_1 = w\times H_2m$. This allows us to set a “weight” for each server, which scales the number of hashes associated with that server. Unfortunately this is not a replacement for the constant scale factor we added in the section above. That scaling needs to be there to set a minimum error margin, which will show up in the servers with the lowest weights.For us, since we want workload to be scaled based on storage, we can use the disk space as the weight, which is exactly what the Pingora team has been doing for years. Elsewhere in the company where workloads are more compute-intensive, weights might be based on CPU or GPU count.What if we add even more hashes???The last problem we need to address is that so far we are working under the assumption that any server can handle any request, but in practice that is not the case. Things like compliance requirements or enabled caching features mean only a subset of servers can handle any particular request. Unfortunately, unlike before, we can’t solve this problem by adding more hashes to the same ring. We have to add completely new rings, and not only that — every combination of features potentially needs its own specific ring!Duplication based on combinations is a classic recipe for exponential explosion. In our case, we have a handful of different features leading to $m2^\text{handful} = \text{dozens}m$ of separate consistent hash rings. So as you have probably guessed by now, the "excessive memory use" (6GB in some cases) that Ivan found was due to an enormous number of hashes to accommodate all the functionality we need and which have to be stored in memory. So what can we do?Storage improvementsOne big improvement came from Zaidoon, who had an insight about our struct for storing hashes in PBR. That struct looks like this:struct Point { impl Point { fn index(&self) -> u16 { \text{Exp}_k = \frac{1}{N}, \text{SD}_k=\sqrt{\frac{(k+1)}{N(kN+1)}-\frac{1}{N^2}} m$$To see how increasing the hash count improves the accuracy, we need to look again at the coefficient of variation.$$m \text{CV}_k=\frac{\text{SD}_k}{\text{Exp}_k}=\sqrt{\frac{N-1}{(N*k+1)}} m$$Plotting $m\text{CV}_km$ shows a potential problem with the “just add more hashes” mentality (other than overusing RAM).You can see each step down in error margin requires (almost) an order of magnitude increase in the number of hashes per server, so adding more hashes yields less and less improvement. Recall that we are using a base of 160 hashes scaled by the server's storage size. To make the math easier, we'll say the weighting factor $m{m_w}m$ for a server is 625, so we get $m{k = 160\times625 = 100{,}000}m$. We can see from the chart above that the last 90,000 hashes we added are buying us a minuscule 0.7% reduction in error. Unfortunately things get even worse from there.The predictions from my beautiful math only work if we think about hashes in a continuous ring, but in practice we use 32-bit numbers for the hashes that have the potential for collisions, and the probability of collisions goes up surprisingly quickly as the number of hashes increases (see the birthday paradox). Collisions matter because in the ideal case, every hash contributes to the volume and distribution of requests handled by the associated server, but a collision means some contributions are randomly dropped, introducing unpredictable error. If we compare some simulated results with 32-bit hashes with the predicted error rate, we can see that for data centers with 2048 servers, the error rate increases: between 10,000 and 100,000 hashes per server.Ultimately, even though this realization feels kind of bad, it’s great news for our plan to reclaim some RAM! Now that we have some math to back it up, we determined that we could decrease the number of hashes we were generating for each server by 90% without incurring any appreciable error, so that is what we set out to do.Migrating without melting originsThere was one more problem: changing the hash ring changes where some cacheable requests go. Even if the new ring is better, switching the whole network at once would effectively invalidate almost all cached content. It would turn a memory optimization into an apocalyptic increase in origin traffic.So we did not make this a single global flip. For a while, PBR carried both versions of the cacheable load balancer in memory: the old ketama ring and the new smaller one. Each request used our normal migration framework to decide which ring should select the backend. That meant the rollout decision was stable per request hash, and it also gave us a clean rollback path. If anything looked wrong, we could send new requests back through the old ring without redeploying PBR.We then rolled the migration out in layers. We started with small validation locations, moved through progressively larger groups of data centers, and only then continued toward the rest of the world. The important part was that we controlled two dimensions independently: how much traffic used the new ring, and where that traffic was allowed to move. A plain global percentage rollout would have spread cache churn everywhere at once. Data-center-scoped rollout kept the blast radius small and made it much easier to tell whether a change was actually safe.During the migration, we watched backend-selection traces, ring-version counters, PBR connection errors, process memory, startup time, cache behavior, and origin traffic. Once the migration reached 100%, we removed the temporary old-ring path, and voila!The chart above shows the comparison of the memory used by PBR the week of the change compared with data from a few weeks before, as well as the result of subtracting one from the other. The sharp drop is the day where the version of PBR with the large (now unused) hash rings was decommissioned forever. Looking at the difference, we get the satisfying result that our changes dropped the used memory by 100TB!Try it yourselfAll the changes we talked about in this post are available now in the pingora-ketama crate in the form of a (for now) unadvertised cargo feature. The v2 ring has the compacted storage format, a faster sorting method, and the ability to scale the base number of hashes per node. Our focus in making these changes had to be on stability and control, so the v1 ring is identical to what pingora ketama has always used, and the library makes it possible to run both simultaneously and decide on a request-by-request basis which to use and when. Beyond trying our literal consistent hashing changes, I would like you to take away from this some inspiration to dig into your own systems to see what “simple” or “obvious” decisions are hiding potential wins, if you’re willing to get into the numbers. You might not be able to solve all your problems with Rust, but math is universal.On this pageDiscuss OnlineRelated tagsDeep DiveEngineeringOpen SourceOptimizationPerformancePingoraRustFollow on Social MediaCloudflareKevin GuthrieMariia IurchenkoSubscribe to receive notifications of new postsEmail addressWe’ll never share your email address.SubscribeThanks for subscribing! Check your inbox to confirm.Getting StartedPlansContact salesPartnersFind a partnerStartupsUnder attack?Domain name searchPublic interestProject GalileoAthenian ProjectCloudflare for CampaignsProject FairshotImpact/ESGResourcesApp innovation reportCloudflare RadarCase studiesStatusSupportEventsBlogSolutionsSSE and SASE platformCloudflare AI CloudAI SecurityFrontend Development PlatformMulti-Tenant Platform DevelopmentWeb Security PlatformCompanyAboutCareersInvestorsPressPress kitGlobal networkComplianceCompliance resourcesTrustGDPRResponsible AITransparency reportReport abuseDevelopersDocumentationLearning centerCommunityStart buildingLogin© 2026 Cloudflare, Inc.|Your privacy choicesReport security issues|Privacy Policy|Terms of use|GDPR|TrademarkSearch is temporarily unavailable.ProductsSolutionsResourcesPricingLogin opens in a new tabDashboard opens in a new tabContact Sales opens in a new tabStart Building opens in a new tab opens in a new tab opens in a new tab opens in a new tabAll CategoriesAIDevelopersRadarProduct NewsSecurityPolicy & LegalZero TrustSpeed & ReliabilityLife at CloudflarePartnersEnglishSwitch Site LanguageEnglishDeutschEspañolEspañol (Latinoamérica)FrançaisItaliano日本語한국어繁體中文简体中文PortuguêsРусскийBahasa IndonesiaภาษาไทยTiếng ViệtPolskiالعربيةעבריתSvenskaNederlandsTürkçeLightDark |
Cloudflare manages an infrastructure of immense scale, necessitating continuous optimization where even marginal improvements yield significant resource savings. This post details how small adjustments to an underlying algorithm, specifically in the Pingora-ketama consistent hashing implementation within the Pingora Backend Router (PBR), resulted in the reclamation of over 100 terabytes of RAM globally. The optimization began with an investigation into excessive memory usage stemming from the consistent hashing mechanism. Consistent hashing is employed to distribute tasks across multiple servers dynamically, allowing for the addition or removal of servers without major disruptions. Conceptually, it maps tasks and servers onto a circular number line. While simple, the distribution of requests across this ring is statistically uneven because hash outputs are random. To quantify this uncertainty and guide optimization, the authors established mathematical measures, including the expected value and standard deviation, which are crucial for understanding the distribution of workload across servers. They introduced the coefficient of variation to measure the error margin, which revealed that the system suffered from significant imbalance, illustrating that servers often handled workload far exceeding their proportional capacity. The team explored several avenues for improvement. One strategy involved increasing the number of hashes per server, which intuitively seemed like a solution to distribute the load more evenly. However, mathematical analysis, specifically calculating the coefficient of variation, demonstrated that this approach yields diminishing returns. While adding more hashes reduces the error margin, the gains diminish rapidly, and practical concerns about the increased probability of hash collisions, exacerbated by the use of 32-bit numbers, introduced unpredictable errors. The mathematical findings suggested that increasing the number of hashes provided negligible error reduction beyond a certain point, leading to a determination to reduce the total number of hashes used. Another optimization involved incorporating server-specific weights, based on factors like disk space, into the hashing process using the ketama algorithm. This approach allows the system to scale the number of associated hashes proportionally to a server's capacity, thereby ensuring that the workload distribution reflects actual resource availability. Furthermore, implementing these changes required addressing memory structure itself. An initial analysis revealed that the memory footprint of storing hash data was unnecessarily large due to alignment rules in the Rust programming language. By restructuring the data storage to use raw byte arrays for hashes and indices, the team managed to reduce the memory consumption for consistent hashing by twenty-five percent. Finally, transitioning the hash ring itself presented a significant challenge, as altering the ring would invalidate existing cached content. To manage this risk, the migration was executed through a layered, progressive rollout across data centers. This method allowed them to independently control the traffic migration and validation checks, ensuring that the rollout was stable per request hash and provided a controlled rollback path. By monitoring backend selection and cache behavior during the migration, the team confirmed the stability of the changes, ultimately decommissioning the large, unused hash rings and achieving the substantial memory reduction. The work demonstrated that applying rigorous mathematical analysis and careful system design could lead to massive aggregate savings. |