Tag

distributed-systems.

11 writings found

Latest Archives

How Meta's ZGateway Proxy Solves the Million-Client Problem

Understanding how a shared proxy tier manages connection meshes, enables cross-client batching, and scales reliability at hyperscale infrastructure.

The Proxy Layer Pattern: Where Infrastructure Gets Smart

How Meta's ZGateway shows why interposing a managed tier between clients and backends solves problems no individual client library can.

Why Proxies Win at Scale: Learning from Meta's ZGateway

How interposing a stateless proxy tier between millions of clients and a shared backend solves reliability and efficiency problems that client libraries cannot.

MetaRoCE: Rethinking Network Transport for AI at Scale

Meta's new MetaRoCE protocol redesigns Ethernet transport for million-GPU clusters, prioritizing edge intelligence over fabric control. What this means for infrastructure engineers.

Aurora DSQL Goes Global: What Multi-Region Consistency Means for You

Aurora DSQL expands to 16 regions with multi-region clusters. I break down why active-active consistency matters for distributed systems and what it changes for developers.

etcd v3.7.0 Released: RangeStream, v2 Store Removal and More

etcd v3.7.0 ships RangeStream RPCs, drops the legacy v2 store bootstrap, overhauls protobuf deps, and brings real CPU savings for Kubernetes clusters.

Cloudflare Meerkat: Building Consensus Without a Leader

Cloudflare's new Meerkat service uses QuePaxa to eliminate leader-dependent consensus bottlenecks across 330+ global data centers. Here's why it matters.

Meta's Adaptive Ranking Model: The Real Cost of Serving Trillion-Parameter Ads

Meta scaled ads recommendations to LLM complexity while keeping latency under a second. Here's why their inference trilemma solution matters beyond advertising.

Meta's RCCLX: Why AMD's GPU Communication Stack Just Got Interesting

Meta open-sources RCCLX with Direct Data Access and FP8 collectives for AMD GPUs. A deep look at what this means for multi-GPU AI workloads.

Meta's RCCLX: Why AMD GPU Communication Just Got Interesting

Meta open-sources RCCLX with Direct Data Access and low-precision collectives, potentially reshaping distributed AI workloads on AMD hardware.

Meta Open Sources RCCLX: AMD Gets Serious Performance Boosts for AI Workloads

Meta's RCCLX brings Direct Data Access and low-precision collectives to AMD GPUs, delivering 10-50% speedups for LLM inference on MI300X hardware.

View all rollups →