Expert Track · Modules 47–73
27 modules.
The depth arc.
The Intermediate track taught you to compose distributed systems. This one takes you underneath them — consensus proofs, storage and query engine internals, hardware, formal verification — and then out to the production disciplines that define staff-level work: observability, chaos, multi-tenancy, DR, FinOps, platform engineering, and privacy.
Haven't finished the earlier tracks? Start with the Intermediate track →
Phase J · Consensus & Correctness — Modules 47–52
M.47
→
M.48
→
M.49
→
M.50
→
M.51
→
M.52
→
Paxos & Consensus
Intermediate said "use a quorum." This shows the two-phase message protocol underneath it — proposal numbers, prepare and accept, Multi-Paxos, and how Raft trades generality for teachability.
Byzantine Fault Tolerance
Crash faults are the honest assumption inside a datacenter. Byzantine faults — nodes that lie, equivocate, or collude — need 3f+1 replicas and PBFT's three phases. Why, and when it's worth the cost.
Time & Causality
Consensus assumed events could be ordered. This asks how. NTP skew, logical and vector clocks, hybrid logical clocks, and TrueTime's explicit uncertainty interval.
Linearizability & Consistency Models
Every vendor claims "strong consistency" and each means something different. The precise definitions — linearizability, serializability, strict serializability, snapshot isolation — and why two of them are orthogonal.
Distributed Transactions
Atomic commit when data spans nodes. Two-Phase Commit and its blocking problem, why 3PC didn't solve it, Percolator's client-driven commits, and Spanner's 2PC-over-Paxos.
Geo-Distributed Transactions
Transactions at continental scale, where the speed of light sets the floor. Spanner's TrueTime and commit-wait, CockroachDB's HLC and read restarts, and Calvin's deterministic pre-ordering.
Phase J · Systems Internals — Modules 53–56
M.53
→
M.54
→
M.55
→
M.56
→
Storage Engine Internals
Below every SQL API sits a storage engine. B+ trees for read-heavy OLTP, LSM trees for write-heavy ingest, compaction strategies, and the write/read/space amplification triangle.
Hardware-Aware Design
Below the storage engine sits silicon. Cache hierarchy and cache lines, false sharing, NUMA locality, branch prediction, NVMe queue depth, and when kernel bypass earns its complexity.
Formal Methods & TLA+
Testing finds bugs by execution; formal methods prove their absence by exhaustive search. TLA+ specifications, invariants, model checking with TLC, and the design bugs AWS found this way.
Custom Network Protocols
Framing, multiplexing, reliability, congestion control — the four decisions in any wire protocol, why head-of-line blocking made QUIC necessary, and when rolling your own is justified.
Phase J · Data Systems at Depth — Modules 57–63
M.57
→
M.58
→
M.59
→
M.60
→
M.61
→
M.62
→
M.63
→
Query Engine Internals
A query engine is a compiler. Parsing to AST, logical rewriting, cost-based physical planning from statistics, join algorithm selection, and vectorized versus compiled execution.
Vector Databases & ANN
Brute force over a billion embeddings is not an option. HNSW navigable graphs, IVF partitioning, product quantization — and the recall, latency, and memory triangle you pick a corner of.
Search Engine Internals
Turning a text query into ranked results in milliseconds. Inverted indexes and posting lists, BM25 scoring, analysis and tokenization, faceting, and multi-stage reranking.
Time-Series Databases
A million points per second, compressed 12× by Gorilla encoding. Delta-of-delta timestamps, XOR float compression, retention and downsampling — and why cardinality is the real ceiling.
Streaming Architectures
Exactly-once end to end, not just in the engine — replayable sources, checkpointed state, transactional sinks. Plus backpressure propagation, watermarks, and reprocessing from a retained log.
Data Lakehouse Architectures
Warehouse semantics on open files. Delta Lake, Iceberg, and Hudi compared — how a metadata layer gives object storage ACID commits, snapshot isolation, time travel, and schema evolution.
Data Mesh & Data Contracts
When the central data team becomes a six-month bottleneck. Domain ownership, data as a product, self-serve platform, federated governance — and the contracts that break producer builds instead of consumer dashboards.
Phase J · ML Systems — Modules 64–65
M.64
→
M.65
→
Distributed ML Training
When a 70B-parameter model fits on no single GPU. Data, tensor, and pipeline parallelism; ZeRO and FSDP memory sharding; collective communication; and 3D parallelism at frontier scale.
Inference Serving at Scale
From 5% GPU utilization to 60%. PagedAttention's virtual-memory approach to KV cache, continuous batching at iteration granularity, speculative decoding, quantization, and prefill/decode disaggregation.
Phase J · Production Engineering — Modules 66–72
M.66
→
M.67
→
M.68
→
M.69
→
M.70
→
M.71
→
M.72
→
Observability at Scale
Observability as an economics problem. Cardinality budgets, head versus tail sampling, exemplars linking a metric spike to one trace, OpenTelemetry standardization, and eBPF instrumentation.
Chaos Engineering & Resilience
From "we hope it's resilient" to verified. Steady-state hypotheses, controlled fault injection, blast radius limits and abort conditions, game days, and running experiments in production safely.
Multi-Tenant SaaS Architecture
Forty-seven thousand tenants sharing infrastructure safely. Silo, pool and bridge isolation models, tenant-aware partitioning, noisy-neighbor quotas, and per-tenant cost attribution.
Disaster Recovery
Two numbers set the architecture: RTO and RPO. Backup-restore through pilot light, warm standby and active-active — plus the restore testing that separates a backup from a capability.
Cost Optimization & FinOps
When the cloud bill rises 40% and nobody knows why. Tagging and showback, unit economics per request and per tenant, commitment coverage, spot strategy, tiering, and egress traps.
Platform Engineering & IDPs
When every team hand-rolls its own manifests and pipelines. Internal developer platforms, golden paths that are easier than the alternative rather than mandated, and DORA metrics as the scorecard.
Data Privacy & Compliance
PII scattered across fifty services you didn't map. Discovery and classification, GDPR/CCPA deletion within deadline across every store and backup, lineage tracking, residency, and consent enforcement. Curriculum finale.
◈ Curriculum complete
73 modules. Beginner to staff+.
73
Modules
3
Tracks
~77h
Total content
Free
Always
One bonus module closes the curriculum: the staff-plus interview toolkit — what changes above senior, how breadth, depth and judgment are scored separately, and how to run a 45-minute expert walkthrough.
M.73 — Staff+ Interview Toolkit →