A live leaf-spine fabric simulation — as a 2D schematic or a full 3D ops console, driven by the same state machine. Select two hosts, watch traffic hash across every equal-cost spine path, then kill a link and see the fabric reconverge onto the surviving spines.
Follow the four steps below in either view — the 2D schematic for clarity, or the 3D console for the full picture (drag to orbit, scroll to zoom, particles show live traffic). Cinematic mode runs the whole demo hands-free.
Training clusters don't generate polite client-server flows. They generate synchronized, full-bandwidth, east-west bursts — and the job moves at the speed of its slowest path.
Collectives like AllReduce and AllGather fire across thousands of GPUs simultaneously. One congested or failed path stalls the entire training step — the network is inside the compute loop, not next to it.
Equal-cost multi-path hashing distributes flows across every spine. Predictable, uniform bandwidth between any two endpoints — no tuning per flow, no manual traffic engineering.
When a link dies, only the local leaf reroutes onto surviving spines. No fabric-wide reconvergence, no topology recalculation storm. The workload keeps moving.
| Topology | 3-stage Clos — 4 spines × 6 leaves × 12 hosts. Every leaf connects to every spine (24 fabric links); hosts dual-homed logically to their leaf. |
|---|---|
| Oversubscription | 2 host ports : 4 fabric uplinks per leaf — non-blocking with headroom. Real AI fabrics target 1:1 for the backend/compute network. |
| Routing model | eBGP-style behavior: each switch is its own ASN, routes advertised up/down the fabric. Path selection reduces to a spine filter — ecmpPaths() below is the whole "protocol." |
| Load sharing | Per-flow 5-tuple ECMP hashing. A single flow stays pinned to one path (no reordering); the aggregate distributes across all equal-cost spines. |
| Failure model | Single link failure → local withdraw → FIB reprogram on the affected leaf → traffic redistributes across surviving spines. Sequenced at realistic control-plane ordering. |
| Rendering | One state machine, two renderers: a 2D SVG schematic and a 3D WebGL console (Three.js, lazy-loaded on demand). All pulse animation is driven from a single clock, so the fabric breathes in sync by construction. |
| Simulated vs. real | Simulated: BGP timers, the packet-loss window, buffer behavior, ECN/PFC. Real: the topology math, path computation, hash distribution logic, and reconvergence sequencing. Model boundaries are stated, not implied. |
// Clos symmetry makes routing trivial — that's the design win. // No Dijkstra, no graph library: reconvergence is a spine filter. function ecmpPaths(src, dst) { const sl = leafOf(src), dl = leafOf(dst); if (sl === dl) return [[src, sl, dst]]; // intra-leaf return SPINES .filter(s => up(`${sl}-${s}`) && up(`${dl}-${s}`)) .map(s => [src, sl, s, dl, dst]); }
One HTML file, one state machine, two renderers. The 2D schematic runs on zero dependencies; the 3D console lazy-loads Three.js only when you ask for it. The concepts carry the weight — the code stays out of the way.
Fifteen years of enterprise WAN/DC networking (CCNP — BGP, MPLS, Cisco, F5, Palo Alto), now deliberately deepened into the fabric layer that AI infrastructure runs on.
Anyone can define ECMP. This shows how traffic actually distributes, what a link failure looks like from the control plane's perspective, and why reconvergence stays local in a Clos design.
The simulation is explicit about what's real (topology math, path logic, sequencing) and what's simplified (timers, loss windows, buffer behavior). Knowing the boundary between model and wire is the operational skill.
The 2D core is one file with no framework and no dependencies; the 3D console loads its engine only on demand. The topology's symmetry makes the routing trivial — recognizing that, instead of importing a graph library, is the point.