Home / Computer Networks
Computer Networks

Clos Network Topologies and VXLAN-EVPN Fabrics: High-Performance Data Center Architecture

Master Clos network topologies and VXLAN-EVPN fabrics: Spine-and-Leaf ECMP, BGP-EVPN control planes, VTEP encapsulation, and distributed Anycast routing.

Aug 25, 2026 • 8 min read

The explosive growth of containerized microservices, distributed databases, and artificial intelligence training clusters has revolutionized data center network engineering. Consequently, implementing clos network topologies and vxlan evpn fabrics has become the industry standard for delivering predictable, non-blocking East-West throughput and programmatic multi-tenancy at hyper-scale.

Historically, enterprise data centers relied on hierarchical three-tier architectures governed by the legacy Spanning Tree Protocol (STP). However, blocking redundant links to prevent broadcast loops severely constrained available network bandwidth and created crippling latency bottlenecks.

In this comprehensive engineering masterclass, we dissect the internal mechanics of modern data center fabrics. We explore folded Clos (Spine-and-Leaf) topologies, Equal-Cost Multi-Pathing (ECMP), VXLAN encapsulation, BGP-EVPN control-plane route types, and distributed Anycast gateway routing.

The Structural Collapse of Traditional Spanning Tree Networks

The classical three-tier network design (Core, Aggregation, and Access) was optimized for North-South client-to-server traffic patterns. In that legacy paradigm, client requests entered the data center, hit a monolithic database, and exited through the perimeter firewall.

However, modern distributed architectures generate massive East-West traffic flows between servers, microservices, and storage nodes. Under a Spanning Tree Protocol (STP) topology, up to 50% of physical switch links are artificially forced into a blocking state to avoid Layer 2 broadcast storms.

Furthermore, when physical links fail, STP recalculation takes seconds to reconverge, causing severe application packet loss and session termination. Therefore, modern network architecture completely eliminates Layer 2 switching at the physical core, replacing it with an all-routed Layer 3 fabric.

Fundamentals of Clos Network Topologies and VXLAN-EVPN Fabrics

Originally formulated by Charles Clos for telephone switching networks, the modern two-tier Spine-and-Leaf Clos topology connects every leaf switch directly to every spine switch. This design ensures that every server is separated from every other server by an identical, deterministic number of network hops.

The architectural diagram below illustrates how an underlying Layer 3 routed fabric supports virtualized Layer 2 and Layer 3 overlays using VXLAN and BGP-EVPN:

+-----------------------------------------------------------------------------------+
|                        SPINE LAYER (LAYER 3 ROUTED CORE)                          |
|                                                                                   |
|         [ Spine Switch 1 ]                  [ Spine Switch 2 ]                    |
|         (BGP Route Reflector)               (BGP Route Reflector)                 |
|               ||      \\                         //      ||                       |
+---------------+--------\\-----------------------//-------+------------------------+
                ||        \\                     //        ||
         Equal-Cost Multi-Pathing (ECMP) Underlay Routing (eBGP)
                ||          \\                 //          ||
+---------------+------------\\---------------//-----------+------------------------+
|               ||            \\             //            ||                       |
|         [ Leaf Switch 1 ]    \\           //       [ Leaf Switch 2 ]              |
|         (VTEP / Gateway)      \\         //        (VTEP / Gateway)               |
|  +-----------------------------------------------------------------------------+  |
|  | OVERLAY: VXLAN Encapsulated Tunnels (UDP 4789) Over Underlay IPv4           |  |
|  | CONTROL PLANE: BGP-EVPN (Type-2 MAC/IP & Type-5 Prefix Routes)              |  |
|  +-----------------------------------------------------------------------------+  |
|               ||                                         ||                       |
+---------------+------------------------------------------+------------------------+
                || ESI-LAG Active/Active Multihoming       ||
                \/                                         \/
+------------------------------------+         +------------------------------------+
| TENANT COMPUTE NODE A              |         | TENANT COMPUTE NODE B              |
| (VM / Container / Bare-Metal)      |         | (VM / Container / Bare-Metal)      |
+------------------------------------+         +------------------------------------+

1. The Non-Blocking Layer 3 Underlay and ECMP

The physical underlay network serves a single, dedicated purpose: transporting IP packets between leaf switches with maximum throughput and minimum latency. The underlay runs standard Interior Gateway Protocols (IGP) or, more commonly, external BGP (eBGP).

Because every leaf connects to all spines, the network leverages Equal-Cost Multi-Pathing (ECMP) across all available links. When a leaf switch transmits data, hashing algorithms distribute flows across multiple spine switches simultaneously.

Consequently, 100% of physical fiber bandwidth remains active without risk of broadcast loops. If a spine switch suffers a hardware failure, remaining spines absorb traffic in microseconds without routing disruptions.

2. The VXLAN Data Plane: Layer 2 over Layer 3 Encapsulation

While the underlay network operates purely on Layer 3 IP routing, modern virtualization platforms require Layer 2 adjacency for virtual machine mobility and clustering. Virtual Extensible LAN (VXLAN - RFC 7348) bridges this gap through packet encapsulation.

VXLAN encapsulates entire Ethernet frames inside standard Layer 4 UDP packets using port 4789. Each virtual segment is tagged with a 24-bit Virtual Network Identifier (VNI), expanding the legacy 4,096 VLAN limitation to over 16 million isolated tenant broadcast domains.

Leaf switches function as VXLAN Tunnel Endpoints (VTEPs), encapsulating local host packets into UDP before transmission across the underlay and decapsulating them at the destination leaf switch.

BGP-EVPN: The Universal Control Plane

Early VXLAN implementations relied on flood-and-learn multicast mechanisms across the underlay to discover remote MAC addresses. This approach generated substantial broadcast noise and wasted network bandwidth.

Ethernet Virtual Private Network with BGP (BGP-EVPN - RFC 7432) replaced flood-and-learn with a scalable, standardized control plane. Leaf switches use MP-BGP to advertise MAC and IP addresses directly to route reflectors.

The BGP-EVPN control plane standardizes five critical route types:

  • Type-1 (Ethernet Auto-Discovery): Advertises multihoming and fast convergence parameters across redundant leaf switches.
  • Type-2 (MAC/IP Advertisement): Publishes host MAC and IP bindings to remote VTEPs, eliminating data-plane ARP broadcast flooding.
  • Type-3 (Inclusive Multicast Ethernet Tag): Sets up ingress replication tunnels for necessary Broadcast, Unknown Unicast, and Multicast (BUM) traffic.
  • Type-4 (Ethernet Segment Route): Enables leaf switches to discover peer multihomed switches sharing the same physical compute node.
  • Type-5 (IP Prefix Route): Advertises routed IP subnets across different Virtual Routing and Forwarding (VRF) tenant instances.

Implementation Guide: BGP-EVPN Configuration on Open Networking Switches

The configuration snippet below demonstrates a production-grade FRRouting (FRR) setup for an open networking leaf switch running SONiC or Linux NOS. It defines the BGP-EVPN address family, VNI mappings, and distributed Anycast routing:

! /etc/frr/frr.conf - Open Network Switch Leaf Configuration
router bgp 65011
 bgp router-id 10.0.0.11
 no bgp default ipv4-unicast
 
 ! Underlay eBGP Peers (Spine Switches)
 neighbor SPINE-FABRIC peer-group
 neighbor SPINE-FABRIC remote-as 65000
 neighbor 10.255.0.1 peer-group SPINE-FABRIC
 neighbor 10.255.0.2 peer-group SPINE-FABRIC

 ! Underlay Address Family
 address-family ipv4 unicast
  network 10.0.0.11/32
  neighbor SPINE-FABRIC activate
 exit-address-family

 ! BGP-EVPN Overlay Address Family
 address-family l2vpn evpn
  neighbor SPINE-FABRIC activate
  neighbor SPINE-FABRIC route-reflector-client
  advertise-all-vni
  advertise-svi-ip
  advertise-default-gw
 exit-address-family
!
! VXLAN Interface Definition
interface vxlan100
 vxlan id 10100
 vxlan local-tunnelip 10.0.0.11
 bridge-access 100
!
! Symmetric IRB Anycast Gateway Interface (VRF Tenant-Production)
interface vlan100
 vrf forwarding Tenant-Prod
 ip address 192.168.10.1/24
 ip address-virtual 00:00:5E:00:01:01 192.168.10.1

The directive ip address-virtual configures a Distributed Anycast Gateway. Every leaf switch in the fabric hosts the exact same default gateway IP and virtual MAC address. Consequently, virtual machines can migrate between physical racks without changing network configurations or experiencing routing hairpinning.

Comparative Matrix: Legacy Spanning Tree vs. Clos Network Topologies and VXLAN-EVPN Fabrics

To summarize the evolution of data center networking, the comparative table below details the performance differences across both models:

Engineering Dimension Legacy 3-Tier STP Architecture Modern Clos Topology with VXLAN-EVPN
Link Utilization 50% active (STP blocks redundant paths). 100% active (ECMP active/active forwarding).
Control Plane Protocol Spanning Tree Protocol (STP / RSTP). Multiprotocol BGP-EVPN (RFC 7432).
Layer 2 Domain Scale Limited to 4,096 VLANs per fabric. Over 16 million VNIs per fabric.
Default Gateway Model Centralized HSRP/VRRP on core switches. Distributed Anycast Gateway on every leaf.
Multihoming Technology Proprietary MLAG / vPC pairing. Standardized ESI (Ethernet Segment Identifier).

High Availability with EVPN Multi-Homing (ESI-LAG)

Traditional data center switches relied on proprietary Multi-Chassis Link Aggregation (MLAG/vPC) protocols to connect dual-homed servers. However, MLAG requires dedicated inter-switch links (peer links) that introduce split-brain failure risks and vendor lock-in.

In contrast, EVPN Multi-Homing utilizes standardized Ethernet Segment Identifiers (ESI). Redundant leaf switches share the same ESI value without requiring physical inter-switch synchronization cables.

When a server transmits data across its LACP bond, both leaf switches forward packets simultaneously using active/active forwarding. Furthermore, if a leaf fails, remote VTEPs switch paths instantly using BGP Type-1 route withdrawal, delivering sub-millisecond failover.

Production Blueprint for Network Architects

To successfully transition an enterprise data center to a high-performance VXLAN-EVPN fabric, engineering teams should follow a four-stage execution roadmap:

  1. Underlay Fabric Provisioning: Configure unnumbered eBGP peering across point-to-point links and enforce Jumbo Frames (MTU 9216 bytes) to accommodate VXLAN encapsulation overhead.
  2. BGP-EVPN Control Plane Activation: Establish route reflectors on spine switches and configure unique Autonomous System numbers per leaf tier (eBGP Underlay / iBGP Overlay).
  3. Symmetric IRB and Tenant VRF Isolation: Deploy Distributed Anycast Gateways and configure symmetric routing VNIs for multi-tenant traffic isolation.
  4. Continuous Telemetry and Streaming Observability: Deploy gNMI and streaming telemetry agents to monitor ECMP flow distribution and BGP session states in real time.

Conclusion: Consolidating Clos Network Topologies and VXLAN-EVPN Fabrics

In conclusion, building clos network topologies and vxlan evpn fabrics represents the state of the art in high-density, automated data center infrastructure.

By replacing brittle Layer 2 broadcast domains with an all-routed ECMP underlay and a programmable BGP-EVPN control plane, network architects deliver massive scale, non-blocking East-West throughput, and bulletproof tenant isolation. Mastering these distributed networking standards is the essential foundation for powering the next generation of cloud-native and AI-driven data centers.

Tags Connectivity · Computer Networks
Enjoyed this read? Share it with someone who also wants to apply technology without the noise.