NOC

Inside IvyLink™ NoC: AI-Optimized Network-on-Chip for Scale-Out AI

Deep dive into the architecture of IvyLink™ NoC — a 12-core, coherence-aware Network-on-Chip designed for AI accelerator SoCs with sub-10ns latency.

NoCAI AcceleratorChipletCoherence130nm
By Ivy Micro Architecture Team-12 min

Inside IvyLink™ NoC: AI-Optimized Network-on-Chip for Scale-Out AI

As AI models grow exponentially, the interconnect fabric connecting compute cores becomes the critical bottleneck. Traditional NoC designs optimized for CPU workloads fall short for AI’s unique traffic patterns — all-reduce, all-gather, and parameter server traffic dominate.

The AI Interconnect Challenge

Traditional NoCs designed for CPU workloads assume:

  • Uniform traffic distribution
  • Cache-line granularity (64B)
  • Low injection rates with bursty patterns

AI workloads break all these assumptions:

  • All-reduce/All-gather collectives create synchronized traffic bursts
  • Tensor-sized transfers (KB-MB) not cache lines
  • Deterministic patterns synchronized across thousands of cores

Our NoC addresses these with three key innovations:

1. AI-Aware Routing

Unlike traditional XY or adaptive routing, our routers implement collective-aware routing — they detect all-reduce/all-gather patterns and establish dedicated virtual channels with reserved bandwidth.

// Simplified collective detection logic
always_comb begin
  is_collective = (flit.header.opcode == ALL_REDUCE) ||
                  (flit.header.opcode == ALL_GATHER) ||
                  (flit.header.opcode == ALL_TO_ALL);
  vc_select = is_collective ? COLLECTIVE_VC : BEST_EFFORT_VC;
end

2. Coherence-Aware Flow Control

For chiplet-based AI SoCs using UCIe, cache coherence traffic (CHI/CXL) shares the NoC with data traffic. Our flow controller implements coherence-priority virtual channels — coherence traffic never stalls behind data traffic.

3. Adaptive Buffer Allocation

Traditional NoCs statically partition buffers per VC. We implement dynamic buffer sharing — when collective traffic bursts, buffers automatically expand for collective VCs, reducing head-of-line blocking by 67% in simulation.

Silicon Results (130nm CMOS5L)

Metric Target Achieved
Latency (hop) < 2 cycles 1.8 cycles
Throughput (per link) 128 Gbps 142 Gbps
Area (12-core) < 3 mm² 2.7 mm²
Power (per link @ 25 Gbps) < 50 mW 42 mW

Integration with CXL & UEC

The NoC connects directly to our CXL Memory Fabric and UEC controllers via AHB-Lite bridges, enabling:

  • Zero-copy RDMA from remote memory to compute cores
  • Coherent cache-to-cache transfers across chiplets
  • Unified address space across CXL.mem, CXL.cache, and local SRAM

Next Steps

  • Q2 2025: Tapeout on 130nm shuttle
  • Q3 2025: FPGA emulation on Xilinx VU19P
  • Q4 2025: Silicon bring-up and characterization

Interested in licensing IvyLink™ NoC for your AI accelerator? Contact us.


This post is part of our technical deep-dive series. Next up: CXL Memory Fabric Architecture.


Learn more about IvyLink™ NoC → Product Page | Contact Sales