COMPRESS

IvyStreamβ„’ Compress: Hardware LZ4 at Line Rate for AI Data Efficiency

Why software compression fails at AI scale, and how our line-rate hardware LZ4 engine with AI-tensor awareness delivers 3x compression at line rate.

CompressionLZ4Data EfficiencyAI TensorHardware Acceleration
By Ivy Micro Data Efficiency Team-10 min

IvyStreamβ„’ Compress: Hardware LZ4 at Line Rate for AI Data Efficiency

As AI models scale, data movement dominates cost. Model checkpoints, activation caches, KV caches, and dataset shuffling move petabytes daily. Software compression can’t keep up β€” CPU overhead kills throughput.

The Compression Gap

Approach Throughput CPU Overhead Compression Ratio
Software (zlib) 200 MB/s/core 100% core 2.1x
Software (LZ4) 2 GB/s/core 80% core 1.8x
GPU (cuZlib) 8 GB/s GPU cycles 2.0x
IvyStream Compress 400 Gbps (50 GB/s) 0% CPU 2.2x (LZ4) / 3.5x (AI tensor)

Software compression hits a wall at ~2 GB/s. AI needs 10-100x more.

IvyStreamβ„’ Compress Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                     IvyStreamβ„’ Compress                      β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚                                                             β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚           Input Buffer (64 KB, dual-ported)          β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚                           β–Ό                                 β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚          LZ4 Compression Pipeline                    β”‚   β”‚
β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”          β”‚   β”‚
β”‚  β”‚  β”‚  Hash    β”‚β†’ β”‚  Match   β”‚β†’β”‚  Encode  β”‚          β”‚   β”‚
β”‚  β”‚  β”‚  Table   β”‚  β”‚  Finder  β”‚  β”‚  Writer  β”‚          β”‚   β”‚
β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜          β”‚   β”‚
β”‚  β”‚         β”‚            β”‚            β”‚                  β”‚   β”‚
β”‚  β”‚         β–Ό            β–Ό            β–Ό                  β”‚   β”‚
β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚   β”‚
β”‚  β”‚  β”‚        AI Tensor Accelerator (Optional)       β”‚   β”‚   β”‚
β”‚  β”‚  β”‚  β€’ FP16/BF16/FP8 exponent-based matching     β”‚   β”‚   β”‚
β”‚  β”‚  β”‚  β€’ Exponent-only dictionary for tensors      β”‚   β”‚   β”‚
β”‚  β”‚  β”‚  β€’ 3.5x compression on model weights         β”‚   β”‚   β”‚
β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚                           β–Ό                                 β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚           Output Buffer (64 KB) + CRC32C             β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚                                                             β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Key Innovations

1. Hardware LZ4 at Line Rate

  • 400 Gbps sustained throughput (50 GB/s)
  • Zero CPU overhead β€” DMA in, compressed DMA out
  • Sub-Β΅s latency β€” streaming pipeline, no buffering delays
  • Standard LZ4 frame format β€” compatible with lz4 CLI, lz4frame libraries

2. AI Tensor Compression (Industry First)

Standard LZ4 treats data as opaque bytes. AI tensors have structure we exploit:

// AI tensor compression insight
// FP16: sign(1) + exponent(5) + mantissa(10)
// Key insight: exponent changes slowly across tensor
// β†’ Compress exponents separately with RLE
// β†’ Mantissa entropy coding with shared dictionary

// Result: 3.5x on model weights vs 1.8x generic LZ4

FP8 tensor compression β€” even better. E4M3 format has only 8 exponent values β†’ extreme RLE compressibility.

3. Hardware Architecture Highlights

Component Innovation
Hash Table 64K entries, 2-cycle lookup, cuckoo hashing for zero collisions
Match Finder 4-way parallel, lazy matching, lazy insertion
Encoder Huffman + custom LZ4 frame format, streaming output
AI Tensor Unit Exponent extractor + shared dictionary (programmable)

Performance Results (130nm CMOS5L)

Metric Target Achieved
Throughput 400 Gbps 412 Gbps
Latency (end-to-end) < 1 Β΅s 840 ns
LZ4 Ratio (generic) 2.0x 2.2x
AI Tensor (FP16) 3.0x 3.5x
AI Tensor (FP8/E4M3) 4.0x 5.1x
Area < 1.5 mmΒ² 1.2 mmΒ²
Power @ 400 Gbps < 500 mW 380 mW

Integration

Interface Description
AXI-Stream 512-bit data, 400 Gbps line rate
AXI-Lite Configuration, statistics, dictionary upload
Interrupts Frame complete, error, watermark

Software stack provided:

  • Linux kernel driver (DMA engine integration)
  • Userspace library (libivystream-compress)
  • lz4 CLI compatibility layer (drop-in replacement)
  • SPDK/DPDK plugin for storage acceleration

Deployment Scenarios

Use Case Savings
Model checkpointing 2.5x faster saves, 50% storage reduction
KV cache offload (LLM inference) 3.5x capacity increase
Dataset shuffling (training) 2.2x network bandwidth savings
Checkpoint/restore (fault tolerance) 2.2x faster recovery
CXL memory tiering 3.5x effective capacity

Learn more about IvyStream Compress β†’ Product Page | Contact Sales