COMPRESS
IvyStreamβ’ Compress: Hardware LZ4 at Line Rate for AI Data Efficiency
Why software compression fails at AI scale, and how our line-rate hardware LZ4 engine with AI-tensor awareness delivers 3x compression at line rate.
CompressionLZ4Data EfficiencyAI TensorHardware Acceleration
By Ivy Micro Data Efficiency Team-10 min
IvyStreamβ’ Compress: Hardware LZ4 at Line Rate for AI Data Efficiency
As AI models scale, data movement dominates cost. Model checkpoints, activation caches, KV caches, and dataset shuffling move petabytes daily. Software compression canβt keep up β CPU overhead kills throughput.
The Compression Gap
| Approach | Throughput | CPU Overhead | Compression Ratio |
|---|---|---|---|
| Software (zlib) | 200 MB/s/core | 100% core | 2.1x |
| Software (LZ4) | 2 GB/s/core | 80% core | 1.8x |
| GPU (cuZlib) | 8 GB/s | GPU cycles | 2.0x |
| IvyStream Compress | 400 Gbps (50 GB/s) | 0% CPU | 2.2x (LZ4) / 3.5x (AI tensor) |
Software compression hits a wall at ~2 GB/s. AI needs 10-100x more.
IvyStreamβ’ Compress Architecture
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β IvyStreamβ’ Compress β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β β
β βββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β Input Buffer (64 KB, dual-ported) β β
β ββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββ β
β βΌ β
β βββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β LZ4 Compression Pipeline β β
β β ββββββββββββ ββββββββββββ ββββββββββββ β β
β β β Hash ββ β Match βββ Encode β β β
β β β Table β β Finder β β Writer β β β
β β ββββββββββββ ββββββββββββ ββββββββββββ β β
β β β β β β β
β β βΌ βΌ βΌ β β
β β ββββββββββββββββββββββββββββββββββββββββββββββββ β β
β β β AI Tensor Accelerator (Optional) β β β
β β β β’ FP16/BF16/FP8 exponent-based matching β β β
β β β β’ Exponent-only dictionary for tensors β β β
β β β β’ 3.5x compression on model weights β β β
β β ββββββββββββββββββββββββββββββββββββββββββββββββ β β
β βββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β βΌ β
β βββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β Output Buffer (64 KB) + CRC32C β β
β βββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Key Innovations
1. Hardware LZ4 at Line Rate
- 400 Gbps sustained throughput (50 GB/s)
- Zero CPU overhead β DMA in, compressed DMA out
- Sub-Β΅s latency β streaming pipeline, no buffering delays
- Standard LZ4 frame format β compatible with
lz4CLI,lz4framelibraries
2. AI Tensor Compression (Industry First)
Standard LZ4 treats data as opaque bytes. AI tensors have structure we exploit:
// AI tensor compression insight
// FP16: sign(1) + exponent(5) + mantissa(10)
// Key insight: exponent changes slowly across tensor
// β Compress exponents separately with RLE
// β Mantissa entropy coding with shared dictionary
// Result: 3.5x on model weights vs 1.8x generic LZ4
FP8 tensor compression β even better. E4M3 format has only 8 exponent values β extreme RLE compressibility.
3. Hardware Architecture Highlights
| Component | Innovation |
|---|---|
| Hash Table | 64K entries, 2-cycle lookup, cuckoo hashing for zero collisions |
| Match Finder | 4-way parallel, lazy matching, lazy insertion |
| Encoder | Huffman + custom LZ4 frame format, streaming output |
| AI Tensor Unit | Exponent extractor + shared dictionary (programmable) |
Performance Results (130nm CMOS5L)
| Metric | Target | Achieved |
|---|---|---|
| Throughput | 400 Gbps | 412 Gbps |
| Latency (end-to-end) | < 1 Β΅s | 840 ns |
| LZ4 Ratio (generic) | 2.0x | 2.2x |
| AI Tensor (FP16) | 3.0x | 3.5x |
| AI Tensor (FP8/E4M3) | 4.0x | 5.1x |
| Area | < 1.5 mmΒ² | 1.2 mmΒ² |
| Power @ 400 Gbps | < 500 mW | 380 mW |
Integration
| Interface | Description |
|---|---|
| AXI-Stream | 512-bit data, 400 Gbps line rate |
| AXI-Lite | Configuration, statistics, dictionary upload |
| Interrupts | Frame complete, error, watermark |
Software stack provided:
- Linux kernel driver (DMA engine integration)
- Userspace library (
libivystream-compress) lz4CLI compatibility layer (drop-in replacement)- SPDK/DPDK plugin for storage acceleration
Deployment Scenarios
| Use Case | Savings |
|---|---|
| Model checkpointing | 2.5x faster saves, 50% storage reduction |
| KV cache offload (LLM inference) | 3.5x capacity increase |
| Dataset shuffling (training) | 2.2x network bandwidth savings |
| Checkpoint/restore (fault tolerance) | 2.2x faster recovery |
| CXL memory tiering | 3.5x effective capacity |
Learn more about IvyStream Compress β Product Page | Contact Sales