KV cache

Tensormesh & AMD Cut GPU Costs with KV Cache Tiering

Tensormesh & AMD Cut GPU Costs with KV Cache Tiering

Tensormesh and AMD have announced a collaboration that uses tiered KV cache management to double AI model density on fixed GPU hardware. Benchmarks show a nearly 7x improvement in time-to-first-token on a 300 GB document set. The approach targets the core economics of enterprise AI inference, not just raw performance.

Tensormesh & AMD Cut GPU Costs with KV Cache Tiering Read More »

MinIO MemKV: Purpose-Built AI Inference Cache Storage

MinIO MemKV: Purpose-Built AI Inference Cache Storage

MinIO has launched MemKV, a purpose-built KV cache storage product targeting NVIDIA’s G3.5 memory tier via NVMe/RDMA. The product promises a 75x improvement in inference time-to-first-token and up to $2M in annual GPU efficiency savings for a typical enterprise deployment. ECI Research examines the business case, technical architecture, and what buyers need to validate before committing.

MinIO MemKV: Purpose-Built AI Inference Cache Storage Read More »