Lightbits Launches Inferra to Expand GPU Memory Capacity for AI Inference
Lightbits Labs has released Inferra, a software engine that extends GPU memory by managing key-value cache data across multiple storage tiers, promising significant performance gains for long-context language model workloads.

Storage software specialist Lightbits Labs has made Inferra generally available, unveiling a new engine that aims to enhance both the cost-effectiveness and speed of AI inference operations. The tool relocates key-value cache information from the constrained high-bandwidth memory found on graphics processors to other storage layers. Lightbits, recognized for developing the NVMe over TCP storage protocol, is targeting the product toward neocloud operators and large enterprises deploying language models that handle extended context windows or process numerous parallel requests.
First unveiled in March, Inferra orchestrates KV cache management across three tiers: GPU high-bandwidth memory, dynamic random-access memory, and Non-Volatile Memory Express storage. The engine employs predictive prefetching techniques to position required data adjacent to the GPU ahead of actual demand.
Key-value caches store intermediate attention information that models generate while interpreting input text. These caches expand as conversations lengthen or documents grow larger, placing pressure on limited GPU memory resources. When the system cannot locate needed cache data in fast memory, it must either fetch from slower storage or recalculate the values, both scenarios causing GPU idle time and slower response generation.
The data is always available for the compute to operate on. We do predictive prefetch, which essentially prevents this stall.
Ramesh Chettuvetty, senior vice president of product and business for AI solutions at Lightbits
According to Lightbits' own testing, Inferra can reduce time to first token by over 100 times in certain long-context scenarios, enable context windows exceeding 10 million tokens on standard hardware, and boost concurrent session density by more than 16 times. The company notes these measurements derive from internal benchmarks and projections that lack independent verification.
Chettuvetty explained that Lightbits analyzes multiple signals within the inference pipeline to anticipate which data the GPU will access next. The company reports achieving cache hit rates approaching 99.9% across most test cases. Beyond core functionality, Inferra incorporates quality-of-service controls, data encryption, tenant isolation mechanisms, and the capacity to relocate cached information when a session transfers between GPU clusters.
CPU-inspired
Arthur Rasmusson, director of AI architecture at Lightbits, drew parallels to memory management strategies that emerged when CPU performance began outpacing memory speeds. "We're doing something inspired by that," Rasmusson stated. "The algorithms are obviously different when you're dealing with LLM inference."
The strategy tackles two interconnected challenges: storage volume and data access velocity. Rather than transferring an entire context into GPU memory simultaneously, Inferra segments it into smaller portions and supplies the GPU with only what it requires at each moment. This approach permits operators to maintain larger caches while accommodating additional users on existing hardware.
What we're providing is sort of a larger bookcase in the sense that we have more ability to store long-term, but also fast retrieval.
Arthur Rasmusson, director of AI architecture at Lightbits
Scenarios involving retrieval-augmented generation, autonomous AI agents, extended prompts, and heavily multiplexed GPU environments should derive the greatest advantage. Rasmusson noted that organizations operating oversized dedicated GPU clusters with minimal concurrent users and contexts that already fit within available memory would see diminished returns.
Separating cache from the GPU itself introduces additional latency concerns. Chettuvetty indicated that Inferra mitigates this overhead through anticipatory data staging, positioning information before the GPU explicitly requests it. Improvements in network and storage speed can reduce the advance prediction requirement and lower the likelihood of misprediction.
Lightbits indicated it is conducting production-phase pilots and initially concentrating on neocloud operators, whose pressure to enhance utilization rates and profit margins makes them quicker to embrace new technologies than major cloud providers. The company is showcasing Inferra at the AI Infra Summit in Santa Clara in the coming week.


