Funding

Positron AI Secures $875M to Bring Consumer-Grade Memory Inference Systems to Market

The Nevada-based chipmaker has closed a Series C funding round led by NEA and other major investors, valuing the company at $5 billion as it prepares to launch inference hardware that replaces expensive HBM memory with more affordable alternatives.

·3 min read
Chipmaker Positron nabs $875M to speed up inference with consumer-grade memory
Chipmaker Positron nabs $875M to speed up inference with consumer-grade memory

Positron AI Inc., which develops inference appliances using consumer-grade memory instead of high-bandwidth alternatives, announced a $875 million Series C funding round today. Leading the investment were NEA, Atreides Management, Valor Equity Partners, Andra Capital, SemiAnalysis Capital, and Jim Clark, co-founder of Silicon Graphics Inc. and Netscape Inc., alongside more than a dozen additional institutional backers. The round values Positron at $5 billion, representing a five-fold jump from its February valuation.

Large language model inference requires constant data movement between computing and memory components on graphics processors, with memory bandwidth—the speed of this data transfer—directly affecting inference performance. Server-grade graphics cards typically rely on HBM memory, which offers the highest bandwidth available, but current supply falls far short of industry demand. Adding to supply constraints, the specialized interconnects needed to integrate HBM with graphics card processors are also in short supply.

Positron, headquartered in Reno, Nevada, has developed a workaround to these supply chain bottlenecks by substituting HBM with LPDDR5X, a memory technology primarily found in smartphones. This approach offers two immediate advantages: LPDDR5X is substantially more available and considerably cheaper than HBM. The trade-off, however, is significantly reduced memory bandwidth, which typically results in slower inference performance.

The company contends it has solved this performance gap through software and hardware optimization. Positron notes that most AI accelerators utilize less than 30 percent of their HBM modules' available bandwidth, whereas its inference systems extract more than 90 percent of LPDDR5X's throughput. This higher utilization rate, according to the company, compensates for the theoretical performance difference between the two memory types.

Titan and Asimov: The Hardware Stack

Positron's primary offering is Titan, an inference appliance equipped with up to 18.4 terabytes of LPDDR5X memory capable of delivering 23.68 terabits per second of bandwidth. According to the company, a single Titan system can run language models with 32 trillion parameters and support a context window of 10 billion tokens.

Inference processing on Titan relies on Asimov, a custom processor built around a systolic array architecture. This design consists of identical computing modules, each with integrated memory for storing LLM weights—the numerical parameters essential to model output generation. The systolic array is complemented by specialized modules for executing activation functions, which control which neural networks participate in each inference operation. Asimov also incorporates CPU cores that serve as a programmable escape hatch, handling tasks outside the chip's optimized capabilities.

Each Titan appliance houses up to eight Asimov chips, and multiple systems can be clustered together to support up to 16,384 accelerators total.

Timeline and Production Plans

Neither Titan nor Asimov have entered production yet. Positron's simulation data suggests a server rack using its silicon could process up to 26 times more tokens per dollar compared to Nvidia Corp.'s Blackwell GB300 NVL72 system.

Chief Executive Officer Mitesh Agrawal stated: "Our focus now is to tape out Asimov, bring Titan to production, and scale manufacturing to meet the demand in front of us. This financing gives us the resources to do exactly that."

Positron plans to tape out Asimov by year-end using Taiwan Semiconductor Manufacturing Co.'s three-nanometer process technology. Volume production is targeted for the second half of 2027, with concurrent scaling of Titan manufacturing. The company will prioritize securing long-term LPDDR5X supply agreements with partners to support this expansion.