Gimlet Labs has raised $300 million in Series B funding as it scales an unusual approach to AI infrastructure: instead of building an inference cloud around one accelerator architecture, the company is assembling heterogeneous data centers in which GPUs, CPUs, near-memory compute and dataflow processors can work together on different portions of the same inference workload. The financing comes only about five months after Gimlet announced an $80 million Series A.
Andreessen Horowitz led the round, joined by Sapphire Ventures, Menlo Ventures, 645 Ventures, Arm, Eclipse, Emergence, Factory, Hudson River Trading, M12, OnePrime Capital, Prosperity7, QuantumLight, Samsung Ventures, Tiger Global Management, Triatomic, Wing Ventures and XTX Markets. Gimlet says that since March it has added billions of dollars in contracted revenue and a gigawatt-scale data center pipeline, while moving toward hundreds of megawatts of managed capacity. The company did not identify the customers associated with those contracts in its Series B announcement.
The larger technical ambition is what Gimlet calls a "multi-silicon cloud" built specifically for inference. Rather than treating an AI accelerator as the unit of infrastructure, Gimlet's software traces and decomposes models into components and schedules those components across different processor architectures according to workload requirements and available capacity. Gimlet claims that breaking a model across different accelerator types can produce 5–10X speedups within the same power footprint, or corresponding throughput gains at a given latency. These are company-reported performance claims. The architecture extends the growing practice of separating prefill and decode, and Gimlet says it is also developing finer-grained forms of disaggregation, including speculative-decode separation and attention/FFN separation.
What remains unknown is whether the approach can hold up at scale. Gimlet's Series B announcement does not disclose the fabric architecture, interconnect speeds, topology or protocols used to tie heterogeneous accelerator pools together. At hundreds of megawatts of managed capacity, that fabric becomes a central architectural question. If multi-silicon inference proves practical at large scale, the AI cloud could evolve from fleets of largely homogeneous accelerators toward more heterogeneous infrastructure, but the company has not yet identified the customers behind its contracted revenue or independently verified its performance claims.