Gimlet Labs Raises $80M for Multi-Chip AI Inference

Gimlet Labs has raised $80M in Series A funding for its multi-silicon inference platform that splits AI workloads across chips from NVIDIA, AMD, and others.

TL;DR
  • Funding: Gimlet Labs raised $80 million in Series A funding led by Menlo Ventures for its multi-silicon AI inference platform.
  • Technology: The startup’s software splits AI workloads across chips from NVIDIA, AMD, Intel, and others, claiming 3x to 10x inference speedups.
  • Traction: Gimlet launched with 8-figure revenues in October 2025 and has since tripled its customer base across frontier AI labs.
  • Market Context: Data center spending is projected to reach nearly $7 trillion by 2030, with current hardware utilization rates at only 15 to 30 percent.

Gimlet Labs on March 23 announced it has raised $80 million in Series A funding for software that splits AI workloads across chips from NVIDIA, AMD, Intel, and others simultaneously. The startup claims 3x to 10x inference speedups at the same cost and power. Menlo Ventures led the round, with Factory, Eclipse Ventures, Prosperity7, and Triatomic also participating.

Led by Stanford adjunct professor and serial entrepreneur Zain Asgar, the startup calls its platform a multi-silicon inference cloud. Rather than locking customers into a single chip vendor, the software automatically maps agentic workloads to well-suited chips without requiring developers to rewrite their code.

According to Asgar, AI applications currently use deployed hardware only 15 to 30 percent of the time, leaving hundreds of billions of dollars in compute sitting idle. As a result, Gimlet was built to close that utilization gap, routing each phase of an AI workload to whichever available processor handles it with the greatest efficiency, regardless of vendor.

“Another way to think about this: you’re wasting hundreds of billions of dollars because you’re just leaving idle resources. Our goal was basically to try to figure out how you can get AI workloads to be 10x more efficient than ever, today.”

Zain Asgar, CEO and Co-founder of Gimlet Labs (via TechCrunch)

How Multi-Silicon Inference Works

Different stages of AI inference place different demands on hardware. “Prefill is compute-bound; decode is memory-bound; and tool calls are network-bound,” Menlo Ventures partner Tim Tully wrote in the firm’s investment blog. A single GPU architecture cannot optimize for all three simultaneously, which is why current data centers leave so much capacity underused.

Building on this analysis, Tully’s thesis is that the hardware diversity already exists across the industry, but what has been missing is a software abstraction layer capable of coordinating it. Menlo described Gimlet’s vision as building the first multi-silicon inference and compute cloud, one that deploys traditional GPUs alongside SRAM-centric silicon purpose-built for memory-intensive workloads, extracting value from hardware that operators have already paid for.

In practice, Gimlet’s software addresses this mismatch by splitting workloads across diverse hardware, slicing the underlying model so that it runs across different chip architectures, routing each portion to an optimal processor. Compute-heavy prefill operations go to GPUs, memory-intensive decode tasks run on high-bandwidth chips, and network-bound tool calls get routed to low-latency accelerators.

gimlet_labs-orchestration flow
Gimlet works by breaking workloads across multi-generation and multi-vendor systems, allowing inference to inherit the performance and cost characteristics of the best possible hardware for each task. (Source: MENLO Ventures)

Performance gains extend specifically to very large frontier models, which place uneven demands across different hardware types. By matching each stage to the right silicon, Gimlet claims to deliver an order of magnitude improvement in performance per watt for customers operating at scale.

Hardware partners announced by Gimlet include NVIDIA, AMD, Intel, ARM, Cerebras, and d-Matrix, ensuring the orchestration layer works across each vendor’s silicon. Meanwhile, d-Matrix, which specializes in in-memory inference compute for data centers, announced jointly with Gimlet that their combined stack delivers 10x speed improvements with notable power efficiency gains for frontier AI workloads. Tully framed the overall opportunity in blunt terms: “the multi-silicon fleet is ready, it’s just missing the software layer to make it work.”

Gimlet offers software and an API through its managed Gimlet Cloud, with an option for customers to deploy on their own data center infrastructure. Its target market is large AI model labs and data center operators rather than general application developers.

More broadly, Asgar described the approach as addressing a fundamental bottleneck: the speed of intelligence has become the constraint, and heterogeneous hardware is what unlocks the next leap in performance for use cases like coding agents. Agentic workloads compound the challenge because they chain together multiple inference calls, tool invocations, and memory retrievals in a single session, creating demand patterns that no single chip type can serve efficiently.

Traction and Funding Details

Gimlet Labs launched with 8-figure revenues in October 2025, an unusual milestone for a company emerging from stealth. In the five months since, according to the company’s press release, it has tripled its customers across frontier labs. Gimlet now counts one of the top-three frontier labs and one of the top-three global hyperscalers among its clients, two of the highest-volume AI workload environments in the industry.

However, landing both a top frontier lab and a top hyperscaler within five months of launch represents enterprise traction at the highest tier of the AI industry. Such accounts validate the platform’s ability to operate at a scale where even small efficiency gains translate into hundreds of millions in annual infrastructure savings.

Asgar said they got “a pretty big swarm of funding” after the October launch. Menlo Ventures led the Series A, with Factory, Eclipse Ventures, Prosperity7, and Triatomic also participating, according to Gimlet’s Series A funding announcement.

According to the company’s October 2025 announcement, Gimlet had previously raised a $12 million seed round led by Factory, bringing total funding to $92 million. Gimlet currently employs 30 people, a lean team by the standards of its ambitions and customer profile.

Moreover, angel investors signal deep industry connections: backers include Sequoia’s Bill Coughran, Stanford Professor Nick McKeown, former VMware CEO Raghu Raghuram, and Intel CEO Lip-Bu Tan. Having an active Intel CEO among its backers is notable for a company whose pitch depends on routing workloads across competing chip architectures.

In addition, Prosperity7, the venture arm of Saudi Aramco, adds a sovereign wealth dimension that reflects growing Middle Eastern interest in AI infrastructure investments. Both a chip company executive and a sovereign wealth fund backing the same startup underscores how broadly the multi-silicon thesis resonates across different categories of strategic investor.

The Scale of the Inference Problem

Investor conviction is grounded in concrete market scale. According to McKinsey, data center spending will reach nearly $7 trillion by 2030, TechCrunch reported, if the current trend of deploying more hardware continues. According to Gimlet’s press release, in 2026 alone the industry is gearing up to spend $650 billion in AI datacenter CapEx, while inference volume has already reached quadrillions of tokens per month and is still growing.

At that scale of capital deployment, software-layer efficiency gains become disproportionately valuable compared to incremental hardware upgrades. Against that backdrop, the low hardware utilization rates Asgar identified represent a structural inefficiency worth hundreds of billions of dollars, one that more single-vendor hardware procurement alone cannot fix. Gimlet’s bet is that a software layer unifying existing heterogeneous fleets can extract far more value from infrastructure already in place.

Sachin Katti, who previously served as Intel Chief Technology and AI Officer, puts it directly:

“The industry is hitting the limits of homogeneous, vertically integrated one-size-fits-all AI infrastructure. Agentic AI applications are inherently heterogeneous, reasoning across multiple models, modalities and data sources.”

Sachin Katti, former Intel Chief Technology and AI Officer (via Gimlet agentic workloads report)

Whether Gimlet can deliver on its 3x to 10x performance claims at scale remains to be proven across a broader customer base. Nevertheless, the founding team brings relevant credentials.

According to the company, Asgar was previously a GPU architect at NVIDIA and an engineering lead at Google AI, giving him firsthand experience designing and deploying the silicon architectures his company now orchestrates across vendor boundaries. Before Gimlet, the cofounders previously worked together at Pixie, an open-source Kubernetes observability tool acquired by New Relic in December 2020, just two months after a $9 million Series A led by Benchmark.

Rapid acquisition demonstrated the team’s ability to build infrastructure software that large enterprises adopt quickly and at production scale. Alongside Asgar, cofounders Michelle Nguyen, Omid Azizi, and Natalie Serrino bring engineering and systems depth that spans the full stack of modern AI infrastructure.

Markus Kasanmascheff
Markus Kasanmascheff
Markus has been covering the tech industry for more than 15 years. He is holding a Master´s degree in International Economics and is the founder and managing editor of Winbuzzer.com.
Subscribe
Notify of
guest
0 Comments
Newest
Oldest Most Voted