Cerebras and Gimlet Labs announced a collaboration on September 28 to combine Cerebras wafer-scale systems with Gimlet's inference cloud and GPU infrastructure. The companies say the design will route different inference phases to the silicon best suited to them.
A large deployment is planned
The companies say they plan to deliver up to 3,000 tokens per second and expect the first Cerebras-powered Gimlet Cloud datacenter later in 2026. Gimlet separately says the planned capacity is 100 megawatts. These are targets and schedules, not operating capacity or independently measured performance.
What is already running
Cerebras' release says the integrated solution is already serving tokens in private deployments. The announcement does not identify public customers, disclose independent throughput results, or show that the planned datacenter is online.
