Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs
Kimi K3, a 2.8‑terabyte language model, produced one token per second on a MacBook Pro. The test streamed data from four SSDs, showing that large‑scale inference can run on consumer hardware with modest speed. The benchmark demonstrates a practical approach to deploying massive models on limited devices, balancing size, speed, and storage in real‑world scenarios.