Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
A developer posted on Hacker News a demonstration of the Qwen3.8‑Flash‑Next language model, which weighs 104 GB, being executed on a Mac equipped with 48 GB of memory. The setup manages to generate roughly twelve tokens per second, showcasing that large‑scale models can run on consumer‑grade hardware with careful optimization, though performance remains modest compared with specialized servers.