Same 27B, same Q4.5, 16x apart. My M5 Max: 18 tok/s. Same model on a 2080 Super via…
This is a AI post classified by Jev as AI infra & evals (a model release), kept by the AI Radar because it carries real work, not commentary.
Same 27B, same Q4.5, 16x apart. My M5 Max: 18 tok/s. Same model on a 2080 Super via OCuLink: 1.1 tok/s. TeksEdge showed a GMKtec EVO-T1 + RTX 2080 Super over OCuLink, PCIe 4.0 x4. 448 GB/s VRAM, 8 GB total. My 27B Q4.5 is 15.2 GB. The card keeps 8, streams 7.2 per token over the cable. PCIe 4.0 x4 is 7.875 GB/s. 7.2 / 7.875 = 0.914 seconds. That is the wall. Per GB the 2080S has 56 GB/s per GB, my box has 2.13. 26x on that ratio. Which is why 7B Q4 at 4 GB would hit 112 tok/s on the card vs 68 here. The combo answers well: what fits 8 GB. What it cannot answer: my 27B. Haven't tested it pe
Posted by 코지베어 🐻 CozyBear (1.2k followers) 19 h ago · 0 likes · 15 views · view the original post on X. Kept by the AI Radar as AI infra & evals.
More AI work like this
- Trains your own large language model from scratch using plain PyTorch. — @tom_doerr
- Here's a useful use-case for @typesafeai for a change. — @iam_zachi
- This is genius. — @daniel_mac8
- Which GPU should you use for embedding workloads? — @runpod
- Big update to the @Alibaba_Qwen Qwen3.8-Flash-Next single DGX Spark recipe! — @jvr0x
- 🔰 The Databricks Advanced Learning Festival runs September 16 through October 14, with… — @databricks
- Q: What is a trace? — @HamelHusain
- Inference scaling part 1. — @rasbt
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.2k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:08 UTC. Full method.