Last night, in a basement in Florida, one conversation ran across two kinds of silicon…
This is a AI post classified by Jev as AI infra & evals (a tool drop), kept by the AI Radar because it carries real work, not commentary.
Last night, in a basement in Florida, one conversation ran across two kinds of silicon at once. Two NVIDIA DGX Sparks read the prompt. The finished cache crossed an RDMA link straight into a Mac Studio's memory. Apple silicon wrote the answer. No file in the middle, no re-reading, one model, one thought — GLM-5.3-Flash with ~380,000 tokens of context in view. That hop should not exist. CUDA and Metal were never meant to touch each other's memory. @ashhart built the bridge that makes them: MCDMA — Metal CUDA Direct Memory Access. Kernel-level RDMA on macOS, written by one person, open source.
Posted by Volatile Markets (1.5k followers) 1 h ago · 15 likes · 371 views · view the original post on X. Kept by the AI Radar as AI infra & evals.
More AI work like this
- Wafer just dropped an AI performance engineering repo that covers: — @Hesamation
- How I AI: spend 9¢ to look at ~1750 PRs and group them into thematic investments. 17,364… — @clairevo
- I have to retract what I posted as the fastest speed of Qwen 3.8 Flash on a single DGX… — @yume_arasaki
- 如何自己训练一个 Jev 一样的判断模型? — @AstroHanRay
- Qwen-Image-2.1 on one Spark: day-0 Comfy, 23/23, H3 kept — @MichaelGannotti
- UPDATE we COOKING ! — @Tech2Wild
- just launched teams on http://classifier.dev! — @michael_chomsky
- we've been exploring how to use Jev to build a better harness — @hwchase17
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 35k posts from 5k X accounts over the last 14 days, 1.5k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 21:11 UTC. Full method.