VLAs? WAMs? Are language and video models the right foundations for robotics?
This is a AI post classified by Jev as AI research (a model release), kept by the AI Radar because it carries real work, not commentary.
VLAs? WAMs? Are language and video models the right foundations for robotics? Introducing Grounded Action Model (GAM): a new paradigm that builds robot foundation models on top of a pretrained 3D grounding model. Ground first. Then learn to act. 🧵👇
Posted by Jiafei Duan (7.4k followers) 1 h ago · 53 likes · 2k views · view the original post on X. Kept by the AI Radar as AI research.
More AI work like this
- New podcast with @datagenproc of @EpochAIResearch digging into the open questions… — @natolambert
- OPENAI'S AI JUST SOLVED A MATH PROBLEM THAT WAS SUPPOSED TO BE IMPOSSIBLE FOR COMPUTERS. — @ArkkDaily
- Aging remains one of the most complex challenges in biological science. In a landmark… — @InSilicoMeds
- 57 interviews before joining OpenAI—and then she published the entire playbook. — @davidarngar
- 很有意思的图啊,MiMo 是想一步一个脚印靠自己爬~ — @Fei2411
- “MiMo-V2.6 Scaling Reinforcement Learning Towards Self-Improvement” — @askalphaxiv
- Hoping someone can explain to me what's going on here. — @lu__jasper
- $10,000 went out to solvers last week. — @yukonresearch
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 41.8k posts from 5k X accounts over the last 14 days, 1.7k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 15:28 UTC. Full method.