Today we are releasing DolphinBench: Mapping the Pareto frontier of agent memory.
This is a AI post classified by Jev as RAG & memory (a launch), kept by the AI Radar because it carries real work, not commentary.
Today we are releasing DolphinBench: Mapping the Pareto frontier of agent memory. Coding, tool use, and long-context retrieval already have strong public benchmarks. Memory still mostly gets graded as a quiz: ask a question, judge the answer, maybe report precision and recall. DolphinBench grades that action after long history, and puts accuracy, cost, and latency on one board. https://dolphinbench.ai/
Posted by mem0 (20.3k followers) 55 min ago · 36 likes · 2.9k views · view the original post on X. Kept by the AI Radar as RAG & memory. Tools mentioned: DolphinBench.
More AI work like this
- The fastest PDF-to-Markdown parser just got faster. ⚡️ — @llama_index
- Keenable is now a web-fetch backend in @cognee_ — @KeenableAI
- Introducing Alexandria. — @firecrawl
- SANS sınav kitabı 8 bin dolar. Evet, yanlış okumadınız, 8 bin dolar :D — @_shadowintel_
- this guy spent 3 days reverse-engineering @noahrshinn's Instinct’s memory. — @DataChaz
- Meet TeleOCR, a new vision language model that turns any image with text into… — @HuggingModels
- you can find a nice interactive visualization of docjev splitting in action here. — @jerryjliu0
- with jev, everyone is understanding the importance of calibrated confidence scores for… — @jerryjliu0
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 42.4k posts from 5k X accounts over the last 14 days, 1.8k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 17:33 UTC. Full method.