DolphinBench
DolphinBench is DolphinBench is an open benchmark from Mem0 that evaluates long-term agent memory through future task completion. Agents with a memory system act on 600 requests across 3 simulated users with years of conversation…. It is ranked #258 on the AI Radar, in RAG & memory, first seen 1 h ago and shared in 1 post (2.9k views).
DolphinBench: Mapping the Pareto frontier of agent memory DolphinBench Discord GitHub DolphinBench · 3 personas · 600 tasks Mapping the Pareto frontier of agent memory DolphinBench compares agent configurations across memory systems, models, providers, and harnesses. Agents complete tasks that depend on past conversations and must recognize which earlier information matters to guide their decisions and actions. View leaderboard → Explore dataset Run and submit Loading results... Why DolphinBench We evaluate agent memory directly through future task completion Most memory benchmarks ask a question and check the answer. DolphinBench gives the…
What people said about DolphinBench on X
Today we are releasing DolphinBench: Mapping the Pareto frontier of agent memory. Coding, tool use, and long-context retrieval already have strong public benchmarks. Memory still mostly gets graded as a quiz: ask a question, judge the answer, maybe report precision and recall. DolphinBench grades that action after…
— @mem0ai, 1 h ago · 36 likes · see the post
Alternatives to DolphinBench
- Firecrawl Alexandria — The knowledge library for superintelligence.
- liteparse — A fast, helpful, and open-source document parser. Contribute to run-llama/liteparse development by creating an account…
- Developer Documentation — If you have needs for large scale doc extraction, come check it out
- agentmemoryleaderboard.ai — 文本、代码、多模态三赛道,以统一 Add / Search 契约检验 Agent 记忆能力。
- pixelrag — https://arxiv.org/abs/2606.28344. The end of web parsing. The beginning of scalable pixel-native search. link:…
- t-mem — T-Mem: Memory That Anticipates, Not Archives — accepted to EMNLP 2026 Main Conference. Trigger-augmented graph memory…
DolphinBench in numbers
- Rank on the AI Radar: #258 of 1771
- Shared in 1 post by 1 account: @mem0ai
- 2.9k views on those posts
- First seen 1 h ago, last shared 1 h ago
- Pricing seen by Jev: open source
- Market: RAG & memory
FAQ
What is DolphinBench?
DolphinBench is an open benchmark from Mem0 that evaluates long-term agent memory through future task completion. Agents with a memory system act on 600 requests across 3 simulated users with years of conversation… It was first shared on X 1 h ago and is ranked #258 on the AI Radar.
Is DolphinBench free?
It is open source.
Who shared DolphinBench?
1 account on X, including @mem0ai, in 1 post totalling 2.9k views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 42.4k posts from 5k X accounts over the last 14 days, 1.8k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 17:33 UTC. Full method.