llama.cpp
llama.cpp is Summary Integration checkpoint: Bonsai 2 27B served at its full 262,144-token window with the MTP draft head, on a 12 GB RTX 4070, at 103 tok/s greedy (three-prompt mean, up from 60 without the dra.... It is ranked #1255 on the AI Radar, in Open & local models, first seen 1 h ago and shared in 1 post (4 views).
cuda: Bonsai 2 27B at the full 262k window with the MTP head on 12 GB Ada: #218 + #215 hybrid PTQ1_0 dispatch, in-place q4_0/q8_0 K/V flash attention (integration checkpoint) by professorpalmer · Pull Request #221 · PrismML-Eng/llama.cpp · GitHub Skip to content Navigation Menu Sign in Appearance settings Search / Sign in Sign up Appearance settings You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. Dismiss alert {{ message }} PrismML-Eng / llama.cpp Public forked…
What people said about llama.cpp on X
A 7.8 point jump off a patch is exactly the kind of number my own box can fake with zero code changes. Same 200 token prompt, same flags, warm, four repeats: 18.29 / 19.49 / 17.95 / 18.04 tok/s. Ninety minutes earlier the same prompt measured 55.4. Until a run_order column says what state the engine was in, patch…
— @cozybearlog, 1 h ago · 0 likes · see the post
Alternatives to llama.cpp
- Intern-S2-397B — 🤖 ModelScope
- Pirate Face — The un-bannable, checksum-verified mirror for open models. Claim your handle before a squatter does.
- bespoke-nimble-9b — We’re on a journey to advance and democratize artificial intelligence through open source and open science.
- atomic.chat — Run AI models locally ->
- deepseek-v4.1-flash-exl3-2x-dgx-sparks — DeepSeek v4.1 Flash EXL3 2.9 bpw for 2x DGX Sparks - MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks
- orcabonsai-27b-uncensored — Open source
llama.cpp in numbers
- Rank on the AI Radar: #1255 of 1345
- Shared in 1 post by 1 account: @cozybearlog
- 4 views on those posts
- First seen 1 h ago, last shared 1 h ago
- Pricing seen by Jev: open source
- Market: Open & local models
FAQ
What is llama.cpp?
Summary Integration checkpoint: Bonsai 2 27B served at its full 262,144-token window with the MTP draft head, on a 12 GB RTX 4070, at 103 tok/s greedy (three-prompt mean, up from 60 without the dra... It was first shared on X 1 h ago and is ranked #1255 on the AI Radar.
Is llama.cpp free?
It is open source.
Who shared llama.cpp?
1 account on X, including @cozybearlog, in 1 post totalling 4 views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 31.7k posts from 4.8k X accounts over the last 14 days, 1.3k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 04:05 UTC. Full method.