Skip to content Navigation Menu Pricing Provide feedback Saved searches Use saved searches to filter your results more quickly Sign up Appearance settings Notifications You must be signed in to change notification settings Fork 19.3k Star 116k Code Issues 687 Pull requests 1.1k Discussions Actions Projects Wiki Security and quality 13 Insights MergedMergedConversation Copy link Copy Markdown Member Overview cont #23398 alt #24270 The copy in can become expensive in some host configurations. This will be refactored properly, but for now a quick patch to avoid the performance hit. Requirements I have read and agree with the contributing guidelines AI usage disclosure: NO Closed 1 task ggerganov deleted the gg/kv-cache-temp-patch-for-sharing branch June 7, 2026 18:43 Copy link Copy Markdown Tested this on an RTX 5090 with Gemma 4 12B QAT + draft-mtp at 64k context. Confirms the perf is restored: structured decode went from 104 to 149 tok/s (+43%), free-text +20%, same draft acceptance. Thanks! Open Labels None yet 2 participants
kv-cache: avoid kv cells copies by ggerganov · Pull Request #24277 · ggml-org/llama.cpp
Skip to content Navigation Menu Pricing Provide feedback Saved searches Use saved searches to filter your results more quickly Sign up Appearance settings Notifications You must be signed in to change notification settings Fork 19.3k Star 1
Skip to content Navigation Menu Pricing Provide feedback Saved searches Use saved searches to filter your results more quickly Sign up Appearance settings Notifications You must be signed in to change notification settings Fork 19.3k Star 1
- This will be refactored properly, but for now a quick patch to avoid the performance hit.
- Confirms the perf is restored: structured decode went from 104 to 149 tok/s (+43%), free-text +20%, same draft acceptance.
What people are saying
Hot takes
Loading takes…
Comments
Discussion · 0
Sign in to comment, like, and save articles.
Sign inLoading comments…

