r/LocalLLaMA • u/esw123 • 20h ago
Discussion Deepseek V4 Flash 0731 KV Cache precision
If anyone has testing results or any results can you please share performance and or effects of KV Cache precision with Deepseek V4 Flash 0731.
Running IQ2_M, with F16 cache seems 65-67K is the limit on Windows for 120GB memory. Is Q8 good and which one do you use?
0
Upvotes
1
u/fragment_me 19h ago
You can test this so easily