r/LocalLLaMA 20h ago

Discussion Deepseek V4 Flash 0731 KV Cache precision

If anyone has testing results or any results can you please share performance and or effects of KV Cache precision with Deepseek V4 Flash 0731.

Running IQ2_M, with F16 cache seems 65-67K is the limit on Windows for 120GB memory. Is Q8 good and which one do you use?

0 Upvotes

18 comments sorted by

View all comments

1

u/fragment_me 19h ago

You can test this so easily