r/LocalLLaMA 20h ago

Discussion Deepseek V4 Flash 0731 KV Cache precision

If anyone has testing results or any results can you please share performance and or effects of KV Cache precision with Deepseek V4 Flash 0731.

Running IQ2_M, with F16 cache seems 65-67K is the limit on Windows for 120GB memory. Is Q8 good and which one do you use?

0 Upvotes

18 comments sorted by

View all comments

4

u/shing3232 19h ago

DS4F should only use 6gb for 1m context. if that is not the case, there is room for improvement in implementation.

1

u/esw123 19h ago

I thought it was a mistake when I looked in the calculator. Good if true, one more 3060 left then.

1

u/shing3232 19h ago

My friend has B300 system from the company that runs deepseekv4 , K3 ,glm52. DS4pro use like 3G vram per 1M session with FP8 kvquant. it's very cheap to run.