Skip to content

llama : enforce the same K and V cache types for DeepSeek V4; enable FA if V cache is quantized - #25871

Open
fairydreaming wants to merge 3 commits into
ggml-org:masterfrom
fairydreaming:force-fa-quant-kv-cache
Open

llama : enforce the same K and V cache types for DeepSeek V4; enable FA if V cache is quantized#25871
fairydreaming wants to merge 3 commits into
ggml-org:masterfrom
fairydreaming:force-fa-quant-kv-cache