Posts tagged “kv-cache”
llama.cpp b11413: How to Test Vulkan Sparse FlashAttention with Quantized KV
Check eligibility for llama.cpp’s Vulkan sparse-attention change, understand the 16× cache threshold, and compare performance with the follow-up fix included.