Posts tagged “vulkan”
llama.cpp b11413: How to Test Vulkan Sparse FlashAttention with Quantized KV
Check eligibility for llama.cpp’s Vulkan sparse-attention change, understand the 16× cache threshold, and compare performance with the follow-up fix included.
llama.cpp b11414: Vulkan Scratch-Buffer Fix and Long-Prompt Checks
llama.cpp b11414 fixes stale Vulkan scratch-buffer reuse. Learn which workloads merit testing, what upstream results show and how to validate long prompts.