Spiritbuun's VBR KV cache squeezes more context into 12GB
A llama.cpp fork from spiritbuun adds Variable Bit Rate KV cache quantization, dynamically degrading cache precision only when VRAM pressure demands it.
A llama.cpp fork from spiritbuun adds Variable Bit Rate KV cache quantization, dynamically degrading cache precision only when VRAM pressure demands it.
A community rig ran NVIDIA's Nemotron Puzzle 75B-A9B in NVFP4 across three power-capped RTX 3090s at 132 t/s decode, and asked why almost nobody ships models shaped for multi-24GB rigs.
Unsloth published NVFP4 quantized checkpoints for Qwen3.6 27B and 35B-A3B that reportedly beat NVIDIA's reference quants on throughput, with FP8 KV cache support for longer contexts.