NInfer: 542 tok/s on Qwen3.6-35B-A3B, single request, one RTX 5090
A solo dev open-sourced NInfer, a from-scratch C++/CUDA engine that sustains 542 tok/s on Qwen3.6-35B-A3B over a full 65K-token decode, single request, on one RTX 5090.
A solo dev open-sourced NInfer, a from-scratch C++/CUDA engine that sustains 542 tok/s on Qwen3.6-35B-A3B over a full 65K-token decode, single request, on one RTX 5090.
Running Qwen3.6 27B with Opencode? The model's appetite for context tokens can gut your usable window fast. Here's what's happening and how to push back.
A Reddit user reportedly ran three concurrent Qwen3.5 122B sessions on a single Mac Studio, serving over 93% of prompt tokens from an on-disk KV cache instead of recomputing on GPU.
Three identical frontier-battery runs separated a useful local 14B model from two subscription agents, then revealed that the test needs a higher ceiling.
Unsloth published NVFP4 quantized checkpoints for Qwen3.6 27B and 35B-A3B that reportedly beat NVIDIA's reference quants on throughput, with FP8 KV cache support for longer contexts.