NInfer: 542 tok/s on Qwen3.6-35B-A3B, single request, one RTX 5090
A solo dev open-sourced NInfer, a from-scratch C++/CUDA engine that sustains 542 tok/s on Qwen3.6-35B-A3B over a full 65K-token decode, single request, on one RTX 5090.
A solo dev open-sourced NInfer, a from-scratch C++/CUDA engine that sustains 542 tok/s on Qwen3.6-35B-A3B over a full 65K-token decode, single request, on one RTX 5090.