543 tok/s: A custom engine makes Qwen 35B fly on a single RTX 5090
NInfer, an open-source C++/CUDA inference engine built from scratch, hits 543 tokens per second on Qwen3.6-35B-A3B with a 65K-token prompt on a single RTX 5090. Custom quantization and hardware-level optimizations push performance far beyond generic engines, marking a win for self-hosting on consumer GPUs where low latency and data sovereignty matter.