Inkling-Small brings 1M-token context to local GPUs, but the real bottleneck is still memory
thinkingmachines has released Inkling-Small, a mixture-of-experts LLM with 276 billion parameters, only 12 billion active, a 1 million token context window, and NVFP4 and GGUF quantizations. AI-RADAR analysis focuses on the hidden challenge of the KV...