Maple-Preview and Ternary Inference: The On-Premise Threshold Lowers
Maple-Preview applies ternary weights to a 20B-parameter model, shrinking memory to ~5 GB while activating only 1B parameters per token via an MoE-like design. This preview hints at local inference on consumer GPUs, but framework support remains imma...