Budget: a dense 9.4B LLM on a single card shifts the bottleneck from GPU to community
An independent researcher has prepared a dense 9.4-billion-parameter model for training on a single card, with distillation from Llama 3 and architectural simplifications. No benchmarks are public, but the signal for on-premise AI is clear: the bottl...