Benchmarks as Targets: Ranking Distortion and On-Premise Risks
Fine-tuning on SWE-bench does not transfer capability to suites like Django or LiveCodeBench. The benchmark becomes a training target and stops measuring general ability. For self-hosted LLM operators, an inflated score distorts VRAM, quantization, a...