We benchmarked 27 open-source LLMs for quality, latency and reliability—and discovered why benchmark configuration can completely change the results.