Tech
A 3B model that beats a 7B: failure-driven orchestration on uncontaminated knowledge
How a fully sovereign QA stack — local Wikipedia index, distilled Qwen2.5-3B reader, and crutches built only from measured failures — went from 33% to 52% on facts no LLM can have memorized, and matched a zero-shot 7B more than twice its size.
TL;DR
Configuration
Post-cutoff-150 accuracy
3B naked (no retrieval)
0.0% [0–2.5]
7B naked
0.0% (by construction — see below)
3B + sovereign chain (start, Sep 29)
33.3%
3B + sovereign cha...
Read the full discussion on Dev.to
This article was aggregated from Dev.to. Click to join the conversation.
View on Dev.to