
Tech
We Tested a 35B LLM Against Typed-Decision Models on 12,000 Real RFQs—Confidence Changed the Winner
91.9%. 89.6%. 78.0%.
Those were the primary-class accuracies of a typed-decision API, a 35B mixture-of-experts LLM, and a 421M open-weight decision model on the same real classification job.
But accuracy was not the result that changed the deployment decision.
The decisive question was:
Can the model tell us when its answer is safe enough to automate—and when a human needs to look?
To find out, I compared Jev , Qwen3.5-35B-A3B , and Laya 421M on 12,000 U.S. federal IT sol...
Read the full discussion on Dev.to
This article was aggregated from Dev.to. Click to join the conversation.
View on Dev.to