
Tech
Inside the eval gate: how we decide a cheaper model can take over a task
I build Operant, so I have a bias. This post first went out on our blog. I post it here because I want critique of the method. My questions are at the end.
Operant can move a repeated task to a cheaper model. It does this only after a gate. The gate replays held-out work on the cheaper model with a learned skill. A judge scores the result against your own past answers. Then a person ratifies the change.
This post explains each step of the gate. The screenshots come from the live console. T...
Read the full discussion on Dev.to
This article was aggregated from Dev.to. Click to join the conversation.
View on Dev.to