Yeah the calibration is really what makes it useful in practice for quick, small decisions. Asking a LLM to give scores to a problem will yield inconsistently scaled/anchored results that changes at a whim.
The blog is pretty heavy on statistics. I'll have to study it more when I have time. Is it essentially bootstrapping results to statistically normalize the answers?
Enough from the Jevbros! You can take half the real estate on the front page of HN, but you can't have the discoverability of Reflective Liquid Crystal Displays
Yeah the calibration is really what makes it useful in practice for quick, small decisions. Asking a LLM to give scores to a problem will yield inconsistently scaled/anchored results that changes at a whim.
The blog is pretty heavy on statistics. I'll have to study it more when I have time. Is it essentially bootstrapping results to statistically normalize the answers?
RLCD, not defined in the article, is Reinforcement Learning for Calibrated Decisions.
Enough from the Jevbros! You can take half the real estate on the front page of HN, but you can't have the discoverability of Reflective Liquid Crystal Displays