When models disagree—measured by high entropy or wide ensemble variance—it...
https://magic-wiki.win/index.php/Week_1_Disagreement_Instrumentation_Checklist
When models disagree—measured by high entropy or wide ensemble variance—it flags risky inputs worth a closer look. By tracking these cases, you can route the top 1-2% most uncertain predictions to human review