Composite decision scores need an explicit policy for disagreement
I want a composite decision score to carry an explicit policy for disagreement, not just a set of weights. Once several judgments collapse into one number, the application has decided which concerns may cancel each other out. That decision belongs to the people responsible for the outcome. Leaving it implicit does not remove policy — it hides policy inside arithmetic.
TypeSafe’s Score documentation makes the mechanism concrete. Jev can evaluate separate questions about bug severity, customer frustration, and report quality; application code combines their answers. Each Score uses an ordered array of descriptive levels and returns a probability-weighted position, the level probabilities, a legend, and confidence. TypeSafe’s tutorial normalizes each position by its highest level index before applying weights. I read those weights as the tutorial author’s priorities, not priorities discovered or validated by the model.
Normalization solves a range problem. It does not establish that a step in frustration is worth the same operational loss as a step in severity. Nor does it make the component judgments equally reliable. A detailed report of a cosmetic defect and a sparse report of a service outage ask different things of a support team. I would use report quality to decide what evidence to collect next, rather than quietly letting poor documentation reduce the urgency of a suspected outage. That is a policy choice worth exposing.
I separate disagreement across dimensions from uncertainty within a dimension. High severity and low frustration can both be correct — a calm customer can report a serious failure. Those answers express competing priorities, not necessarily contradictory evidence. Within one Score, a distribution split between distant levels means something different from a distribution concentrated on an intermediate level. The mean can hide that difference. I want the original component distributions retained so that a ranking cannot erase the reason a case deserved attention.
My disagreement policy would therefore specify precedence before aggregation. A credible signal of a blocking incident would enter a dedicated triage path under an explicitly tested rule. Missing evidence on a consequential dimension would request clarification or review, rather than receive a neutral value that makes the total look complete. Ordinary cases would proceed to weighted ranking only after those rules had been evaluated. Permissions would remain outside that ranking altogether — a favorable score cannot compensate for absent authority.
I would also resist weighting each component by its confidence as an automatic repair. TypeSafe describes confidence as a statistic derived from the returned probability distribution. It is not a second observation, and it is not automatically the empirical probability that the judgment is correct. Downweighting an uncertain severity assessment could suppress exactly the report that needs investigation. Calibration requires labeled outcomes for a defined event; validating the composite additionally requires evidence that its ordering and decision branches serve the intended operational objective.
Before adopting that policy, I would compare weighted ranking alone with the proposed precedence rules on held-out, labeled cases. I would inspect missed urgent cases, unnecessary escalations, review demand, and outcomes in disagreement-heavy slices. The decision record would retain component answers, rubric and model versions, weights, and the rule that selected the branch. These are proposed tests, not results: the local Jev research ran no model experiments, and I have not established that this design improves triage.
I would accept a simple weighted average for a reversible, low-stakes queue where every dimension is genuinely compensable and held-out evaluation supports the ranking. In that bounded setting, the disagreement policy can be short: trade-offs are allowed, and an operator can reorder the queue before any consequential action.
Outside that boundary, I want the application to explain what it does when its inputs pull in different directions. A composite score should summarize an already chosen decision policy. It should not be the place where the organization discovers, after an incident, which concern it accidentally allowed to disappear.