Decision-model replay needs versioned state, questions, and labels
I treat decision-model replay as a data-versioning problem — preserve the state, the question, and the meaning of every label, or the rerun answers a different question. A model version alone cannot identify the decision that happened. If a team changes what “urgent” means while keeping the same field name, yesterday’s verdict and today’s verdict are no longer directly comparable. The missing artifact is a versioned decision record, not another dashboard.
Jev makes this boundary unusually visible. It is TypeSafe’s structured-decision model, not a routing framework, an orchestration runtime, or a policy engine. Routing is one application of its judgments; the surrounding application still decides what to execute. TypeSafe’s Choice documentation says that option names and descriptions reach the model, while question identifiers do not. I therefore regard a label description as executable decision configuration, not explanatory metadata that someone can tidy without review.
I would start the record with the exact state supplied to the model, alongside its source references and the version of the transformation that assembled it. A link to a mutable customer record is insufficient. A later lookup can import a corrected address, a new complaint, or a changed account status into an apparently historical test. I want to distinguish a change in the evidence from a change in the judgment. Redaction and compaction belong in that lineage because they determine which evidence survives.
Next I would preserve the complete question definition: wording, primitive, option names, descriptions, and ordered rubric where applicable. TypeSafe’s published jaggedness examples show that a yes/no Choice and a Noul can produce different probabilities for similar propositions. Those are vendor-reported examples, not results I reproduced. Their architectural implication is direct: changing the formulation is an evaluation change. A threshold tuned against one formulation does not become valid for another because both outputs fit between zero and one.
Labels need a second kind of versioning: the reference answers used to grade the replay. The available options describe what the model can say; adjudicated labels describe what evaluators consider correct. I would store annotation guidance, adjudication history, and dataset membership separately from predictions. If “refund request” expands to include implicit dissatisfaction, I want both the old grading contract and the revised one. Otherwise an apparent model improvement could be a quiet relabeling of the test set.
I would also retain the requested model identifier, the model reported in the response, the adapter version, the full returned distribution, and the policy revision that consumed it. TypeSafe documents moving aliases and recommends pinning a version when thresholds depend on it. These records support two different investigations: replay the recorded answer through policy, or rerun the frozen input through a candidate model. Neither investigation should reissue the original side effect. Historical permission is evidence about an earlier action, not permission to perform it again.
My proposed comparison would change one component at a time and report disagreements by scenario, including missing context, overlapping labels, and adversarial state. I would compare semantic correctness, review rates, and costly errors rather than collapse everything into one accuracy figure. I have not run these experiments. The local Jev research establishes documented interfaces and limitations; it does not supply independent calibration measurements or production qualification. Preserving inputs makes such tests interpretable, not automatically successful.
For a disposable, read-only triage experiment whose decisions will never be audited or compared across revisions, I would accept a fixed fixture set instead of a production decision ledger. That exception ends when someone uses historical results to justify a model upgrade, tune an action threshold, or explain a consequential outcome.
I want replay to answer what changed, not merely produce another answer. State preserves the evidence, questions preserve the task, and labels preserve the grading contract. Versioning all three turns a model comparison into an attributable engineering decision.