Classifier label changes are downstream data-contract changes
I treat a classifier’s label set as a downstream data contract, not a collection of prompt strings. Once a label selects a queue, feeds a metric, or becomes a training target, changing its meaning changes the interface those consumers depend on. The dangerous release is the one that preserves every JSON field while quietly redrawing the categories. Nothing fails to parse; the business starts counting and acting on something different.
Jev makes this boundary unusually visible. TypeSafe’s current Choice documentation says that option names and descriptions both reach the model; the question identifier does not. The response includes a selected choice, probabilities across the supplied options, and confidence. I read that as two contracts joined together: the criteria tell the model what distinctions to make, and the returned labels tell application code what those distinctions mean. Editing a description can therefore change production behaviour without changing an enum.
Consider a support classifier whose billing category covers charges, invoices, and payment problems. A team splits refund status into its own category and calls the change additive. For an existing consumer, it may be subtractive: billing no longer contains the same population. A dashboard can report fewer billing incidents without any improvement in service. A routing rule can stop sending refund cases to the people who previously owned them. I would require the producer to explain that redistribution before accepting the new label.
This is where I extend data contracts beyond field definitions. My classifier contract would name the taxonomy owner, stable business identifiers, inclusion and exclusion criteria, supported combinations, and the meaning of an unmatched result. It would distinguish a model-facing option name from an application-owned identifier and a display caption. Those distinctions let a team change presentation without accidentally changing inference. They also make merges, splits, and deprecations explicit operations rather than edits hidden inside a configuration diff.
The choice of classifier belongs in that contract too. The Jev research distinguishes a single TypeSafe Choice from GLiClass’s documented multi-label and hierarchical classification interfaces. Selecting one category and returning several applicable categories are different promises. I would not replace one with the other behind an unchanged field called labels. Nor would I assume a dotted hierarchy enforces parent-child consistency. Consumers need to know whether categories compete, coexist, or require deterministic validation before they can interpret the result.
I would make migration evidence consumer-specific. Run the old and proposed definitions against the same adjudicated examples, inspect which cases cross boundaries, then trace those movements into queues, reports, and policy branches. Test mixed requests and cases outside the taxonomy, not just clean examples of each category. TypeSafe recommends an explicit other option when a Choice set does not cover every input. That option needs an owned destination. These are proposed release checks, not results from this research: no Jev inference experiment or integration benchmark was run.
Historical data needs a migration decision as well. I would retain the original label and taxonomy version, publish an explicit mapping where one is defensible, and keep reclassified history distinguishable from original observations. A category split cannot recover a missing historical distinction merely by renaming rows. The evaluation gate should also separate better classification from easier definitions. Comparing headline accuracy across changed target meanings can reward the taxonomy edit rather than the model.
I would exempt a purely presentational rename from semantic reevaluation when the model-visible criteria, stored identifiers, routing rules, and analytical groupings are all unchanged. That is a caption change, and checking those invariants is enough.
For everything that changes what a category contains, I want a producer-owned compatibility verdict and a consumer migration path. A valid structured response proves the interface can be read. It does not prove the reader is still being told the same thing. The release unit is the meaning downstream systems consume—not the line of configuration someone happened to edit.