Cheap routing decisions must survive the cost of fallback
I judge a routing decision by the cost of the verified task it helps finish, not by the price of making the choice. A cheap selector can send work down an expensive path: a weak answer, a retry, a stronger model, then a person repairing the result. If those steps erase the saving, the router has improved the appearance of the invoice rather than the economics. This is what I mean when I say cost is an architectural property.
Jev makes this distinction worth examining because its interface is narrower than a generative model’s. It returns typed decisions rather than prose. That makes it a candidate for selecting a model, not an orchestration framework that owns execution and recovery. RouteLLM is a framework for serving and evaluating learned routers between stronger and weaker models. These are different layers. Comparing their headline figures as interchangeable products would skip the integration that determines the bill.
TypeSafe’s current model documentation lists Jev at $0.042 per million input tokens, with output tokens free. Its launch article reports workflow results of 193.6 times faster and 444.6 times cheaper, while saying these likely sit toward the high end of real-world gains. The reference answers average frontier-model predictions, and the comparison asks generative models for probability-bearing structured decisions. Those are vendor-reported results under a particular evaluation design. I have not run a local benchmark, and I would not translate those ratios into savings for an entire agent workflow.
My accounting unit would be total operating cost across a representative task cohort divided by verified successful tasks. The numerator includes routing, embeddings where required, every downstream attempt, validation, repair, human review, and attributable serving overhead. Failed and abandoned tasks still contribute their costs. The denominator counts completed tasks meeting a predefined acceptance contract, not responses returned or routes selected. I would compare policies on the same workload under matched quality requirements, with completion rate reported separately so that rejecting difficult work cannot quietly improve the result.
Latency belongs beside that calculation, not hidden inside an invented currency conversion. I would report end-to-end completion latency and its tail, including backoff and review queues; monetize delay only where the business has a defensible valuation. A cheap first attempt followed by serial fallback can be slower than sending the task directly to a stronger model. Conversely, a wrong answer accepted without repair may look both fast and cheap. Evals are the contract that stops this accounting from rewarding undetected failure.
The documented LangChain integration exposes a concrete boundary: its experimental ModelRouterMiddleware classifies the latest human message and uses the chosen model for every model call in that run. It does not promise automatic reconsideration when a tool result reveals a harder problem. I would therefore specify when escalation happens, what state moves with it, and how many attempts the task may consume. Provider unavailability and a semantically inadequate answer need different recovery paths. Failure-mode design matters more here than an impressive initial selection rate.
I would also keep the routing threshold attached to observed outcomes. RouteLLM warns that routing proportions change with the query distribution; choosing a target share of strong-model calls is not proof of successful completion. The trace should connect the initial choice to retries, fallback, review, and acceptance. Trajectory replay makes the expensive path visible. Without that connection, the selector gets credit for savings while another service or team absorbs the repair bill.
There is a narrow case where I would accept a much lighter evaluation: stable, low-stakes routing with cheap deterministic acceptance checks and bounded, side-effect-free fallback. There, a small representative comparison may establish useful savings without elaborate review-cost modelling. Outside that boundary, I want the decision layer to earn its place on the completed-task ledger. Selecting cheaply is an implementation property; finishing correctly at lower total cost is the economic result.