Decision Models: The Non-Generative Layer Slotting Under Agent Pipelines
A new class of small, structured-output-only classifiers is being wired into agent graphs for routing, scoring, and judging — decoupling 'decide' from 'generate' in ways that change both cost and evaluation coverage.
A quiet shift ran through the week ending September 26, 2026: several independent releases converged on the idea that the decision steps in an agent graph — routing, classification, scoring, judging — should not be handled by the same frontier LLM that does the generation. Instead, a distinct class of small, structured-output-only models is being slotted underneath the agent, and the economics are changing what teams can afford to evaluate.
The generative/decisional split is now a product category
Three separate LangChain releases in the same week point at the same architectural seam. Jev, TypeSafe AI’s ‘decision model’, is pitched explicitly as a System-One replacement for frontier LLMs on narrow routing and classification steps inside LangGraph, with claimed 200x speed and 400x cost improvements on those specific tasks. LangSmith Gateway now hosts SemIf, a 4B open-source decision model that returns structured classifications and scores rather than generated text. And the TypeSafeClassifier runnable exposes typed yes/no, choice, and score primitives as first-order components in agent middleware.
The framing across all three is the same: these are not smaller LLMs competing on the generative frontier. They are a different shape of model — one that produces distributions over a fixed schema and nothing else. That constraint is what makes them fast and cheap enough to run on every hop of an agent graph rather than only at decision points expensive enough to justify a Claude or GPT call.
This is the same move that happened in web serving a decade ago: nginx displaced Apache from the outer layer not because it was better at everything, but because it did a narrower thing (routing, TLS, static content) faster and cheaper, and passed the hard work upstream. Model routing has been discussed as a cost play; decision models make it a type distinction rather than a tier distinction.
Judging every trace changes what evaluation means
The more consequential effect shows up in evaluation. LangChain reports that using Jev as an LLM judge matched a human reviewer on every call at $0.34 total versus $28.17 for Claude Sonnet 4.6, with up to 913x lower variance and sub-half-second latency. Whether those exact numbers hold across domains is beside the point; the shift is that scoring every production trace becomes financially trivial rather than a sampling problem.
Sampling has been the tacit constraint behind most production evaluation work. You could not afford to run a frontier LLM judge over 100% of traces, so you sampled, and your dashboards were statistical estimates. If a purpose-built classifier can grade every trace at a rounding-error cost, the evaluation surface stops being a sampled signal and becomes a dense one. That in turn feeds directly into the trace-mining loops other teams are building — dense per-trace scores are exactly the signal fine-tuning pipelines need, and LangChain’s parallel launch of LangSmith Fine-Tuning is not a coincidence.
The practical consequence: LLM-as-judge stops being a periodic offline evaluation and becomes an online gate. Every tool call, every subagent handoff, every retrieval result can carry a score. Dashboards shift from ‘here’s our pass rate on last week’s sample’ to ‘here’s the distribution of judge scores across every trace we ran, live.‘
Structured-output constraints are the actual differentiator
It is worth being precise about what these models are and are not. They are not distilled frontier LLMs answering short questions. Their output space is constrained by construction to a schema — yes/no, choice-of-N, bounded score — and they are trained and served with that constraint baked in. That is why the variance numbers are so low: there is no long-tail free-form output to disagree about.
This matters because a lot of practitioner pain in agent systems comes from asking a general-purpose LLM to produce structured output as a side task. Retry loops, schema-repair passes, and JSON-mode fallbacks are all workarounds for using the wrong tool. The structured output patterns literature has largely been about coaxing generative models into behaving decisionally. A model that only does the decisional thing sidesteps that class of failure entirely, and pushes the generative model back to what it is actually good at.
The corollary is that not every ‘small model’ call in your graph is a candidate. If a step genuinely needs to produce prose — a search query, a plan, a code diff — a decision model is the wrong substrate. The split is along output shape, not along task difficulty.
What this reorganizes in the harness
Agent harnesses have been trending toward middleware architectures where tool calls, routing, and safety checks all sit as interceptable stages. Managed Deep Agents formalizes middleware for tool-call interception; the Jev-in-harness webinar walks through inserting decision models at exactly those seams — confidence thresholds, safety controls, model routing, judge stages for RL verifiers. What was previously a prompt-and-parse pattern (call the LLM, parse JSON, branch) becomes a typed component with its own latency and cost budget.
For practitioners auditing an existing agent, the useful exercise is to walk the graph and tag each LLM call as either decisional (output is one of a bounded set of things) or generative (output is open-ended text). The decisional calls are the candidates. In most production agents this includes: tool selection when the tool set is small, retrieval reranking, safety pre-checks, sub-agent routing, done/not-done judgments in ReAct loops, and post-hoc quality scoring. That inventory is usually longer than teams expect — often a majority of calls — and it is exactly the workload that a frontier model is overqualified for.
The deeper implication is for agent architecture as a whole. Cost pressure has been pushing teams toward tiered routing between frontier and mid-tier LLMs. Decision models introduce a third tier that isn’t on the LLM spectrum at all — and the boundary of what belongs in the LLM layer versus what belongs in a specialized classifier is where the next round of architectural work is happening.
This article was generated with AI assistance and reviewed by the editors before publication.