Adding an AI model can make an application cheaper to run.
If a specialised model takes enough work off a more expensive language model, total costs fall. It can also be considerably faster: in our trial, Jev responded in 0.32 seconds, compared with 2.44 seconds for the LLM we tested.
We are exploring this principle while developing our new application. Some parts use AI for recurring assessments. The cost of each assessment is small, but across thousands of requests, both cost and processing time add up.
The potential benefit comes from how the work is divided. A specialised model handles well-defined questions. A general-purpose language model remains available for cases that need further interpretation. The benefit depends largely on how much work the first step can handle on its own.
And on quality. A cheaper assessment offers little value if someone then spends extra time correcting mistakes.
A different model for recurring work
The introduction of Jev by TypeSafe prompted us to test this. Jev is a decision model: it receives context and focused questions, and returns choices, scores and probabilities.
It still interprets language, but does not need to generate free-form text to return an assessment. For software that mainly needs a category or score, that can be an efficient approach.
There are now several candidates, including Clef by Cloudflare. We wanted to compare different options and see how they perform on our own examples.
Putting it to the test
We started with one part of the application. We asked Jev, Clef Flash and Claude Haiku to assess 48 practical examples twice. For Haiku, we used a variant of our existing prompt. That resulted in 288 model calls.
The trial ran separately from the existing workflow. We collected the responses for comparison; no changes were applied.
These were the measured response times and estimated API costs:
| Model | Median response time | Estimated API cost per 1,000 calls |
|---|---|---|
| Jev 1.13.0 | 0.32 seconds | $0.20 |
| Clef Flash | 0.68 seconds | $0.40 |
| Haiku 4.5 with our prompt variant | 2.44 seconds | $3.56 |
Measured on 3 October 2026. Costs are calculated from reported usage and the applicable rates. All responses are included, including those rejected by our validation.
Jev was the fastest and cheapest candidate in this trial. However, the models produced different outputs: the decision models answered classification questions, while the LLM produced a more detailed proposal with explanatory text. These figures compare individual calls within our test setup. They do not yet demonstrate savings across the complete workflow.
The economics of combining models
For the business case, what matters most is what happens when we combine these models.
Suppose we combine the models like this:
Jev assesses
Every case is assessed by Jev first.
The application checks
Is there enough information, do the answers agree, and is further assessment needed?
The LLM when needed
Only some cases then go on to the LLM.
Using our measured costs gives the following hypothetical scenarios:
| Workflow for 1,000 cases | Estimated API cost |
|---|---|
| Everything handled by the tested LLM variant | $3.56 |
| Everything through Jev first; 50% then go to the LLM | $1.98 |
| Everything through Jev first; 25% then go to the LLM | $1.09 |
In these scenarios, the extra Jev calls pay for themselves by reducing the number of LLM calls. If everything still goes to the LLM, costs actually rise to approximately $3.76. Cases passed on also incur the processing time of the first step.
An extra model can reduce total costs if it takes enough work off the more expensive model.
These scenarios are not measured savings. We assume that a follow-up LLM assessment costs the same as in our trial. Human review, infrastructure and possible reassessments are not included. Different prompts, batching or reusing requests can also change the balance.
What we are taking into the next trial
The first trial also provided practical lessons for the design of our application:
Fixed rules stay in the software.
A model sometimes recognised information correctly, but then failed to adequately follow the associated business rule.
Each part has a clear role.
The decision model classifies, the software applies rules, and an LLM or a person handles unclear or more complex cases.
Explanations can often use templates.
Established facts and assessments form a clear proposal. The user retains final approval.
Responses need to be usable.
Our validation rejected 30 of the 96 Haiku responses because of their structure. That does not establish 30 incorrect assessments, but it does give us a reason to improve the integration.
The model remains interchangeable.
We use fixed questions and response structures so new candidates can go through the same test. Switching requires further testing and calibration.
The next trial will cover 300–1,000 human-reviewed examples, with a separate holdout set for the final evaluation. We will compare the complete workflow on quality, cost, speed and corrections needed. The current trial contains too few verified answers to identify a reliable winner on quality.
We have not yet established how much work can be handled without an LLM. Because confirmed context was missing, all decision-model results in this trial were routed to additional human review. The next step is to measure how much of the speed and cost advantage we can realise while maintaining quality.
To follow developments, see the community benchmark Decision Index on Hugging Face. A Rubik's Cube demo comparing different types of models also provides a light-hearted illustration of the discussion about dividing tasks. We test the value of the combination against our own practical examples.
Are you building an application with many recurring AI assessments?
We would be happy to help design a practical trial that compares speed, cost and quality together.


