Cut Your Jev Bill Linearly with ContextPress
Last Updated on September 22, 2026 by Editorial Team
Author(s): Taha Azizi
Originally published on Towards AI.
Why Jev Is Different
Jev is TypeSafe’s first System One model. It returns typed decisions, not text. On OpenRouter it costs $0.042 per million input tokens and $0 for output.
Your cost is the sum of all tokens in your prompt. Jev’s usage.cost is linear in usage.input_tokens, so token savings equal dollar savings.

Why It Matters
Jev is cheap. The problem starts when you use it more. Decisions are cheap, so you run them a lot. Volume grows, prompts grow with it, and the bill follows.
At my measured mix, a raw prompt averages about 2,450 input tokens. That works out to about $103 per million decisions. It is a small number per call but a real number per month once Jev sits in a high-volume pipeline.
Since output is free, the prompt is the only thing you can shrink. That is where ContextPress comes in.
What I Tested
In my previous article, Cut Your LLM Token Costs for Free with ContextPress, I benchmarked the three presets on 222 real conversation items and measured token savings and critical fact loss. This time I wanted to see what happens on a real bill, with a real model, on labelled decisions.
The question is: What fraction of Jev’s billed input does ContextPress remove, and does Jev still answer correctly?
How I Measured It
I used two public labelled datasets: BoolQ (RAG, yes/no) and QuALITY (long files, 4-way multiple choice). That is 200 items and 800 Jev calls, each item billed raw and again at low, medium and high.
The metric is the invoice save: one minus compressed input tokens over raw input tokens. Confidence intervals come from a 10,000-resample bootstrap on the 120 items where I kept the paired bills.
Numbers Tell The Story

The raw bill was $0.01238. After ContextPress it was $0.01193 on low, $0.00692 on medium and $0.00419 on high.
RAG on medium saves 28.33% and files on medium save 51.69%. Retrieved Wikipedia is already tight, so there is less to remove. Long articles have more to give. Give ContextPress real turns and it has room to work.
Did the Cheaper Prompt Change Jev’s Answers?
Cost is only half of the picture. Compression removes tokens, so I checked whether Jev still reached the same decisions.

low matched raw accuracy exactly, and it saves about 4% of the bill. The medium drop of 3.33 points is not statistically significant on 120 items. That does not prove there is no loss, only that this sample cannot separate it from noise. Most of it came from long files, where the drop was 6.7 points. However, RAG on medium showed no drop. The 8.33 point drop is statistically significant, and it tells us that high is too aggressive to keep the facts Jev needs.
Choose by Budget
The right preset depends on what you can afford to lose.
low: no accuracy drop and about 4% cost savings. Choose it when accuracy comes first and the savings are a bonus.medium: a much bigger cost cut, with a drop of about 3 points in this study. Choose it when volume is high and you can accept a small accuracy trade-off. Test it on your own labels first, especially for long files.high: avoid it unless the budget is very limited and the task tolerates lost facts.
Cut Even More by Stacking
ContextPress is always lowering the cost but in Jev’s case it is more direct and obvious. Here are some instructions to keep your bills under control.
- Retrieve less, then compress.
mediumremoves 28% of a RAG bill, you save 28% on your Jev’s token cost. - Use Jev as a cheap first pass. Escalate low-confidence cases to a larger model. ContextPress shrinks the prompt for both.
- Fit more in the window. Jev has a 32K/64K context, and compression lets you pack more evidence into it.
At my measured mix, medium brings the bill from about $103 to about $58 per million items, and low brings it to about $98/$99.

Try It Yourself
pip install contextpress
from contextpress import ContextManager
cm = ContextManager(type="rag_doc", compression="low") # or "medium"
packed = cm.compress(chunk_turns + [question]) # many turns, question last
Use rag_doc for retrieved chunks and files, and put the question last so recency knows what to keep.
Caching Note
Jev’s bill had no cache discount in this study, so compression is safe. If you add a frontier fallback, compress only the new, uncached part and keep a stable prefix untouched. Recompressing the whole history breaks prefix caching. I cover this trade-off in more detail in the appendix of the previous article.
Try it on your own bills and read usage.cost yourself. Please like, subscribe, visit the GitHub page and feel free to contribute.
Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.
Published via Towards AI
Towards AI Academy
We Build Enterprise-Grade AI. We'll Teach You to Master It Too.
15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.
Start free — no commitment:
→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day
→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages
Our courses:
→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.
→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.
→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.
Note: Article content contains the views of the contributing authors and not Towards AI.