Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Read by thought-leaders and decision-makers around the world. Phone Number: +1-650-246-9381 Email: pub@towardsai.net
228 Park Avenue South New York, NY 10003 United States
Website: Publisher: https://towardsai.net/#publisher Diversity Policy: https://towardsai.net/about Ethics Policy: https://towardsai.net/about Masthead: https://towardsai.net/about
Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Founders: Roberto Iriondo, , Job Title: Co-founder and Advisor Works for: Towards AI, Inc. Follow Roberto: X, LinkedIn, GitHub, Google Scholar, Towards AI Profile, Medium, ML@CMU, FreeCodeCamp, Crunchbase, Bloomberg, Roberto Iriondo, Generative AI Lab, Generative AI Lab VeloxTrend Ultrarix Capital Partners Denis Piffaretti, Job Title: Co-founder Works for: Towards AI, Inc. Louie Peters, Job Title: Co-founder Works for: Towards AI, Inc. Louis-François Bouchard, Job Title: Co-founder Works for: Towards AI, Inc. Cover:
Towards AI Cover
Logo:
Towards AI Logo
Areas Served: Worldwide Alternate Name: Towards AI, Inc. Alternate Name: Towards AI Co. Alternate Name: towards ai Alternate Name: towardsai Alternate Name: towards.ai Alternate Name: tai Alternate Name: toward ai Alternate Name: toward.ai Alternate Name: Towards AI, Inc. Alternate Name: towardsai.net Alternate Name: pub.towardsai.net
5 stars – based on 497 reviews

Frequently Used, Contextual References

TODO: Remember to copy unique IDs whenever it needs used. i.e., URL: 304b2e42315e

Resources

Free: 6-day Agentic AI Engineering Email Guide.
Learnings from Towards AI's hands-on work with real clients.
Jev and RLCD: Architecture, Open-Source Models, and Practical Agent Use Cases
Latest   Machine Learning

Jev and RLCD: Architecture, Open-Source Models, and Practical Agent Use Cases

Last Updated on October 6, 2026 by Editorial Team

Author(s): Rajesh K

Originally published on Towards AI.

Jev and RLCD: Architecture, Open-Source Models, and Practical Agent Use Cases

A practical guide to typed AI decisions, calibrated probabilities, and the software that turns them into useful workflows.

An agent receives a request: “Investigate the failed deployment and explain what changed.”

Before it produces an answer, the system must make several decisions. Which capability should handle the request? Does it have enough context? Which tool should run next? When should it stop gathering evidence?

These decisions are small individually, but their quality shapes the entire workflow. A wrong route can waste several tool calls. A premature stopping decision can produce an unsupported explanation.

Jev makes this decision layer worth examining explicitly. This article explores its interface, the meaning of RLCD, two independent open-source implementations, and practical ways to evaluate them.

Table of contents

  1. What is Jev? What is RLCD?
  2. Jev architecture: the documented interface
  3. RLCD architecture: an inspectable open-source training loop
  4. Practical use cases and examples
  5. Calling hosted Jev from Python
  6. Open-source models and their architectures
  7. Running a local decision model
  8. Comparisons and trade-offs
  9. Evaluating a decision layer
  10. Further reading

1. What is Jev? What is RLCD?

Jev is TypeSafe AI’s decision model. TypeSafe describes it as a System One model: it reads supplied context and answers typed questions with values and probabilities, rather than writing an open-ended reply. Its current interface accepts text, including structured text such as JSON. It does not generate code or explanations. See the official System One documentation.

The core idea is decisions for code. A chat response needs to be interpreted; a typed decision supplies a value that software can use in a branch, route or scoring rule. The application declares the question and the permitted answer space.

Three primitives define its answer space:

Jev and RLCD: Architecture, Open-Source Models, and Practical Agent Use Cases

For a support ticket, Choice can select among billing, technical and sales. Score can place the customer’s frustration on a defined scale. Noul can assess the proposition that the ticket is urgent. These are separate judgments: a technical category does not, by itself, establish urgency. Numbers shown in an illustrative example are not measured accuracy or guarantees.

Choice and Score include distributions and a confidence field; Noul returns a probability between zero and one. These are documented in TypeSafe’s introduction.

RLCD means Reinforcement Learning for Calibrated Decisions. TypeSafe describes it as training for decisions whose stated probabilities reflect observed outcomes. The AI primer explains the objective, but does not provide a reproducible implementation of Jev’s internal training method.

Calibration is measured across predictions. Among comparable events assigned an 80% probability, roughly 80% should occur. That does not guarantee the outcome of a particular decision.

One subtlety matters: Jev’s confidence is a statistic derived from the distribution. For Choice it rescales the highest probability relative to an even split among the options. It is not simply the probability that the selected answer is correct. See the confidence documentation.

2. Jev architecture: the documented interface

It helps to separate three things: the model’s API, its internal architecture, and the application surrounding it.

The documented API evaluates multiple questions independently against the same state. Applications can mix the three primitives in one request. Questions should be narrow; separate judgments are composed through code. This describes an interface contract, not a complete account of the model’s layers or training machinery. Source: TypeSafe introduction.

For an agent application, I would use the following architecture:

This is a proposed application architecture, not a diagram of Jev’s internal network.

Each component has a distinct responsibility:

  • The state builder selects relevant context and identifies available capabilities.
  • The decision model supplies bounded judgments.
  • Policy code checks authorization, budgets and allowed actions.
  • The harness manages execution, state, retries and tool calls.
  • Verification checks whether the result supports completion.

LangChain’s Jev harness guide illustrates model routing and tool-risk middleware. These are examples of placing decisions inside the execution loop, rather than treating Jev as a replacement chat model.

3. RLCD architecture: an inspectable open-source training loop

RLCD is a training objective, not one universally specified neural architecture.

For a concrete implementation, consider eve-rlcd. It scores declared options through a constrained letter-token readout. Its policy samples an option, receives feedback about whether that option was correct, and uses the reward:

reward = outcome − probability_assigned_to_sampled_option

Here, outcome is one for a correct selection and zero otherwise. The project connects this REINFORCE estimator to the Brier-score objective under bandit feedback. This is the repository’s method, not a disclosed formula for Jev’s proprietary training.

Conceptual training loop for the independent eve-rlcd implementation.

Why use a proper scoring rule? A useful probability model should be rewarded for representing uncertainty accurately. Always pushing the winning answer toward certainty can make an action policy decisive while making its reported probabilities less useful.

For evaluation, the multiclass Brier score can be written as:

Brier = sum over options of (predicted probability − observed indicator)²

Lower is better. Accuracy evaluates the selected option; Brier evaluates the distribution. Neither metric alone establishes performance on an unfamiliar deployment domain.

4. Practical use cases and examples

Ten uses for Jev

A decision model fits where an application needs a bounded judgment before choosing its next action. The following ten categories follow the supplied presentation; the questions and workflow examples are practical interpretations, not measured deployment results.

For corpus map-reduce, imagine reviewing customer feedback. The map step asks whether each message describes a billing problem and scores its urgency. The reduce step counts categories or summarizes score distributions with ordinary code. Independent documents can be processed concurrently; request size and service limits still apply.

For a live UI, a request such as “show unresolved billing tickets” can map to an allowed filter state. Jev selects a bounded option; the frontend validates the selection and updates its components. This example concerns interface decisions, rather than generating arbitrary UI code.

Architecture: rules, agents, and hybrid software

Rule-based software, agents, and hybrid AI-powered software differ in who controls the workflow. Rules encode branches directly. An agent uses a generative model to select steps as state changes. A hybrid application keeps explicit control flow and calls Jev for bounded decisions at selected checkpoints. It can combine ordinary conditions, model judgments and tools within the same workflow. The model supplies judgments; application code controls allowed transitions and execution.

Figure 3. Three control-flow patterns, conceptually adapted from the supplied architecture slide in Rob Shocks’ Jev breakdown. Branching code is explicit; a model’s decision can still be wrong.

Architecture: one help-desk state, several typed questions

A help-desk application can ask about priority, team, deadline and the next tool against the same ticket state. These questions can share one request while application code combines their answers into the next step.

Figure 4. Proposed help-desk architecture. Policy checks and fallback belong to the application; the diagram does not describe Jev’s neural layers.

Write on Medium

The examples below extend these patterns to other workflows. They are proposed designs, not results from a deployment I have measured.

Example A: Routing an incident investigation

A request says: “A batch job failed with S878. Explain the error and check whether recent changes contributed.”

I would split the request into decisions about the next step: document lookup, operational investigation, clarification, or escalation. The state should include the supplied evidence, environment identifiers and available read-only tools.

If logs are absent, the application can gather them before asking a domain agent to investigate. The domain agent performs the investigation; the router does not establish the root cause.

Example B: Choosing an agent harness

Suppose an agent factory offers two harnesses with different configured skills. Describe their actual capabilities in the routing criteria, including supported tools and execution limits.

Harness names such as Hermes or CUGA are not sufficient criteria by themselves. The router should judge the capabilities actually installed and available.

Example C: Checking evidence in RAG

After retrieval, ask separate questions about relevance, coverage and whether a specific claim is supported by the supplied passages.

If the passage discusses storage exhaustion but the answer asserts that a deployment caused it, an evidence check should flag that gap. The application can retrieve change records or narrow the answer.

A model’s support judgment is still a prediction. Keep the source passages and citations available for verification.

Example D: Selecting the next tool

For a code investigation, candidate actions might be “search symbols,” “read a file,” “run a permitted check,” or “request clarification.” Restrict candidates to actions available in the current state.

The harness validates arguments and tool results. Authorization remains a deterministic check; a model’s high probability does not grant permission.

Example E: Assessing ticket completeness

Define an ordered rubric: missing key identifiers, partially specified, sufficient for investigation. Judge that dimension separately from urgency.

A ticket can be urgent and incomplete. Separating the two avoids hiding conflicting judgments inside one overall score.

5. Calling hosted Jev from Python

The official quickstart documents typesafe-sdk, TypeSafeClient and the system_one call. Set TYPESAFE_API_KEY through your environment before running this adapted example.

pip install typesafe-sdk
from typesafe_sdk import Choice, Noul, TypeSafeClient
client = TypeSafeClient()
response = client.system_one(
state={
"request": "Explain S878 and investigate this job failure.",
"evidence": "No job identifier or incident logs supplied.",
},
questions={
"next_step": Choice(
instructions="Choose the next step supported by this context.",
criteria={
"documentation": "Explain the error from documentation.",
"investigation": "Investigate a specified incident with evidence.",
"clarification": "Request missing incident details.",
},
),
"has_job_id": Noul(
instructions="Does the supplied context identify the failed job?"
),
},
)
route = response.answers["next_step"]
print(route.choice)
print(route.probabilities)
print(route.confidence)
print(response.answers["has_job_id"].noul)

Treat this as a decision probe. Do not assume that it will select a particular route until you run it. For reproducible evaluation, record the resolved model version, criteria and input alongside each output.

6. Open-source models and their architectures

Two inspectable projects provide different ways to explore the idea.

Both are independent projects. Neither claims to reveal Jev’s internal training algorithm.

In eve-rlcd, the decision-only export retains the transformer body and a small letter-token decision head. It requires the repository’s loader, rather than a stock text-generation pipeline. The inference documentation describes prefilling the shared state once and batching question suffixes against that cache.

OpenJev-RLCD follows a different path. Its paper scores answer distributions after sampled reasoning. Its repository describes a two-stage recipe: first calibrate the readout, then reinforce rationales with a proper-score reward and a KL anchor. Because this path generates rationales, its latency profile should not be assumed equivalent to a direct decision readout.

7. Running a local decision model

For a first local experiment, use the released eve-rlcd export. The model card documents the download and typed question interface.

git clone https://github.com/anthony-maio/eve-rlcd
cd eve-rlcd
uv sync
uv run hf download anthonym21/qwen3-0.6b-rlcd-decision \
--local-dir qwen3-rlcd-decision

Save the following as local_demo.py inside the repository, then run uv run python local_demo.py.

import json
from rlcd.decide import ChoiceQ, Decider, NoulQ
decider = Decider.load("qwen3-rlcd-decision", device="cpu")
answers = decider.ask(
"Customer reports API errors after a deployment. No billing issue.",
[
ChoiceQ(
"Which department should handle this ticket?",
["BILLING", "INFRASTRUCTURE", "SECURITY", "PRODUCT_SUPPORT"],
),
NoulQ("Does this ticket describe an API failure?"),
],
)
print(json.dumps(answers, indent=2))

The loader’s device argument supports selecting CPU. This example has been checked against the source interface; it is not a measured CPU benchmark. The repository’s environment pins a CUDA PyTorch build, so uv sync is not a minimal CPU-only installation.

The model card describes the model as a research demonstration. Its published calibration evidence concerns its evaluation distribution, not arbitrary enterprise traffic. Do not interpret local deployment as proof of production suitability.

For OpenJev-RLCD, start by rebuilding the published result tables:

git clone https://github.com/ZimmyGao/openjev-rlcd
cd openjev-rlcd
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
bash scripts/make_tables.sh

The repository says table reconstruction requires no GPU; its training experiments used a 48 GB GPU. See its reproduction instructions before attempting training.

8. Comparisons and trade-offs

These are architecture comparisons, not a head-to-head benchmark.

Ordinary classifiers can also be calibrated. A decision API is valuable because it defines a useful software contract; it does not make classification a newly invented problem.

RLCD should also be compared with simpler training baselines. If complete labels are available, supervised training with calibration may be the appropriate starting point. Bandit learning becomes relevant when feedback identifies whether a chosen action worked without identifying every alternative’s outcome.

Do not reuse thresholds blindly across implementations. Jev and eve-rlcd expose different confidence definitions. Normalize adapters around explicit option probabilities and validate their meaning for each provider. See TypeSafe confidence and the eve model card.

An independent September 2026 Jev evaluation preprint reports strong results across several classification datasets, but also limitations on noisy labels and rubric judgments. It reports that some binary tasks benefited from threshold tuning. Those findings justify testing a candidate; they do not establish results for an enterprise incident workflow.

9. Evaluating a decision layer

I would begin with a labelled set of representative requests, including unsupported tasks and ambiguous cases. Separate training or calibration data, threshold-selection data, and the final test set.

Measure the complete workflow:

For a local CPU experiment, record CPU type, thread count, memory, input length and concurrency. Compare cold and warm runs. Do not transfer a reported GPU latency number to CPU inference.

A provider adapter should preserve the selected value, full probability distribution, model version and criteria version. Add timeouts, unsupported-option handling and a fallback route before connecting the router to an operational harness.

My working hypothesis is that focused decision models can make agent control flow easier to inspect and improve. The evidence to seek is concrete: fewer wrong routes, better handling of missing context, and lower cost per successfully completed task.

10. Further reading

Official documentation, source code and research papers provide further detail on the interface and independent open-source implementations.

Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.

Published via Towards AI


Towards AI Academy

We Build Enterprise-Grade AI. We'll Teach You to Master It Too.

15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.

Start free — no commitment:

→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day

→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages

Our courses:

→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.

→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.

→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.

Note: Article content contains the views of the contributing authors and not Towards AI.