LitJev: the zero-training Jev reproduction on Qwen, tested

Artificial Intelligence 8 min
Abstract editorial still of a local decision layer: probability bars rising from an open-weight model cube on a dark
LitJev turns off-the-shelf Qwen checkpoints into a local /v1/systemone decision API with no training and no generated answer text.

LitJev is an independent, zero-training attempt to reproduce Jev’s decision interface on Hugging Face Qwen checkpoints. It accepts typed questions and returns probability distributions without generating an answer in prose. Its public API follows the Jev schema, but its model, probabilities, confidence values and performance are not those of TypeSafe’s official Jev.

That distinction does most of the work here. LitJev offers a practical local surface for experiments and client development. It does not reveal or recreate TypeSafe’s private RLCD training process.

What LitJev actually is

LitJev is an independent research project by Zhengxu Yu. According to the LitJev README, it puts a Jev-shaped decision API in front of off-the-shelf Qwen models. You provide a shared state, one or more questions and the permitted options. The server returns typed JSON rather than a paragraph that an application must parse.

No fine-tuning is required for the main decision API. There is no answer-text generation in that path either. The project reads scores from the language model’s output head, turns those scores into distributions and assembles the response in code.

The README says LitJev supports the Qwen3.x text and vision family across model sizes. Its default and most-tested checkpoint is Qwen/Qwen3.8-27B, run on a single H100 with 80 GB of memory. That is the project’s stated setup, not a benchmark I reproduced. Vision checkpoints are needed for screenshot-based decisions, and the README does not guarantee support for other model families.

The repository also includes a playground, benchmark modules, optional calibration and an experimental System Two endpoint. Those extras should not blur the central claim: the standard /v1/systemone route is a zero-training readout over a loaded Qwen checkpoint.

The original code is available under Apache-2.0. Adapted jevlike examples retain their MIT license. The README also carries an unusually useful disclaimer: LitJev is not affiliated with or endorsed by TypeSafe AI, is based on a hypothesis formed from public information, and is not the official Jev implementation.

For the broader context, my earlier Jev System One overview explains why typed decisions are different from ordinary chat completion. The AI section collects the rest of this series.

How does the readout work?

Two-step diagram: shared state prefill into a model cube, then option-word readout as probability bars.
LitJev answers without generating text: prefill the shared state, then score each option’s own words.

The mechanism is compact, although it deserves more precision than “Qwen picks an option.”

First comes prefill. A shared state opens every prompt. With the SGLang backend, that prefix can be cached once. With the default Transformers backend, LitJev batches the option rows.

Each question then presents its options as plain text ending in Answer:. For every candidate, LitJev teacher-forces the words of that option and adds up their token log-probabilities. The resulting value is the option score. A softmax converts all option scores for that question into a probability distribution, and Python code builds the typed response.

LitJev scores the option words themselves by default, not arbitrary labels such as A, B or C. The project says this choice avoids priors attached to label tokens. A coded mode remains available through --readout coded for anyone who wants letter-code readout.

Questions in the same request do not see one another. The shared state is common, but one answer cannot quietly influence the next question. According to the README, the API accepts up to ten questions in one request.

This design explains both the appeal and the caveat. It gets a distribution without asking the model to write and then parsing that writing. Yet a softmax output is not automatically a trustworthy real-world probability. The LitJev README says its probabilities are not calibrated by default. A displayed 0.85 therefore should not be treated as an observed 85 percent success rate without validation on the task that matters.

The public Jev primitives fit this response shape:

  • choice selects among named options and returns a probability distribution.
  • score places the input on a defined scale and returns probabilities around the score.
  • noul returns a value from 0 to 1.

TypeSafe’s official documentation says hosted Jev also returns confidence for Choice and Score. If Noul is unfamiliar, see my guide to Choice, Score and Noul.

LitJev vs hosted Jev: same schema, different model

Split panel: local GPU decision API on the left, keyed cloud gateway on the right, shared schema bridge between them.
Same public /v1/systemone contract; different model, auth, and calibration story.

LitJev copies the public request and response contract closely enough to be useful for compatible clients. Both surfaces accept model, state and questions in a POST /v1/systemone request, then return answers and usage. Contract compatibility is the useful part. Numerical equivalence is not promised.

Detail LitJev local Jev hosted
URL http://127.0.0.1:8000/v1/systemone https://api.typesafe.ai/v1/systemone
Authentication None on the local default Authorization: Bearer API_KEY
model value litjev or the loaded checkpoint ID jev-latest
Underlying model A Qwen checkpoint loaded by the user TypeSafe’s official Jev
Probability behavior Not calibrated by default Official Jev behavior and confidence contract

The last two rows are where careless comparisons go wrong. TypeSafe describes Jev as its System One model, trained with RLCD to answer structured Choice, Score and Noul questions. LitJev does not ship TypeSafe’s private weights or reproduce that training. It applies a documented readout method to a general Qwen checkpoint.

So “same schema” means an application can target a familiar JSON shape. It does not mean a local LitJev result should match jev-latest, or that confidence and performance carry across. The LitJev README explicitly says internals, confidence values and performance are not identical.

There are also several hosted routes around the official model, including OpenRouter, Vercel and Cloudflare integrations. I mapped those options in the Jev ecosystem guide. The choice is less about syntax than operational needs: local control and inspection on one side, official model behavior and managed infrastructure on the other.

How to run LitJev locally

I read the repository instructions for this section; I did not run LitJev on the machine used to prepare this post. The quick start in the README uses Git, uv and the project’s locked environment:

git clone https://github.com/zhengxuyu/litjev.git
cd litjev
uv run --locked litjev --model Qwen/Qwen3.8-27B

The server then exposes its playground at http://127.0.0.1:8000/ and the decision endpoint at http://127.0.0.1:8000/v1/systemone. The first request can take minutes while the checkpoint downloads and loads. LitJev is not on PyPI at the time covered by the README, so the repository is the installation source.

Its basic request shape is:

{
  "model": "litjev",
  "state": "Customer writes: my order arrived broken and I need it replaced today.",
  "questions": {
    "intent": {
      "type": "choice",
      "instructions": "What does the customer want?",
      "criteria": {
        "refund": "Money back",
        "replace": "A replacement",
        "info": "Just information"
      }
    },
    "escalate": {
      "type": "noul",
      "instructions": "Should a human take over?"
    }
  }
}

The exact values returned will depend on the loaded checkpoint and prompt. Values printed in repository examples are illustrative, not measurements from this post.

Transformers is the default backend. The README gives these commands for SGLang:

uv pip install sglang
uv run litjev --model Qwen/Qwen3.8-27B --backend sglang

Hardware is not a footnote. The named 27B default was tested by the project on an H100 80 GB, which is far beyond a typical laptop GPU. Smaller Qwen checkpoints may lower the entry cost, but model size can change decision quality. The repository claims support across sizes, not equal results across sizes.

When LitJev helps, and when it does not

LitJev makes sense when the method itself is part of the work. A team can inspect the prompts, readout code and distributions locally. It can develop a schema-compatible client without an API key, keep a prototype inside a controlled environment, or study how direct option scoring behaves on different Qwen checkpoints.

It may also suit an offline lab with suitable GPU capacity. For screenshot decisions, a supported Qwen vision checkpoint offers a path that keeps the image and inference on local infrastructure.

Hosted Jev is the more direct choice when the requirement is the official model rather than an open reproduction. That includes work depending on TypeSafe’s RLCD-trained behavior, its confidence outputs, managed availability or an official production path. Running a large checkpoint also transfers model downloads, GPU memory, serving and monitoring to the local operator. “No API key” does not mean “no operating cost.”

Calibration needs its own test plan in either case, but it is especially important here because LitJev labels its default probabilities as uncalibrated. Before attaching an automated action to a threshold, collect task-specific examples, compare predicted distributions with outcomes and choose thresholds from that evidence. A visually precise decimal can still be wrong.

My practical reading is narrow. LitJev is a clear, inspectable experiment around Jev’s public interface and a useful adapter for Qwen. It is not evidence that a stock checkpoint has acquired the behavior of TypeSafe’s private system. Keeping those two statements together makes the project more interesting, not less.

Sources

Oğuzhan Koçaklı

Oğuzhan Koçaklı writes and advises on AI engineering, agents, GenAI products, and applied ML. Daily digests and deep dives in EN + TR at oguzhan.co.

All posts

1 comment

Leave a Reply

Your email address will not be published. Required fields are marked *