Engineering7 min read

Open-Weight Jev Alternatives: The First Two Weeks

Within three days of Jev's launch, open models speaking the same /v1/systemone API appeared on GitHub and Hugging Face. What they are, how they differ, and when self-hosting one makes sense.

A snapshot, not a verdict

Jev launched on September 15, 2026. By September 18, at least five open-weight models had been released that answer the same Choice, Score and Noul questions over the same wire format. Every one of them is less than two weeks old as of this writing.

This post describes that ecosystem as it stands at the end of September. Expect it to go stale quickly. Projects will be renamed, abandoned or retrained, and the benchmark numbers below will move. Treat it as a map of what exists, not a ranking.

One note on what these are not. They are not general-purpose open models like Qwen, Gemma or Kimi being prompted to act like Jev. Each one is a model trained specifically to produce typed decisions, usually by attaching a scoring head to an existing base model and fine-tuning it. The base models are old. The decision training and the Jev-compatible server on top are new.

The part that makes switching cheap

Nearly every project here implements TypeSafe's POST /v1/systemone request and response shape. That means the official SDK works unchanged: you point it at your own server by setting TYPESAFE_BASE_URL, and the rest of your code does not know the difference.

This is the practical payoff of keeping your integration thin, which we recommended in What Jev Actually Costs. If Jev reprices, changes its rate limits, or you need data to stay on your own hardware, the fallback now has a concrete shape. Trying one against a sample of your traffic is an afternoon of work, not a migration.

Three architectures

The projects fall into three groups, and the group tells you more about how a model will behave than its name does.

Encoders (Laya, Von). A bidirectional encoder like ModernBERT reads the state and the options, and the answer is read from a softmax over option positions. These models are small (roughly 300-400M parameters), run on CPU, and are very fast. The tradeoff is knowledge: they know less about the world than a multi-billion-parameter decoder, and it shows on questions with many options or questions that need facts not in the state.

Fine-tuned decoders (Kev, Decider). An existing Qwen3.5 base model gets a LoRA adapter and a scoring head, then is trained on decision data. They come in ladders from under 1B to 27-35B parameters. The larger sizes get closest to Jev's accuracy, and they need a real GPU to do it.

Diffusion (OpenJev). A discrete diffusion model masks one answer slot per question and reads the probability distribution straight from a forward pass. This is the closest in spirit to how Jev describes itself: answers come from probabilities over a fixed output space, so they cannot go off-schema.

The models

Model Builder Base Size License Created
Laya Convai Innovations ModernBERT-large / mmBERT 421M / 322M Apache 2.0 Sep 18
Kev Jared Palmer Qwen3.5 0.8B-27B Apache 2.0 Sep 17
Decider Mapika Qwen3.5 0.8B-35B MoE Apache 2.0 Sep 16
Von wfzyx ModernBERT-large 395M Apache 2.0 Sep 18
OpenJev razorback16 DiffusionGemma 26B-A4B 26B (4B active) Apache 2.0 Sep 18

All five say they implement /v1/systemone.

Laya

The most popular so far, at roughly 28,000 GitHub stars. It ships an English model on ModernBERT-large and a multilingual one on mmBERT, plus a laya-serve HTTP server. Recent releases added support for long documents of up to 8,192 tokens.

Its model card reports two results that are worth reading together. On a 2,000-item typed-decisions set, Laya scores 0.766 against Jev's 0.727, at about 33ms per question against Jev's 236-276ms. On Banking77, which has 77 possible labels, Laya scores 0.425 against Jev's 0.870. That is the encoder tradeoff in two numbers: fast and competitive on small option sets, much weaker when a Choice has dozens of options.

Kev

Kev comes in four sizes (0.8B, 4B, 9B and 27B), built from LoRA adapters on Qwen base models. The project recommends 4B as the starting point, and 0.8B runs on a laptop. The README says it was trained on public and synthetic data, not on Jev outputs, and it supports fine-tuning on your own labelled examples.

On the project's own held-out set, Kev-27B scores 0.848 against Jev's 0.857. Knowledge-heavy questions are where it falls behind: on MMLU, Kev-9B scores 0.74 against Jev's 0.90. The 27B model needs an 80GB GPU.

Decider

Also on Qwen3.5, from 0.8B up to a 35B mixture-of-experts model with 3B active parameters, plus a 2B vision variant. The project calls itself an independent reproduction and says it was not distilled from Jev. Its labels came from a local Qwen3.5-27B teacher. It reports competitive results on language and tool-use questions and weaker results on knowledge questions (0.51 against Jev's 0.69).

Von

A single 395M ModernBERT model trained on about 63,000 operational decision items. It masks attention between options so that reordering the options cannot change the answer, which is a property worth having. It ships von-sdk for Python and TypeScript.

Von's README is candid about the tradeoff. On JevBench it roughly halves Jev's latency and comes in at a fraction of the cost, with calibration close to Jev's (75.7 against 76.3). On the hard tier, it scores 37.3% against Jev's 74.1%.

OpenJev

The diffusion option. It runs DiffusionGemma with 26B total and 4B active parameters, and serves both /v1/systemone and an OpenAI-compatible /v1/chat/completions from the same weights. The project reports about 31ms for a three-question request on an RTX PRO 6000 at a single request, and about 57 requests per second at 64 concurrent requests.

Read the benchmarks carefully

Every number above comes from the project's own README or model card. None of it has been independently reproduced yet. There are also two reasons the numbers do not compare with each other:

  • Different benchmarks. Laya, Kev, Decider and Von each report against different evaluation sets. A 0.848 on one and a 0.766 on another tell you nothing about which model is better.
  • Different Jev scores. Jev itself scores 0.727 on Laya's set, 0.857 on Kev's, and 74.1% on Von's hard tier. The comparison that matters is each model against Jev on the same set, and those gaps are what we quoted.

The pattern across all of them is consistent, though. Open models are close to Jev on straightforward classification with few options, and fall behind on questions with many labels and on questions that need world knowledge. That matches what you would expect from smaller models, and it matches the calibration advice in Confidence and Calibration: whatever you run, measure it on your own data and set thresholds from that.

When to self-host

Self-hosting one of these is worth it when:

  • Data can't leave your infrastructure. This is the strongest reason, and the one hosted Jev cannot address.
  • Your questions are narrow. A Noul or a Choice over five options is where the small encoders hold up. Test that on your own traffic before trusting it.
  • You need a hedge. Having a Jev-compatible fallback that you have already run against a sample of real traffic turns a pricing or availability change into a config change.
  • You want to fine-tune. Kev and Decider both support training on your own labelled examples.

Stay on hosted Jev when:

  • Your Choices have many options, or answers depend on general knowledge rather than what's in the state.
  • You don't want to run GPUs. Jev's pricing is hard to beat once you count the hardware for a 27B model, and the encoder models that do run cheaply on CPU are the least accurate.
  • You need stability. These projects are days old. Pinning jev-1.13.0 gives you fixed behavior. Pinning a commit of a two-week-old repo gives you a snapshot nobody else is maintaining.

A reasonable middle path is to keep Jev as the default and run one open model in shadow on a sample of requests. Log both answers, compare them, and you will know whether the open model is good enough for your task before you depend on it.

Are you ready to become a Jev expert?

All interactive courses, 8 mini-projects, the playground, and 1,000 live Jev credits. $19.99 once, with updates as Jev ships new versions.