Engineering7 min read

Option Order Changes Jev's Answers. Here's How to Average It Out

Reordering a Choice's options can move its probabilities by 20 points or more and flip some answers. What the community measured, when it matters, and a cheap fix that runs in one request.

The finding

A Choice question sends its options as a map. Maps have an order, and Jev reads it. Put the same options in a different order and you can get a different probability distribution back, and sometimes a different answer.

This was one of the first things people measured after launch, and the results line up across several independent tests of jev-1.13.0:

  • Probabilities move, even when the answer doesn't. In finnhll/jev-eval, reordering the criteria map on an ambiguous input moved the winning probability from 0.62 to 0.48 and confidence from 0.43 to 0.22. The chosen option stayed the same. The number a confidence gate reads did not.
  • Some answers flip. Archestra tested 100 real agent tool calls from Claude Code sessions. Reordering options never flipped a two-way label, but on three-way labels, shifting or reversing the order changed the outcome on 5 to 7 calls per 100.
  • With no right answer, the first option wins. In pobooo/jev-dice, asking Jev to "roll a die" by picking a face returned the first-listed option on all 2,000 identical requests. When one option was clearly correct, order had little effect.
  • Order can look like something else. pawarbi/jev-bias-audit asked which of two candidates would make a better president, a man or a woman, and got "man" at 67%. Most of that turned out to be position: whichever option was listed first gained about 0.37. Without checking the reversed order, a position effect reads like a demographic bias.

The pattern is consistent. When the state clearly supports one option, order barely matters. When the input is ambiguous, or when no option is really correct, order fills the gap.

Why this matters more than it sounds

Most integrations don't read choice. They read confidence or probabilities and compare them against a threshold. That's the whole point of a calibrated model, and it's what we recommend in Confidence and Calibration.

Option order moves exactly those numbers, on exactly the inputs where your threshold is doing its work. A clear case at 0.95 stays above your threshold in any order. A borderline case at 0.62 can drop to 0.48 depending on how you happened to write your options, and the decision whether to auto-approve or escalate changes with it.

It also means two teams asking the "same" question can get different calibration curves, and a refactor that reorders a dictionary can silently change your escalation rate.

The fix: ask in more than one order and average

Treat option order as noise and average it out. Ask the same Choice with its options in two or more orders, line the probabilities up by option name, and average them. This is what the pijev package does, and several projects have added the same thing. Averaging has a useful guarantee: the averaged distribution's Brier score and log loss are never worse than the average of the individual orderings. You don't have to know which order is "right".

The key detail is that this costs very little with Jev. All the orderings go in one request, as separate questions against the same state:

  • The state is sent and billed once, no matter how many questions ride along.
  • Questions are evaluated in parallel, so latency barely moves.
  • Question keys are never sent to the model, so you can name the copies anything.

The only extra cost is the tokens for the repeated question text.

from typesafe_sdk import Choice, TypeSafeClient
 
def orderings(criteria: dict, k: int = 2) -> list[dict]:
    """Original order, reversed, then cyclic shifts up to k orderings."""
    items = list(criteria.items())
    orders = [items, items[::-1]]
    for shift in range(1, len(items)):
        if len(orders) >= k:
            break
        orders.append(items[shift:] + items[:shift])
    return [dict(o) for o in orders[:k]]
 
def order_averaged_choice(client, state, instructions, criteria, k=2):
    questions = {
        f"q_{i}": Choice(instructions=instructions, criteria=c)
        for i, c in enumerate(orderings(criteria, k))
    }
    response = client.system_one(state=state, questions=questions)
 
    averaged = {
        option: sum(response.answers[key].probabilities[option] for key in questions) / len(questions)
        for option in criteria
    }
    choice = max(averaged, key=averaged.get)
    return choice, averaged
 
with TypeSafeClient() as client:
    choice, probs = order_averaged_choice(
        client,
        state="I was charged twice for the same order and the app keeps crashing when I open receipts.",
        instructions="Which team should handle this?",
        criteria={
            "billing": "Charges, refunds, invoices, payment problems",
            "technical": "Crashes, errors, features not working",
            "shipping": "Delivery status, delays, lost packages",
        },
        k=3,
    )

A few practical notes:

  • Two orders get you most of the way. Original plus reversed cancels the strongest effect, which is first position. Add cyclic shifts if you have many options or the decision is expensive to get wrong.
  • Recompute confidence from the averaged probabilities. Don't average the per-order confidence fields. Compute your gate from the averaged distribution, either the peak or the margin between first and second place, and set thresholds against averaged numbers on your own labelled data.
  • Nouls have no order. A Noul is a single proposition, so there's nothing to reorder. Averaging only applies to Choice (and to Score, if you want to reverse a rubric, but then you have to map the levels back).
  • Keep the question count in mind. Three orderings of five Choices is fifteen questions. That's fine for Jev, and it doesn't touch the 32k limit, which only counts the state plus the single longest question. It does count toward the 64k total per request.

Cheaper habits that also help

Averaging treats the symptom. These reduce how much there is to average:

  • Give ambiguity somewhere to go. When pawarbi's audit added a "can't tell" option, Jev picked it on every one of 72 ambiguous calls instead of falling back on position. finnhll's tests found that without an "other" option, two of three out-of-taxonomy inputs came back with confidence of 0.93 or higher on a wrong answer. An explicit fallback option is the single most useful thing you can add to a Choice.
  • Write options that separate cleanly. Order effects show up when the state doesn't clearly favour one option. Descriptions that say what distinguishes each option from the others, as covered in Writing Choice Questions, leave less room for position to decide.
  • Keep option lists short and relevant. 123Satyajeet123/jev-wide found that adding unrelated candidates shifts the balance between the options you care about. If you have 40 categories, a first Choice over groups followed by a narrower one often beats one big list.
  • Fix your order in code. Even if you don't average, make the order deterministic, for example sorted by option name. Then a refactor that rebuilds a dictionary can't quietly change your production behaviour.

Test it on your own questions

Order sensitivity depends on the task, so measure it on yours. Take a few hundred labelled examples, send each Choice in the original and reversed order, and look at two numbers: how often the chosen option flips, and how far the top probability moves on the cases near your threshold. If flips are rare and the shift is small, a fixed order is fine. If not, average.

This is also worth rerunning when you move off jev-1.13.0. Position effects are the kind of thing a new model version can improve or make worse, and your thresholds were tuned against the old behaviour.

Are you ready to become a Jev expert?

All interactive courses, 8 mini-projects, the playground, and 1,000 live Jev credits. $19.99 once, with updates as Jev ships new versions.