Fundamentals9 min read

Writing Choice Questions That Hold Up

How to write options Jev can tell apart, why overlapping options quietly wreck your confidence values, and how to find and fix them.

What a Choice actually does

A Choice asks Jev to pick one option from a set you define, up to 255 of them. You get back the winning option, a probability for every option, and a confidence value derived from how concentrated those probabilities are.

Two properties of that answer drive everything in this guide:

  • The probabilities sum to 1 across your options. Jev does not score each option on its own. It divides one unit of probability between them.
  • Both the option names and their descriptions are sent to the model. The descriptions are not documentation for your teammates. They are the definition Jev judges against.

The first property is why a Choice always picks something, even when nothing fits. The second is why the words you put in criteria matter more than anything else in the request.

A running example

Say you are building an expense tool. Employees upload a receipt and type a line about it, and you want each expense in the right category. A first draft usually looks like this:

{
  "state": "Dinner with the client after the site visit in Denver.",
  "model": "jev-latest",
  "questions": {
    "category": {
      "type": "choice",
      "instructions": "Which expense category does this belong to?",
      "criteria": {
        "travel": "Travel expenses",
        "meals": "Food and drink",
        "entertainment": "Client entertainment",
        "transport": "Getting around",
        "software": "Software and tools",
        "subscriptions": "Recurring subscriptions"
      }
    }
  }
}

It looks reasonable. It has at least four overlaps in it, and they will cost you on every expense that touches them.

Why overlapping options are the main problem

Two options overlap when a real input fits both of them. In the draft above:

  • A client dinner is meals, entertainment, and, because it happened on a trip, arguably travel.
  • An Uber is transport, and also travel when it is on a trip.
  • A GitHub plan is software and also a subscription.

Because the probabilities have to sum to 1, Jev cannot say "this is fully meals and fully entertainment". It has to split the mass. An expense that any accountant would categorize instantly comes back looking something like this (illustrative values):

{
  "choice": "meals",
  "probabilities": {
    "meals": 0.46, "entertainment": 0.38, "travel": 0.12,
    "transport": 0.02, "software": 0.01, "subscriptions": 0.01
  },
  "confidence": 0.35
}

That low confidence is not the model being unsure about the expense. Confidence is computed from the distribution, and the distribution is split because your options told it to split. This causes four separate problems.

Confidence stops meaning anything. If you gate on confidence, and you should, every expense that lands in an overlap gets sent to review. Your review queue fills up with the easy cases, and the hard cases are hidden among them.

No threshold fixes it. Lowering the gate lets these through, but it also lets through the genuinely ambiguous inputs you wanted a human to look at. You cannot tell the two apart by confidence, because the overlap has made them look the same. The only fix is in the criteria.

The winner flips on near-identical inputs. When two options sit at 0.46 and 0.38, a small change in wording moves the top spot. "Dinner with the client" goes to meals, "client dinner" goes to entertainment. Your reports show two categories for the same kind of spend, and nobody can explain why.

Your data drifts silently. Nothing errors. Every answer is a valid option with a valid distribution. The damage shows up months later as a category total that does not match what finance expected.

The three shapes of overlap

Most overlaps fall into one of three patterns, and each has a different fix.

Synonyms

Two options that mean nearly the same thing: software and subscriptions, billing and payments, bug and error. Nobody intended the distinction, it just crept in as the list grew.

Fix: merge them. If the distinction genuinely matters to you, write down the rule that separates them and put it in both descriptions. If you cannot write the rule, the distinction does not exist, and Jev will not find it either.

Nesting

One option is a subset of another: transport inside travel, airfare inside transport. Every input that fits the child also fits the parent, so the two always split the mass.

Fix: pick one level of granularity for the whole set. If you need both levels, ask two Choices in the same request, one broad and one specific. They evaluate in parallel against the same state, so the second one costs you almost nothing.

Mixed axes

The subtlest one. The options answer different questions. travel describes when an expense happened. meals and transport describe what was bought. A dinner on a trip is legitimately both, because it sits on two axes at once.

Fix: split the axes into separate questions. One Choice for what was bought, and one Noul for whether it happened on a business trip:

{
  "category": {
    "type": "choice",
    "instructions": "What was purchased in this expense?",
    "criteria": {
      "meals": "Food or drink, for the employee alone",
      "client_entertainment": "Food, drink, or events where a client or prospect was hosted",
      "transport": "Any ride, flight, train, rental car, fuel, or parking",
      "lodging": "Hotels and other overnight stays",
      "software": "Software, SaaS tools, and their subscriptions",
      "other": "Anything that does not fit the categories above"
    }
  },
  "on_business_trip": {
    "type": "noul",
    "instructions": "The expense was incurred while travelling away from the employee's home office"
  }
}

Your code now knows both facts, and "travel spend" becomes a query over on_business_trip rather than a category that competes with everything else.

A variant of this is multi-label inputs. If an input can genuinely belong to several options at once, and that is the correct answer rather than an ambiguity, a Choice is the wrong primitive. It will always force a split. Ask one Noul per option instead. Each Noul is judged on its own, with no sum-to-1 constraint, so an input can score high on all of them.

Writing options Jev can tell apart

With overlap handled, the rest is craft.

Describe the boundary, not the center. "Food and drink" defines meals in isolation. "Food or drink, for the employee alone" tells Jev where meals stops and client entertainment starts. Write each description with its nearest neighbour in mind, and say which side of the line the tricky cases fall on.

Put tie-break rules in the descriptions. When two options genuinely touch, say who wins. "Transport: any ride, flight, or train, including those taken during a business trip". The model reads the description, so the rule has to be there, not in a comment in your code.

Name options for meaning. Option names are sent to the model along with the descriptions. cat_3 wastes that signal. client_entertainment reinforces the description.

Keep options at the same granularity. A set with meals, transport, and uber_rides invites the nesting problem. If one option is much narrower than the others, it is probably a child of one of them.

Add a catch-all. If your inputs can fall outside your options, and real inputs almost always can, add an other and give it a route in your code. Without it, the mass goes to whichever option is least wrong, often with enough confidence to pass your gate. Toggle the catch-all below and watch where the probability goes.

Choice · 3 options

“Which team should handle this?”

department

billing
0.16
technical
0.38
sales
0.46

Response

"department": { "type": "choice", "choice": "sales", "confidence": 0.19, "probabilities": { "billing": 0.16, "technical": 0.38, "sales": 0.46 } }

Your code

match answers["department"].choice:
case "billing": route_to(billing_queue)
case "technical": route_to(engineering)
case "sales": route_to(sales_team)

Try the job question with and without other. Without it, Jev still has to pick one of your three teams, so it lands on sales with a low confidence. The distribution always sums to 1 across the options you defined, so a missing option does not make the answer go away. It pushes it somewhere wrong.

Write the full question in instructions. Question keys are not sent to the model, so "category" tells Jev nothing. "What was purchased in this expense?" does.

Do not encode an order. If your options are low, medium, and high, you want a Score, not a Choice. A Choice treats them as unrelated labels. It has no idea that confusing low with high is worse than confusing low with medium.

Finding overlaps in production

Overlaps you did not anticipate will show up in your logs, as long as you log the full probabilities object and not just the winning choice. Look for pairs of options that keep sharing the top two spots with a small gap between them:

from collections import Counter
 
confusions = Counter()
 
for probabilities in logged_answers:  # the probabilities dict from each Choice answer
    first, second = sorted(probabilities, key=probabilities.get, reverse=True)[:2]
    if probabilities[first] - probabilities[second] < 0.3:
        confusions[tuple(sorted((first, second)))] += 1
 
for pair, count in confusions.most_common(5):
    print(pair, count)

A pair that appears once in a while is normal ambiguity. A pair that dominates the list is a boundary your descriptions have not drawn. Read twenty of those inputs yourself. If you can categorize them instantly, the problem is the criteria, not the model.

When you want the overlap on purpose

Sometimes you really do want a fine-grained Choice and a coarse decision from it. Because the probabilities in one Choice answer sum to 1, you can add sibling options together in code:

p = response.answers["category"].probabilities
 
food_and_drink = p["meals"] + p["client_entertainment"]
 
if food_and_drink > 0.8:
    apply_meal_policy(expense_id)

That gives you a clean, confident signal for the broad group even when Jev is split on which of the two it is. Do this with options from the same Choice only. Probabilities from different questions come from different distributions, and adding them means nothing.

A checklist before you ship a Choice

  • Could a real input fit two of these options? If so, merge them, split the axes, or write the tie-break into the descriptions.
  • Does each description say where it ends and its nearest neighbour begins?
  • Are all the options at the same level of detail?
  • Is there an other, and does your code route it somewhere?
  • Is there an order hiding in the options? Use a Score.
  • Can an input legitimately be several of these at once? Use one Noul per option.
  • Are you logging the full distribution, so the next overlap shows up in your data before it shows up in a report?

Are you ready to become a Jev expert?

All interactive courses, 9 mini-projects, the playground, and 1,000 live Jev credits. $19.99 once, with updates as Jev ships new versions.