How to set Jev confidence thresholds with your own data
Evaluate Jev confidence using a runnable Python script. Measure accepted errors, coverage and fallback before choosing a threshold for your application.
Set a Jev confidence threshold by measuring errors among accepted decisions and the share of traffic accepted on labeled data. Do not choose 0.8 or 0.9 because it looks reassuring. A threshold only becomes meaningful when the question, model version, input population and fallback behavior are defined.
This tutorial gives you an offline evaluation script, a synthetic fixture with known answers, and a path from existing Jev logs to a held-out test. It is designed for a Choice classification question, such as routing requests among support queues. It does not make live API calls, measure Jev accuracy, or validate automatic high-impact decisions. Start with the Jev API guide if you have not yet collected responses.
Understand which number you are thresholding
The official API distinguishes an answer’s probabilities from its confidence. For a Choice with n options, the documented confidence formula is (n × p_max − 1) / (n − 1). With three options and a largest probability of 0.8, confidence is 0.7. A threshold on one field is therefore not interchangeable with the same threshold on the other. TypeSafe’s confidence reference.

Original documentation screenshot captured October 2, 2026. This is the field definition, not evidence that a threshold works on your data.
Score uses a different calculation that accounts for distances between ordered levels. Noul returns a yes-probability without a separate confidence field. Do not flatten all three types into a shared column called “certainty” and apply one cutoff. This tutorial’s CSV expects the returned Choice confidence for one fixed question.
Two concepts then need separate measurement. Calibration asks whether predicted probabilities correspond to observed outcomes across groups of cases. Selective evaluation asks how often your accepted decisions are wrong at a chosen cutoff. A selective error rate can look good even when a probability estimate is poorly calibrated, particularly when almost all traffic falls back.
Define acceptance before looking at the results
Suppose your application routes a ticket to billing, technical or other. A correct classification is agreement with an independently assigned label. If a ticket is incomplete, contradictory or outside those categories, the gold-label process must define what should happen. Otherwise you will reward the model for guessing an answer the reviewers cannot justify.
Write down a maximum acceptable error rate, a minimum amount of accepted traffic, a latency budget and the errors that require separate review. These are application decisions, not numbers supplied by Jev. A reversible queue assignment and an irreversible action should not inherit the same acceptance rule merely because both return a label.
Use a label guide with examples and have reviewers resolve disagreements before evaluating model outputs. Where possible, keep reviewers blind to the prediction. If they see the model’s answer first, the resulting “ground truth” can quietly become agreement with the model rather than independent evidence.
Separate development, validation and final testing
Use development examples to improve the question and category definitions. Use a validation split to compare candidate thresholds. Lock both the question and threshold before running the final held-out test. Repeatedly adjusting a cutoff after seeing the final test results turns that set into another validation set.
Split by conversation, customer document or template family when related rows could leak across sets. Include normal traffic and important boundary cases, but report their results separately if the boundary cases are deliberately oversampled. Record language and input-length groups so that a good average cannot hide an unsupported subgroup.
There is no universal sample size that makes deployment safe. In particular, zero errors among a handful of accepted cases proves little. As a rough statistical illustration, with independent trials and zero observed errors, the “rule of three” gives an approximate 95% upper error bound of 3 / accepted_count. It is a planning approximation, not a certificate: correlated cases, selection bias and distribution shifts can invalidate the interpretation. For formal inference, choose a suitable interval and sampling design with your evaluation owner.
Prepare a CSV from a fixed question
Download and unzip the Jev decision kit. You need Python 3; the scripts use only its standard library. No key or network access is required.
The supplied file has five columns:
| Column | Meaning | Validation rule |
|---|---|---|
id | Unique example identifier | Required and unique |
gold | Independently assigned category | Required, including requests that failed |
predicted | Returned Choice label | Required when status is ok |
confidence | Returned Choice confidence | Finite number from 0 to 1 when status is ok |
status | ok or error | Errors remain in total traffic and always fall back |
Keep the actual response model, request ID, full probability map, question version and source-data split in a separate trace manifest linked by id. The small evaluator does not enforce those metadata fields. Do not merge several model versions or questions into one CSV and then describe the result as a single validated policy.
Here is part of the supplied hand-written synthetic fixture, not a Jev response:
id,gold,predicted,confidence,status
case-1,billing,billing,0.98,ok
case-2,technical,technical,0.94,ok
case-3,billing,technical,0.91,ok
For your own file, copy fields from saved responses without rounding before thresholding. Retain failed requests as error rows with empty prediction and confidence. Dropping timeouts artificially increases apparent coverage.
Run the evaluator and verify its arithmetic
From the extracted directory, run:
python3 evaluate.py synthetic.csv
python3 evaluate.py synthetic.csv --thresholds 0.8 0.9
We executed the first command locally against the eight supplied synthetic cases. The output is a test of the calculator’s logic, not measured Jev performance:
| Threshold | Accepted / all cases | Accepted errors | Coverage | Error rate among accepted |
|---|---|---|---|---|
| 0.50 | 6 / 8 | 2 | 75% | 33.33% |
| 0.80 | 5 / 8 | 1 | 62.5% | 20% |
| 0.90 | 3 / 8 | 1 | 37.5% | 33.33% |
| 0.95 | 1 / 8 | 0 | 12.5% | 0% |
Coverage is accepted / all eligible cases; accepted error rate is wrong accepted / accepted. When no cases are accepted, the script returns null for error rate rather than pretending the policy achieved perfect accuracy.
Notice that moving from 0.80 to 0.90 makes the observed error rate worse in this fixture. A higher threshold does not guarantee a monotonic improvement on a finite sample. The 0.95 row has no errors but accepts only one case, which is insufficient evidence for reliability. These deliberate examples help catch two common misreadings before you use real data.
Select a policy, then test the whole pipeline
On validation data, inspect the candidate rows and the actual accepted errors. Eliminate policies that violate your error constraints. Among the remaining policies, assess coverage, fallback capacity, cost and latency. If none satisfy the requirements, the result is “do not automate this question yet,” not “choose the least bad number.”
Then freeze the selected threshold and evaluate on the held-out data once. Also evaluate the fallback: a low-confidence decision sent to another model or a human is not automatically a successful result. Score the complete workflow, including unavailable handlers and cases that never finish.
A conservative policy skeleton is:
if response_error or label_not_allowed or confidence_invalid:
fallback()
elif confidence < validated_threshold:
fallback()
else:
enforce_permissions_and_business_rules()
route_to_handler()
This is application pseudocode, not an executable Jev SDK example. Reject missing, nonfinite or out-of-range confidence before comparing it: a NaN value must not slip through a numeric cutoff. Authorization, input validation and business rules remain in your code. A high confidence score must not bypass them. When a routed action has side effects, use the same idempotency and confirmation controls that your ordinary workflow requires.
Diagnose the failures instead of hiding them
| Observation | What to inspect next |
|---|---|
| CSV rejected | Duplicate IDs, missing labels, nonfinite values or unsupported status strings; fix the export before calculating metrics |
| High-confidence wrong labels | Ambiguous category meanings, contradictory option names, missing state or out-of-scope inputs |
| Good overall result, bad language subgroup | Separate language labels and validation; do not assume English performance transfers |
| Very low coverage | Question granularity, incomplete context and actual uncertainty; do not lower the threshold without measuring new errors |
| Model routing succeeds but task fails | Handler execution, arguments, tool errors and fallback outcome |
| Results change after an update | Model response version, question hash, preprocessing, option set and traffic mix |
For the broader interpretation of published results, see what Jev benchmarks can establish. That article explains why benchmark calibration cannot approve your particular production cutoff.
Keep the threshold tied to a version
Pin the evaluated model version when reproducibility matters, and log the version reported by the response. Re-evaluate when you change a question, add an option, translate the state, alter preprocessing or move to another model version. Even an unchanged threshold can become a different policy if its inputs change.
Begin with shadow decisions, then a limited reversible rollout with an explicit rollback owner. Monitor accepted errors using newly labeled samples, total coverage, fallback load and complete-task outcomes. Save the previous policy so you can return to it without reconstructing the experiment.
Finally, feed measured coverage into the routing cost calculator. A cutoff is useful when it satisfies the task’s quality rule and leaves a workable application—not when it produces an impressive-looking decimal.
Frequently Asked Questions
- Does Jev confidence 0.9 mean the answer is 90% correct?
- No. API confidence summarizes the answer distribution. You must measure correctness against independent labels for your question, population and model version.
- Does the included script call Jev or cost money?
- No. It reads a local CSV using Python's standard library. Its supplied fixture is hand-written synthetic test data, not Jev output.


