Jev decision kit — Ofox, 2026-10-02
Python 3, standard library only. No API calls or credentials.
python3 evaluate.py synthetic.csv
python3 cost.py
synthetic.csv is HAND-WRITTEN TEST DATA, NOT A JEV RESPONSE OR PERFORMANCE RESULT.
For real evaluation use human-labeled, grouped validation/test splits and record the actual model and question version separately. Do not tune thresholds on your final test set.
Default TypeSafe direct input rate is a dated 2026-10-02 snapshot. Downstream costs are hypothetical, not Ofox prices. Replace all rates and traffic assumptions with your own billing and traces.

CSV header must exactly match: id,gold,predicted,confidence,status
Use one fixed Choice question and model version per file; record the allowed category set externally. The tool checks schema and values, not your label ontology or dataset split. IDs must be unique, fields must have no surrounding whitespace, and error rows must have empty prediction/confidence. UTF-8 with or without a BOM is supported.
The evaluator rejects empty data and malformed CSV rather than reporting zero risk. No accepted cases means error_rate is null. Compare subgroup files independently; this small tool does not compute statistical intervals or evaluate fallback quality.
The cost formula uses a common fallback cost for both the direct and fallback branch. If fallback traffic is systematically longer, or a cheap branch later escalates, replace costs with the appropriate branch averages or extend the model. break_even_acceptance above 1 means no feasible acceptance share yields savings under these assumptions; null means strong <= cheap. Quality and cost per successful completion require separate outcome records.
Local verification covers the supplied arithmetic, boundary thresholds, malformed files and invalid cost arguments; it does not verify Jev model performance.
