Build your own decision model
Nish Tahir's tutorial, published 10 October 2026, uses the open Qwen3-1.7B model to answer multiple-choice questions. It masks every vocabulary token except the five options, A through E, so the model answers in one forward pass. On a holdout of 1,221 CommonsenseQA questions, the base model answered 725 correctly; after finetuning, 762. A GitHub repo includes scripts for building the dataset, evaluation, finetuning and calibration. The project copies commercial "System one" models. Jev, Clef and Laya are examples of this growing class of small, low-cost classifiers built for agents. Normal generation needs one pass for each output token. Tahir points out that raw option probabilities likely measure next-token confidence, not whether the answer is correct. The raw model was 98.6% confident on answers it got right only 70% of the time. Fitting a single temperature of about 3.80 fixed this. For AI companies, this means agent steps like routing, triage and approval checks can run on a 1.7B open model for the cost of one pass. Confidence scores still need calibrating before they trigger any automated action.