Most LLM calls in production are not conversations. They are routing decisions: which support queue does this ticket belong in, is this comment abusive, which intent did the user just express. A new open-source project called Jeff targets exactly that workload with fine-tunes of Qwen3.5 (0.8B and 2B) and Gemma 4 E2B that return a calibrated probability for each option from a single forward pass. No generated text, nothing to parse: about 22 ms per decision on an RTX PRO 6000 and 28 ms on an Apple M4 Max running MLX.
You describe a situation and list the options in plain words. A refund request gets routed by listing the teams and letting the model pick; a support snippet gets an anger score. The request format matches Jev, a larger commercial decision model from TypeSafe, so code written against Jev’s API can point at a local Jeff server instead. Three question types are supported: choice with up to 255 options, yes/no returned as a probability, and a score on a scale you describe. Several independent questions can ride in one request and get answered together.
Zero-shot means your categories were never in the training data
The pitch is that options can be anything: support queues, user intents, moderation labels, voice commands, game moves. Your categories do not need to appear in the training set because you describe them in the request and the model picks. That is what makes it a decision model rather than a classifier, where every label change means retraining.
If zero-shot accuracy is not good enough, a short fine-tune on your own examples moves things dramatically. The project’s voice-navigation fine-tune took held-out accuracy from 31.7% to 95.8% in under half an hour on one GPU. Calibration is handled the boring way: cross-entropy over the option letters during training, then one fitted temperature at the end, with checkpoints selected on a development set rather than on the benchmark panel.
Benchmarks: close to the big model on the useful tasks
Across 4,599 questions from five public benchmarks plus JevBench’s public hard tier, Jeff-Qwen3.5-2B scores 83.1 overall against Jev’s published 83.0. The 0.8B lands at 79.1 and the Gemma 4 E2B fine-tune at 81.6. Financial PhraseBank is effectively solved at 96 or higher across all three. BBH lags far behind the big model, 64 to 68 versus Jev’s 94.3, which the authors are upfront about: at 0.8B to 2B parameters these models make fast, calibrated judgement calls between options you describe. They do not reason through multi-step problems, and no training changes that at this size.
Two fine points from the caveat section deserve attention before anyone wires this in. Benchmark scores did not predict game play in their tests: the untrained Gemma 4 E2B beat the untrained Qwen models on benchmarks yet played their test games worst, right often but not reliably. And prompt phrasing matters more than expected; Jev’s own Doom prompt, a raw bearing number plus an aiming rule, failed for every Jeff model, while options that stated consequences in words worked. It is also English and text only.
Getting started is a three-command affair
The quick start is deliberately boring. Clone, run uv sync, download a checkpoint from Hugging Face with one command, and start the server with a checkpoint path and a port. On Apple silicon there is an MLX backend behind an extra install flag, and it is the faster option on a Mac. The server exposes a single endpoint that takes a state description and a question map, and returns a probability per option plus the chosen answer and a confidence figure. That shape makes it easy to drop behind an existing feature flag and A/B against whatever heuristic or API call you use today.
Trained entirely on local hardware
The whole pipeline ran on a workstation. The 0.8B trains in about two hours on one RTX PRO 6000, the 2B in about three and a half, and all synthetic training data was written by an open model, Qwen3.8-Flash-Next, on two DGX Sparks. No cloud GPUs, and no closed-model output in the training data; a closed model was used only to spot-check the quality of a sample of the synthetic data. Code is MIT, weights are Apache 2.0, and the project is a fork of Denis Yarats’ AutoJev recipe, explicitly unaffiliated with TypeSafe. It has pulled 503 stars in the roughly two days since release, which for a niche inference project is fast.
Where this lands
The interesting claim is economic. If a 0.8B model handles 95% of your routing calls at 22 ms on hardware you already own, the large model only gets invoked when the small one is unsure or the task genuinely needs reasoning. That is the speculative-decoding pattern applied to classification instead of text generation, and it is the same instinct behind the wave of small local models we have covered before. For teams paying per-call API prices for high-volume classification today, downloading three checkpoints and running the quick-start server is a cheap experiment with a measurable answer within an afternoon: point your existing traffic labels at it, compare against your current setup, and keep whatever wins.