AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

The Jeff project announced v1.1 on September 29, saying its Qwen3.5-based 0.8B and 2B models can now choose among up to 254 options. The project reports fast local inference and stronger scores on some classification benchmarks, while its results remain weaker than larger models on reasoning-heavy tests.

The independent Jeff project released version 1.1 on September 29, saying its Qwen3.5-based decision models can now select among up to 254 options, compared with 26 in version 1.0. The project describes Jeff as a small model for returning probabilities over choices in a single pass, with reported response times of 22 milliseconds on an RTX PRO 6000 and 28 milliseconds on an Apple M4 Max.

Jeff is designed for developers to give a situation and describe possible answers in ordinary language. The model returns a probability for each option, the selected option and a confidence measure. Its request format matches Jev, a separate product made by TypeSafe; Jeff’s makers say the project is not affiliated with or endorsed by TypeSafe. The software also supports yes-or-no questions and numeric ratings, and can answer several independent questions in one request.

The September 29 release note says both Jeff-Qwen3.5-0.8B and Jeff-Qwen3.5-2B gained the expanded choice limit and improved calibration. On the project’s long-list test, the 0.8B model’s score rose from 40% to 95%. The 2B model’s overall benchmark score moved from 83.1% in version 1.0 to 82.0% in version 1.1. The project says version 1.0 remains available on Hugging Face. Its Gemma 4 model is listed as version 1.0 and supports up to 26 choices.

For the updated Qwen models, Jeff reports an overall score of 79.1% for the 0.8B model and 82.0% for the 2B model across five public benchmarks. Those figures cover 4,599 questions. The project’s table lists higher scores than published Jev figures on several classification and grounding tasks, including Financial PhraseBank and RAGTruth. However, it says the published Jev and AutoJev results were measured on different samples of the same benchmarks, so the comparison is not a controlled head-to-head evaluation.

At a glance
announcementWhen: Version 1.1 announced September 29, 2026
The developmentJeff released v1.1 of its small, locally runnable decision models, expanding the Qwen models’ choice limit from 26 to 254 options.

Where Small Decision Models Fit

Jeff’s stated use is to handle bounded choices inside local software: routing support requests, identifying user intents, applying moderation labels, interpreting voice commands or selecting a game move. The project says categories can be described at request time, even if they were not represented in training. If the model performs well enough for a particular task, a developer may be able to avoid generating free-form text and parsing it afterward.

The reported latency and training setup may interest teams that want to run classification on their own hardware. The project says the 0.8B model trained in about two hours and the 2B model in about three and a half hours on one RTX PRO 6000 workstation GPU. Its authors also report a voice-navigation fine-tune that raised held-out accuracy from 31.7% to 95.8% in under half an hour on one GPU. These are project-reported results; performance on another dataset or device may differ.

The trade-off is capability. Jeff’s authors say its small models can make fast judgments but do not match the reasoning ability of Jev, which runs on a much larger model. In the benchmark table, Jeff’s scores on reasoning-heavy BBH, JudgeBench and JevBench hard are below Jev’s published figures. That distinction matters for developers deciding whether a quick classifier is appropriate or whether a task needs a more capable system.

How Jeff Was Built and Tested

The project says the models are fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification, and that its training code starts from the open-source AutoJev recipe. It reports that training took place on local hardware: one RTX PRO 6000 workstation GPU, with synthetic training data written by an open model on two DGX Sparks. The authors say a closed model was used only to spot-check a sample of that synthetic data, not to produce training data. Testing was done on a MacBook.

Jeff’s published evaluation includes five public benchmarks and a separate 105-item JevBench hard tier. The project’s authors caution that games are an imperfect zero-shot test because game states differ from typical unstructured data. In its described setup, the code presents the state and legal moves in words, including what each move would lead to, but does not identify the correct move. The project says each game result covers 20 episodes with seed 1234; the supplied report excerpt does not include the full results table.

Although Jeff uses Jev’s request format, that compatibility describes an interface, not a shared product or endorsement. The project presents Jeff as an independent implementation and makes model checkpoints available on Hugging Face. Its serving instructions include CPU and NVIDIA GPU options, plus an Apple silicon backend for Qwen models.

“Jeff returns a calibrated probability for each option from a single forward pass.”

— The Jeff project, describing its response format

Limits of the Reported Results

The latency figures, benchmark scores and fine-tuning result come from the project; the supplied report does not include an independent replication. The benchmark table notes that Jev’s published scores use different samples, which limits direct comparison. It is also not clear from the report excerpt how the long-list test was constructed, how calibration was measured, or whether the same results hold across hardware configurations and real deployments.

The report gives no full results for its three game tests in the supplied material, and it does not establish how often users will need task-specific fine-tuning. The project says a short fine-tune can improve performance on its voice-navigation example, but that result does not establish expected gains for other applications. Developers would need to evaluate accuracy, confidence calibration and error costs against their own data and use cases.

Testing Jeff in Real Applications

The immediate next step for developers is to evaluate the v1.1 checkpoints on their own choices and hardware. The project provides serving instructions and model downloads, while retaining version 1.0 for users who need the earlier release. No later release date or independent evaluation is specified in the announcement. The project’s next concrete evidence will come from more detail on the tests and results it has published, alongside evaluations by users applying Jeff to their own tasks.

Key Questions

What changed in Jeff v1.1?

The project says the Qwen3.5 0.8B and 2B models can choose among up to 254 options, up from 26, and have improved calibration. The 0.8B model rose from 40% to 95% on the project’s long-list test.

How fast is Jeff?

The project reports about 22 milliseconds per decision on an RTX PRO 6000 and 28 milliseconds on an Apple M4 Max. These are project measurements on named hardware, not a guarantee for other systems.

Is Jeff affiliated with Jev?

No. Jeff uses the same request format, but its project says it is not affiliated with or endorsed by TypeSafe, the makers of Jev.

Does Jeff outperform Jev?

The project reports higher Jeff scores than published Jev figures on some classification and grounding benchmarks, but lower results on several reasoning-heavy tests. It also says the published Jev figures used different benchmark samples, limiting direct comparison.

Can Jeff run without cloud GPUs?

The project says it trained and tested the models using local hardware, and provides serving options for CPU, NVIDIA GPUs and Apple silicon for supported models. That describes the project’s setup; deployment requirements depend on the chosen model and device.

Source: hn

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

GPT-6 Astra On OpenRouter

OpenRouter integrates GPT-6 Astra, marking a significant step in open AI deployment. Details remain limited, but interest is surging among developers.

US reportedly allows 10 Chinese companies to buy NVIDIA’s coveted H200 AI chips

The US reportedly permits 10 Chinese companies to buy NVIDIA’s H200 AI processors, marking a potential shift in export controls amid ongoing tensions.

The Learning-by-Doing Wall In AI: China’s Gradual Progress Explained

An analysis of China’s recent advances in chipmaking, including domestic lithography tools and the challenges of scaling production capabilities.

Improving GPT‑5.6 Sol In ChatGPT, Expanding GPT‑5.6 Luna Access For Free Users

OpenAI improves GPT-5.6 Sol integration in ChatGPT and broadens free user access to GPT-5.6 Luna, signaling increased availability and performance.