Technical report / Unee 0.3

How Unee works, and how it compares.

Unee is a small AI model that makes decisions and writes answers, on an ordinary computer or right in the browser. This page explains how it works in plain English, then shows how it measures up against Jev, Laya and other decision models, with every number traced to a benchmark run.

Back to Unee

At a glance

Three decision models, side by side.

UneeJevLaya
WeightsOpen (Apache 2.0)ClosedOpen (Apache 2.0)
Where it runsYour computer, server or browserTypeSafe's hosted APIYour computer, server or browser
AnswersA probability for every optionA probability for every optionA probability for every option
Writes textStreams chat and explanationsNoNo
Answers from your documentsBuilt inNoNo
Input lengthAbout 3,000 tokens (as trained)About 32,000 tokens1,024 tokens
Size0.75B or 1.9B parametersNot published421M parameters
Cost per decisionFree on your hardwarePaid per callFree on your hardware

Jev and Laya facts come from their public documentation and model cards. Laya’s context figure is for its typed-decisions checkpoint.

How it works

Four ideas, explained simply.

1. A decision is one quick look, not an essay.

Big chatbots answer by writing word after word. Unee answers a decision the way you tick a box: it reads your text, your question and the options once, then gives every option a probability. That is why it is fast, and why it can never answer something that is not on your list.

  1. Your text“I was charged twice for order #4821.”
  2. Your question and optionsWhich team? Billing, technical or shipping.
  3. Unee reads it onceOne pass through the model, about a tenth of a second on a laptop graphics card.
  4. A probability for every optionBilling 98%, every other option under 1%. Your app acts when Unee is sure and asks a person when it is not.

2. It practised on checked examples.

A larger open model wrote and checked the practice examples, and Unee learned from them. No benchmark’s test questions were used.

  1. Practice examplesDecisions across 40 kinds of task in 45 languages, written and checked by Qwen3.5-9B.
  2. Tricky pairsNear-twins where one small change flips the right answer.
  3. Best of several runsPractice runs are averaged into one model, which smooths out each run’s mistakes.

3. A tiny date helper does the calendar maths.

Small models are bad at counting days in one glance, and so is Jev. Before Unee reads your text, a simple program works out the gaps between any dates it finds, in English and 18 other languages, and adds them as facts.

  1. Your text“Delivered on 1 May 2026. Return asked for on 5 June 2026.”
  2. The date helper adds“5 June 2026 is 35 days after 1 May 2026.”
  3. Unee just compares“Is it within 30 days?” becomes an easy question. This alone added about 3 points of accuracy.

4. It answers from your own documents.

For a chat assistant, Unee finds the right part of your help pages and answers from it. If the answer is not there, it says so instead of making something up.

  1. Your documentsFAQ, policies, product pages, prices.
  2. Short passagesSplit into bite-sized pieces, in any script.
  3. Best matchesThe pieces closest to the question are picked, in milliseconds.
  4. A streamed answerWord by word, in the user’s language, or “I don’t know” when the documents do not say.

Accuracy and cost

On the frontier: nothing cheaper is more accurate.

DecideBench gives 400 tricky decisions to every model and records how often it is right and what a million decisions would cost. The dashed line joins the models that nothing beats on both at once. Both Unee sizes sit on it.

DecideBench accuracy against cost per million decisions30%40%50%60%70%80%90%100%$3$10$30$100$300Cost per million decisions (log scale)AccuracyDeepSeek-V4-Flash: 99.8%DeepSeek-V4-FlashDeepSeek-V4.1-Flash: 99.2%DeepSeek-V4.1-FlashGLM-5.3-Flash: 99.2%GLM-5.3-FlashJev: 98%Jevimajev-4b: 95%imajev-4bClef 27B: 94.8%Clef 27BBespoke-Nimble-9B: 94%Bespoke-Nimble-9BTEV: 92.8%TEVTEV (Together): 92.8%TEV (Together)Qwen3-8B, no thinking: 90.5%Qwen3-8B, no thinkingJevK5 v0.3: 88.8%JevK5 v0.3Decider-4B: 88.5%Decider-4BClef-Flash 9B: 85.8%Clef-Flash 9BJeff Gemma4-E2B: 80.8%Jeff Gemma4-E2BKev-4B: 78.2%Kev-4BKev-9B: 72.8%Kev-9BJeff Qwen3.5-0.8B: 71.2%Jeff Qwen3.5-0.8BJeff Qwen3.5-2B: 70%Jeff Qwen3.5-2BDecider-2B: 63.2%Decider-2BQwen2-1.5B, Arize: 60.2%Qwen2-1.5B, ArizeLaya 421M: 59.8%Laya 421MGLiNER2.5-Decide 340M: 57%GLiNER2.5-Decide 340MCLM-v0.1-8B: 41%CLM-v0.1-8BJulia-1 144M: 34.8%Julia-1 144MUnee 0.8B: 83%Unee 0.8BUnee 2B: 87%Unee 2B
  • Unee
  • Jev
  • Laya
  • Other models
  • Frontier
DecideBench v1.1. Other entries from the benchmark's published table. Unee runs on its default llama.cpp server with 4-bit weights; its cost uses DecideBench's own method for self-hosted models (graphics-card hours times hourly price, at 4 requests at a time), with throughput measured on a laptop RTX 4070 and priced at the $0.81/hour NVIDIA L4 rate the benchmark uses. That is a stand-in, not a measurement on an L4.

Accuracy and size

More right answers per gigabyte.

Smaller models are cheaper to run and fit on more devices. Among models whose size is published, nothing smaller than Unee is more accurate, and every model that beats Unee 2B is at least twice its size.

DecideBench accuracy against model size30%40%50%60%70%80%90%100%0.1B0.3B1B3B10B30BModel size, billions of parameters (log scale)Accuracyimajev-4b: 95%imajev-4bClef 27B: 94.8%Clef 27BBespoke-Nimble-9B: 94%Bespoke-Nimble-9BQwen3-8B, no thinking: 90.5%Qwen3-8B, no thinkingDecider-4B: 88.5%Decider-4BClef-Flash 9B: 85.8%Clef-Flash 9BKev-4B: 78.2%Kev-4BKev-9B: 72.8%Kev-9BJeff Qwen3.5-0.8B: 71.2%Jeff Qwen3.5-0.8BJeff Qwen3.5-2B: 70%Jeff Qwen3.5-2BDecider-2B: 63.2%Decider-2BQwen2-1.5B, Arize: 60.2%Qwen2-1.5B, ArizeLaya 421M: 59.8%Laya 421MGLiNER2.5-Decide 340M: 57%GLiNER2.5-Decide 340MCLM-v0.1-8B: 41%CLM-v0.1-8BJulia-1 144M: 34.8%Julia-1 144MUnee 0.8B: 84%Unee 0.8BUnee 2B: 88%Unee 2B
  • Unee
  • Laya
  • Other models
  • Smallest for its accuracy
DecideBench v1.1 accuracy against model size, for entries whose size is stated in their name, plus Unee measured with the benchmark's protocol. Sizes in billions of parameters; Gemma4-E2B is left out because E2B is an effective, not a total, size.

Where each model wins

Strong at intent and triage, still learning agent safety.

The 400 decisions come in eight kinds. Unee matches Jev on understanding what a customer wants, and is weakest on reviewing AI agent actions and moderation.

DecideBench accuracy by task familyShould an AI agent's action go ahead?Unee 2BUnee 2B, Should an AI agent's action go ahead?: 86%86%Unee 0.8BUnee 0.8B, Should an AI agent's action go ahead?: 78%78%JevJev, Should an AI agent's action go ahead?: 100%100%LayaLaya, Should an AI agent's action go ahead?: 28%28%Which tool should an agent use?Unee 2BUnee 2B, Which tool should an agent use?: 90%90%Unee 0.8BUnee 0.8B, Which tool should an agent use?: 88%88%JevJev, Which tool should an agent use?: 98%98%LayaLaya, Which tool should an agent use?: 52%52%Does the evidence support the claim?Unee 2BUnee 2B, Does the evidence support the claim?: 86%86%Unee 0.8BUnee 0.8B, Does the evidence support the claim?: 82%82%JevJev, Does the evidence support the claim?: 96%96%LayaLaya, Does the evidence support the claim?: 76%76%How should a post be moderated?Unee 2BUnee 2B, How should a post be moderated?: 82%82%Unee 0.8BUnee 0.8B, How should a post be moderated?: 80%80%JevJev, How should a post be moderated?: 98%98%LayaLaya, How should a post be moderated?: 52%52%Does a return follow the policy?Unee 2BUnee 2B, Does a return follow the policy?: 80%80%Unee 0.8BUnee 0.8B, Does a return follow the policy?: 74%74%JevJev, Does a return follow the policy?: 96%96%LayaLaya, Does a return follow the policy?: 50%50%How happy is the reviewer?Unee 2BUnee 2B, How happy is the reviewer?: 88%88%Unee 0.8BUnee 0.8B, How happy is the reviewer?: 82%82%JevJev, How happy is the reviewer?: 98%98%LayaLaya, How happy is the reviewer?: 72%72%What does the customer want?Unee 2BUnee 2B, What does the customer want?: 100%100%Unee 0.8BUnee 0.8B, What does the customer want?: 98%98%JevJev, What does the customer want?: 100%100%LayaLaya, What does the customer want?: 90%90%How urgent is this ticket?Unee 2BUnee 2B, How urgent is this ticket?: 92%92%Unee 0.8BUnee 0.8B, How urgent is this ticket?: 90%90%JevJev, How urgent is this ticket?: 98%98%LayaLaya, How urgent is this ticket?: 58%58%
DecideBench v1.1 accuracy by task family (50 decisions each). Jev and Laya from the benchmark's published per-family table.

An independent check

On a benchmark Unee never trained on.

JevBench is a separate benchmark with 1,682 checked decisions. Unee 0.1 ran it once, zero-shot, and clearly beat Laya, with Jev clearly ahead. Later versions were not rerun: some of their practice data now overlaps JevBench’s checked questions, so their score would not be independent.

ModelOverallPick oneYes or noRate on a scale
Unee 0.1, 2B70.7%63.2%80%70.4% exact
Unee 0.1, 0.8B63.9%57.5%69%69.1% exact
Jev88.8%---
gpt-oss-20b88.4%---
bart-mnli54.3%---
Laya53.9%---

JevBench leaderboard, 232-row gold sample (not the same rows as Unee's 1,682). Unee 0.1 was run on all 1682 gold decisions, zero-shot, once. After this run, new practice tasks were added for the kinds of question JevBench showed as weak (long option lists, two-step reasoning, arithmetic), so a JevBench score for later versions would no longer be fully independent. Later versions are checked on S1MB and MASSIVE instead.

Beyond English

The same requests, in 51 languages.

MASSIVE is a public set of everyday requests to a voice assistant, such as “wake me up at seven” or “will it rain tomorrow”, written by native speakers in 51 languages. We asked each model to pick which of 18 parts of the assistant should handle 60 of them, the same requests in every language, so the languages can be compared fairly. Jev and Laya have not been run on these requests, so the comparison is with the model Unee is built on.

ModelEnglishAverage over 51 languagesAverage over the 12 priority languages
Unee 0.3, 2B86.7%69.7%77.2%
Unee 0.3, 0.8B91.7%63.3%74.4%
Qwen3.5 0.8B, before Unee training43.3%28.5%35.8%
Accuracy by language on MASSIVEEnglishUnee 2BUnee 2B, English: 86.7%86.7%Unee 0.8BUnee 0.8B, English: 91.7%91.7%ArabicUnee 2BUnee 2B, Arabic: 71.7%71.7%Unee 0.8BUnee 0.8B, Arabic: 56.7%56.7%HindiUnee 2BUnee 2B, Hindi: 68.3%68.3%Unee 0.8BUnee 0.8B, Hindi: 73.3%73.3%SpanishUnee 2BUnee 2B, Spanish: 78.3%78.3%Unee 0.8BUnee 0.8B, Spanish: 78.3%78.3%Chinese (Simplified)Unee 2BUnee 2B, Chinese (Simplified): 88.3%88.3%Unee 0.8BUnee 0.8B, Chinese (Simplified): 90%90%FrenchUnee 2BUnee 2B, French: 76.7%76.7%Unee 0.8BUnee 0.8B, French: 75%75%PortugueseUnee 2BUnee 2B, Portuguese: 73.3%73.3%Unee 0.8BUnee 0.8B, Portuguese: 73.3%73.3%BengaliUnee 2BUnee 2B, Bengali: 68.3%68.3%Unee 0.8BUnee 0.8B, Bengali: 68.3%68.3%RussianUnee 2BUnee 2B, Russian: 78.3%78.3%Unee 0.8BUnee 0.8B, Russian: 78.3%78.3%UrduUnee 2BUnee 2B, Urdu: 71.7%71.7%Unee 0.8BUnee 0.8B, Urdu: 65%65%IndonesianUnee 2BUnee 2B, Indonesian: 85%85%Unee 0.8BUnee 0.8B, Indonesian: 81.7%81.7%GermanUnee 2BUnee 2B, German: 83.3%83.3%Unee 0.8BUnee 0.8B, German: 73.3%73.3%JapaneseUnee 2BUnee 2B, Japanese: 83.3%83.3%Unee 0.8BUnee 0.8B, Japanese: 80%80%
MASSIVE scenario test set, the same 60 requests in each language. English is shown first for reference; the other languages are the ones Unee 0.3 gave extra training weight.

Answers from your documents

Does it stick to your pages?

We wrote small help centres for made-up companies Unee never saw, some in other languages, and asked 116 questions. Some are answered in the pages and some are not. A larger model judged every reply, and strictly: an answer that leaves out part of a policy counts as wrong.

ModelSays “I don’t know” when the pages do not cover itFully right when the pages have the answer
Unee 0.3, 2B67.9%42%
Unee 0.3, 0.8B53.6%34.1%
Qwen3.5 0.8B, before Unee training17.9%28.4%

Every model gets the same pages through the same built-in search. Unee says so when your pages do not have the answer, where the model it is built on mostly guessed. The judge is Qwen3.5-9B, the model Unee learned from, so treat these as a guide rather than a final verdict.

Strict mode checks the answer before it is sent.

In strict mode Unee writes the whole answer, then checks it sentence by sentence against the pages and removes what it cannot support. The same answers were judged with and without that check.

ModelReplies with a made-up fact, no checkWith strict mode
Unee 0.3, 2B10.3%6%
Unee 0.3, 0.8B14.7%11.2%

Strict mode costs a few points of correct answers, and it lowers made-up facts without ruling them out: the checker is the same small model. It checks at a threshold of 0.5 here, and the threshold is yours to set.

A third benchmark

S1MB: over a hundred public tests in one score.

S1MB is a community benchmark that gathers public datasets for the three kinds of decision Unee makes: pick one, yes or no, and rate on a scale. Each test is scored so that 0 means no better than a trivial guess, such as always giving the most common answer, and 100 means perfect. The headline is the average of the three kinds.

ModelTask averagePick oneYes or noRate on a scale
Jev 1.1359.59---
Open-Jev-9B52.10---
bekko-system-one-v0-400m50.60---
Unee 0.3, 2B44.4152.6445.6035.00
Unee 0.3, 0.8B35.4541.4242.4322.51
Kev-0.8b18.68---
von16.21---
Laya13.36---
JevForge-0.8B1.92---

S1MB leaderboard, full English suite, snapshot of 2026-09-30. Unee was run by us with S1MB’s own evaluator on 137 scored tests. S1MB’s test rows were never used in training, but Unee 0.3 did learn from the public training splits of the datasets S1MB draws on. Checking every S1MB test input against all of Unee’s training text finds 6.9% that resemble it, almost all in tests generated from templates, with 48 exact copies. Leaving out every test with any overlap (101 remain), the task average is 45.69 for 2B and 36.74 for 0.8B, no lower than the full scores, so they do not rest on the overlap. The few long-document tests (contracts, research papers) did not fit the four-at-once setup, so they were run again one at a time with a longer context window. Per-kind scores are shown for Unee only, as the leaderboard snapshot lists the average.

General ability

What it keeps from its base.

Unee is trained for decisions and answers from your documents, not general knowledge. These widely used tests show how much general ability it keeps. They were run with EleutherAI’s evaluation harness (lm-eval 0.4.13) on every item, with the model’s thinking switched off.

TestUnee 0.3, 0.8BUnee 0.3, 2B
MMLU: 57 school and professional subjects (5 worked examples)49.3%58.1%
ARC-Challenge: grade-school science, the harder set43.5%49.9%
HellaSwag: finishing an everyday situation51.3%59.6%
TruthfulQA: avoiding common misconceptions32.4%29.9%
GSM8K: maths word problems (5 worked examples)29.7%49%
IFEval: every format instruction in a prompt followed39.7%52.7%
IFEval: share of single instructions followed51.2%62.6%

Qwen publishes none of these for the base models except IFEval: 52.1% for 0.8B and 61.2% for 2B, sampled with some randomness and without saying which IFEval score it is. Unee’s per-instruction score is close to those and its whole-prompt score is lower, so read that comparison as a rough guide. We did not run the base models ourselves.

Honest confidence

Sure answers you can trust, unsure ones you can spot.

A probability is only useful if it is honest. Most of Unee’s answers come with high confidence, and there it is honest: answers given at about 96% are right about 95% of the time. Its less confident answers are right less often than they claim, which is exactly why your app should act on its own above a threshold and hand the rest to a person.

Unee calibration on DecideBench20%20%30%30%40%40%50%50%60%60%70%70%80%80%90%90%100%100%How sure the model said it wasHow often it was rightUnee 2B: said 47%, right 47% (15 answers)Unee 2B: said 54%, right 58% (12 answers)Unee 2B: said 66%, right 77% (22 answers)Unee 2B: said 75%, right 71% (34 answers)Unee 2B: said 85%, right 86% (63 answers)Unee 2B: said 95%, right 97% (249 answers)Unee 0.8B: said 46%, right 47% (17 answers)Unee 0.8B: said 55%, right 56% (36 answers)Unee 0.8B: said 66%, right 77% (35 answers)Unee 0.8B: said 75%, right 78% (37 answers)Unee 0.8B: said 86%, right 85% (114 answers)Unee 0.8B: said 94%, right 96% (161 answers)
  • Unee 2B
  • Unee 0.8B
  • Perfect calibration
DecideBench v1.1, all 400 decisions, grouped by how confident the top answer was; bands with fewer than 10 answers are not drawn. Expected calibration error: Unee 2B 0.026, Unee 0.8B 0.025 (0 is perfect).

Speed on real hardware

Fast on a graphics card, usable on an old laptop.

Times for one decision with worked examples in the prompt (the benchmark’s setup), measured on a recent laptop (Intel Core Ultra 9, RTX 4070). Without examples, decisions are up to about twice as fast. Older, slower processors will take longer.

Where it runsUnee 0.8BUnee 2B
Laptop graphics card (RTX 4070)95.7 ms111.2 ms
Processor only, 4 threads1.2 s2.6 s
Processor only, 2 threads2.1 s4.5 s
Inside a web browser (WebGPU)0.4 to 0.9 sNot packaged yet

Median time per decision, one request at a time, llama.cpp with 4-bit weights. For comparison, DecideBench lists Jev at 639 ms through its hosted API and Laya at 97 ms on its own server, measured on different hardware.

Honest limits

What Unee is not good at yet.

  • Large hosted models are more accurate. Jev and the biggest models still get more decisions right. Unee’s advantage is that it is small, free to run and runs anywhere.
  • One quick look has limits. Decisions that need several steps of reasoning, sums or very long lists of options are harder for it.
  • It is not a general knowledge assistant. On 60 open-ended questions, a larger model judging blind preferred the answers of the plain Qwen model Unee is built on. Unee’s answers are shorter and, like any small model, it gets facts wrong. Its chat is meant for answering from your own documents.
  • The explanation is the model’s own account. When Unee explains a decision, that is useful, but it is not proof of how the decision was made.
  • Not for high-stakes calls on its own. Medical, legal, credit or hiring decisions need a person in the loop.