Product 002 AI model Early access Version 0.3

Unee

A small AI model that lives inside your app.

Unee makes decisions, answers from your own knowledge and streams its replies, on an ordinary computer or right in the browser. No API, no account, no cost per call.

  • Two sizes: 0.8B and 2B
  • Runs offline
  • No API key
  • Open source

Input

Hi, I was charged twice for order #4821 this morning. Please send the extra $39 back to my card.

Which team handles it?

  1. Billing98%
  2. Shipping<1%
  3. Sales<1%
  4. Technical<1%

WhyThe user is reporting a payment error and requesting a refund.

Asks for money back?

  1. Yes97%
  2. No3%

How urgent?

  1. Today49%
  2. Right now35%
  3. Within a few days10%
  4. Can wait a week6%

Real responses from Unee 0.3 (0.8B) running on a laptop CPU. To keep the cards short some answers are left out, including two wrong yes or no answers about the mixed review; every answer is in the recording in the Unee repository. Like any model it can be wrong, so test it on your own data.

Unee 0.8B in the browser, downloaded once and then cached.
469 MB
One decision by Unee 0.8B on a laptop GPU (median, single request, 4-bit).
96 ms
Servers, API keys or per-call fees needed in the browser.
0
Model sizes, from a browser tab to a server.
2

What it does

One small model that decides and talks.

Most small models do one thing. Unee decides, talks, looks things up and summarises, so one download covers the whole conversation.

Calibrated decisions today. Grounded chat that is still improving. Unee 0.3 is our first step towards a small model you can trust with both.

  1. 01

    Decide

    Yes or no, pick one option, or score on a scale. Every option gets a calibrated probability, so your app knows how sure Unee is and when to ask a person.

  2. 02

    Chat

    Streams its reply word by word, the way large assistants do, on your own hardware. Ask for a reason and it explains a decision right after making it.

  3. 03

    Answer from your knowledge

    Give it your help centre, policies and product details. It finds the right passages, answers from them, and says so when they do not cover the question.

  4. 04

    Summarise

    Turns a long thread from email, chat and call notes into a few clear points: who is involved, what they want, what is done and what is still open.

Confidence

It tells you how sure it is.

Every answer comes with a probability, and the probabilities hold up. Act on the answers Unee is sure about, and send the rest to a person.

Unee 2B

Right when it is sure
97.2%
Right overall
88.0%

On 62% of decisions, Unee 2B gave its answer a probability of 90% or more. Those answers were right 97.2% of the time.

Unee 0.8B

Right when it is sure
96.3%
Right overall
84.0%

On 40% of decisions, Unee 0.8B gave its answer a probability of 90% or more. Those answers were right 96.3% of the time.

DecideBench v1.1, 400 decisions per model, measured by UNEEVERSE.

Where it fits

Every message in. Clear next steps out.

  • Customer support and CRM

    Route every message, flag urgency and refund requests, answer from your help centre and summarise a customer's history across every channel.

  • Your website

    An assistant that answers from your own pages and runs in the visitor's browser, so a busy day costs nothing extra.

  • AI agents

    A fast check before an agent runs a command, spends money or sends an email: approve, ask a person, or block.

  • Moderation and routing

    Spam, abuse, intent and priority, answered together in one call over the same message.

  • Private and offline apps

    Text never has to leave the device, which keeps sensitive data where it belongs.

  • Long threads

    Summarise email, chat and call threads in the language they were written in, and pull out what needs doing.

Two sizes

Pick the one that fits your hardware.

Unee 0.8B

For browsers, everyday laptops and modest servers.

Size
0.75B parameters
Download
469 MB in the browser, 529 MB for llama.cpp (4-bit)
Runs on
Browser with WebGPU, any CPU, any GPU
DecideBench
84.0%

Unee 2B

For servers and stronger PCs. The most accurate Unee.

Size
1.9B parameters
Download
1.27 GB for llama.cpp (4-bit)
Runs on
Any CPU, any GPU
DecideBench
88.0%

How it compares

Honest numbers, side by side.

Jev is the most accurate of the three. Unee is ahead of Laya on both leaderboards, runs on your own computer or in the browser, is free to use, and is the only one of the three that can talk back.

Smaller, and still ahead.

Parameters are a model's size. Fewer means a smaller download, less memory and a cheaper machine. Unee is the small one in most of these comparisons.

1.9B parameters

Unee 2B

DecideBench
Ahead of four larger models: Clef-Flash 9B, Kev-4B, Kev-9B and CLM-v0.1-8B.
S1MB
Ahead of two larger models: Decider 4B and kev-9b.

0.75B parameters

Unee 0.8B

DecideBench
Ahead of four larger models: Kev-4B, Kev-9B, Decider-2B and CLM-v0.1-8B.
S1MB
Ahead of two larger models: Open-Jev-2B and Jeff Qwen3.5 2B.

A larger model counts only when its stated size is at least a quarter bigger. Larger models still lead the table: Jev and imajev-4b are both more accurate than Unee.

DecideBench leaderboard

DecideBench v1.1 accuracy for Unee and other decision models
ModelParametersAccuracyPair accuracyRunsNext to Unee
DeepSeek-V4-FlashNot stated99.8%99.5%Hosted API
JevNot stated98.0%96.0%Hosted API
imajev-4b4B95.0%90.5%Self-hosted
Clef 27B27B94.8%89.5%Hosted API
Bespoke-Nimble-9B9B94.0%88.0%Self-hosted
TEVNot stated92.8%86.0%Self-hosted
Qwen3-8B, no thinking8B90.5%81.5%Self-hosted
JevK5 v0.3Not stated88.8%78.5%Self-hosted
Decider-4B4B88.5%78.5%Self-hosted
Unee 2B1.9B88.0%76.5%Any PC
Clef-Flash 9B9B85.8%73.0%Hosted APILarger, behind Unee 2B
Unee 0.8B0.75B84.0%70.0%Any PC or browser
Jeff Gemma4-E2BNot stated80.8%64.5%Self-hosted
Kev-4B4B78.2%65.5%Self-hostedLarger, behind Unee 0.8B
Kev-9B9B72.8%53.5%Self-hostedLarger, behind Unee 0.8B
Jeff Qwen3.5-0.8B0.8B71.2%49.5%Self-hosted
Decider-2B2B63.2%41.0%Self-hostedLarger, behind Unee 0.8B
Laya typed-decisions421M59.8%36.5%Self-hosted
GLiNER2.5-Decide340M57.0%30.5%Self-hosted
CLM-v0.1-8B8B41.0%11.0%Self-hostedLarger, behind Unee 0.8B

What you get

Unee 2B is 10 points behind Jev on DecideBench, 88.0% against 98.0%, and it runs on your own computer or in the browser with no fee per call.

Unee compared with Jev and Laya. The strongest option on each row is marked.
Unee 2BUnee 0.8BJevLaya
DecideBench accuracy88.0%84.0%98.0% (strongest on this row)59.8%
S1MB score44.4135.4559.59 (strongest on this row)15.00
Size (smaller is lighter)1.9B0.75BNot published421M (strongest on this row)
Runs on your own hardwareYes (strongest on this row)Yes (strongest on this row)No, hosted APIYes (strongest on this row)
Streams text and explanationsYes (strongest on this row)Yes (strongest on this row)NoNo
One decision111 ms, laptop GPU96 ms, laptop GPU (strongest on this row)639 ms, through its API97 ms, L4 GPU (strongest on this row)
Cost per million decisionsAbout $23 on a rented GPU$0 in the browser, about $19 on a rented GPU (strongest on this row)$32.26$5.45
Open weightsYes (strongest on this row)Yes (strongest on this row)NoYes (strongest on this row)

Bold marks the strongest on each row. One decision marks the two fastest, which are a millisecond apart on different hardware.

  1. DecideBench v1.1, 400 contrastive decisions, as published in the benchmark's README on 1 October 2026. Unee was measured by UNEEVERSE with the benchmark's own harness and prompts, and is not on the official leaderboard yet.
  2. S1MB: 137 public decision benchmarks, scored so that 0 is a trivial guess and 100 is perfect. The Jev and Laya rows are the live leaderboard on 6 October 2026. Unee trained on the public training splits behind S1MB and never on its test set. On the 101 benchmarks with no overlap at all with its training text, it scores 45.69 (2B) and 36.74 (0.8B).
  3. Unee 0.1 was also checked once on JevBench, which it never trained or tuned on: 2B 70.7%, 0.8B 63.9%. The JevBench leaderboard lists Jev at 88.8% and Laya at 53.9% on a 232-item sample.
  4. You may have read that Laya beats Jev. That comes from Laya's own model card: 0.766 accuracy against 0.727 for Jev on Laya's typed-decisions test of 400 cases, a test that checkpoint was trained for. On the two independent leaderboards here, Jev is ahead. Unee has not been run on Laya's test.
  5. Sources: DecideBench v1.1 (github.com/choyiny/decidebench), the S1MB leaderboard (huggingface.co/spaces/hotchpotch/S1MB-leaderboard) and the Laya typed-decisions model card. Unee's own results are in its repository under bench/results.
  6. Speeds are not like for like. Unee was timed on a laptop RTX 4070 with no network, Jev through its API, and Laya on an L4 graphics card by DecideBench.
  7. Cost follows DecideBench's method for self-hosted models: graphics card time with four requests in flight, at $0.81 an hour. Unee's figure uses the laptop's throughput at that price, so it is a stand-in, not a measurement on an L4. Jev's and Laya's figures are from the DecideBench leaderboard.

Strict mode

Checked before it is sent.

Small models sometimes add a fact that is not in your documents. With strict mode on, Unee checks its own answer first.

  1. 01

    Write

    Unee writes the whole answer from your documents.

  2. 02

    Check

    Its decision side reads the answer sentence by sentence. Numbers, times, email addresses and links must appear in your documents exactly, and every fact is verified against them.

  3. 03

    Send

    Sentences it cannot support are removed, and a short note says the rest could not be confirmed.

A real reply from Unee 2B in our test set, which uses made-up companies.

what is the cost for adding 5 new employees and do you offer payment plans?

The cost for adding 5 new employees is $225, calculated as the base fee of $49 per employee plus $4.50 each.Removed: 225 does not appear in the documents.

Yes, we offer payment plans available in three tranches: 40% upfront, 30% at 30 days, and 30% at 60 days.

I couldn't confirm the rest from the knowledge I have. Would you like me to pass your question to a person?

Replies containing a made-up fact

Unee 2B
10.3% without6.0% with strict mode
Unee 0.8B
14.7% without11.2% with strict mode

Strict mode lowers the risk. It is not a guarantee: the checker is the same small model, and it removes a few correct sentences too. Measured on 116 questions with a larger model as judge.

Get started

A few lines, then it is yours.

  • An npm package for the browser and Node.js
  • A Python package and a one-command server
  • An HTTP API that accepts Jev requests unchanged
  • OpenAI-compatible streaming chat
  • Tools for AI agents through MCP

Unee 0.8B runs in the visitor's browser on WebGPU, or in Node.js on the processor. Nothing is sent to a server.

  1. Install the package and its runtime.

    npm i @uneeverse/unee @huggingface/transformers
  2. Load the model once. The browser downloads 469 MB and then keeps it.

    import { Unee } from "@uneeverse/unee";
    
    const unee = await Unee.load("uneeverse/unee-0.8b");
  3. Ask for a decision. Every option comes back with a probability.

    const res = await unee.decide({
      state: "I was charged twice for order #4821",
      questions: {
        team: { type: "choice", instructions: "Which team handles this?",
                criteria: { billing: "Charges and refunds", technical: "Bugs" } },
      },
    });
  4. Or answer from your own documents, checked before it is shown.

    for await (const piece of unee.stream(messages, { knowledge: helpCentre, strict: true }))
      chat.append(piece);

The HTTP API also accepts Jev requests unchanged.

What it needs.

Unee 0.8BUnee 2B
Download469 MB in the browser, 529 MB for llama.cpp1.27 GB for llama.cpp
Memory while runningAbout 1.0 to 1.2 GBAbout 2.0 to 2.3 GB
Graphics cardNot neededNot needed
One decision, processor only0.64 s1.4 s
One decision, laptop graphics card96 ms111 ms
One decision, in the browser0.4 to 0.9 sNot packaged for the browser
BrowserChrome or Edge, Safari 26 or later, Firefox 141 or later, with WebGPUNot packaged for the browser

Measured on one laptop (Intel Core Ultra 9 185H, 32 GB of memory, RTX 4070 Laptop with 8 GB), using the 4-bit files on llama.cpp. Processor times use four threads and a request with no worked examples; with one worked example per option they are 1.2 s and 2.6 s. Treat these as a guide: an older processor takes longer.

Settings, explained.

Open a setting to see what it controls, what you can give it, and what it will not take.

knowledgeThe documents Unee answers from.Chat default: None

What it controls

What Unee is allowed to answer from. Without it, Unee chats from what it learned in training. With it, Unee answers from your text and says so when the text does not cover the question.

How it works

Unee splits the text into passages of about 120 words, picks the passages whose words best match the last two questions, and answers from those. Nothing is kept between requests.

What you can give it

  • Plain text, as many pieces as you like.
  • A title with its text, such as one help-centre article.
  • An Open Knowledge Format folder, read by the loader.
  • Text in any language Unee handles.

What it will not take

  • Files, PDFs, images or web links. Turn them into text first.
  • A database or search index. Unee searches the text you send by itself.
Example
"knowledge": [
  { "title": "Refunds", "text": "Refunds are paid within 14 days of a return." },
  "Shipping: UK only, 2 to 3 working days."
]
knowledge_kHow many passages go into each answer.Chat default: 4

What it controls

The number of passages Unee reads before it answers. The default is 4. In the browser package it is called knowledgeK.

How it works

More passages give Unee more to work from, and more unrelated text to be distracted by. Change it only if answers are missing something your documents do say.

What you can give it

  • A whole number, 1 or more.
  • A higher number when answers are spread across several pages. Unee reads more, and each answer takes longer.
  • A lower number for speed, when each question has one obvious passage.

What it will not take

  • More passages than you sent. If you send four or fewer, all of them are used.
Example
"knowledge_k": 6
strictCheck every sentence against your documents before sending.Chat default: Off

What it controls

Whether Unee checks its own answer before anyone sees it. It needs knowledge: with no documents there is nothing to check against, and the request runs as ordinary chat.

How it works

Each sentence is sorted into a fact, an 'I do not know', or neither. Facts must pass two checks: numbers, times, email addresses and links have to appear in your documents exactly, and the sentence has to be supported by them. What fails is removed, and the note is added.

What you can give it

  • true, to switch it on with the defaults.
  • threshold: a number from 0 to 1 for how sure a sentence must be to stay. The default is 0.5. Higher removes more.
  • note: the sentence Unee adds when it removed something. The default is in English, so write your own for other languages.

What it will not take

  • Any other keys. Only threshold and note are read.
  • Rewriting. Strict mode keeps a sentence or removes it; it never changes the wording.
  • Word-by-word streaming. The answer arrives whole, because it is checked first.
Example
"strict": { "threshold": 0.7,
            "note": "I could not confirm the rest. Shall I pass this to a person?" }
explainAdd a one-sentence reason to a decision.Decisions default: Off

What it controls

Whether a decision comes back with a short reason. The decision is made first, in one pass; the reason is written after it.

How it works

Reasons take longer than decisions, because they are written word by word. Ask for them where a person will read them.

What you can give it

  • true on any question you want explained.

What it will not take

  • A guarantee. The reason is Unee's explanation, not proof of how it decided.
Example
"team": { "type": "choice", "instructions": "Which team handles this?",
          "criteria": { "billing": "Charges and refunds", "technical": "Bugs" },
          "explain": true }

Server options

OptionDefaultWhat it does
--port8000Where the API listens.
--threadsChosen by llama.cppProcessor threads to use.
--slots4Requests handled at the same time.
--ctx16384Context length, shared across those requests.
--gpu-layers0Layers to place on a graphics card. Use 99 for all of them.

Put it to work

A worked example: a support inbox.

Where Unee earns its place today. Decide, act when it is sure, ask a person when it is not, then draft the reply.

  1. Decide

    One call answers several questions about a message: which team, how urgent, and whether it asks for money back.

  2. Act on the sure ones

    When the chosen option has a probability of 90% or more, route the message automatically.

  3. Ask a person for the rest

    Everything else goes to a person, with Unee's probabilities attached.

  4. Draft the reply

    Unee drafts an answer from your help centre with strict mode on. A person approves it before it goes out.

support_inbox.py
from unee import Client

unee = Client("http://127.0.0.1:8000")
teams = {"billing": "Charges and refunds",
         "technical": "Bugs and outages",
         "shipping": "Deliveries"}

answer = unee.choice(message, "Which team handles this?", teams)
if max(answer["probabilities"].values()) >= 0.9:
    route(message, answer["choice"])      # sure: act
else:
    hand_off(message, answer)             # not sure: a person decides

draft = "".join(unee.chat(message, knowledge=help_centre, strict=True))
queue_for_approval(draft, unee.strict_report)

This is a pattern, not a customer story. route, hand_off and queue_for_approval stand for your own code.

Know the limits

What Unee is not, yet.

  • Larger hosted models are more accurate

    Jev scores 98.0% on DecideBench against 88.0% for Unee 2B.

  • Chat needs a person checking

    About half of its answers from documents are judged fully correct, and some contain a fact the documents do not state.

  • English is strongest

    Check the per-language results in the technical report before relying on another language.

  • Not for high-stakes decisions

    Medical, legal, credit and employment decisions need a person.

Roadmap

What you can expect.

Unee is in early access: usable today, and still improving. Plans can change, so nothing here is a promise of a date.

Available today

  • Decisions with a confidence score

    Yes or no, pick one, or rate on a scale, with a probability for every option.

  • Chat from your own documents

    Streaming answers from the knowledge you give it, with strict mode to check them.

  • Two sizes, three places to run

    In a browser tab, on your own computer, or on a server.

In progress

  • More accurate decisions

    Fewer wrong calls, and a larger share of answers Unee is sure about.

  • Chat you can lean on more

    Better answers from your documents and fewer mistakes in open conversation.

Coming next

  • A stronger answer check

    Strict mode that catches more of what your documents do not support.

  • A larger model

    A bigger size for servers, for work that needs more accuracy.

Open source

Yours to use, change and ship.

Created by Muneef Mumthas at UNEEVERSE.

Open weights and code under Apache 2.0.

Built on Qwen3.5 by the Qwen team, released under Apache 2.0.

Launch list

Hear what Unee does next.

Join the UNEEVERSE launch list to hear about new versions of Unee, new products and updates.