Ask a large language model "is this email urgent?" and it will happily write you a paragraph about how urgency is a nuanced concept that depends on context. That paragraph costs money, takes a second or two to arrive, and buries the one bit of information you actually needed under a lot of throat-clearing. Multiply that by every message, every review, every line item an AI employee looks at in a day, and you have a system that thinks carefully about everything and quickly about nothing.
Jev is built for the opposite problem. It's a model from a company called TypeSafe, and it does not generate text at all. You hand it a situation and a question, and it hands back an answer: yes or no, which one of these, or where on this scale, each one with a number attached saying how sure it is. No paragraph, no punctuation to parse, nothing to argue with. We started using it inside Sanaf this month for exactly the moments where a full, deliberate answer is overkill and a fast, confident judgment is exactly what's needed.
The two ways to think, and why one of them got skipped
The psychologist Daniel Kahneman split human thinking into two systems: System One is fast, automatic, and runs without effort, the kind of judgment that tells you a stranger's face looks kind before you've thought about why. System Two is slow, deliberate, and effortful, the kind of thinking you do when you're doing long division or writing a difficult email. Most days, System One is doing the driving and System Two only gets called in for the hard problems.
Chat-style AI models are almost all System Two. Every question, however small, gets the full deliberate treatment: read the context, reason about it, compose a considered response. That's the right tool when the question is genuinely hard. It's a strange amount of machinery for "does this message need a reply today, or can it wait?"
TypeSafe calls Jev a "System One" model on purpose. It's the fast, intuitive layer that was missing. It looks at a situation and returns a snap judgment, calibrated so that "80% sure" actually means it's right about eight times out of ten, not a model being politely uncertain. It can't write you a sentence explaining its reasoning. What it can do is look at a hundred things and tell you, in under a second and for a fraction of a cent, which ten of them are worth a second look.
Three kinds of questions, and nothing else
Jev only answers three shapes of question, and the discipline is the point.
A yes or no, with a probability instead of a flat answer. Not "is this urgent" as a paragraph, but a number between 0 and 1: 0.92 sure, not 0.51 sure. The difference matters, because a system that treats "probably" and "almost certainly" as the same word can't tell you when to double-check its work.
A pick one of these, with a confidence spread across every option instead of just the winner. Given a message and a list of "is this a booking, a cancellation, or a question," it doesn't just say "booking." It says how much of the probability landed on booking versus how much landed on the runner-up, which is the difference between a clear call and a coin flip wearing a label.
A where does this sit on a scale, the same way you'd rate how urgent something feels from "can wait a week" to "drop everything." Not a guess dressed up as a number, an actual position along a described range, with the same kind of confidence behind it as the other two.
That's the whole vocabulary. Ask it to write a reply, summarize a thread, or invent a number, and it will refuse, because generating new text is exactly the job it was built to leave to something else. It's a specialist, and the specialty is judgment, not composition.
The specialty pays off because you can ask a whole batch of these questions about one situation at once, instead of one at a time. TypeSafe's own testing found that thirteen questions asked together about the same document came back about twelve times cheaper and ten times faster than asking them one by one, with no change in the answers. The situation only has to be read once. Everything you might want to know about it can be asked in that same breath, and the code on the other end only has to look at the answers that turned out to matter.
Where Sanaf actually uses it today
Picture the front desk of a busy office. Not the person who solves your problem, the person who looks up when you walk in, sizes up in about two seconds whether this is a quick question or something that needs the specialist down the hall, and points you the right way. That's the job Jev does inside a Sanaf turn, and it's live in production right now, not a slide in a deck.
Every time somebody sends Sanaf a message, something has to decide how much thinking the reply actually needs. "Thanks, that covers it" and "we need to unwind last week's shipment, rebook the courier, and tell the customer why it's late" are not the same amount of work, but until recently they both got routed to the same model, reasoning as hard about the first as the second. Now a Jev question runs first, sizing up how much the message really needs: a direct answer from what's already on screen, one quick lookup, or a longer investigation across a few tools. Simple messages get answered by a smaller, faster model. Anything that looks even slightly complicated, or that Jev isn't confidently sure about, stays on the model built to handle it. The rule only ever moves work down to the fast lane when Jev is genuinely confident, and any doubt at all keeps it on the careful path, because getting a hard question wrong costs a lot more than a few cents saved on an easy one.
The same trick shows up a layer earlier. An AI employee that's picked up a few dozen skills over time faces its own small triage problem on every message: which of these, if any, actually applies here? Reading through the whole list first, every time, is exactly the kind of System Two overhead that adds up over thousands of messages. A Jev question looks at the message and the skill list together and points at the likely match before any deliberate thinking starts, so the model can go straight to using the right approach instead of first working out what the right approach is.
Neither of those is a hypothetical. They're the two places, out of the handful we've built or are actively testing, where the fast judgment layer is doing real work on real conversations today.
What it will never decide
There's a line we won't move, and it's worth being explicit about it, because "the AI made a judgment call" is exactly the sentence that should make you nervous if the judgment involves your money or your customers.
Jev is a probability, not a permission slip. Whether Sanaf is allowed to send an email, spend money, delete something, or take any action that can't be quietly undone is decided by a separate, fixed set of rules, the same rules every time, checked by ordinary code rather than a confidence score. A judgment from Jev can raise a flag that a request looks riskier than its label suggests. It can never lower the bar and let something through on its own say-so. If a client hasn't approved an action, no amount of "I'm 99% sure this is fine" changes that. The fast layer is allowed to make the assistant quicker. It is never allowed to make it looser.
The same caution applies to arithmetic and dates. Jev doesn't count, and it doesn't do calendar math, on purpose, because a fast intuitive judgment about "does this feel urgent" is a different kind of thing than "is this actually overdue by three days," and blurring the two is how you end up with a system that's confidently wrong about something a calculator would get right. Anything that involves counting, adding up, or comparing timestamps stays in ordinary code, where it belongs.
The part that's still ahead of us
There's a bigger idea we're testing rather than shipping yet, and it's worth saying so plainly rather than letting it sound finished. A lot of the work an AI employee is hired to do isn't one big decision, it's a long list of small ones: which of last night's two hundred customer reviews actually need a personal reply, which dozen bank transactions don't obviously match anything on file, which of forty broken links on a website are worth fixing this week. Running a full, careful thought on every single item in a list like that is slow and, at scale, genuinely expensive. Running a fast judgment over the whole list first, and only spending real effort on the handful that clear the bar, is the shape of the idea.
We're not turning that on until we've checked it against something real: does a fast judgment actually agree with what a person would have picked, on an honest sample where nobody on our team hand-picked the "right" answer in advance. That measurement work is happening now. What we won't do is build a system that quietly decides two hundred things weren't worth your time and shows you none of it. If a triage layer like that ever ships, showing what it set aside, not just what it flagged, is the whole point, not an afterthought.
Why this is worth knowing about, even if you never touch a model directly
You don't need to know what Jev is to use an AI employee. You do get something out of knowing it exists: a hint about how a system that answers instantly and one that thinks carefully can live in the same product without either one pretending to be the other. The fast layer is what makes an AI employee feel responsive instead of ponderous on the easy stuff. The careful layer, and the rules that never bend for a probability, are what make it safe to hand the hard stuff to at all.
That's the whole design, really: know which kind of thinking a moment calls for, and never let the fast kind make a decision only the slow kind is allowed to make. See how the roster of jobs works, or start a conversation if you want to see the judgment calls happening on your own inbox instead of on paper.