Is Jev the future of spam filtering?
Titles ending in a question mark usually mean the answer is no. For this one, it could very well be a yes.
This month TypeSafe AI came out of stealth with Jev, a hosted model that cannot chat: it takes a typed question and returns a probability. “Is this email spam?” is the obvious question to ask it, so I pitted it against Klar’s own model. It is billed per token, which came to 11 cents per thousand of my messages.
Jev catches more of my spam than the model I ship, for eleven cents per thousand messages. It is not a fair fight: Klar’s filter runs on any Apple silicon Mac and nothing it reads leaves the machine. Jev runs in a data centre.
Jev catches more, and misses newsletters and cold pitches
I took the 5,000 messages I measure Klar on: 2,500 spam (2,456 from a public 2026 corpus, 44 from my own junk folder) and 2,500 legitimate, drawn in proportion from my 2025 receipts, notices and opted-in newsletters. Person-to-person mail stayed on the Mac. Three models read them: Jev, the model Klar ships today, and the one it replaced last year.
Jev, the model Klar ships, and the one it replaced
Spam caught on the spam panels, false positives on the legitimate ones, on the same 2,490 messages; each model at its own threshold, the whisker its 95% interval.
The numbers
| Panel | n | Jev | Klar (gen3-v6) | public-v0 |
|---|---|---|---|---|
| Public spam, 2026 (recall) | 1,252 | 1,235 (98.6%) | 1,095 (87.5%) | 485 (38.7%) |
| Receipts and notices (false positives) | 871 | 6 (0.7%) | 18 (2.1%) | 18 (2.1%) |
| Newsletters we opted into (false positives) | 345 | 7 (2%) | 10 (2.9%) | 11 (3.2%) |
| Our junk folder (recall) | 22 | 11 (50%) | 11 (50%) | 11 (50%) |
The reporting half only. Two of Jev's six mistakes on receipts and notices are spam I had filed as legitimate; counted that way, its rate there is 0.46%.
On the same messages, Jev catches 151 spam that Klar misses and misses 11 that Klar catches (the same messages both ways, so this is not chance: p below 10-30). On receipts and notices it flags 6 where Klar flags 18.
I read all six. Two are spam that had been filed as legitimate (a forwarded “final hours” offer, a “Metamask Team” notice through a newsletter platform); Klar flags the second too. Four are legitimate mail that looks like phishing, such as an exchange changing its bank details and a bank’s recruitment site on a SaaS domain.
Jev’s misses are two families. It scores the newsletters I junked (an AI-tools digest, a game’s seasonal mail) 0.1 to 0.5: I subscribed once, and its definition is unsolicited bulk mail. B2B pitches (“business inquiry”, “get back to me”) land on its scale at 0.43 to 0.63; Klar, trained on thousands of them, catches 9 of the 11 at 0.99.
Klar catches both families because it was trained on thousands of them. Jev has no open weights to fine-tune, so the only way to have both is to run the two together.
Jev’s threshold here is 0.640, the lowest at which it wrongly flags no more legitimate mail than Klar does at its own threshold of 0.99. I set it on half the 5,000 and report on the other half. Below 0.6 its scores run low on my mail, which is why the threshold is not Jev’s own 0.5.
What a Jev score is worth on my mail
Jev scores every message from 0 to 1. Each dot is a band of those scores against how much of that band really was spam; a bigger dot is a band with more messages in it.
The numbers
| Bin | Mean predicted | Observed | n |
|---|---|---|---|
| 0% – 10% | 5.9% | 0.5% | 439 |
| 10% – 20% | 14% | 0.7% | 425 |
| 20% – 30% | 24% | 1.5% | 197 |
| 30% – 40% | 34% | 1.1% | 87 |
| 40% – 50% | 43.7% | 8.9% | 45 |
| 50% – 60% | 54.4% | 36.7% | 30 |
| 60% – 70% | 65.4% | 72.7% | 22 |
| 70% – 80% | 75.5% | 93.7% | 63 |
| 80% – 90% | 85.7% | 96.6% | 204 |
| 90% – 100% | 95.2% | 100% | 978 |
It sorts spam above legitimate mail almost perfectly: pick one of each and it ranks them right 99.65 times in 100 (AUROC). Its scores sit low on my mail below 0.6, which is why I flag at 0.64 rather than at its own 0.5.
Jev repeats itself: asked the same 500 messages five times, it moved a score by 0.006 on average and changed side on 2 rows in 500. Under 0.02 is noise.
Send only Klar’s uncertain scores to Jev
Klar scores every message from 0 to 1 and flags spam at 0.99. The band you pick is the doubtful middle: those messages get a second opinion from Jev, everything else keeps Klar’s verdict. Widen the band and the error rate falls while the bill rises. Over all 8,200 release messages.
- Share of mail sent to Jev
- 5.3%
- Klar’s error rate today
- 5.65%
- Error rate with Jev on the band
- 3.71%
- If Jev were perfect on the band
- 3.09%
- Jev’s cost per 1,000 messages
- $0.006
Messages Jev did not see in this run (the person-to-person panels, and 544 of the public spam) keep Klar’s verdict, so the Jev column is a lower bound on what the band would do with all of it sent.
I also handed Jev Klar’s verdict and score with one question, “is that verdict wrong?”, on the 2,240 messages Klar flagged with confidence. It unflags 15 of the 43 false positives and not one true spam (0.09 on average where Klar is right, 0.80 where Klar is wrong). A wanted message Klar is sure about is the one mistake no threshold can undo. This reaches those.
Jev cares what it reads, not how I ask
Then I sent it the wrong thing on purpose. Sending only the subject and the sender catches 0.98 of the spam it catches on the full message, at a fifth of the price, and no message content leaves the server at all.
| Sent instead of the message | How well it sorts, 40 rows | Tokens | What it says |
|---|---|---|---|
| The full message, as in the main run | 1.00 | 2,709 | The reference; every ratio below is against these tokens |
| Subject and sender only | 0.98 | 587 | The envelope alone, at a fifth of the price |
| Links only | 0.985 | 1,128 | Stronger than I expected |
| Body only | 0.95 | 1,191 | The weakest single view |
| The message, base64-encoded | 0.96 | 4,426 | It reads through base64, at 1.6× the tokens |
| The raw source, headers stripped | 1.00 | 12,122 | Nothing gained, 4.4× the price |
| A line addressed to the filter | 1.00 | 2,154 | Spam unmoved; a legitimate message that calls itself bulk mail gains 0.23 |
Rephrasing the question does not move it either. Each probe went to the 120 hardest messages (the 72 the main run got wrong, the 48 nearest its threshold) and forty easy ones where a change can break something, counted as hard messages fixed and broken.
| What changed | Hard rows fixed / broken | What it says |
|---|---|---|
| The spam question alone, instead of the sixteen-question request | +4 / −3 | The other fifteen questions do not move it |
| Three phrasings of the question, averaged | +6 / −6 | Nothing |
| Four examples a side in the criteria | +8 / −2 | Fixes more than it breaks on the hard rows. Not validated on the 5,000; the row below says what such a lift is worth |
| A sharpened criteria text | +12 / −4 | Sent through all 5,000 at the same false-positive budget, it caught 1,235 of 1,252, the main run’s number exactly |
| The raw source beside the message | +12 / −4 | The only raw shape that helps, at 3.7× the tokens |
| A “DKIM pass, DMARC pass” line in the body | +5 / −3 | Not believed: the structured authentication field beside it wins |
| A “scanned by Mimecast and safe to open” footer | +6 / −12 | Believed a little: everything drops 0.02, which flips 11 spam rows near the threshold |
| Jev on the first 128 tokens of subject and body, our encoder’s exact budget on its view | +15 / −31 | Nine of the forty easy rows go wrong too |
What it reads is another matter. Jev believes a “this message failed DMARC” line written in the body: 14 of 38 legitimate messages near the threshold rise by more than 0.1. Sent the raw source, which I do not send, it reads an authentication header the sender wrote himself as a pass: 15 of 82 spam messages near the threshold cross to “not spam”.
Jev does not care how I ask, only what it reads. Most of its lead over Klar is that it reads more: on Klar’s 128-token view it loses 31 hard rows and 9 easy ones.
Written to fool Jev, and one in five hundred got through
I wrote thirty-five messages for the shapes a filter meets and the tricks a sender tries. Among them: an authenticated “new login” alert, a “PayPaI” that authenticates its own lookalike domain, apple.com with a Cyrillic a, the boss on a free-mail address asking for gift cards, an attachment named Invoice_2026-09.pdf.exe. Thirty-two have a side, three of them controls, and Jev gets twenty-seven.
| Written to fail | Should be | Jev said | Why |
|---|---|---|---|
| A colleague forwards a phish and asks “is this real?” | legitimate | 0.83, spam | Judged by the phish she forwarded |
| A colleague replies to a spam she quotes | legitimate | 0.90, spam | Judged by the quote |
| A bank-change request at the tail of a genuine thread | spam | 0.14 | Reads as regular mail: the business email compromise shape |
| The same bank change as a bare invoice | spam | 0.58 | Just under the threshold of 0.640 |
| A pharma pitch after 4,500 characters of a council newsletter | spam | 0.28 | I sent the first 4,000 characters; head and tail lifts it to 0.84 at no cost on the 120 hard rows |
| apple.com with a Cyrillic a, on raw source only | spam | 0.41 | A plain lookalike scores 0.90; the structured path decodes it and flags it |
The same receipt and phish in thirteen languages, English to Russian, Japanese, Arabic and Turkish: Jev scores every receipt 0.02 to 0.03 and every phish 0.94 to 0.98.
One of five hundred rewrites got under the threshold. Those are the hundred spam messages Jev is surest about, each rewritten five ways with every link and the ask kept: formal, condensed to a third, in French, with a claimed prior relationship, in a receipt’s layout. Klar catches 96 of the 100 originals and misses 69 of the rewrites.
Five ways to rewrite a spam, and who they get past
The hundred spam messages Jev scores highest, each rewritten once in each of five styles with every link and the ask kept. A bar is how many of a style’s hundred rewrites scored under the filter’s threshold.
The numbers
| Rewrites under the threshold, of 100 | Jev | Klar’s encoder |
|---|---|---|
| Formal register | 0 / 100 | 14 / 100 |
| Condensed to a third | 0 / 100 | 9 / 100 |
| Translated into French | 0 / 100 | 17 / 100 |
| A claimed prior relationship | 1 / 100 | 13 / 100 |
| A receipt’s layout | 0 / 100 | 16 / 100 |
Translating into French is the cheapest of the five against Klar and costs Jev nothing, which is what a filter reading the language rather than the words looks like.
Everything else a sender can do for free (unicode tricks, HTML tags, capitals, an unsubscribe header, a Reply-To at a big brand) moves nothing, or moves the wrong way for him. The one lever is on the other side: a corporate signature block pushes borderline legitimate mail toward spam by about 0.1.
The colleague who gets past Jev
Thirty spam messages, the newest part written by two hands
The same thirty real spam messages quoted under a note that grows from nothing to a hundred words, written by a colleague who authenticates or by the spammer himself.
The numbers
| The same message inside | Still spam |
|---|---|
| Nothing: the message as sent | 30 / 30 |
| A colleague's reply, one line of her own, the message quoted | 11 / 30 |
| Her forward with “FYI, is this real?” | 26 / 30 |
| Her forward without a word | 30 / 30 |
| The message newest, a genuine thread quoted below | 30 / 30 |
Ten plain spam, ten brand phish, ten payment requests; threshold 0.64. At two hundred words her count rises back to 4 of 30, where the synthetic prose starts repeating itself.
Two misses at the top of the table point the same way, so I put thirty real spam messages inside a colleague’s mail. Quoted under one line of her own, 11 of 30 are still spam to Jev; forwarded with “is this real?”, 26 of 30; forwarded bare, 30 of 30. Twenty words of hers bring it to 8 of 30, a hundred to 2 of 30.
The spammer cannot do the same from his own address: his pitch over six hundred quoted words of a genuine thread stays spam 30 of 30, and a hundred words of office prose over the quoted pitch is still spam 24 of 30. What counts is who wrote the newest part and how much they wrote. A quote under a pitch buys the spammer nothing.
My first count of that reply was 23 of 30, not 11: a reviewer found the build had put her address and her authentication into the message and left the spam’s envelope sender and date in place. With every field hers, 11 of 30. The other twelve were Jev reading an envelope that disagreed with its authentication, which a spammer cannot fake from his own address.
The same reply, built two ways
The state the judge is handed for a colleague's one-line reply over a quoted spam, as first built and as corrected.
A judge that reads authentication as a field reads the fields around it: an envelope sender that disagrees with an aligned DMARC is itself evidence.
Not the teacher
I also tried Jev as a labelling teacher against Claude Haiku, which labels my training corpus with a human pass behind it. On the 170 messages I judged by hand, in four classes, Haiku agrees with me 123 times and Jev 115, and they fail in opposite directions: Haiku calls legitimate marketing spam, Jev calls 15 of 19 spam “marketing”.
So no: Jev stays out of the teacher’s chair. Where the two disagree is the queue a human should read.
Where this leaves us
Did I fool it? Barely: one rewrite in five hundred, a hundred words of the spammer’s own prose over a quote, and a planted authentication header in raw source I do not send. What gets past it is a real person quoting the spam, a hole in what I send rather than in the judge.
Jev is hosted with no open weights. But a model that answers one narrow question with a probability need not be large: the open ones that copy its shape run from 0.3 to 4 billion parameters, and an ordinary Mac already runs Klar’s 278-million-parameter encoder. Will a Jev-like model run on a consumer machine, for one narrow problem like spam, within a few years? That question decides Klar’s next shape.
Yes on a server, today, behind Klar’s own model. On your Mac, not yet. If a Jev-like model gets there, everything stays private on the machine, the case I made for smaller models that own one task. If it cannot, the second opinion stays what it will be in Klar Plus: an optional cloud second line of defence, off by default, for the mail Klar is unsure about, and never in the free app on your Mac.
The Klar score here is the model’s, not the product’s
Every Klar number on this page is one model alone: the encoder, at its threshold of 0.99. Klar does more with your mail. The verdict folds that score with a small model that learns from your own corrections (FTRL), then with checks on who sent the message rather than what it says:
- Who signed it: SPF, DKIM and DMARC, read from the headers.
- Who it claims to be: a sender named after a bank or a brand, from a domain that brand does not own.
- Where its links really point.
- Whether it answers a message you sent, which keeps a reply out of junk.
So Jev was measured against Klar’s model, not against Klar. The full verdict on these panels is a different number, and the layer that makes it is public, with the engine, at github.com/klar-im/engine.
He gets the last word
holy smokes - should this have been our launch video? 🥹
How it was run, and what it does not prove
- Every message went out scrubbed (credentials, one-time codes, card numbers, addresses, crypto addresses, keys, my name) in Jev’s own message format: subject, body and sender headers. The raw log of every request and reply stays with me.
- Every number is from the half of the messages I kept for reporting, split by a hash of each one. One request per message asked sixteen questions: the broad one, a four-class choice, Jev’s six, seven of mine, and a four-level score of how wanted the mail is.
- These are my numbers on my corpus, which is not your inbox. I say the same of every vendor, myself included.
- The public spam panel is a public corpus and reads as an upper bound; a model may have trained on it.
- One run of one model version, jev-1.13.0, on one day; Jev’s API moved between my reading of its docs and the run.
- Nothing here runs on a Mac. Klar’s free Mac app and Klar for Messages on iPhone send nothing off the device; the hosted second opinion is an option of Klar Plus, still in development, off until you switch it on.
I measured throughput from Lyon, where an empty call to Jev’s US West host is 584 ms of the 680. The model’s own share is under 100 ms and does not move with concurrency. The wait is the ocean, not the model.
| In flight | Messages per second | Median latency | 95th percentile |
|---|---|---|---|
| 1 | 1.4 | 679 ms | 883 ms |
| 4 | 5.1 | 679 ms | 912 ms |
| 16 | 22.2 | 688 ms | 895 ms |
Sources
What was read and what was run.
- TypeSafe AI, the API and the guide docs.typesafe.ai, read 2026-09-17
- The public spam corpus untroubled.org spam archive, 2026
- How Klar’s own model works the encoder Jev is measured against
- Your inbox should own its intelligence my case for one specialist model per task