← Klar blog

Is Jev the future of spam filtering?

Titles ending in a question mark usually mean the answer is no. For this one, it could very well be a yes.

This month TypeSafe AI came out of stealth with Jev, a hosted model that cannot chat: it takes a typed question and returns a probability. “Is this email spam?” is the obvious question to ask it, so I pitted it against Klar’s own model. It is billed per token, which came to 11 cents per thousand of my messages.

Jev catches more of my spam than the model I ship, for eleven cents per thousand messages. It is not a fair fight: Klar’s filter runs on any Apple silicon Mac and nothing it reads leaves the machine. Jev runs in a data centre.

Jev catches more, and misses newsletters and cold pitches

I took the 5,000 messages I measure Klar on: 2,500 spam (2,456 from a public 2026 corpus, 44 from my own junk folder) and 2,500 legitimate, drawn in proportion from my 2025 receipts, notices and opted-in newsletters. Person-to-person mail stayed on the Mac. Three models read them: Jev, the model Klar ships today, and the one it replaced last year.

Jev, the model Klar ships, and the one it replaced

Spam caught on the spam panels, false positives on the legitimate ones, on the same 2,490 messages; each model at its own threshold, the whisker its 95% interval.

0%25%50%75%100%Public spam, 2026 (recall, n = 1,252)98.6% · Jev87.5% · Klar (gen3-v6)38.7% · public-v0Receipts and notices (false positives, n = 871)0.7% · Jev2.1% · Klar (gen3-v6)2.1% · public-v0Newsletters we opted into (false positives, n = 345)2% · Jev2.9% · Klar (gen3-v6)3.2% · public-v0
The numbers
PanelnJevKlar (gen3-v6)public-v0
Public spam, 2026 (recall)1,2521,235 (98.6%)1,095 (87.5%)485 (38.7%)
Receipts and notices (false positives)8716 (0.7%)18 (2.1%)18 (2.1%)
Newsletters we opted into (false positives)3457 (2%)10 (2.9%)11 (3.2%)
Our junk folder (recall)2211 (50%)11 (50%)11 (50%)

The reporting half only. Two of Jev's six mistakes on receipts and notices are spam I had filed as legitimate; counted that way, its rate there is 0.46%.

On the same messages, Jev catches 151 spam that Klar misses and misses 11 that Klar catches (the same messages both ways, so this is not chance: p below 10-30). On receipts and notices it flags 6 where Klar flags 18.

I read all six. Two are spam that had been filed as legitimate (a forwarded “final hours” offer, a “Metamask Team” notice through a newsletter platform); Klar flags the second too. Four are legitimate mail that looks like phishing, such as an exchange changing its bank details and a bank’s recruitment site on a SaaS domain.

Jev’s misses are two families. It scores the newsletters I junked (an AI-tools digest, a game’s seasonal mail) 0.1 to 0.5: I subscribed once, and its definition is unsolicited bulk mail. B2B pitches (“business inquiry”, “get back to me”) land on its scale at 0.43 to 0.63; Klar, trained on thousands of them, catches 9 of the 11 at 0.99.

Klar catches both families because it was trained on thousands of them. Jev has no open weights to fine-tune, so the only way to have both is to run the two together.

Jev’s threshold here is 0.640, the lowest at which it wrongly flags no more legitimate mail than Klar does at its own threshold of 0.99. I set it on half the 5,000 and report on the other half. Below 0.6 its scores run low on my mail, which is why the threshold is not Jev’s own 0.5.

What a Jev score is worth on my mail

Jev scores every message from 0 to 1. Each dot is a band of those scores against how much of that band really was spam; a bigger dot is a band with more messages in it.

0%0%25%25%50%50%75%75%100%100%5.9% → 0.5%, n = 43914% → 0.7%, n = 42524% → 1.5%, n = 19734% → 1.1%, n = 8743.7% → 8.9%, n = 4554.4% → 36.7%, n = 3065.4% → 72.7%, n = 2275.5% → 93.7%, n = 6385.7% → 96.6%, n = 20495.2% → 100%, n = 978Predicted P(spam)Observed spam rate
The numbers
BinMean predictedObservedn
0% – 10%5.9%0.5%439
10% – 20%14%0.7%425
20% – 30%24%1.5%197
30% – 40%34%1.1%87
40% – 50%43.7%8.9%45
50% – 60%54.4%36.7%30
60% – 70%65.4%72.7%22
70% – 80%75.5%93.7%63
80% – 90%85.7%96.6%204
90% – 100%95.2%100%978

It sorts spam above legitimate mail almost perfectly: pick one of each and it ranks them right 99.65 times in 100 (AUROC). Its scores sit low on my mail below 0.6, which is why I flag at 0.64 rather than at its own 0.5.

Jev repeats itself: asked the same 500 messages five times, it moved a score by 0.006 on average and changed side on 2 rows in 500. Under 0.02 is noise.

Send only Klar’s uncertain scores to Jev

Klar scores every message from 0 to 1 and flags spam at 0.99. The band you pick is the doubtful middle: those messages get a second opinion from Jev, everything else keeps Klar’s verdict. Widen the band and the error rate falls while the bill rises. Over all 8,200 release messages.

From a score of
Up to
Share of mail sent to Jev
5.3%
Klar’s error rate today
5.65%
Error rate with Jev on the band
3.71%
If Jev were perfect on the band
3.09%
Jev’s cost per 1,000 messages
$0.006

Messages Jev did not see in this run (the person-to-person panels, and 544 of the public spam) keep Klar’s verdict, so the Jev column is a lower bound on what the band would do with all of it sent.

I also handed Jev Klar’s verdict and score with one question, “is that verdict wrong?”, on the 2,240 messages Klar flagged with confidence. It unflags 15 of the 43 false positives and not one true spam (0.09 on average where Klar is right, 0.80 where Klar is wrong). A wanted message Klar is sure about is the one mistake no threshold can undo. This reaches those.

Jev cares what it reads, not how I ask

Then I sent it the wrong thing on purpose. Sending only the subject and the sender catches 0.98 of the spam it catches on the full message, at a fifth of the price, and no message content leaves the server at all.

Sent instead of the messageHow well it sorts, 40 rowsTokensWhat it says
The full message, as in the main run1.002,709The reference; every ratio below is against these tokens
Subject and sender only0.98587The envelope alone, at a fifth of the price
Links only0.9851,128Stronger than I expected
Body only0.951,191The weakest single view
The message, base64-encoded0.964,426It reads through base64, at 1.6× the tokens
The raw source, headers stripped1.0012,122Nothing gained, 4.4× the price
A line addressed to the filter1.002,154Spam unmoved; a legitimate message that calls itself bulk mail gains 0.23

Rephrasing the question does not move it either. Each probe went to the 120 hardest messages (the 72 the main run got wrong, the 48 nearest its threshold) and forty easy ones where a change can break something, counted as hard messages fixed and broken.

What changedHard rows fixed / brokenWhat it says
The spam question alone, instead of the sixteen-question request+4 / −3The other fifteen questions do not move it
Three phrasings of the question, averaged+6 / −6Nothing
Four examples a side in the criteria+8 / −2Fixes more than it breaks on the hard rows. Not validated on the 5,000; the row below says what such a lift is worth
A sharpened criteria text+12 / −4Sent through all 5,000 at the same false-positive budget, it caught 1,235 of 1,252, the main run’s number exactly
The raw source beside the message+12 / −4The only raw shape that helps, at 3.7× the tokens
A “DKIM pass, DMARC pass” line in the body+5 / −3Not believed: the structured authentication field beside it wins
A “scanned by Mimecast and safe to open” footer+6 / −12Believed a little: everything drops 0.02, which flips 11 spam rows near the threshold
Jev on the first 128 tokens of subject and body, our encoder’s exact budget on its view+15 / −31Nine of the forty easy rows go wrong too

What it reads is another matter. Jev believes a “this message failed DMARC” line written in the body: 14 of 38 legitimate messages near the threshold rise by more than 0.1. Sent the raw source, which I do not send, it reads an authentication header the sender wrote himself as a pass: 15 of 82 spam messages near the threshold cross to “not spam”.

Jev does not care how I ask, only what it reads. Most of its lead over Klar is that it reads more: on Klar’s 128-token view it loses 31 hard rows and 9 easy ones.

Written to fool Jev, and one in five hundred got through

I wrote thirty-five messages for the shapes a filter meets and the tricks a sender tries. Among them: an authenticated “new login” alert, a “PayPaI” that authenticates its own lookalike domain, apple.com with a Cyrillic a, the boss on a free-mail address asking for gift cards, an attachment named Invoice_2026-09.pdf.exe. Thirty-two have a side, three of them controls, and Jev gets twenty-seven.

Written to failShould beJev saidWhy
A colleague forwards a phish and asks “is this real?”legitimate0.83, spamJudged by the phish she forwarded
A colleague replies to a spam she quoteslegitimate0.90, spamJudged by the quote
A bank-change request at the tail of a genuine threadspam0.14Reads as regular mail: the business email compromise shape
The same bank change as a bare invoicespam0.58Just under the threshold of 0.640
A pharma pitch after 4,500 characters of a council newsletterspam0.28I sent the first 4,000 characters; head and tail lifts it to 0.84 at no cost on the 120 hard rows
apple.com with a Cyrillic a, on raw source onlyspam0.41A plain lookalike scores 0.90; the structured path decodes it and flags it

The same receipt and phish in thirteen languages, English to Russian, Japanese, Arabic and Turkish: Jev scores every receipt 0.02 to 0.03 and every phish 0.94 to 0.98.

One of five hundred rewrites got under the threshold. Those are the hundred spam messages Jev is surest about, each rewritten five ways with every link and the ask kept: formal, condensed to a third, in French, with a claimed prior relationship, in a receipt’s layout. Klar catches 96 of the 100 originals and misses 69 of the rewrites.

Five ways to rewrite a spam, and who they get past

The hundred spam messages Jev scores highest, each rewritten once in each of five styles with every link and the ask kept. A bar is how many of a style’s hundred rewrites scored under the filter’s threshold.

05101520Formal registerJev: 0 / 1000Klar’s encoder: 14 / 10014Condensed to a thirdJev: 0 / 1000Klar’s encoder: 9 / 1009Translated into FrenchJev: 0 / 1000Klar’s encoder: 17 / 10017A claimed prior relationshipJev: 1 / 1001Klar’s encoder: 13 / 10013A receipt’s layoutJev: 0 / 1000Klar’s encoder: 16 / 10016JevKlar’s encoderRewrites under the threshold, of 100
The numbers
Rewrites under the threshold, of 100JevKlar’s encoder
Formal register0 / 10014 / 100
Condensed to a third0 / 1009 / 100
Translated into French0 / 10017 / 100
A claimed prior relationship1 / 10013 / 100
A receipt’s layout0 / 10016 / 100

Translating into French is the cheapest of the five against Klar and costs Jev nothing, which is what a filter reading the language rather than the words looks like.

Everything else a sender can do for free (unicode tricks, HTML tags, capitals, an unsubscribe header, a Reply-To at a big brand) moves nothing, or moves the wrong way for him. The one lever is on the other side: a corporate signature block pushes borderline legitimate mail toward spam by about 0.1.

The colleague who gets past Jev

Thirty spam messages, the newest part written by two hands

The same thirty real spam messages quoted under a note that grows from nothing to a hundred words, written by a colleague who authenticates or by the spammer himself.

010203002050100A colleague replying: 0 → 30 / 30A colleague replying: 20 → 8 / 308A colleague replying: 50 → 6 / 306A colleague replying: 100 → 2 / 302A colleague replyingThe spammer himself: 0 → 30 / 30The spammer himself: 20 → 30 / 3030The spammer himself: 50 → 30 / 3030The spammer himself: 100 → 24 / 3024The spammer himselfWords of the newest part's own textStill spam, of 30
The numbers
The same message insideStill spam
Nothing: the message as sent30 / 30
A colleague's reply, one line of her own, the message quoted11 / 30
Her forward with “FYI, is this real?”26 / 30
Her forward without a word30 / 30
The message newest, a genuine thread quoted below30 / 30

Ten plain spam, ten brand phish, ten payment requests; threshold 0.64. At two hundred words her count rises back to 4 of 30, where the synthetic prose starts repeating itself.

Two misses at the top of the table point the same way, so I put thirty real spam messages inside a colleague’s mail. Quoted under one line of her own, 11 of 30 are still spam to Jev; forwarded with “is this real?”, 26 of 30; forwarded bare, 30 of 30. Twenty words of hers bring it to 8 of 30, a hundred to 2 of 30.

The spammer cannot do the same from his own address: his pitch over six hundred quoted words of a genuine thread stays spam 30 of 30, and a hundred words of office prose over the quoted pitch is still spam 24 of 30. What counts is who wrote the newest part and how much they wrote. A quote under a pitch buys the spammer nothing.

My first count of that reply was 23 of 30, not 11: a reviewer found the build had put her address and her authentication into the message and left the spam’s envelope sender and date in place. With every field hers, 11 of 30. The other twelve were Jev reading an envelope that disagreed with its authentication, which a spammer cannot fake from his own address.

The same reply, built two ways

The state the judge is handed for a colleague's one-line reply over a quoted spam, as first built and as corrected.

First build: the envelope left behindFromCamille Rey, example.comDMARCpass, alignedReturn-Paththe spam's, parcel-notify.topDatethe spam's own dateThanks, noted. I will have a look tomorrow.> On Sat, Parcel Desk wrote:> Your parcel is waiting. Pay the fee at> parcel-notify.top/t within 24 hours.23 / 30still spam, of 30Coherent: every field hersFromCamille Rey, example.comDMARCpass, alignedReturn-Pathhers, example.comDatea day after the spamThanks, noted. I will have a look tomorrow.> On Sat, Parcel Desk wrote:> Your parcel is waiting. Pay the fee at> parcel-notify.top/t within 24 hours.11 / 30still spam, of 30

A judge that reads authentication as a field reads the fields around it: an envelope sender that disagrees with an aligned DMARC is itself evidence.

Not the teacher

I also tried Jev as a labelling teacher against Claude Haiku, which labels my training corpus with a human pass behind it. On the 170 messages I judged by hand, in four classes, Haiku agrees with me 123 times and Jev 115, and they fail in opposite directions: Haiku calls legitimate marketing spam, Jev calls 15 of 19 spam “marketing”.

So no: Jev stays out of the teacher’s chair. Where the two disagree is the queue a human should read.

Where this leaves us

Did I fool it? Barely: one rewrite in five hundred, a hundred words of the spammer’s own prose over a quote, and a planted authentication header in raw source I do not send. What gets past it is a real person quoting the spam, a hole in what I send rather than in the judge.

Jev is hosted with no open weights. But a model that answers one narrow question with a probability need not be large: the open ones that copy its shape run from 0.3 to 4 billion parameters, and an ordinary Mac already runs Klar’s 278-million-parameter encoder. Will a Jev-like model run on a consumer machine, for one narrow problem like spam, within a few years? That question decides Klar’s next shape.

Yes on a server, today, behind Klar’s own model. On your Mac, not yet. If a Jev-like model gets there, everything stays private on the machine, the case I made for smaller models that own one task. If it cannot, the second opinion stays what it will be in Klar Plus: an optional cloud second line of defence, off by default, for the mail Klar is unsure about, and never in the free app on your Mac.

The Klar score here is the model’s, not the product’s

Every Klar number on this page is one model alone: the encoder, at its threshold of 0.99. Klar does more with your mail. The verdict folds that score with a small model that learns from your own corrections (FTRL), then with checks on who sent the message rather than what it says:

  • Who signed it: SPF, DKIM and DMARC, read from the headers.
  • Who it claims to be: a sender named after a bank or a brand, from a domain that brand does not own.
  • Where its links really point.
  • Whether it answers a message you sent, which keeps a reply out of junk.

So Jev was measured against Klar’s model, not against Klar. The full verdict on these panels is a different number, and the layer that makes it is public, with the engine, at github.com/klar-im/engine.

He gets the last word

holy smokes - should this have been our launch video? 🥹

Diogo Almeida @CompleteSkeptic · 23 September 2026

How it was run, and what it does not prove

  • Every message went out scrubbed (credentials, one-time codes, card numbers, addresses, crypto addresses, keys, my name) in Jev’s own message format: subject, body and sender headers. The raw log of every request and reply stays with me.
  • Every number is from the half of the messages I kept for reporting, split by a hash of each one. One request per message asked sixteen questions: the broad one, a four-class choice, Jev’s six, seven of mine, and a four-level score of how wanted the mail is.
  • These are my numbers on my corpus, which is not your inbox. I say the same of every vendor, myself included.
  • The public spam panel is a public corpus and reads as an upper bound; a model may have trained on it.
  • One run of one model version, jev-1.13.0, on one day; Jev’s API moved between my reading of its docs and the run.
  • Nothing here runs on a Mac. Klar’s free Mac app and Klar for Messages on iPhone send nothing off the device; the hosted second opinion is an option of Klar Plus, still in development, off until you switch it on.

I measured throughput from Lyon, where an empty call to Jev’s US West host is 584 ms of the 680. The model’s own share is under 100 ms and does not move with concurrency. The wait is the ocean, not the model.

In flightMessages per secondMedian latency95th percentile
11.4679 ms883 ms
45.1679 ms912 ms
1622.2688 ms895 ms

Sources

What was read and what was run.

All posts

Reclaim your inbox

Spam filter for Apple Mail. Free on the App Store.

Download on the App Store