Jev model test results

Test Jev on news relevance.

Jev caught every relevant article in all three trials and cost less than the other models. Run it beside the current model first, without using its answer. Don't use it for the other tasks yet.

Jev 1.13.0 3 trials People reviewed the answers September 20, 2026
News relevance 1.00 Recall in all 3 Jev trials
Jev trial cost $0.046 522 relevance requests
Current production cost $2.033 7 fixed News reports
Live model changes 0 Current models still decide
01 / What changes

Run one live test

Run Jev beside the current model for news relevance. Keep its answer out of the product until a separate review.

Try this next

News relevance only

Jev recorded 100% recall and 81.46-82.00% precision in the three trials. Run it on the same articles as the current model, but keep using the current model's answer. Record disagreements, speed, usage, cost, skipped requests, and the exact inputs.

i
Work item: news-aggregator-vt6.8

Before Jev can decide live results, test it again on a fixed reviewed sample. It must find every relevant article, keep precision at or above 80%, and account for every request.

Job relevance, tier, sentiment

The models missed the required accuracy, prompts were too long, or there were no human-reviewed answers to test against.

Grouping, ranking, daily themes

Choosing between two items doesn't prove that a model can group everything, rank a full list, or write a useful report.

Confidence

Four reviewed examples aren't enough to justify changing the live model.

02 / Best test

News relevance

The same 174 reviewed cases were run three times. Precision and recall show the lowest and highest result.

Jev

typesafe-jev / jev-1.13.0
Precision81.46-82.00%
Recall100.00%
$0.0460Total cost
522Requests
824.7 msSlow requests (p95)

Gemini 2.5 Flash

openrouter-gemini-2.5-flash
Precision80.69-80.95%
Recall95.12-96.75%
$0.2663Total cost
522Requests
857 msSlow requests (p95)

Gemini 3.8 Flash

openrouter-gemini-3.8-flash
Precision90.40-90.48%
Recall91.87-92.68%
$2.0442Total cost
522Requests
5,485.85 msSlow requests (p95)
ModelCorrect by trialRequestsTokens in / outCost
Jev146, 147, 146 / 1745221,095,921 / 11,484$0.046028682
Gemini 2.5 Flash142, 141, 140 / 174522865,962 / 2,610$0.266313600
Gemini 3.8 Flash153, 152, 152 / 174522884,232 / 368,266$2.044171500
03 / Job Finder

Job relevance

The main test checked all six current hiring rules together on 121 reviewed cases. Every model missed at least one required accuracy check.

Jev

jev-1.13.0
Correct by trial104-106 / 121
$0.1212Total cost
1,530Requests
4,300.9 msSlow requests (p95)

Gemini 2.5 Flash

google-gemini-2.5-flash
Correct by trial104-106 / 121
$0.8058Total cost
1,538Requests
4,640 msSlow requests (p95)

Gemini 3.8 Flash

google-gemini-3.8-flash
Correct by trial104-106 / 121
$4.3303Total cost
1,569Requests
32,881.2 msSlow requests (p95)
04 / News task

News grouping

The test asked whether 23 article pairs belonged together. It did not build complete groups.

Jev

typesafe-jev
Correct by trial20-21 / 23
$0.1167Cost per 1,000
69Requests
839.8 msSlow requests (p95)

Gemini 2.5 Flash

openrouter-gemini-2.5-flash
Correct by trial21 / 23
$0.6631Cost per 1,000
69Requests
906.5 msSlow requests (p95)

Gemini 3.8 Flash

openrouter-gemini-3.8-flash
Correct by trial19-20 / 23
$3.3287Cost per 1,000
69Requests
4,695.3 msSlow requests (p95)
05 / News task

Ranking

The test asked which article should come first in 28 pairs. It did not rank a full report.

Jev

typesafe-jev
Correct by trial20-22 / 28
$0.3321Cost per 1,000
84Requests
1,057.3 msSlow requests (p95)

Gemini 2.5 Flash

openrouter-gemini-2.5-flash
Correct by trial12 / 28
$1.8752Cost per 1,000
84Requests
1,174.2 msSlow requests (p95)

Gemini 3.8 Flash

openrouter-gemini-3.8-flash
Correct by trial25-27 / 28
$6.1984Cost per 1,000
84Requests
4,863.3 msSlow requests (p95)
06 / News task

Tier

The test turned two yes-or-no answers into a tier for 50 cases. It did not create complete article assessments.

Jev

typesafe-jev
Correct by trial23 / 50
$0.6031Cost per 1,000
150Requests
1,083.1 msSlow requests (p95)

Gemini 2.5 Flash

openrouter-gemini-2.5-flash
Correct by trial25 / 50
$2.6005Cost per 1,000
150Requests
1,420.4 msSlow requests (p95)

Gemini 3.8 Flash

openrouter-gemini-3.8-flash
Correct by trial27-30 / 50
$11.0901Cost per 1,000
150Requests
5,005.5 msSlow requests (p95)
07 / News task

Confidence

Only four reviewed cases were available. That is enough to describe these runs, not enough to choose a model.

Jev

typesafe-jev
Correct by trial4 / 4
$0.0887Cost per 1,000
12Requests
745.6 msSlow requests (p95)

Gemini 2.5 Flash

openrouter-gemini-2.5-flash
Correct by trial3 / 4
$0.4620Cost per 1,000
12Requests
844.4 msSlow requests (p95)

Gemini 3.8 Flash

openrouter-gemini-3.8-flash
Correct by trial4 / 4
$4.2838Cost per 1,000
12Requests
3,816.7 msSlow requests (p95)
08 / News task

Daily theme

The test checked 27 pairs of possible themes. It did not write or review a complete daily report.

Jev

typesafe-jev
Correct by trial20 / 27
$0.0315Cost per 1,000
81Requests
694.8 msSlow requests (p95)

Gemini 2.5 Flash

openrouter-gemini-2.5-flash
Correct by trial20-21 / 27
$0.1631Cost per 1,000
81Requests
916.8 msSlow requests (p95)

Gemini 3.8 Flash

openrouter-gemini-3.8-flash
Correct by trial17 / 27
$1.6034Cost per 1,000
81Requests
3,125.4 msSlow requests (p95)
09 / News task

Sentiment

No one had reviewed and fixed the correct sentiment labels before this research. There is no fair model comparison to show.

Jev

typesafe-jev

No reviewed labels were available, so there is no measured result for this test.

-Accuracy
-Cost
0Test requests

Gemini 2.5 Flash

openrouter-gemini-2.5-flash

No reviewed labels were available, so there is no measured result for this test.

-Accuracy
-Cost
0Test requests

Gemini 3.8 Flash

openrouter-gemini-3.8-flash

No reviewed labels were available, so there is no measured result for this test.

-Accuracy
-Cost
0Test requests
10 / Input size

Some prompts are too long

Jev allows 64,000 tokens per request and 32,000 tokens for the saved state plus the longest question. The character check below is only a rough warning.

0 / 21,139

Grouping

Every production request we checked was below the size limit.

190 / 8,367

Ranking

190 requests were too long. We haven't tested how well the fallback works on those cases.

197 / 228

Tier model calls

Most requests were too long. Combining questions saved tokens in one test, but we don't know whether it changes the answers.

Size check
70,000
state characters, not tokens
If the input is too long or Jev rejects it: send the original full input to the current model. Don't cut up or remove source material. Don't send the same rejected request to Jev again.
11 / Missing answers

What we still don't know

People reviewed the correct answers before testing. A model's own answer was never treated as correct just because the model produced it.

What was directly tested?

Job relevance tested all six current hiring rules together. News relevance tested one yes-or-no question on 174 articles.

We didn't test each hiring rule separately. The News test also didn't test a complete report.

Which tests were incomplete?

Grouping, ranking, tier, confidence, and daily themes were broken into small yes-or-no questions. That shows how the models compare on those questions, but not whether they can complete the real task.

What could not be evaluated?

No one had reviewed the correct sentiment for the articles or groups. We also had no reviewed answers for structured data or written summaries. Keep the current models for those tasks.

How was production cost counted?

We followed every model call used to build seven fixed report versions. They contain 1,656 separate model outputs. Every call was counted once and none were missing.

Attributed and unique physical cost are both exactly $2.032867430.

Source data

This page summarizes the fixed test results below. It contains no passwords, model responses, reports, or article text.

artifacts/jev-research-v1/jev-research-snapshot-v1.json SHA-256: 8a889bef56d63c071803fe476041a0c5c4414bafc10e7bc6e474d059208e779f