Jev model test results

Test Jev on news relevance.

Jev caught every relevant article in all three trials and cost less than the other models. Run it beside the current model first, without using its answer. Don't use it for the other tasks yet.

Jev 1.13.0 3 trials People reviewed the answers September 20, 2026
News relevance 1.00 Recall in all 3 Jev trials
Jev trial cost $0.046 522 relevance requests
Current production cost $2.033 7 fixed News reports
Live model changes 0 Current models still decide
01 / What changes

Run one live test

Run Jev beside the current model for news relevance. Keep its answer out of the product until a separate review.

Try this next

News relevance only

Jev was the only model that passed the rules set before this test. Run it on the same articles as the current model, but keep using the current model's answer. Record disagreements, speed, usage, cost, skipped requests, and the exact inputs.

i
Work item: news-aggregator-vt6.8

Before Jev can decide live results, test it again on a fixed reviewed sample. It must find every relevant article, keep precision at or above 80%, and account for every request.

Don't change

Job relevance, tier, sentiment

The models missed the required accuracy, prompts were too long, or there were no human-reviewed answers to test against.

Keep testing

Grouping, ranking, daily themes

Choosing between two items doesn't prove that a model can group everything, rank a full list, or write a useful report.

Need more examples

Confidence

Four reviewed examples aren't enough to justify changing the live model.

02 / Best test

News relevance

The same 174 reviewed cases were run three times. Precision and recall show the lowest and highest result.

Didn't pass

Gemini 2.5 Flash

openrouter-gemini-2.5-flash
Precision80.69-80.95%
Recall95.12-96.75%
$0.2663Total cost
857 msSlow requests (p95)
Didn't pass

Gemini 3.8 Flash

openrouter-gemini-3.8-flash
Precision90.40-90.48%
Recall91.87-92.68%
$2.0442Total cost
5,485.85 msSlow requests (p95)
ModelCorrect by trialRequestsTokens in / outCostBest trade-off
Jev146, 147, 146 / 1745221,095,921 / 11,484$0.046028682Yes
Gemini 2.5 Flash142, 141, 140 / 174522865,962 / 2,610$0.266313600No
Gemini 3.8 Flash153, 152, 152 / 174522884,232 / 368,266$2.044171500No
!
Job relevance did not produce a winner

All three models failed at least one required check on the main 121-case test. They passed the smaller 25-case ATS test, but that isn't enough. Keep the current model and separately test a version that asks each hiring rule as its own question.

03 / Incomplete tests

Other News tasks

These tests asked small yes-or-no questions. They compare the models, but don't show that a model can finish the full job.

TaskModelCorrect by trialCost / 1,000Typical / slow msSame answerBest overall
GroupingJev21, 21, 20 / 23$0.1167646 / 839.822 / 23Yes
GroupingGemini 2.521, 21, 21 / 23$0.6631670 / 906.523 / 23Yes
GroupingGemini 3.819, 19, 20 / 23$3.32871,658 / 4,695.322 / 23No
RankingJev21, 22, 20 / 28$0.3321816.5 / 1,057.326 / 28Yes
RankingGemini 2.512, 12, 12 / 28$1.8752678.5 / 1,174.228 / 28No
RankingGemini 3.827, 27, 25 / 28$6.19842,467.5 / 4,863.324 / 28Yes
TierJev23, 23, 23 / 50$0.60311,015.5 / 1,083.150 / 50Yes
TierGemini 2.525, 25, 25 / 50$2.6005769.5 / 1,420.3550 / 50Yes
TierGemini 3.830, 27, 29 / 50$11.09013,238 / 5,005.547 / 50Yes
ConfidenceJev4, 4, 4 / 4$0.0887625 / 745.64 / 4Yes
ConfidenceGemini 2.53, 3, 3 / 4$0.4620562.5 / 844.354 / 4No
ConfidenceGemini 3.84, 4, 4 / 4$4.28383,492 / 3,816.654 / 4No
Daily themeJev20, 20, 20 / 27$0.0315622 / 694.827 / 27Yes
Daily themeGemini 2.521, 20, 20 / 27$0.1631551 / 916.824 / 27No
Daily themeGemini 3.817, 17, 17 / 27$1.60341,926 / 3,125.427 / 27No
i
What "best overall" means

No other tested model was better on accuracy, cost, and slow-request speed at the same time. This still doesn't show that the model can do the full production task.

04 / Input size

Some prompts are too long

Jev allows 64,000 tokens per request and 32,000 tokens for the saved state plus the longest question. The character check below is only a rough warning.

All fit
0 / 21,139

Grouping

Every production request we checked was below the size limit.

Some too long
190 / 8,367

Ranking

190 requests were too long. We haven't tested how well the fallback works on those cases.

Usually too long
197 / 228

Tier model calls

Most requests were too long. Combining questions saved tokens in one test, but we don't know whether it changes the answers.

Size check
70,000
state characters, not tokens
If the input is too long or Jev rejects it: send the original full input to the current model. Don't cut up or remove source material. Don't send the same rejected request to Jev again.
05 / Missing answers

What we still don't know

People reviewed the correct answers before testing. A model's own answer was never treated as correct just because the model produced it.

What was directly tested?

Job relevance tested all six current hiring rules together. News relevance tested one yes-or-no question on 174 articles.

We didn't test each hiring rule separately. The News test also didn't test a complete report.

Which tests were incomplete?

Grouping, ranking, tier, confidence, and daily themes were broken into small yes-or-no questions. That shows how the models compare on those questions, but not whether they can complete the real task.

What could not be evaluated?

No one had reviewed the correct sentiment for the articles or groups. We also had no reviewed answers for structured data or written summaries. Keep the current models for those tasks.

How was production cost counted?

We followed every model call used to build seven fixed report versions. They contain 1,656 separate model outputs. Every call was counted once and none were missing.

Attributed and unique physical cost are both exactly $2.032867430.

Source data

This page summarizes the fixed test results below. It contains no passwords, model responses, reports, or article text.

artifacts/jev-research-v1/jev-research-snapshot-v1.json SHA-256: 8a889bef56d63c071803fe476041a0c5c4414bafc10e7bc6e474d059208e779f