Test Jev on news relevance.
Jev caught every relevant article in all three trials and cost less than the other models. Run it beside the current model first, without using its answer. Don't use it for the other tasks yet.
Run one live test
Run Jev beside the current model for news relevance. Keep its answer out of the product until a separate review.
News relevance only
Jev recorded 100% recall and 81.46-82.00% precision in the three trials. Run it on the same articles as the current model, but keep using the current model's answer. Record disagreements, speed, usage, cost, skipped requests, and the exact inputs.
Before Jev can decide live results, test it again on a fixed reviewed sample. It must find every relevant article, keep precision at or above 80%, and account for every request.
Job relevance, tier, sentiment
The models missed the required accuracy, prompts were too long, or there were no human-reviewed answers to test against.
Grouping, ranking, daily themes
Choosing between two items doesn't prove that a model can group everything, rank a full list, or write a useful report.
Confidence
Four reviewed examples aren't enough to justify changing the live model.
News relevance
The same 174 reviewed cases were run three times. Precision and recall show the lowest and highest result.
Jev
Gemini 2.5 Flash
Gemini 3.8 Flash
| Model | Correct by trial | Requests | Tokens in / out | Cost |
|---|---|---|---|---|
| Jev | 146, 147, 146 / 174 | 522 | 1,095,921 / 11,484 | $0.046028682 |
| Gemini 2.5 Flash | 142, 141, 140 / 174 | 522 | 865,962 / 2,610 | $0.266313600 |
| Gemini 3.8 Flash | 153, 152, 152 / 174 | 522 | 884,232 / 368,266 | $2.044171500 |
Job relevance
The main test checked all six current hiring rules together on 121 reviewed cases. Every model missed at least one required accuracy check.
Jev
Gemini 2.5 Flash
Gemini 3.8 Flash
News grouping
The test asked whether 23 article pairs belonged together. It did not build complete groups.
Jev
Gemini 2.5 Flash
Gemini 3.8 Flash
Ranking
The test asked which article should come first in 28 pairs. It did not rank a full report.
Jev
Gemini 2.5 Flash
Gemini 3.8 Flash
Tier
The test turned two yes-or-no answers into a tier for 50 cases. It did not create complete article assessments.
Jev
Gemini 2.5 Flash
Gemini 3.8 Flash
Confidence
Only four reviewed cases were available. That is enough to describe these runs, not enough to choose a model.
Jev
Gemini 2.5 Flash
Gemini 3.8 Flash
Daily theme
The test checked 27 pairs of possible themes. It did not write or review a complete daily report.
Jev
Gemini 2.5 Flash
Gemini 3.8 Flash
Sentiment
No one had reviewed and fixed the correct sentiment labels before this research. There is no fair model comparison to show.
Jev
No reviewed labels were available, so there is no measured result for this test.
Gemini 2.5 Flash
No reviewed labels were available, so there is no measured result for this test.
Gemini 3.8 Flash
No reviewed labels were available, so there is no measured result for this test.
Some prompts are too long
Jev allows 64,000 tokens per request and 32,000 tokens for the saved state plus the longest question. The character check below is only a rough warning.
Grouping
Every production request we checked was below the size limit.
Ranking
190 requests were too long. We haven't tested how well the fallback works on those cases.
Tier model calls
Most requests were too long. Combining questions saved tokens in one test, but we don't know whether it changes the answers.
What we still don't know
People reviewed the correct answers before testing. A model's own answer was never treated as correct just because the model produced it.
What was directly tested?
Job relevance tested all six current hiring rules together. News relevance tested one yes-or-no question on 174 articles.
We didn't test each hiring rule separately. The News test also didn't test a complete report.
Which tests were incomplete?
Grouping, ranking, tier, confidence, and daily themes were broken into small yes-or-no questions. That shows how the models compare on those questions, but not whether they can complete the real task.
What could not be evaluated?
No one had reviewed the correct sentiment for the articles or groups. We also had no reviewed answers for structured data or written summaries. Keep the current models for those tasks.
How was production cost counted?
We followed every model call used to build seven fixed report versions. They contain 1,656 separate model outputs. Every call was counted once and none were missing.
Attributed and unique physical cost are both exactly $2.032867430.
This page summarizes the fixed test results below. It contains no passwords, model responses, reports, or article text.
artifacts/jev-research-v1/jev-research-snapshot-v1.json
SHA-256: 8a889bef56d63c071803fe476041a0c5c4414bafc10e7bc6e474d059208e779f