Test Jev on news relevance.
Jev caught every relevant article in all three trials and cost less than the other models. Run it beside the current model first, without using its answer. Don't use it for the other tasks yet.
Run one live test
Run Jev beside the current model for news relevance. Keep its answer out of the product until a separate review.
News relevance only
Jev was the only model that passed the rules set before this test. Run it on the same articles as the current model, but keep using the current model's answer. Record disagreements, speed, usage, cost, skipped requests, and the exact inputs.
Before Jev can decide live results, test it again on a fixed reviewed sample. It must find every relevant article, keep precision at or above 80%, and account for every request.
Job relevance, tier, sentiment
The models missed the required accuracy, prompts were too long, or there were no human-reviewed answers to test against.
Grouping, ranking, daily themes
Choosing between two items doesn't prove that a model can group everything, rank a full list, or write a useful report.
Confidence
Four reviewed examples aren't enough to justify changing the live model.
News relevance
The same 174 reviewed cases were run three times. Precision and recall show the lowest and highest result.
Jev
Gemini 2.5 Flash
Gemini 3.8 Flash
| Model | Correct by trial | Requests | Tokens in / out | Cost | Best trade-off |
|---|---|---|---|---|---|
| Jev | 146, 147, 146 / 174 | 522 | 1,095,921 / 11,484 | $0.046028682 | Yes |
| Gemini 2.5 Flash | 142, 141, 140 / 174 | 522 | 865,962 / 2,610 | $0.266313600 | No |
| Gemini 3.8 Flash | 153, 152, 152 / 174 | 522 | 884,232 / 368,266 | $2.044171500 | No |
All three models failed at least one required check on the main 121-case test. They passed the smaller 25-case ATS test, but that isn't enough. Keep the current model and separately test a version that asks each hiring rule as its own question.
Other News tasks
These tests asked small yes-or-no questions. They compare the models, but don't show that a model can finish the full job.
| Task | Model | Correct by trial | Cost / 1,000 | Typical / slow ms | Same answer | Best overall |
|---|---|---|---|---|---|---|
| Grouping | Jev | 21, 21, 20 / 23 | $0.1167 | 646 / 839.8 | 22 / 23 | Yes |
| Grouping | Gemini 2.5 | 21, 21, 21 / 23 | $0.6631 | 670 / 906.5 | 23 / 23 | Yes |
| Grouping | Gemini 3.8 | 19, 19, 20 / 23 | $3.3287 | 1,658 / 4,695.3 | 22 / 23 | No |
| Ranking | Jev | 21, 22, 20 / 28 | $0.3321 | 816.5 / 1,057.3 | 26 / 28 | Yes |
| Ranking | Gemini 2.5 | 12, 12, 12 / 28 | $1.8752 | 678.5 / 1,174.2 | 28 / 28 | No |
| Ranking | Gemini 3.8 | 27, 27, 25 / 28 | $6.1984 | 2,467.5 / 4,863.3 | 24 / 28 | Yes |
| Tier | Jev | 23, 23, 23 / 50 | $0.6031 | 1,015.5 / 1,083.1 | 50 / 50 | Yes |
| Tier | Gemini 2.5 | 25, 25, 25 / 50 | $2.6005 | 769.5 / 1,420.35 | 50 / 50 | Yes |
| Tier | Gemini 3.8 | 30, 27, 29 / 50 | $11.0901 | 3,238 / 5,005.5 | 47 / 50 | Yes |
| Confidence | Jev | 4, 4, 4 / 4 | $0.0887 | 625 / 745.6 | 4 / 4 | Yes |
| Confidence | Gemini 2.5 | 3, 3, 3 / 4 | $0.4620 | 562.5 / 844.35 | 4 / 4 | No |
| Confidence | Gemini 3.8 | 4, 4, 4 / 4 | $4.2838 | 3,492 / 3,816.65 | 4 / 4 | No |
| Daily theme | Jev | 20, 20, 20 / 27 | $0.0315 | 622 / 694.8 | 27 / 27 | Yes |
| Daily theme | Gemini 2.5 | 21, 20, 20 / 27 | $0.1631 | 551 / 916.8 | 24 / 27 | No |
| Daily theme | Gemini 3.8 | 17, 17, 17 / 27 | $1.6034 | 1,926 / 3,125.4 | 27 / 27 | No |
No other tested model was better on accuracy, cost, and slow-request speed at the same time. This still doesn't show that the model can do the full production task.
Some prompts are too long
Jev allows 64,000 tokens per request and 32,000 tokens for the saved state plus the longest question. The character check below is only a rough warning.
Grouping
Every production request we checked was below the size limit.
Ranking
190 requests were too long. We haven't tested how well the fallback works on those cases.
Tier model calls
Most requests were too long. Combining questions saved tokens in one test, but we don't know whether it changes the answers.
What we still don't know
People reviewed the correct answers before testing. A model's own answer was never treated as correct just because the model produced it.
What was directly tested?
Job relevance tested all six current hiring rules together. News relevance tested one yes-or-no question on 174 articles.
We didn't test each hiring rule separately. The News test also didn't test a complete report.
Which tests were incomplete?
Grouping, ranking, tier, confidence, and daily themes were broken into small yes-or-no questions. That shows how the models compare on those questions, but not whether they can complete the real task.
What could not be evaluated?
No one had reviewed the correct sentiment for the articles or groups. We also had no reviewed answers for structured data or written summaries. Keep the current models for those tasks.
How was production cost counted?
We followed every model call used to build seven fixed report versions. They contain 1,656 separate model outputs. Every call was counted once and none were missing.
Attributed and unique physical cost are both exactly $2.032867430.
This page summarizes the fixed test results below. It contains no passwords, model responses, reports, or article text.
artifacts/jev-research-v1/jev-research-snapshot-v1.json
SHA-256: 8a889bef56d63c071803fe476041a0c5c4414bafc10e7bc6e474d059208e779f