Skip to content
latentSource

Jev ranked jobs better. MiniLM found about as many strong matches.

I compared Jev with three local embedding models on 237 job postings. Jev led on graded relevance, but a stricter definition of a good match changed the conclusion.

View agent and demonstration
·8 min read
AI evaluationSemantic searchJevEmbeddings
Share
Better rankings, similar strong matches

Recorded demonstration. Unmute the player to listen.

I wanted to know whether Jev could find better jobs than ordinary embeddings. I already had a working search interface, real postings, and a model returning relevance judgments. That makes for a decent demo. It doesn't tell me whether I should pay for the model call.

So I put three local embedding models behind the same interface, stored their vectors in PostgreSQL with pgvector, and ran the comparison. Then I added an independent LLM judge because reading more than a thousand request/job pairs by hand wasn't going to happen.

Jev won the graded ranking metric. When I counted only strong matches in the first five results, MiniLM had 33 and Jev had 32.

I find that more useful than a clean victory. It tells me which part of the result list improved, and which claim I still can't make.

What I actually compared

The live benchmark froze 237 eligible job postings and ran 48 predefined searches through four methods. Each search ran twice for timing. Quality measurements use only the first repeat, so a second attempt can't quietly replace an inconvenient result.

The embedding models were sentence-transformers/all-MiniLM-L6-v2, BAAI/bge-small-en-v1.5, and snowflake/snowflake-arctic-embed-s. All three ran locally on CPU. Their search paths used exact cosine similarity in pgvector. There was no approximate nearest-neighbor index to introduce another source of retrieval error.

Descriptions were encoded in 600-character chunks with 100-character overlap, then combined with normalized mean pooling. That's one document representation, with tradeoffs of its own. I haven't shown that it's the best configuration for any of these models.

The Jev pipeline first retrieved up to 30 candidates through lexical matching and aliases, then judged the first 1,000 characters of each candidate's search text. The embedding methods searched the whole eligible corpus using their document representations.

That difference matters. This is a comparison of the pipelines I built. Candidate retrieval can prevent Jev from seeing a good job at all, and shorter text can hide a requirement that the embedding representation includes. I can't take the final numbers and attribute every difference to model intelligence.

TypeSafe describes Jev as a model for typed judgments that software can consume directly. Its Choice primitive, for example, returns a selected option, probabilities, and confidence. In this search application, I used Jev to evaluate candidate relevance. That was the behavior under test.

I made the judge blind to the search method

For each query, I pooled the top ten results from all four methods and added two randomly selected postings where available. A job returned by several methods was evaluated once for that query. This produced 1,399 unique request/job pairs.

GPT-6.1 Sol evaluated the request against the full posting and explicit job metadata. It saw neither the retrieval method nor its ranking or score. The rubric stayed fixed, and the model returned a grade with a short reason, evidence quotes, and any unmet or unknown requirements.

Grade 2 meant a strong match: the work fit and every explicitly requested requirement had support in the posting. Grade 1 meant a partial match or missing requested information. Grade 0 meant unrelated work or an explicit conflict.

A remote Python role with no stated country eligibility can be a partial match for someone asking to work from Mexico. It shouldn't become a strong match because the judge feels optimistic about the employer. Silence about a requirement is still silence.

All 1,399 pairs received accepted judgments. Two requests hit HTTP 520 errors and succeeded on retry. The runner rejected invalid output instead of turning failed calls into negative labels.

I still have one model's opinion of relevance. Blind judging removes an obvious source of favoritism; it doesn't remove the judge's mistakes. These labels are provisional. The report includes the individual judgments and posting evidence so those decisions can be inspected.

Two metrics, two different conclusions

The ranking metric was pooled nDCG@10. It rewards useful results near the top and gives a strong match more credit than a partial one. We used gains of 3 for grade 2, 1 for grade 1, and 0 for grade 0, with a discount for lower positions. Each query's score is normalized against the best ordering within its judged pool.

Forty-five queries had at least one relevant job in that pool. Those queries entered the ranking average. The other three were excluded from nDCG, but remained in the strict precision calculation below.

MethodPooled nDCG@10, 45 queriesStrong-match Precision@5, all 48 queriesMedian search time
Jev0.63013.3%235 ms
MiniLM0.48713.8%29 ms
BGE0.38012.5%33 ms
Arctic0.36712.5%33 ms

The first column makes Jev look good. The second makes me slow down.

For strict precision, I counted only grade-2 results in the first five positions. Forty-eight searches give each method 240 possible positions. An unfilled position counts as a miss. The counts for Jev and MiniLM were:

Assessment in the first five resultsJevMiniLM
Strong matches3233
Partial matches10781
Strong plus partial matches139114

These are query/job appearances, not distinct jobs. A posting can be relevant to more than one query.

Jev's larger count came from partial matches. It produced more related results overall, while the count of clearly supported matches near the top was almost identical. A difference of one strong match doesn't establish that MiniLM is better either. It does prevent me from claiming that Jev found more jobs satisfying every requirement.

The live ranking difference is descriptive. I haven't established statistical significance or tested how stable it is under a second independent judge.

One query makes the tradeoff visible

The request was: "I want to interview users and design and test product experiences and prototypes."

Jev's top ten contained ten partial matches, with no strong or irrelevant results. MiniLM returned one strong match, three partial matches, and six irrelevant results.

If I'm browsing for nearby opportunities, I might prefer Jev's list. If I'm checking whether a system found a job that clearly supports everything I asked for, MiniLM found one and Jev didn't.

Those are different product goals. Combining them into the word "accuracy" hides a decision I need to make as the person building the application. For strict requirements, I would report strong-match precision separately and show which conditions remain unknown.

The synthetic test was easier to impress

Before the live evaluation, I ran a controlled benchmark with 72 fictional postings and 48 queries. The facts were explicit, and the expected matches came from those facts rather than an LLM judge. Sixteen queries were used for development and 32 were held out, with different job topics in the two groups.

Jev's deployed pipeline scored 1.000 on held-out nDCG@10. BGE scored 0.986. The paired bootstrap interval for that difference included zero.

That was a useful check that the system could handle the constructed cases. The live results were much less tidy. BGE, which was close to Jev on the synthetic ranking task, fell behind MiniLM on the real postings.

Short, explicit descriptions made the controlled test easier to label. Real postings contain missing facts, vague requirements, and work that doesn't fit a clean category. I want both tests, and I want their conclusions kept separate.

What the evaluation cost

The judge consumed 2,634,345 input tokens and 187,458 output tokens. The output total already includes 15,268 reasoning tokens.

I placed explicit cache boundaries after the shared rubric and the job description. The changing search request came afterward. Requests for a particular posting ran sequentially, so later evaluations could reuse the posting prefix; different postings ran concurrently.

The API reported 2,485,742 cached input tokens, or 94.4% of input. At published standard GPT-6.1 Sol rates, the recorded usage cost about $2.48. The same token counts without caching would cost about $7.14. The estimate includes cache-write charges, and excludes any unreported usage from the two failed requests. It isn't an invoice.

That's the judging cost. It doesn't include hosting, document embedding, or the earlier Jev search calls. In the live search run, MiniLM's median latency was about 29 ms against Jev's 235 ms. A local model with an existing index has a substantial advantage there. Whether Jev's improvement is worth that latency depends on which kinds of matches the application values.

What I would build next

For this job-search implementation, I would keep MiniLM as a serious baseline. Jev has evidence in its favor on graded relevance, and I want to test a retrieval-plus-Jev configuration on real jobs with equal candidate and text coverage. I wouldn't remove the cheaper path based on these results.

The experiment also leaves me interested in a different use of Jev: choosing the next action in a changing situation. I've written a concept called ReplayOps, a small incident-response simulator. An alert arrives, the model chooses a bounded action, the simulator changes state, and the next decision must account for what happened.

The same "errors increased after deployment" message could justify rollback when it is reversible, or investigation when the current release included an incompatible schema change. A restart loop should accumulate a visible cost. An uncertain model should be able to ask for evidence or escalate.

That concept has no model results yet. It would need a competent rules baseline and a conventional LLM, with success judged by the simulator's outcomes. If Jev can resolve more unseen incidents at an acceptable error rate and lower latency or cost, that would be evidence for a useful role in a decision loop.

For now, the measured result is narrower: Jev improved the overall relevance of this job-search pipeline. MiniLM found about as many strong matches in the first five results and returned them much faster. I'll keep that distinction in the next evaluation.

Evidence and limits

The full report contains per-query counts, rankings summarized against the labels, and links to every judgment. Token accounting records cache reads, writes, and pricing assumptions.

The 48 intentions were predefined for the benchmark, not sampled from actual user traffic. The same intentions were used in the controlled and live tracks. Pooled labels cannot establish corpus-wide recall or prove there were no relevant jobs outside the pool. One judge and one small corpus are insufficient to declare a universal winner.

I preserved the raw search snapshots, structured judgments, attempt records, evaluation code, and conclusions in a checksum-verified archive. The original complete HTTP response envelopes weren't captured. Reconstructed request bodies are labeled as reconstructed. That is the evidence available for this run.