Jev vs Laya: Testing Decision Models on Configuring Search

David Tippett,•Engineering
Fig.09 — Decision models
CaitlinKaitlyn

Should this field match by sound?

Jev35/35correct
0.16s
Laya29/35correct
1.10s
Claude Opus 535/35correct
2.27s
Jev vs Laya · Relevan EngineeringRelevan
Median time per request and fields configured correctly: Jev 35 of 35 in 0.16 seconds, Laya 29 of 35 in 1.10 seconds, Claude Opus 5 35 of 35 in 2.27 seconds.

At Relevan’s core is a tuning machine we call ASRE (Automated Search Relevance Engine). This is our secret sauce for tuning the relevance of each customer’s indexes. Not every customer needs every relevance feature, so ASRE’s core job is deciding which features to enable, disable, or tune to make search more relevant.

When TypeSafe released Jev, we immediately realized that it could seriously simplify how ASRE picks which features to enable. Right now we use complicated heuristics, or even LLM calls that can fail or have output parsing issues. With Jev (and similar models), you get a discrete output that matches the exact data shape you need.

To test it, we decided to use a straightforward case. We have a phonetic analyzer that we can enable for customers. The analyzer helps when people search for things they can only spell phonetically. Take, for example, the name Caitlin, which has over a hundred recognizable spellings . As long as you can spell it phonetically, Relevan can find it.

We tested three models on this task, and we measured two things: how often the model was correct, and how long each answer took. The eval is open source at relevan-dev/decision-model-eval  so you can reproduce or modify it for your own experiments.

The test

We created 10 sample indexes that are similar to some we’ve seen at Relevan. Each index has a mapping, a short description, and some sample values for each field. They include recruiter candidates, medical providers, English parish registers, and application logs. There’s a nice mix of fields that can benefit from phonetic matching and ones that don’t.

For each index, a search engineer wrote a reasonable configuration, which we use as the label. Then we asked two questions about each of the 35 text fields:

  1. Should this field get phonetic matching? The model answers with a probability.
  2. Which encoder should it use: double_metaphone, metaphone, or soundex?

Here are the three models we tested:

  • Laya, an open-weights model that answers typed questions. We ran it on a CPU on our own machine.
  • Jev, the hosted model from TypeSafe. Laya and Jev use the same request format, so both got the exact same requests.
  • Claude Opus 5, through the Claude API. Claude gets the same state and questions as a prompt, and it replies with structured output that matches a schema built from those questions. We set the effort to low, since this is a short classification task.

We sent the requests one at a time, and each model got one warm-up request before we started the clock.

The results

LayaJevClaude Opus 5
Enable decision correct29 of 3535 of 3535 of 35
Precision71%100%100%
Recall100%100%100%
Encoder correct9 of 1515 of 1515 of 15
Whole index config correct4 of 1010 of 1010 of 10
Brier score (0 is perfect)0.1730.0200.010
Median time per request1.10 s0.16 s2.27 s
95th percentile time1.14 s0.24 s3.42 s
Time for the full test38 s6 s84 s

Jev and Claude both nailed it. Every field, every encoder, all ten indexes. Claude’s probabilities were a little sharper than Jev’s. Fields that needed phonetic matching came back close to 1, and fields that didn’t came back close to 0. That’s what the Brier score captures. Lower is better, and 0 means every probability was spot on.

Funny enough, the first mistake we found was ours. In an earlier run, Jev and Claude both picked double_metaphone for the English parish registers. Our label said metaphone, since every name in that index is English. We assumed the narrower encoder would mean fewer false matches, but we’d never actually tested that. Double Metaphone handles English names just fine, so we changed our rule: use double_metaphone unless there’s a compatibility requirement or a measured improvement. All the results in this post use the updated labels.

The real difference is speed. Jev answered in about 160 ms, which is fast enough to run while a customer is still creating their index. Claude took about 2.3 seconds per answer. That’s still fine for a background job, and we didn’t have to train or host anything. But it’s roughly 14 times slower than Jev for the same answers.

Laya struggled. It caught every field that needed phonetic matching, but it also turned it on for six fields that didn’t, like summary, email, and cuisine. Most of its probabilities landed between 0.5 and 0.75, which isn’t much better than a coin flip. It also picked soundex for four indexes that had no reason to use it.

Architecting with decision models

One request per field is the obvious approach, but it isn’t the only one. We wanted to know if grouping the questions differently would get us better answers or fewer requests, so we tried four setups.

1. Per field(default)
one call per field
enable?encoder?
phonetic config for the index
2. Encoder first--chain
one call per index
encoder?
encoder applies to every field
one call per field
enable?
phonetic config for the index
3. Index gate--gate
one call per index
any names?
no → every field off, no field calls
yes
one call per field
enable?encoder?
phonetic config for the index
4. Gate and encoder first--gate --chain
one call per index
any names?encoder?
no → every field off, no field calls
yes
one call per field
enable?
phonetic config for the index
The four architectures we tested. Each box is one request to the model, and the tags are the typed questions it answers.
  1. Per field. Each field gets one request that asks both questions.
  2. Encoder first. The encoder usually applies to the whole index, so one request per index picks it. Then each field request only asks whether to enable it.
  3. Index gate. One request per index asks whether any field holds names. If not, we skip every field in that index.
  4. Gate and encoder first. One index request asks the gate question and the encoder question together.
ArchitectureRequestsLaya: enable, encoder, configsJev and ClaudeJev totalClaude total
Per field3529 of 35, 9 of 15, 4 of 10all correct5.9 s84 s
Encoder first4529 of 35, 8 of 15, 3 of 10all correct7.4 s111 s
Index gate3827 of 35, 7 of 12, 3 of 10all correct6.4 s102 s
Gate and encoder first3827 of 35, 7 of 12, 3 of 10all correct6.7 s121 s

A quick note on the table: when the gate closes an index, its fields never get an encoder question. That’s why Laya’s encoder score is out of 12 instead of 15 for the two gated setups.

For Jev and Claude, the setup didn’t matter. They gave the same answers every time, and only the number of requests changed. Claude’s total time also jumped around between runs, with its median time per request ranging from 2.3 s to 3.1 s.

With Laya, every alternative made things worse. Encoder first spread soundex across entire indexes instead of just a few fields. The gate correctly shut off the application logs, which fixed one wrong field. But it also shut off three indexes that do hold names, and Laya’s recall dropped from 100% to 80%.

The gate didn’t save us any time either. It only pays off when a lot of indexes don’t need phonetic matching. In our set, eight of ten indexes hold names, so the gate added more index requests than it saved in field requests. This is probably something we should have modeled by keeping the distribution of phonetic enabled indexes the same as in production but that’s work for a later date.

Shaping the state

Laya’s biggest improvements didn’t come from the architecture. They came from how we worded the request. We made two changes during development (against an earlier version of the labels), and both made a big difference.

Describe when to use it, not what it does. Our first encoder options described what each encoder does, like “short English surnames only; it keeps the first letter and is coarse.” That’s a great description for a search engineer, but the model has nothing in the index description to match it against. Laya picked soundex for 27 of 28 fields. So we rewrote each option as the situation that calls for it: “another system already runs a soundex index over the same records, and this index must return the same results.” Laya’s encoder accuracy jumped from 21% to 71%.

Show the model sample values. Our index request originally included only the field names. With just that, Laya picked the right encoder 57% of the time. Adding a few sample values per field bumped that to 79%, and whole configs correct went from 5 to 6 of 10.

Here’s what a single field request looks like:

{ "index": "restaurant-listings", "index_contents": "Restaurant listings for a city guide. Diners search for a place a friend recommended out loud, so they rarely have the spelling.", "field": "restaurantName", "declared_type": "text", "search_capabilities": ["searchable", "autocompletable"], "analyzer_language": "none", "sample_values": ["Pizzeria Bianco", "Le Bernardin", "Sqirl"] }

And here’s the index request, including the sample values that made the difference:

{ "index": "restaurant-listings", "index_contents": "Restaurant listings for a city guide. Diners search for a place a friend recommended out loud, so they rarely have the spelling.", "fields": { "restaurantName": "Pizzeria Bianco, Le Bernardin, Sqirl", "neighbourhood": "Kreuzberg, Shoreditch, Ravenswood", "cuisine": "italian, vietnamese, seafood", "phone": "+1 212 554 1515" } }

What this means for Relevan

Every feature in ASRE comes down to two questions: should this be on for this customer, and how should it be set? Today, answering that takes heuristics, LLM calls we have to babysit, and a search engineer double-checking the result. Phonetic matching was our first test, and honestly, it went better than we expected.

The part I’m most excited about is the probabilities. We already tune with a human in the loop. A well-calibrated probability tells that person which decisions are easy and which ones aren’t. With a model like Jev making the first pass, they can skip the fields the model is sure about and spend their time on the handful it isn’t.

Next up, we want to put our labels to the test with real spoken-name queries and count how many false matches each encoder actually produces. This is how we can ultimately be certain that we’re producing a net-positive effect.

Small note: the day before we were set to publish this, Contrastive-LM (a collaboration between Nvidia and Stanford) released CLM-8B . As of this writing, its performance wasn’t good enough to include in this post, but we’ll be keeping our eyes open for new open models.

Want to try this with your own use cases? Clone decision-model-eval , write one function that answers the typed questions, and see how it stacks up against Laya, Jev, and Claude (and other models we add).

And if you’d rather have search that tunes itself, try our free tier, or book a call to see what ASRE can do for your data.