Can a 400M decision model replace a local LLM for job alerts?
Decision models pick an answer from a fixed list and give a probability for each one, instead of writing text. Several new ones came out while I was building my job-alert pipeline, so I tested them against the general LLM the pipeline started with. A 400-million-parameter decision model, asked three narrow questions about each posting (what kind of role it is, whether anything rules me out, and how well it fits my resume), now matches the LLM and runs about 60 times faster, with a reason for every decision.
vs 97 s for the local LLM
random pick: 7%
measured by relabelling a sample
Why I tried decision models
My job search pulls postings from 12 job boards and has to decide, for each one, whether it deserves an email alert, a save for later, or nothing. I started with a general local LLM writing a 1–10 match score. It worked, but slowly, and its reasons were free text I couldn't count on.
Each check in my pipeline has a fixed question and a fixed set of answers, which is what decision models are made for. Several appeared since I started this project. I set two rules for trying them: everything runs on my own machine, because every check includes my resume, and it costs nothing. That ruled out hosted ones such as Jev.
Gemma (gemma4:e4b)
Through Ollama. Reads my resume and the posting and writes a 1–10 score with its reasoning. The starting point.
Nimble
Answers yes/no and multiple-choice questions through Ollama. Used for one question: "Is this a software engineering role?"
Clef-Flash
Cloudflare's decision model, tried in Nimble's place on the same question.
GLiNER2.5-Decide
A small open decision model from Fastino, used as downloaded.
Laya
A ModernBERT-based decision model on Hugging Face (convaiinnovations/laya), fine-tuned on my own labelled jobs. I asked it two ways, shown in the next section: one question, then three.
Rules
Country, years of experience asked for, title words such as "senior", contract and hourly-pay wording.
One broad question, then three narrow ones
One question. The Laya I started with was fine-tuned on a single multiple-choice question, "Should the candidate be considered for this role?", with three answers:
- notify: strong fit, right level, good alignment with the candidate's profile
- log: possible fit, but borderline or uncertain; worth a later look
- skip: weak fit, wrong level, or a disqualifying mismatch
It alerts when it answers notify and is at least 70% sure. That one answer has to cover three separate judgements: is this my kind of role, does anything rule me out, and how well do my skills match.
Three questions. Laya's documentation shows a version trained on several question types at once, so I split the judgement and asked in order. A posting that fails a step is skipped, and the run log records why; the grey boxes below are those log lines.
Which stack is this role?
Seven choices: backend, frontend, full stack, AI, data, infrastructure or other, read from the posting alone. My search targets backend, full-stack, AI and data roles, so those go on; frontend, infrastructure and other roles are skipped. The accepted list is a setting.
Is there a dealbreaker?
Clearance, citizenship, a senior level, the wrong kind of role, from the posting alone.
How well does it fit?
A 1–10 score from my resume and the posting. Laya returns a probability for each score; their weighted average is the fit.
500 unseen postings, judged two ways
I picked 500 postings at random from four days of my search. A larger LLM read each with my full resume and marked 35 as good jobs. No model had seen any of them. Every setup is scored on two questions: alerts worth opening (of the alerts it sent, how many were good jobs) and good jobs caught (of the 35). Both are measured against the labelling LLM, not my own judgement.
You can't max out both. Alert on every posting and you catch all 35 good jobs, but only 7% of the alerts are worth opening. Alert on almost nothing and most alerts are good, but most good jobs are missed. So every table below shows both scores, plus how many alerts were sent.
To know how high "good" can go, I had the labelling LLM rate 120 of the postings again, including all 35 good jobs. It agreed with itself on about 95% of them, but the borderline postings that decide precision are where it wavered, so even a second pass of the labeller reaches only about 70% of alerts worth opening. That is the realistic ceiling.
Each model alone: none is good enough
| Setup, on its own | Alerts | Worth opening | Good jobs caught |
|---|---|---|---|
| Gemma, thinking on (random 60 of the 500) | 28 of 60 | 21% | 6 of 6 |
| Three-question Laya, fit score 6.07 or more | 91 | 17.6% | 16 of 35 |
| Nimble, asked "should I get an alert?" | 109 | 17.4% | 19 of 35 |
| One-question Laya, notify at any confidence | 268 | 10.1% | 27 of 35 |
| GLiNER2.5-Decide | 499 | 6.8% | 34 of 35 |
| Random pick | – | 7% | – |
Alone, every model sends far too many alerts. GLiNER2.5-Decide beats Laya on its maker's own benchmark, yet it said "alert" to 499 of the 500 postings. What helped was putting cheap rules in front of them.
Rules first, then one narrow question, then the fit
The rules removed almost two thirds of the postings and only 1 good job. On their own they already help a lot: one-question Laya behind the rules alone, alerting at any confidence, sends 97 alerts and 26.8% are worth opening (26 good jobs caught), against 10.1% with no rules in front. Nimble's single yes/no question removed 29 more postings, none of them good jobs. Clef-Flash in Nimble's place did the same (41% of alerts worth opening against 42%), so I kept the smaller download. The final model then only has to judge 156 plausible postings:
Each final model decides an alert in its own terms. Gemma writes a whole-number match score from 1 (no match) to 10 (exceptional); my pipeline has always alerted at 8 or more, and the table also shows 9 or more. Three-question Laya gives a fit score with decimals on the same 1–10 scale, the average of the probabilities it gives each score, and alerts at 6.07, a line fitted on held-out training postings before this test.
| Final model, behind rules and Nimble | Alerts | Worth opening (95% range) | Good jobs caught | Per posting |
|---|---|---|---|---|
| Three-question Laya: alert when its fit score is 6.07 or more | 35 | 42.9% (28–59%) | 15 of 35 | 1.6 s |
| One-question Laya: alert when it answers notify, at least 70% sure | 43 | 41.9% (28–57%) | 18 of 35 | 1.5 s |
| Gemma: alert when its match score is 8 or more | 102 | 31.4% (23–41%) | 32 of 35 | 97 s |
| Gemma: alert when its match score is 9 or more | 45 | 35.6% (23–50%) | 16 of 35 | 97 s |
Behind the same rules, the two designs make different trades. With a match score of 8 or more, Gemma catches almost every good job, 32 of 35, but sends three times as many alerts and fewer than a third are worth opening. At 9 or more, it sends about as many alerts as Laya and catches about as many good jobs, with fewer worth opening: 35.6% against 42–43%. With 35 good jobs the ranges overlap, so at the same number of alerts the small decision model matches the LLM, in 1.6 seconds per posting instead of 97.
Seconds per posting on a 24 GB M3 MacBook
Typical time per posting, log scale. Laya answers in one forward pass per question; the LLM writes its reasoning token by token.
What the three questions changed
Before the split, I retrained the one-question model on new labels several times, with the same three-answer question and later a plain yes/no question (next section). None came close to the original model.
The three-question version brought a retrained model level with the original one, and every skip now says why. Because each question can be tested on its own, I also tried replacing Laya's stack-role and dealbreaker answers with the labels' own, a perfect version of both. Alerts did not change at all: the rules and Nimble had already removed those jobs. So the only part left to improve is the fit score.
Why a 1–10 fit score, not "would this candidate fit? yes or no"
A yes/no fit question sounds simpler, and I tried it twice.
As its own model. One of the retrains before the split asked Laya a single yes/no question with my resume and the posting: "Is this job worth an alert?" It reached 24% of alerts worth opening behind the rules, against about 42% for the original model, and its "yes" probabilities never went above 0.39, so it could not tell a strong fit from a borderline one.
Inside the three-question model. The fit question gives a probability for each score from 1 to 10. Adding the probabilities for 7 to 10 gives the chance that the job is a good fit, which is a yes/no answer read from the same output. Both readings were measured:
| Fit read as | Ranking quality on held-out postings | Alerts | Worth opening | Good jobs caught |
|---|---|---|---|---|
| The decimal score, alert at 6.07 or more | 0.551 | 35 | 42.9% | 15 of 35 |
| The chance of a 7–10 (yes/no), alert at 51% or more | 0.528 | 30 | 40.0% | 12 of 35 |
Ranking quality is average precision: how well the reading puts good jobs ahead of the rest, before any cut-off is chosen. The benchmark columns are behind the rules and Nimble, each reading at its own cut-off, fitted on held-out postings before the benchmark.
The decimal score comes out slightly ahead on both, though with 35 good jobs the difference is within noise. It also ranked good jobs better before the benchmark, which is why it was chosen. Three reasons the score is the better thing to train on:
- It carries more information. The labels are 1–10 scores. Training on the score teaches the model that a 6 is nearly good and a 2 is far off; a yes/no label calls both "no".
- "Yes" is rare. Only about 1 posting in 5 is a good job, and a model trained on a rare "yes" learns to stay cautious, which is likely why the yes/no model's probabilities stayed low.
- The score model gives the yes/no answer for free. The 7–10 sum above is that answer; a yes/no model cannot give the score back.
Five things about decision models in practice
- Ask one specific question.Nimble asked "Is this a software role?" lost no good jobs. Asked "Should I get an alert?", one alert in six was worth opening.
- Order beats model choice.Every model did far better behind cheap rules than on its own, and swapping Nimble for Clef-Flash changed nothing.
- Test on your own task, not the maker's benchmark.A model that leads its own benchmark can still be unusable on a different task, as GLiNER2.5-Decide was here.
- Split a fuzzy judgement into narrow questions.A small model fine-tuned on three questions matched my best model and explains every skip.
- Measure the ceiling and the best case before tuning.Relabelling a sample put the ceiling near 70%; a perfect-answer test showed which part can still move it.