Job alerts

Can a 400M decision model replace a local LLM for job alerts?

Decision models pick an answer from a fixed list and give a probability for each one, instead of writing text. Several new ones came out while I was building my job-alert pipeline, so I tested them against the general LLM the pipeline started with. A 400-million-parameter decision model, asked three narrow questions about each posting (what kind of role it is, whether anything rules me out, and how well it fits my resume), now matches the LLM and runs about 60 times faster, with a reason for every decision.

Time per posting1.6 sthree-question Laya
vs 97 s for the local LLM
Alerts worth opening42.9%behind the rules, on 500 postings
random pick: 7%
Realistic ceiling~70%what the labels themselves allow
measured by relabelling a sample
Why decision models

Why I tried decision models

My job search pulls postings from 12 job boards and has to decide, for each one, whether it deserves an email alert, a save for later, or nothing. I started with a general local LLM writing a 1–10 match score. It worked, but slowly, and its reasons were free text I couldn't count on.

Each check in my pipeline has a fixed question and a fixed set of answers, which is what decision models are made for. Several appeared since I started this project. I set two rules for trying them: everything runs on my own machine, because every check includes my resume, and it costs nothing. That ruled out hosted ones such as Jev.

General LLM · local

Gemma (gemma4:e4b)

Through Ollama. Reads my resume and the posting and writes a 1–10 score with its reasoning. The starting point.

Decision model · 9B · local

Nimble

Answers yes/no and multiple-choice questions through Ollama. Used for one question: "Is this a software engineering role?"

Decision model · local

Clef-Flash

Cloudflare's decision model, tried in Nimble's place on the same question.

Decision model · 340M · local

GLiNER2.5-Decide

A small open decision model from Fastino, used as downloaded.

Decision model · 400M · fine-tuned

Laya

A ModernBERT-based decision model on Hugging Face (convaiinnovations/laya), fine-tuned on my own labelled jobs. I asked it two ways, shown in the next section: one question, then three.

Plain code

Rules

Country, years of experience asked for, title words such as "senior", contract and hourly-pay wording.

Two ways to ask Laya

One broad question, then three narrow ones

One question. The Laya I started with was fine-tuned on a single multiple-choice question, "Should the candidate be considered for this role?", with three answers:

It alerts when it answers notify and is at least 70% sure. That one answer has to cover three separate judgements: is this my kind of role, does anything rule me out, and how well do my skills match.

Three questions. Laya's documentation shows a version trained on several question types at once, so I split the judgement and asked in order. A posting that fails a step is skipped, and the run log records why; the grey boxes below are those log lines.

1

Which stack is this role?

Seven choices: backend, frontend, full stack, AI, data, infrastructure or other, read from the posting alone. My search targets backend, full-stack, AI and data roles, so those go on; frontend, infrastructure and other roles are skipped. The accepted list is a setting.

skip · skipped, stack role frontend
2

Is there a dealbreaker?

Clearance, citizenship, a senior level, the wrong kind of role, from the posting alone.

skip · skipped, likely dealbreaker (72%)
3

How well does it fit?

A 1–10 score from my resume and the posting. Laya returns a probability for each score; their weighted average is the fit.

alert · fit 7.02/10, at or above the alert cut-off
The test

500 unseen postings, judged two ways

I picked 500 postings at random from four days of my search. A larger LLM read each with my full resume and marked 35 as good jobs. No model had seen any of them. Every setup is scored on two questions: alerts worth opening (of the alerts it sent, how many were good jobs) and good jobs caught (of the 35). Both are measured against the labelling LLM, not my own judgement.

You can't max out both. Alert on every posting and you catch all 35 good jobs, but only 7% of the alerts are worth opening. Alert on almost nothing and most alerts are good, but most good jobs are missed. So every table below shows both scores, plus how many alerts were sent.

To know how high "good" can go, I had the labelling LLM rate 120 of the postings again, including all 35 good jobs. It agreed with itself on about 95% of them, but the borderline postings that decide precision are where it wavered, so even a second pass of the labeller reaches only about 70% of alerts worth opening. That is the realistic ceiling.

Round 1

Each model alone: none is good enough

Setup, on its ownAlertsWorth openingGood jobs caught
Gemma, thinking on (random 60 of the 500)28 of 6021%6 of 6
Three-question Laya, fit score 6.07 or more9117.6%16 of 35
Nimble, asked "should I get an alert?"10917.4%19 of 35
One-question Laya, notify at any confidence26810.1%27 of 35
GLiNER2.5-Decide4996.8%34 of 35
Random pick–7%–

Alone, every model sends far too many alerts. GLiNER2.5-Decide beats Laya on its maker's own benchmark, yet it said "alert" to 499 of the 500 postings. What helped was putting cheap rules in front of them.

Round 2

Rules first, then one narrow question, then the fit

500postings35 good jobs
185after the rules34 good jobs
156after Nimble's "software role?"34 good jobs
~35–43alerts from the final model15–18 good jobs

The rules removed almost two thirds of the postings and only 1 good job. On their own they already help a lot: one-question Laya behind the rules alone, alerting at any confidence, sends 97 alerts and 26.8% are worth opening (26 good jobs caught), against 10.1% with no rules in front. Nimble's single yes/no question removed 29 more postings, none of them good jobs. Clef-Flash in Nimble's place did the same (41% of alerts worth opening against 42%), so I kept the smaller download. The final model then only has to judge 156 plausible postings:

Each final model decides an alert in its own terms. Gemma writes a whole-number match score from 1 (no match) to 10 (exceptional); my pipeline has always alerted at 8 or more, and the table also shows 9 or more. Three-question Laya gives a fit score with decimals on the same 1–10 scale, the average of the probabilities it gives each score, and alerts at 6.07, a line fitted on held-out training postings before this test.

Final model, behind rules and NimbleAlertsWorth opening (95% range)Good jobs caughtPer posting
Three-question Laya: alert when its fit score is 6.07 or more3542.9% (28–59%)15 of 351.6 s
One-question Laya: alert when it answers notify, at least 70% sure4341.9% (28–57%)18 of 351.5 s
Gemma: alert when its match score is 8 or more10231.4% (23–41%)32 of 3597 s
Gemma: alert when its match score is 9 or more4535.6% (23–50%)16 of 3597 s

Behind the same rules, the two designs make different trades. With a match score of 8 or more, Gemma catches almost every good job, 32 of 35, but sends three times as many alerts and fewer than a third are worth opening. At 9 or more, it sends about as many alerts as Laya and catches about as many good jobs, with fewer worth opening: 35.6% against 42–43%. With 35 good jobs the ranges overlap, so at the same number of alerts the small decision model matches the LLM, in 1.6 seconds per posting instead of 97.

Speed

Seconds per posting on a 24 GB M3 MacBook

Rulesunder 0.01 s
Laya, one question1.5 s
Laya, three questions1.6 s
GLiNER2.5-Decide (CPU)5.9 s
Nimble, one yes/no question7.5 s
Nimble, full decision12.9 s
Gemma, thinking off29 s
Gemma, thinking on97 s

Typical time per posting, log scale. Laya answers in one forward pass per question; the LLM writes its reasoning token by token.

Fine-tuning Laya

What the three questions changed

Before the split, I retrained the one-question model on new labels several times, with the same three-answer question and later a plain yes/no question (next section). None came close to the original model.

The three-question version brought a retrained model level with the original one, and every skip now says why. Because each question can be tested on its own, I also tried replacing Laya's stack-role and dealbreaker answers with the labels' own, a perfect version of both. Alerts did not change at all: the rules and Nimble had already removed those jobs. So the only part left to improve is the fit score.

Fit as yes or no

Why a 1–10 fit score, not "would this candidate fit? yes or no"

A yes/no fit question sounds simpler, and I tried it twice.

As its own model. One of the retrains before the split asked Laya a single yes/no question with my resume and the posting: "Is this job worth an alert?" It reached 24% of alerts worth opening behind the rules, against about 42% for the original model, and its "yes" probabilities never went above 0.39, so it could not tell a strong fit from a borderline one.

Inside the three-question model. The fit question gives a probability for each score from 1 to 10. Adding the probabilities for 7 to 10 gives the chance that the job is a good fit, which is a yes/no answer read from the same output. Both readings were measured:

Fit read asRanking quality on held-out postingsAlertsWorth openingGood jobs caught
The decimal score, alert at 6.07 or more0.5513542.9%15 of 35
The chance of a 7–10 (yes/no), alert at 51% or more0.5283040.0%12 of 35

Ranking quality is average precision: how well the reading puts good jobs ahead of the rest, before any cut-off is chosen. The benchmark columns are behind the rules and Nimble, each reading at its own cut-off, fitted on held-out postings before the benchmark.

The decimal score comes out slightly ahead on both, though with 35 good jobs the difference is within noise. It also ranked good jobs better before the benchmark, which is why it was chosen. Three reasons the score is the better thing to train on:

What I learned

Five things about decision models in practice