TL;DR
- I reran my job matching benchmark on the same 241 jobs with the same prompt and the same GPT-6 Astra reference, adding gpt-6-luna and gpt-6-sol.
- gpt-6-luna at medium reasoning effort reaches 0.96 rank agreement with Astra, level with gpt-5.6-sol (0.96) and above Jev (0.93) and gpt-5.6-luna (0.95).
- It costs $0.55 per 1,000 jobs. That is 2.5x Jev, but about half the price of gpt-5.6-luna and a fortieth of gpt-5.6-sol.
- Correction to the first post. It used out-of-date GPT-5.6 prices. At OpenAI's current list prices gpt-5.6-luna costs $1.04 per 1,000 jobs, not $5.18, so Jev's price advantage over it is about 5x, not 24x.
- Medium effort matters for gpt-6-luna. At low effort it drops to 0.93 for a saving of only $0.06 per 1,000 jobs. For the GPT-5.6 models, extra effort bought nothing.
- Using Jev as a front filter now saves 30 to 40% of a very small bill, and loses some of the best jobs at the tighter settings. For OutRung I would now run gpt-6-luna on every job and keep Jev for small typed judgement calls.
Four days ago I published a benchmark of TypeSafe’s Jev against a set of GPT-5.6 models. Jev ranked jobs about as well as the GPT-5.6 models for a fraction of the price, so I proposed a split design in which Jev filters every job and an LLM only sees the survivors.
Then gpt-6-luna and gpt-6-sol arrived. gpt-6-luna is listed on OpenAI’s pricing page at $0.10 per million input tokens and $0.50 per million output tokens, half of what gpt-5.6-luna costs. That makes it only two to two and a half times the price of Jev on this task, which erases most of the gap my split design was built to exploit. So I ran the benchmark again.
In short, gpt-6-luna at medium reasoning effort is now the best value on this benchmark by a wide margin. It matches the best GPT-5.6 model’s agreement with Astra for about a fortieth of the price.
A correction first. The first post originally priced the GPT-5.6 models from an out-of-date price list, and I have since corrected it. At OpenAI’s current list prices gpt-5.6-luna costs $1.04 per 1,000 jobs, not $5.18, so Jev’s advantage over it is about 5x, not 24x. The rankings and the filter simulation are unaffected, but the savings from the filter were smaller than I first claimed. Every cost in this post uses OpenAI’s official list prices.
Same benchmark, two new models
Nothing else changed. The 241 jobs are the same ones the pipeline scored for my profile, every LLM gets the same short neutral prompt that asks for a 0-100 score and a one-sentence reason, and GPT-6 Astra at medium effort is still the reference. The metric is still Spearman rank correlation with Astra, which measures whether two methods put the jobs in the same order.
I ran gpt-6-luna and gpt-6-sol at both low and medium reasoning effort, and gpt-6-luna at low effort twice to see how much a model disagrees with itself between runs. The Jev, embedding, reranker, and GPT-5.6 results are reused from the first run. They are the same scores on the same jobs, so the comparison is like for like. There is no gpt-6-terra deployment yet, so the GPT-6 middle tier is missing.
How closely each method agrees with Astra
gpt-6-luna at medium effort reaches 0.96. That is level with gpt-5.6-sol (0.96) and slightly above gpt-5.6-luna (0.95) and Jev (0.93). gpt-6-sol lands at 0.95, which is a little surprising for the larger model, but on 241 jobs that gap is too small to mean much.
The top group was already tight in the first run, and it is tighter now. Six models sit between 0.93 and 0.96, and on 241 jobs most of those differences are within noise. The two gpt-6-luna low-effort runs agree with each other at 0.97, which is about the ceiling you can expect from any LLM on this data. Once a model is in the mid 0.9s, its agreement with Astra is limited mostly by randomness, not by ability.
Reasoning effort matters this time
In the first run, medium effort bought nothing. gpt-5.6-luna stayed at 0.95 and gpt-5.6-terra fell slightly. gpt-6-luna behaves differently.
| Model | Effort | Agreement with Astra | Cost per 1,000 jobs | Median latency |
|---|---|---|---|---|
| gpt-6-luna | low | 0.93 | $0.49 | 1.2s |
| gpt-6-luna | medium | 0.96 | $0.55 | 1.9s |
| gpt-6-sol | low | 0.95 | $9.97 | 2.1s |
| gpt-6-sol | medium | 0.95 | $10.48 | 2.2s |
At low effort gpt-6-luna is level with Jev. At medium effort it moves to the top of the table for an extra six cents per 1,000 jobs and under a second of extra latency. It spends about 190 output tokens per job at medium against about 80 at low, but output is so cheap on this model that the extra thinking barely shows on the bill. gpt-6-sol gains almost nothing from medium effort, which matches what I saw with the GPT-5.6 models.
What the scores look like
As before, each dot is one job. Astra’s score runs across and the selected model’s score runs up. The slider keeps each side’s top X% so that models with different habits on the scale are compared fairly.
The most striking thing about gpt-6-luna at medium effort is its calibration. Across all 241 jobs its average score is within a hundredth of a point of Astra’s on the 0-10 scale. Jev marks about 1.5 points more strictly than Astra, gpt-5.6-luna about 1 point more generously, and gpt-6-sol about 0.9 points more strictly. None of that matters for ranking, but it matters a lot if the number is ever shown to a user, because an 8 from gpt-6-luna means roughly what an 8 from Astra means.
At the top 20%, gpt-6-luna picks 41 of Astra’s 49 jobs, the same as Jev. gpt-6-sol picks 44, the best of any model I have tested, and gpt-5.6-luna 39. If what you care about is agreeing on the very best jobs rather than the whole order, gpt-6-sol has an edge, although at 19 times the price of gpt-6-luna.
Does it keep the best jobs?
The same safety test as last time. Sort all 241 jobs by each model’s score, keep only the top X%, and count how many of Astra’s 25 best jobs survive.
gpt-6-luna keeps all 25 from the top 20% of jobs up. Jev keeps 20 at that setting and gpt-5.6-luna 21. At the harshest setting, the top 10%, gpt-6-sol is clearly ahead with 20 of 25, while gpt-6-luna keeps 16 and Jev 15. With a target set this small, one job is four percentage points, so treat the gaps at 10% as indicative.
What it costs
These are measured token counts from the run, priced at OpenAI’s list rates and normalised to 1,000 candidate-job pairs. Every LLM saw the same 4,536 input tokens per job on average, so the differences come from price per token and how much each model writes.
- Embedding
- Reranker
- Jev
- GPT-6 Luna
- GPT-5.6 Luna
- GPT-6 Sol
- GPT-5.6 Terra
- GPT-5.6 Sol
| Model | Agreement with Astra | Cost per 1,000 jobs | Multiple of Jev | Median latency |
|---|---|---|---|---|
| Jev | 0.93 | $0.22 | 1x | 0.3s |
| gpt-6-luna (medium) | 0.96 | $0.55 | 2.5x | 1.9s |
| gpt-5.6-luna (low) | 0.95 | $1.04 | 5x | 2.1s |
| gpt-5.6-terra (low) | 0.93 | $9.83 | 45x | 2.0s |
| gpt-6-sol (medium) | 0.95 | $10.48 | 48x | 2.2s |
| gpt-5.6-sol (medium) | 0.96 | $21.37 | 97x | 3.0s |
| GPT-6 Astra (medium) | reference | $49.88 | 227x | 3.7s |
The chart makes the shift easy to see. In the first run the cheapest LLM that could match Jev was gpt-5.6-luna at about five times Jev’s price, and a filter was the obvious way to exploit that gap. gpt-6-luna now sits between them, at half the price of gpt-5.6-luna and higher up than both. It is also the fastest LLM I have tested, although Jev still answers in about a sixth of the time.
Is the Jev filter still worth it?
This was the practical point of the first post. Score every job with Jev, pass only the top X% to the LLM, and save the rest of the bill. Here is the same table with gpt-6-luna at medium effort as the LLM behind the filter.
| Jobs passed to gpt-6-luna | Cost per 1,000 jobs | Saving | Astra’s top 10% kept |
|---|---|---|---|
| All of them, no filter | $0.55 | 0% | 25 of 25 (100%) |
| Top 50% by Jev | $0.50 | 10% | 25 of 25 (100%) |
| Top 30% by Jev | $0.39 | 30% | 24 of 25 (96%) |
| Top 20% by Jev | $0.33 | 40% | 20 of 25 (80%) |
| Top 10% by Jev | $0.28 | 50% | 15 of 25 (60%) |
With gpt-5.6-luna behind it at current prices, the 30% setting saves about half the bill, $0.53 against $1.04. With gpt-6-luna it saves 30%, which is 16 cents per 1,000 jobs, and it still loses one of Astra’s best 25. The 50% setting that lost nothing now saves 5 cents. Jev’s own call has become a large share of the total, so a harsher filter is the only way to save anything, and a harsher filter is exactly what drops good jobs.
So for scoring, I would now drop the filter. Running gpt-6-luna on every job costs 55 cents per 1,000, gives a score and a written reason in one call, and needs one vendor instead of two. The explanation problem that made me keep an LLM in the first post disappears, because the model that scores is the model that explains. A split design only pays for itself again at a volume where tens of cents per thousand jobs add up, and OutRung is nowhere near that.
That is also why I am moving the model behind OutRung’s job scoring to gpt-6-luna.
Where Jev still fits
Jev did not get worse. TypeSafe built it as a System One model that returns typed answers instead of text, and it is still the cheapest capable scorer here, and it still answers in 0.3 seconds against 1.2 to 3.7 seconds for the LLMs, which matters for anything interactive. The location probe from the first post also still stands. A typed yes or no question like “is Royston within a daily commute of Cambridge, UK?” costs a fraction of a cent, returns a usable probability instead of text to parse, and got the two overseas Cambridges right where embeddings got them badly wrong.
What changed is the case for Jev as a cost guard in front of an LLM. That case rested on the price gap between Jev and the cheapest LLM that could match it. The first post put that gap at 24x, at current prices it was really about 5x, and with gpt-6-luna it is 2.5x. I would still reach for Jev for many small typed judgement calls per job, such as commute, seniority, visa, and whether a role is really remote, where a full LLM call for each one would add up.
Caveats
The biggest new caveat is family resemblance. gpt-6-luna and gpt-6-sol come from the same generation as Astra, so some of their agreement may be shared habits rather than better judgement. I checked this by using gpt-6-sol at medium effort as the reference instead of Astra. gpt-6-luna stays near the top at 0.93, gpt-5.6-sol reaches 0.94, gpt-5.6-luna 0.90, and Jev drops to 0.87. The order of the groups holds, but Jev’s gap to the LLMs widens under a different reference. A reference from a different vendor, or real outcomes such as applications and interviews, would be a stronger test.
Everything else from the first post still applies. All 241 pairs come from one candidate profile, the prompt is a short neutral one rather than OutRung’s production rubric, and agreement with Astra is not the same as a good outcome for the person reading the list. Prices are list prices at the time of writing, and at these levels a single price change moves the conclusion, so I will rerun this when prices or models move again.
Related questions
Nothing in the setup. Same 241 jobs, same candidate profile, same neutral prompt, and GPT-6 Astra is still the reference. The only change is two new models, gpt-6-luna and gpt-6-sol, each run at low and medium reasoning effort.

About the author
Tian
Tian is an AI professional, builder, and the founder of OutRung. Holding a PhD in deeptech, Tian navigated the frustrating modern job market first-hand before transitioning into the AI space. OutRung was built to share the exact strategies that made that transition successful. Tian's goal is to help everyday job seekers use AI to find their ideal roles efficiently, without needing to be computer experts themselves.



