1# Chapter 14: eval, choosing an LLM with a test instead of a hunch 2 3There are a dozen Workers AI models that can call tools, at prices that 4differ by more than tenfold. Which one should write Jev's questions? The 5rule (the owner, 2026-10-02): **the cheapest model that writes a correct Jev 6tool call every time, with Claude Haiku's price as the ceiling.** Not the 7smartest, not the newest: the cheapest that passes. 8 9So this program asks each candidate, cheapest first, to do the job on a set 10of cases, and stops at the first that gets them all right. It uses the 11Worker's own three steps (`llm::request`, `llm::parse`, `ask::check`), so a 12pass here is a pass on the site. 13 14> **Aside.** The cases have opinions. "How likely is X" counts as a 15> yes-or-no question, because the probability of yes *is* the likelihood. It 16> was first written as a how-much, and failed a model for being right. The 17> case was fixed, not the model. 18 19 nix develop .#owner 20 21Every run **spends Workers AI neurons from the same free daily allowance as the live site** (about 350 per model per run of 10,000 a day, which resets at 00:00 UTC), so it refuses to run without `--spend`. Run it once, when asked to, never in a loop: ten runs in one afternoon (2026-10-05) used the whole allowance and the site could not draft questions until the reset. 22 lmjtfy-eval --spend # cheapest first, stop at the first that passes every case 23 lmjtfy-eval --all --spend # every candidate 24 lmjtfy-eval @cf/qwen/qwen3-30b-a3b-fp8 --spend # one model, printing each reply 25 26The candidates and their prices are `llm::CANDIDATES` (chapter 8). The cases 27are `CASES` in `src/main.rs`: four yes-or-no questions, four with a best 28answer, three of degree, one "how likely", one with quotes and symbols, and 29one that asks two things. A case passes when every tool call is well formed, 30Jev's protocol accepts it, and the tools called are the ones the case 31expects. A run spends real neurons from the account's daily 10,000: 14 32requests a model. 33 34## Result, 2026-10-02 35 36| Model | Passed | Neurons a call | Seconds a call | 37| --- | --- | --- | --- | 38| `@cf/ibm-granite/granite-4.0-h-micro` | 11 / 14 | 2.4 | 3.6 | 39| `@cf/qwen/qwen3-30b-a3b-fp8` | 14 / 14 | 16.0 | 3.1 | 40 41granite called `jev_choice` for "should I rewrite it in rust", left 42`yes_means` out of a `jev_noul`, and made one call for the two-part question. 43qwen3 is the Worker's `LLM_MODEL`. At 16 neurons a call the free allocation 44covers about 625 questions a day. The more expensive candidates were not run. 45 46## The facts: `lmjtfy-eval facts` 47 48The nine questions Jev is asked first (chapter 7) decide everything: what 49is refused, what Jev answers alone, what reaches the LLM, what goes on the 50feed. Their wording is behaviour, and this checks it. 51 52 op-env-run -- lmjtfy-eval facts 53 54`src/facts.rs` has 25 inputs, each with what the facts should be: whole, 55answerable 56or not, which kinds it may be read as, whether it is several questions, 57whether it needs a scale of its own, whether it is fit to show. The Worker's 58own request (`ask::wanted` with every Jev fact) goes to Jev through 59`jev-http`, so the estate's shared spend ledger admits it, and `ask::learned` 60reads the answers as the Worker does. 61 62Run it after changing a fact question's wording (`packages/ask`), the split 63threshold (`ask::ALSO`) or the Jev model. A run costs under a tenth of a 64cent. The key is `LMJTFY_TYPESAFE_API_KEY`, from `op-env-run`. 65 66`lmjtfy-eval halves` is the other half of `whole`: eight questions cut off 67part way, as the page sends them while they are typed, each of which must 68be read as not whole. It is a run of its own because the estate's ledger 69lets thirty requests through a minute, and anything else asking Jev on this 70machine shares them: a run that is throttled says so on each line it could 71not send, and is run again. 72 73Result, 2026-10-03, with `whole` added: 25 / 25 of the inputs (over two 74runs, the ledger having throttled each), about 830 tokens a request; 7 / 8 75of the halves ("can peng" is read as whole). 2026-10-02: 25 / 25. 76 77## In this folder 78 79| Path | What | 80| --- | --- | 81| [src/](src/) | The two evals. | 82| [Cargo.toml](Cargo.toml) | The program: `llm`, `ask` and `rules`, `ureq` for Cloudflare's REST API, and `jev-http` for the facts eval. | 83 84← Previous: [Chapter 13, tools/](../) · Up: [tools](../) · Next: [Chapter 14½, eval/src/](src/) →