lmjtfy.git / tools / eval / README.md
1# Chapter 14: eval, choosing an LLM with a test instead of a hunch
2
3There are a dozen Workers AI models that can call tools, at prices that
4differ by more than tenfold. Which one should write Jev's questions? The
5rule (the owner, 2026-10-02): **the cheapest model that writes a correct Jev
6tool call every time, with Claude Haiku's price as the ceiling.** Not the
7smartest, not the newest: the cheapest that passes.
8
9So this program asks each candidate, cheapest first, to do the job on a set
10of cases, and stops at the first that gets them all right. It uses the
11Worker's own three steps (`llm::request`, `llm::parse`, `ask::check`), so a
12pass here is a pass on the site.
13
14> **Aside.** The cases have opinions. "How likely is X" counts as a
15> yes-or-no question, because the probability of yes *is* the likelihood. It
16> was first written as a how-much, and failed a model for being right. The
17> case was fixed, not the model.
18
19    nix develop .#owner
20
21Every run **spends Workers AI neurons from the same free daily allowance as the live site** (about 350 per model per run of 10,000 a day, which resets at 00:00 UTC), so it refuses to run without `--spend`. Run it once, when asked to, never in a loop: ten runs in one afternoon (2026-10-05) used the whole allowance and the site could not draft questions until the reset.
22    lmjtfy-eval --spend      # cheapest first, stop at the first that passes every case
23    lmjtfy-eval --all --spend        # every candidate
24    lmjtfy-eval @cf/qwen/qwen3-30b-a3b-fp8 --spend   # one model, printing each reply
25
26The candidates and their prices are `llm::CANDIDATES` (chapter 8). The cases
27are `CASES` in `src/main.rs`: four yes-or-no questions, four with a best
28answer, three of degree, one "how likely", one with quotes and symbols, and
29one that asks two things. A case passes when every tool call is well formed,
30Jev's protocol accepts it, and the tools called are the ones the case
31expects. A run spends real neurons from the account's daily 10,000: 14
32requests a model.
33
34## Result, 2026-10-02
35
36| Model | Passed | Neurons a call | Seconds a call |
37| --- | --- | --- | --- |
38| `@cf/ibm-granite/granite-4.0-h-micro` | 11 / 14 | 2.4 | 3.6 |
39| `@cf/qwen/qwen3-30b-a3b-fp8` | 14 / 14 | 16.0 | 3.1 |
40
41granite called `jev_choice` for "should I rewrite it in rust", left
42`yes_means` out of a `jev_noul`, and made one call for the two-part question.
43qwen3 is the Worker's `LLM_MODEL`. At 16 neurons a call the free allocation
44covers about 625 questions a day. The more expensive candidates were not run.
45
46## The facts: `lmjtfy-eval facts`
47
48The nine questions Jev is asked first (chapter 7) decide everything: what
49is refused, what Jev answers alone, what reaches the LLM, what goes on the
50feed. Their wording is behaviour, and this checks it.
51
52    op-env-run -- lmjtfy-eval facts
53
54`src/facts.rs` has 25 inputs, each with what the facts should be: whole,
55answerable
56or not, which kinds it may be read as, whether it is several questions,
57whether it needs a scale of its own, whether it is fit to show. The Worker's
58own request (`ask::wanted` with every Jev fact) goes to Jev through
59`jev-http`, so the estate's shared spend ledger admits it, and `ask::learned`
60reads the answers as the Worker does.
61
62Run it after changing a fact question's wording (`packages/ask`), the split
63threshold (`ask::ALSO`) or the Jev model. A run costs under a tenth of a
64cent. The key is `LMJTFY_TYPESAFE_API_KEY`, from `op-env-run`.
65
66`lmjtfy-eval halves` is the other half of `whole`: eight questions cut off
67part way, as the page sends them while they are typed, each of which must
68be read as not whole. It is a run of its own because the estate's ledger
69lets thirty requests through a minute, and anything else asking Jev on this
70machine shares them: a run that is throttled says so on each line it could
71not send, and is run again.
72
73Result, 2026-10-03, with `whole` added: 25 / 25 of the inputs (over two
74runs, the ledger having throttled each), about 830 tokens a request; 7 / 8
75of the halves ("can peng" is read as whole). 2026-10-02: 25 / 25.
76
77## In this folder
78
79| Path | What |
80| --- | --- |
81| [src/](src/) | The two evals. |
82| [Cargo.toml](Cargo.toml) | The program: `llm`, `ask` and `rules`, `ureq` for Cloudflare's REST API, and `jev-http` for the facts eval. |
83
84← Previous: [Chapter 13, tools/](../) · Up: [tools](../) · Next: [Chapter 14½, eval/src/](src/) →