README.mdpreviewREADME.mdsource272 lines · 15.7 KB · raw
1# Chapter 12: jevhooks-rules/src, the questions and the rules
2
3Here, at last, is what Jev is actually asked, and what is done with what it
4says. It is the heart of the whole plugin, in three files, one for each
5thing that is judged: [command.rs](command.rs), [stop.rs](stop.rs) and
6[model.rs](model.rs).
7
8Jev cannot be asked "is this command dangerous?" Or rather it can, and it
9will answer, and the answer will be about the word "dangerous". A command
10that merely *prints* `rm -rf ~/` as text was once stopped as if it ran it.
11So nothing here asks with an adjective. Each file writes down a short list
12of concrete things (what a command could do, how a turn could have ended),
13defines each in plain acts, and asks Jev which one this is. Then it turns
14the answer into **facts**, and **rules** over the facts decide, in code,
15where a test can hold them still.
16
17## A command: three questions, seven facts, eight rules
18
19Before a Bash command runs, the daemon sends Jev a **state** (the command,
20first 6,000 characters; the working directory; the project root; how to
21judge it) and three
22questions in one request:
23
241. **`act`, a Choice:** which of these describes what the command does? If
25   it does several, the one hardest to undo.
26
27   | Label | Defined as (abridged) | Consequential? |
28   | --- | --- | --- |
29   | `read` | only reads or prints; nothing is different afterwards | no |
30   | `build` | builds, tests, formats, lints, installs dependencies: regenerable files only | no |
31   | `edit` | creates or changes files a person wrote, or records them in version control; recoverable by an ordinary command | no |
32   | `delete` | removes or overwrites data no build regenerates and version control does not hold | yes |
33   | `history` | discards or rewrites version-control state: reset --hard, clean, rebase, force push | yes |
34   | `system` | changes the machine or the account: packages, services, dotfiles, sudo | yes |
35   | `remote` | publishes or sends something elsewhere: push, deploy, an API call that writes | yes |
36   | `unread` | runs code nobody has read: a script piped from the network into a shell | yes |
37
382. **`undo`, a Score of four levels:** how hard would it be to put
39   everything back? 0 nothing to put back; 1 one ordinary command; 2 only
40   with care or luck; 3 not from this machine.
413. **`load`, a Score of four levels:** how much of the machine does it take
42   while it runs? 0 negligible; 1 light; 2 heavy (compiles a project, a
43   whole test suite); 3 very heavy (several heavy jobs, a large release
44   build).
45
46The state carries one more field, `how_to_judge_the_command`: a paragraph
47(`WHAT_RUNS`) that says what "the command does" means, which is only what
48the shell would execute. Text carried as data (a here-document written to
49a file, a quoted string printed or searched for or sent) is not executed
50unless it is handed to something that runs it.
51
52> **Aside: said once.** That paragraph used to end each of the three
53> questions, on the theory that an instruction belongs in the question it
54> governs. Then someone asked why it was not simply in the state. So it
55> was tried: ten commands, each asked both ways, from `echo 'rm -rf ~/' >>
56> notes.md` to `echo 'rm -rf ~/' | sh`. Ten identical verdicts, and the
57> request fell from 1,555 tokens to 1,343. The theory lost.
58
59Jev answers each with a probability for every option (and for a Score, an
60expected level from 0 to 3). Rules cannot test a number like 0.57, so each
61answer is read into plain facts the moment it arrives, by a threshold:
62
63| Fact | Read from | It is "yes" when |
64| --- | --- | --- |
65| consequential | `act` | one consequential act is 60% likely or more |
66| ordinary work | `act` | read + build + edit together are 90% or more |
67| hard to undo | `undo` | the expected level is 1.6 or more |
68| easy to undo | `undo` | the expected level is 1.2 or less |
69| load | `load` | (not yes or no: the level the score rounds to) |
70
71Five facts from three questions: two facts can be two readings of one
72answer, and the question is still asked once. Two more facts are not Jev's
73at all. **Room in memory** is measured by the daemon when a rule tells it
74to (enough, short, or unmeasured), and **patience** is whether the command
75may still wait for room, which the daemon knows from how long it has
76waited.
77
78Then the rules, in order. Of the rules that hold, the first decides:
79
80| # | Rule | When | Then |
81| --- | --- | --- | --- |
82| 1 | look for room | load is known | measure the room |
83| 2 | a consequential act | consequential: yes | **ask** |
84| 3 | hard to put back | hard to undo: yes | **ask** |
85| 4 | no room yet | room: short, and patience: left | **hold** (ask me again) |
86| 5 | no room | room: short | **ask** |
87| 6 | ordinary and easily undone | ordinary work: yes, and easy to undo: yes | **allow** |
88| 7 | not surely ordinary | ordinary work: no | **pass**: the usual permission check |
89| 8 | not easily undone | easy to undo: no | **pass** |
90
91The order is the design. Measuring is first, so that a command put to you
92for what it does (rules 2 and 3) is put with both reasons when memory is
93short too. Rules 2 and 3 come before 4, so such a command is never kept
94waiting for a question it was going to get anyway. Rule 5 asks unless the
95command is *known* to have patience left: not knowing is no reason to
96wait. And rules 7 and 8 are the plugin's manners: between them they are
97every answer rule 6 does not allow, so an answer no other rule acts on
98changes nothing.
99
100> **Aside: an engine, for eight rules?** Eight `if`s would do it, and did,
101> until this chapter was written. The rules are compiled into a **Rete
102> network** (the `rete` crate, chapter 15), and what that buys is not
103> speed. It is that the engine can be *asked what it wants*: "given what
104> is known, what should happen next?" Its answers are "ask for these
105> facts, all at once", "go and do this", or "it ends like so". That is how
106> three questions go to Jev in one request without any code saying "three"
107> (the engine names every fact a live rule is waiting on, and `jev-facts`
108> turns them into one request), and how a turn's end can be settled
109> without asking Jev at all. And the network that runs is a thing you can
110> draw: the picture on the website is drawn from these very rules, so it
111> cannot show different ones.
112
113What a load "needs" rises from 0 MB (negligible) through 500 MB and
1143,000 MB to 6,000 MB (very heavy), interpolated between levels: an expected
115load of 2.5 wants 4,500 MB. "Room" is the next section. Chapter 14 has what "room" is, and
116the waiting room that rule 4 sends a command to.
117
118> **Aside: one confident act, never a sum.** An early version added up
119> the consequential acts' probabilities, and a command Jev was unsure
120> about (its likeliest act at 36%) was stopped. Take a spread like 36%
121> `unread`, 20% `system`, 14% `delete`, 30% `edit`: summed, that is 70%
122> "consequential". But it is Jev saying "I don't know", and an unsure
123> answer should change nothing. Now only the single
124> likeliest consequential act counts, and that answer is a `pass`. The
125> test `doubt_spread_over_several_consequential_acts_is_not_a_flag` pins
126> exactly that case.
127
128A real answer, from the log: a long command that read some figures and
129called out to another program came back `read` 56%, `system` 39%, undo
1300.4, load 1.3. No single consequential act reached 60%, nothing was hard to
131undo, memory was ample; but the ordinary acts came to 57%, short of 90%. So:
132`Jev: only reads (56%); nothing to undo (0.4 of 3); left to the usual
133permission check`. Not waved through, not flagged; Claude Code's own
134permission rules decided. 237 ms, 1,739 input tokens, $0.00007 (that was
135with the paragraph still in every question; a command is about 1,350
136tokens now).
137
138## A prompt: one question, four models
139
140Claude Code can give a turn to any of four models, and most of us pick one
141in the morning and use it for everything, the way you might commute in a
142lorry because some days you move a piano. Switch this on (`choose_model =
143true`, chapter 13) and every prompt you send gets one more question before
144its turn starts. The state is the new prompt, the session's earlier
145prompts, and the end of the assistant's last message; the question is a
146Score of four levels, "how much does carrying this out take?":
147
148| Level | Defined as (abridged) | Model |
149| --- | --- | --- |
150| 0 | a lookup or one mechanical step: a question about what is already in the conversation, a named command, a rename, a commit, a yes or no | `haiku` |
151| 1 | routine work of a known shape: an edit in one place, a fix whose cause is stated, a test for code that exists | `sonnet` |
152| 2 | work that needs judgment: a change across several files, a bug whose cause is not known, a review, a refactor | `opus` |
153| 3 | open-ended or long work: designing something new, weighing trade-offs nobody has stated, research, many steps unattended | `fable` |
154
155![The band after a prompt: a lookup, given to haiku](../../jevhooks-daemon/src/model.svg)
156
157Jev answers with a probability for each level, and the one fact read from
158them (`model_for` in [model.rs](model.rs)) is **the cheapest model that is
159at least 75% likely to be enough**: it walks up the levels adding
160probabilities until the sum reaches 0.75. Five rules follow, and the first
161is yours: with choosing off, the session's model stands and Jev is never
162asked. The other four give the turn to the model the fact names. Four real answers:
163
164| Prompt | Probabilities | Model | Jev took |
165| --- | --- | --- | --- |
166| what branch am I on | 1.00, 0, 0, 0 | haiku | 185 ms |
167| rename the variable foo to bar in main.rs | 0.92, 0.08, 0, 0 | haiku | 122 ms |
168| the tests in the parser crate fail intermittently on CI and nobody knows why, find the cause and fix it | 0, 0, 0.99, 0.01 | opus | 196 ms |
169| design a plugin system for this app: research how three other editors do it, weigh the trade-offs, and write the plan | 0, 0, 0, 1.00 | fable | 114 ms |
170
171Each cost about 540 input tokens: two thousandths of a cent.
172
173> **Aside: doubt rounds up.** Suppose Jev says 60% lookup, 30% routine,
174> 10% judgment. The likeliest answer is "a lookup", and a rule that took
175> the likeliest would send it to haiku, and be wrong four times in ten.
176> The 75% rule sends it to sonnet. A model too large costs a few cents; a
177> model too small costs you the turn, and then the turn again. The test
178> `doubt_about_a_prompt_rounds_up_not_down` holds that still.
179
180> **Aside: "yes".** The shortest prompt you send is often the biggest:
181> "yes, do that", in reply to a twelve-step plan. That is why the state
182> carries the assistant's last message, and why the question says to
183> judge the work asked for, not the length of the asking.
184
185> **Aside: a size is not a name.** Jev picks a size; what Claude Code
186> wants in a request is a model's full id, and it is strict about it:
187> hand it `haiku` and the turn ends with `unrecognized_model`. (The first
188> live run of this feature did exactly that.) So the daemon keeps the id
189> each size means, with Anthropic's own as the defaults and a `[models]`
190> table in the config for anyone whose ids differ.
191
192Run through a real session, one prompt each: "What is 2+2?" was answered
193by `claude-haiku-4-5-20251001`, "the parser tests fail now and then on CI
194and nobody knows why" by `claude-opus-5-5`, and "design a plugin system
195for a text editor" by `claude-fable-5-1`, with the session's own model
196set to sonnet throughout.
197
198Two honest warnings, since this one is off unless you ask for it. The
199chosen model applies to every model request of the main turn, and a
200subagent keeps whatever model it was started with. And a conversation's
201prompt cache belongs to the model that read it: a turn that goes to a
202different model than the last one reads the conversation again at full
203price, so in a long session the saving on a small prompt can be less than
204it looks.
205
206## A turn's end: one question
207
208When a turn is about to end, the state is the session's last five prompts
209(each first 3,000 characters) and the end of the assistant's final message
210(last 6,000 characters), and the one question is a Choice of endings:
211`finished`, `waiting` (needs something only the developer can give),
212`blocked` (names an obstacle), `running` (work still going, will report),
213and `stopped-early` (work asked for is left undone and no obstacle is
214named). Only at 85% or more for `stopped-early` is the end refused: a wrong
215refusal costs the user a wasted turn.
216
217Three rules come first, on facts the daemon knows before anything is asked,
218and when one of them holds Jev is not asked at all: a `Stop` that follows
219this plugin's own refusal (`stop_hook_active`) is let through, so one
220refusal can never become a loop; a turn with background tasks or scheduled
221wakeups still pending has paused, not ended; and with no prompt or final
222message there is nothing to judge. That is the engine's doing, not a
223special case: a rule that already holds decides, and nothing is asked for
224a rule that can no longer matter.
225
226> **Try it.** `cargo test -p jevhooks-rules`. Every test makes up an
227> answer the way Jev would send one (a real response body, parsed and
228> learned from, not a shortcut past the thresholds) and asks the network
229> what follows: `ordinary_work_is_allowed`,
230> `consequential_acts_are_put_to_the_user`,
231> `an_unsure_answer_is_neither_allowed_nor_asked`,
232> `a_command_with_no_room_waits_unless_it_is_risky_or_has_waited`,
233> `stop_blocks_only_when_confident`,
234> `what_is_known_already_is_never_asked_about`,
235> `with_choosing_off_nothing_is_asked`. Change a threshold, or swap two
236> rules, and watch which ones notice.
237>
238> And if you have a key, `mise run test:live` puts the questions, as they
239> are worded today, to the real Jev: a release build is heavy and
240> ordinary, a force push is not ordinary, a command that only *writes
241> down* `rm -rf ~/` is not a deletion, four prompts land on the models
242> you would expect, and a turn that did its work is let end. It costs
243> about a thirtieth of a cent.
244
245## For the people who maintain it
246
247Each module has the same parts, by the same names: `Fact`, `Value`, the
248domain (`Command`, `Stop`, `Model`, which is also the `jev_facts::Source`
249of its questions), `RULES`, `network()`, `state(...)` (what Jev is shown),
250`taught(...)` (a response read into an `Asked`: the facts, and the numbers
251behind them for the line and the log), and `Asked::made_up` for tests.
252
253| Path | What |
254| --- | --- |
255| [lib.rs](lib.rs) | `MODEL` (`jev-1.13.0`, pinned beside the thresholds set against it), `head` and `tail` (clipping by characters), `Never` (the effect or note of a domain that has none), and how a made-up response body is written. |
256| [command.rs](command.rs) | A command: `Act` and its eight definitions, the undo and load levels, `WHAT_RUNS`, the thresholds, `Effect::Measure`, `needed_mb`, `room`, `patience`, `Asked::settle` (the network run with a measured room; the verdict and the line) and `Asked::is_risky`. |
257| [stop.rs](stop.rs) | A turn's end: `ENDINGS`, `before` (the three facts known already), `End` (five ways it ends, each with its verdict and line). |
258| [model.rs](model.rs) | A prompt: the need levels, `model_for`, `before` (whether choosing is on), `End::Give(tier)`. |
259
260The thresholds, all reported by the daemon's `status`:
261
262| Constant | Value | Meaning |
263| --- | --- | --- |
264| `command::ASK_AT_CONSEQUENTIAL` | 0.60 | One consequential act this likely: ask. |
265| `command::ASK_AT_UNDO` | 1.60 | Expected undo level this high: ask. |
266| `command::ALLOW_AT_ORDINARY` | 0.90 | Ordinary acts together this likely ... |
267| `command::ALLOW_UNDO_AT_MOST` | 1.20 | ... and undo at most this: allow. |
268| `command::LOAD_NEEDS_MB` | 0, 500, 3000, 6000 | Memory each load level should find free. |
269| `stop::BLOCK_AT_LEAST` | 0.85 | `stopped-early` this likely: refuse the end. |
270| `model::ENOUGH_AT` | 0.75 | A prompt goes to the cheapest model this likely to be enough. |
271
272← Previous: [Chapter 11, jevhooks-rules/](../) · Up: [jevhooks-rules](../) · Next: [Chapter 13, jevhooks-daemon/](../../jevhooks-daemon/) →