Chapter 12: jevhooks-rules/src, the questions and the rules
Here, at last, is what Jev is actually asked, and what is done with what it says. It is the heart of the whole plugin, in three files, one for each thing that is judged: command.rs, stop.rs and model.rs.
Jev cannot be asked "is this command dangerous?" Or rather it can, and it
will answer, and the answer will be about the word "dangerous". A command
that merely prints rm -rf ~/ as text was once stopped as if it ran it.
So nothing here asks with an adjective. Each file writes down a short list
of concrete things (what a command could do, how a turn could have ended),
defines each in plain acts, and asks Jev which one this is. Then it turns
the answer into facts, and rules over the facts decide, in code,
where a test can hold them still.
A command: three questions, seven facts, eight rules
Before a Bash command runs, the daemon sends Jev a state (the command, first 6,000 characters; the working directory; the project root; how to judge it) and three questions in one request:
-
act, a Choice: which of these describes what the command does? If it does several, the one hardest to undo.Label Defined as (abridged) Consequential? readonly reads or prints; nothing is different afterwards no buildbuilds, tests, formats, lints, installs dependencies: regenerable files only no editcreates or changes files a person wrote, or records them in version control; recoverable by an ordinary command no deleteremoves or overwrites data no build regenerates and version control does not hold yes historydiscards or rewrites version-control state: reset --hard, clean, rebase, force push yes systemchanges the machine or the account: packages, services, dotfiles, sudo yes remotepublishes or sends something elsewhere: push, deploy, an API call that writes yes unreadruns code nobody has read: a script piped from the network into a shell yes -
undo, a Score of four levels: how hard would it be to put everything back? 0 nothing to put back; 1 one ordinary command; 2 only with care or luck; 3 not from this machine. -
load, a Score of four levels: how much of the machine does it take while it runs? 0 negligible; 1 light; 2 heavy (compiles a project, a whole test suite); 3 very heavy (several heavy jobs, a large release build).
The state carries one more field, how_to_judge_the_command: a paragraph
(WHAT_RUNS) that says what "the command does" means, which is only what
the shell would execute. Text carried as data (a here-document written to
a file, a quoted string printed or searched for or sent) is not executed
unless it is handed to something that runs it.
Aside: said once. That paragraph used to end each of the three questions, on the theory that an instruction belongs in the question it governs. Then someone asked why it was not simply in the state. So it was tried: ten commands, each asked both ways, from
echo 'rm -rf ~/' >> notes.mdtoecho 'rm -rf ~/' | sh. Ten identical verdicts, and the request fell from 1,555 tokens to 1,343. The theory lost.
Jev answers each with a probability for every option (and for a Score, an expected level from 0 to 3). Rules cannot test a number like 0.57, so each answer is read into plain facts the moment it arrives, by a threshold:
| Fact | Read from | It is "yes" when |
|---|---|---|
| consequential | act | one consequential act is 60% likely or more |
| ordinary work | act | read + build + edit together are 90% or more |
| hard to undo | undo | the expected level is 1.6 or more |
| easy to undo | undo | the expected level is 1.2 or less |
| load | load | (not yes or no: the level the score rounds to) |
Five facts from three questions: two facts can be two readings of one answer, and the question is still asked once. Two more facts are not Jev's at all. Room in memory is measured by the daemon when a rule tells it to (enough, short, or unmeasured), and patience is whether the command may still wait for room, which the daemon knows from how long it has waited.
Then the rules, in order. Of the rules that hold, the first decides:
| # | Rule | When | Then |
|---|---|---|---|
| 1 | look for room | load is known | measure the room |
| 2 | a consequential act | consequential: yes | ask |
| 3 | hard to put back | hard to undo: yes | ask |
| 4 | no room yet | room: short, and patience: left | hold (ask me again) |
| 5 | no room | room: short | ask |
| 6 | ordinary and easily undone | ordinary work: yes, and easy to undo: yes | allow |
| 7 | not surely ordinary | ordinary work: no | pass: the usual permission check |
| 8 | not easily undone | easy to undo: no | pass |
The order is the design. Measuring is first, so that a command put to you for what it does (rules 2 and 3) is put with both reasons when memory is short too. Rules 2 and 3 come before 4, so such a command is never kept waiting for a question it was going to get anyway. Rule 5 asks unless the command is known to have patience left: not knowing is no reason to wait. And rules 7 and 8 are the plugin's manners: between them they are every answer rule 6 does not allow, so an answer no other rule acts on changes nothing.
Aside: an engine, for eight rules? Eight
ifs would do it, and did, until this chapter was written. The rules are compiled into a Rete network (theretecrate, chapter 15), and what that buys is not speed. It is that the engine can be asked what it wants: "given what is known, what should happen next?" Its answers are "ask for these facts, all at once", "go and do this", or "it ends like so". That is how three questions go to Jev in one request without any code saying "three" (the engine names every fact a live rule is waiting on, andjev-factsturns them into one request), and how a turn's end can be settled without asking Jev at all. And the network that runs is a thing you can draw: the picture on the website is drawn from these very rules, so it cannot show different ones.
What a load "needs" rises from 0 MB (negligible) through 500 MB and 3,000 MB to 6,000 MB (very heavy), interpolated between levels: an expected load of 2.5 wants 4,500 MB. "Room" is the next section. Chapter 14 has what "room" is, and the waiting room that rule 4 sends a command to.
Aside: one confident act, never a sum. An early version added up the consequential acts' probabilities, and a command Jev was unsure about (its likeliest act at 36%) was stopped. Take a spread like 36%
unread, 20%system, 14%delete, 30%edit: summed, that is 70% "consequential". But it is Jev saying "I don't know", and an unsure answer should change nothing. Now only the single likeliest consequential act counts, and that answer is apass. The testdoubt_spread_over_several_consequential_acts_is_not_a_flagpins exactly that case.
A real answer, from the log: a long command that read some figures and
called out to another program came back read 56%, system 39%, undo
0.4, load 1.3. No single consequential act reached 60%, nothing was hard to
undo, memory was ample; but the ordinary acts came to 57%, short of 90%. So:
Jev: only reads (56%); nothing to undo (0.4 of 3); left to the usual permission check. Not waved through, not flagged; Claude Code's own
permission rules decided. 237 ms, 1,739 input tokens, $0.00007 (that was
with the paragraph still in every question; a command is about 1,350
tokens now).
A prompt: one question, four models
Claude Code can give a turn to any of four models, and most of us pick one
in the morning and use it for everything, the way you might commute in a
lorry because some days you move a piano. Switch this on (choose_model = true, chapter 13) and every prompt you send gets one more question before
its turn starts. The state is the new prompt, the session's earlier
prompts, and the end of the assistant's last message; the question is a
Score of four levels, "how much does carrying this out take?":
| Level | Defined as (abridged) | Model |
|---|---|---|
| 0 | a lookup or one mechanical step: a question about what is already in the conversation, a named command, a rename, a commit, a yes or no | haiku |
| 1 | routine work of a known shape: an edit in one place, a fix whose cause is stated, a test for code that exists | sonnet |
| 2 | work that needs judgment: a change across several files, a bug whose cause is not known, a review, a refactor | opus |
| 3 | open-ended or long work: designing something new, weighing trade-offs nobody has stated, research, many steps unattended | fable |
Jev answers with a probability for each level, and the one fact read from
them (model_for in model.rs) is the cheapest model that is
at least 75% likely to be enough: it walks up the levels adding
probabilities until the sum reaches 0.75. Five rules follow, and the first
is yours: with choosing off, the session's model stands and Jev is never
asked. The other four give the turn to the model the fact names. Four real answers:
| Prompt | Probabilities | Model | Jev took |
|---|---|---|---|
| what branch am I on | 1.00, 0, 0, 0 | haiku | 185 ms |
| rename the variable foo to bar in main.rs | 0.92, 0.08, 0, 0 | haiku | 122 ms |
| the tests in the parser crate fail intermittently on CI and nobody knows why, find the cause and fix it | 0, 0, 0.99, 0.01 | opus | 196 ms |
| design a plugin system for this app: research how three other editors do it, weigh the trade-offs, and write the plan | 0, 0, 0, 1.00 | fable | 114 ms |
Each cost about 540 input tokens: two thousandths of a cent.
Aside: doubt rounds up. Suppose Jev says 60% lookup, 30% routine, 10% judgment. The likeliest answer is "a lookup", and a rule that took the likeliest would send it to haiku, and be wrong four times in ten. The 75% rule sends it to sonnet. A model too large costs a few cents; a model too small costs you the turn, and then the turn again. The test
doubt_about_a_prompt_rounds_up_not_downholds that still.
Aside: "yes". The shortest prompt you send is often the biggest: "yes, do that", in reply to a twelve-step plan. That is why the state carries the assistant's last message, and why the question says to judge the work asked for, not the length of the asking.
Aside: a size is not a name. Jev picks a size; what Claude Code wants in a request is a model's full id, and it is strict about it: hand it
haikuand the turn ends withunrecognized_model. (The first live run of this feature did exactly that.) So the daemon keeps the id each size means, with Anthropic's own as the defaults and a[models]table in the config for anyone whose ids differ.
Run through a real session, one prompt each: "What is 2+2?" was answered
by claude-haiku-4-5-20251001, "the parser tests fail now and then on CI
and nobody knows why" by claude-opus-5-5, and "design a plugin system
for a text editor" by claude-fable-5-1, with the session's own model
set to sonnet throughout.
Two honest warnings, since this one is off unless you ask for it. The chosen model applies to every model request of the main turn, and a subagent keeps whatever model it was started with. And a conversation's prompt cache belongs to the model that read it: a turn that goes to a different model than the last one reads the conversation again at full price, so in a long session the saving on a small prompt can be less than it looks.
A turn's end: one question
When a turn is about to end, the state is the session's last five prompts
(each first 3,000 characters) and the end of the assistant's final message
(last 6,000 characters), and the one question is a Choice of endings:
finished, waiting (needs something only the developer can give),
blocked (names an obstacle), running (work still going, will report),
and stopped-early (work asked for is left undone and no obstacle is
named). Only at 85% or more for stopped-early is the end refused: a wrong
refusal costs the user a wasted turn.
Three rules come first, on facts the daemon knows before anything is asked,
and when one of them holds Jev is not asked at all: a Stop that follows
this plugin's own refusal (stop_hook_active) is let through, so one
refusal can never become a loop; a turn with background tasks or scheduled
wakeups still pending has paused, not ended; and with no prompt or final
message there is nothing to judge. That is the engine's doing, not a
special case: a rule that already holds decides, and nothing is asked for
a rule that can no longer matter.
Try it.
cargo test -p jevhooks-rules. Every test makes up an answer the way Jev would send one (a real response body, parsed and learned from, not a shortcut past the thresholds) and asks the network what follows:ordinary_work_is_allowed,consequential_acts_are_put_to_the_user,an_unsure_answer_is_neither_allowed_nor_asked,a_command_with_no_room_waits_unless_it_is_risky_or_has_waited,stop_blocks_only_when_confident,what_is_known_already_is_never_asked_about,with_choosing_off_nothing_is_asked. Change a threshold, or swap two rules, and watch which ones notice.And if you have a key,
mise run test:liveputs the questions, as they are worded today, to the real Jev: a release build is heavy and ordinary, a force push is not ordinary, a command that only writes downrm -rf ~/is not a deletion, four prompts land on the models you would expect, and a turn that did its work is let end. It costs about a thirtieth of a cent.
For the people who maintain it
Each module has the same parts, by the same names: Fact, Value, the
domain (Command, Stop, Model, which is also the jev_facts::Source
of its questions), RULES, network(), state(...) (what Jev is shown),
taught(...) (a response read into an Asked: the facts, and the numbers
behind them for the line and the log), and Asked::made_up for tests.
| Path | What |
|---|---|
| lib.rs | MODEL (jev-1.13.0, pinned beside the thresholds set against it), head and tail (clipping by characters), Never (the effect or note of a domain that has none), and how a made-up response body is written. |
| command.rs | A command: Act and its eight definitions, the undo and load levels, WHAT_RUNS, the thresholds, Effect::Measure, needed_mb, room, patience, Asked::settle (the network run with a measured room; the verdict and the line) and Asked::is_risky. |
| stop.rs | A turn's end: ENDINGS, before (the three facts known already), End (five ways it ends, each with its verdict and line). |
| model.rs | A prompt: the need levels, model_for, before (whether choosing is on), End::Give(tier). |
The thresholds, all reported by the daemon's status:
| Constant | Value | Meaning |
|---|---|---|
command::ASK_AT_CONSEQUENTIAL | 0.60 | One consequential act this likely: ask. |
command::ASK_AT_UNDO | 1.60 | Expected undo level this high: ask. |
command::ALLOW_AT_ORDINARY | 0.90 | Ordinary acts together this likely ... |
command::ALLOW_UNDO_AT_MOST | 1.20 | ... and undo at most this: allow. |
command::LOAD_NEEDS_MB | 0, 500, 3000, 6000 | Memory each load level should find free. |
stop::BLOCK_AT_LEAST | 0.85 | stopped-early this likely: refuse the end. |
model::ENOUGH_AT | 0.75 | A prompt goes to the cheapest model this likely to be enough. |
← Previous: Chapter 11, jevhooks-rules/ · Up: jevhooks-rules · Next: Chapter 13, jevhooks-daemon/ →