Chapter 12: jevhooks-rules/src, the questions and the rules

Here, at last, is what Jev is actually asked, and what is done with what it says. It is the heart of the whole plugin, in three files, one for each thing that is judged: command.rs, stop.rs and model.rs.

Jev cannot be asked "is this command dangerous?" Or rather it can, and it will answer, and the answer will be about the word "dangerous". A command that merely prints rm -rf ~/ as text was once stopped as if it ran it. So nothing here asks with an adjective. Each file writes down a short list of concrete things (what a command could do, how a turn could have ended), defines each in plain acts, and asks Jev which one this is. Then it turns the answer into facts, and rules over the facts decide, in code, where a test can hold them still.

A command: three questions, seven facts, eight rules

Before a Bash command runs, the daemon sends Jev a state (the command, first 6,000 characters; the working directory; the project root; how to judge it) and three questions in one request:

  1. act, a Choice: which of these describes what the command does? If it does several, the one hardest to undo.

    LabelDefined as (abridged)Consequential?
    readonly reads or prints; nothing is different afterwardsno
    buildbuilds, tests, formats, lints, installs dependencies: regenerable files onlyno
    editcreates or changes files a person wrote, or records them in version control; recoverable by an ordinary commandno
    deleteremoves or overwrites data no build regenerates and version control does not holdyes
    historydiscards or rewrites version-control state: reset --hard, clean, rebase, force pushyes
    systemchanges the machine or the account: packages, services, dotfiles, sudoyes
    remotepublishes or sends something elsewhere: push, deploy, an API call that writesyes
    unreadruns code nobody has read: a script piped from the network into a shellyes
  2. undo, a Score of four levels: how hard would it be to put everything back? 0 nothing to put back; 1 one ordinary command; 2 only with care or luck; 3 not from this machine.

  3. load, a Score of four levels: how much of the machine does it take while it runs? 0 negligible; 1 light; 2 heavy (compiles a project, a whole test suite); 3 very heavy (several heavy jobs, a large release build).

The state carries one more field, how_to_judge_the_command: a paragraph (WHAT_RUNS) that says what "the command does" means, which is only what the shell would execute. Text carried as data (a here-document written to a file, a quoted string printed or searched for or sent) is not executed unless it is handed to something that runs it.

Aside: said once. That paragraph used to end each of the three questions, on the theory that an instruction belongs in the question it governs. Then someone asked why it was not simply in the state. So it was tried: ten commands, each asked both ways, from echo 'rm -rf ~/' >> notes.md to echo 'rm -rf ~/' | sh. Ten identical verdicts, and the request fell from 1,555 tokens to 1,343. The theory lost.

Jev answers each with a probability for every option (and for a Score, an expected level from 0 to 3). Rules cannot test a number like 0.57, so each answer is read into plain facts the moment it arrives, by a threshold:

FactRead fromIt is "yes" when
consequentialactone consequential act is 60% likely or more
ordinary workactread + build + edit together are 90% or more
hard to undoundothe expected level is 1.6 or more
easy to undoundothe expected level is 1.2 or less
loadload(not yes or no: the level the score rounds to)

Five facts from three questions: two facts can be two readings of one answer, and the question is still asked once. Two more facts are not Jev's at all. Room in memory is measured by the daemon when a rule tells it to (enough, short, or unmeasured), and patience is whether the command may still wait for room, which the daemon knows from how long it has waited.

Then the rules, in order. Of the rules that hold, the first decides:

#RuleWhenThen
1look for roomload is knownmeasure the room
2a consequential actconsequential: yesask
3hard to put backhard to undo: yesask
4no room yetroom: short, and patience: lefthold (ask me again)
5no roomroom: shortask
6ordinary and easily undoneordinary work: yes, and easy to undo: yesallow
7not surely ordinaryordinary work: nopass: the usual permission check
8not easily undoneeasy to undo: nopass

The order is the design. Measuring is first, so that a command put to you for what it does (rules 2 and 3) is put with both reasons when memory is short too. Rules 2 and 3 come before 4, so such a command is never kept waiting for a question it was going to get anyway. Rule 5 asks unless the command is known to have patience left: not knowing is no reason to wait. And rules 7 and 8 are the plugin's manners: between them they are every answer rule 6 does not allow, so an answer no other rule acts on changes nothing.

Aside: an engine, for eight rules? Eight ifs would do it, and did, until this chapter was written. The rules are compiled into a Rete network (the rete crate, chapter 15), and what that buys is not speed. It is that the engine can be asked what it wants: "given what is known, what should happen next?" Its answers are "ask for these facts, all at once", "go and do this", or "it ends like so". That is how three questions go to Jev in one request without any code saying "three" (the engine names every fact a live rule is waiting on, and jev-facts turns them into one request), and how a turn's end can be settled without asking Jev at all. And the network that runs is a thing you can draw: the picture on the website is drawn from these very rules, so it cannot show different ones.

What a load "needs" rises from 0 MB (negligible) through 500 MB and 3,000 MB to 6,000 MB (very heavy), interpolated between levels: an expected load of 2.5 wants 4,500 MB. "Room" is the next section. Chapter 14 has what "room" is, and the waiting room that rule 4 sends a command to.

Aside: one confident act, never a sum. An early version added up the consequential acts' probabilities, and a command Jev was unsure about (its likeliest act at 36%) was stopped. Take a spread like 36% unread, 20% system, 14% delete, 30% edit: summed, that is 70% "consequential". But it is Jev saying "I don't know", and an unsure answer should change nothing. Now only the single likeliest consequential act counts, and that answer is a pass. The test doubt_spread_over_several_consequential_acts_is_not_a_flag pins exactly that case.

A real answer, from the log: a long command that read some figures and called out to another program came back read 56%, system 39%, undo 0.4, load 1.3. No single consequential act reached 60%, nothing was hard to undo, memory was ample; but the ordinary acts came to 57%, short of 90%. So: Jev: only reads (56%); nothing to undo (0.4 of 3); left to the usual permission check. Not waved through, not flagged; Claude Code's own permission rules decided. 237 ms, 1,739 input tokens, $0.00007 (that was with the paragraph still in every question; a command is about 1,350 tokens now).

A prompt: one question, four models

Claude Code can give a turn to any of four models, and most of us pick one in the morning and use it for everything, the way you might commute in a lorry because some days you move a piano. Switch this on (choose_model = true, chapter 13) and every prompt you send gets one more question before its turn starts. The state is the new prompt, the session's earlier prompts, and the end of the assistant's last message; the question is a Score of four levels, "how much does carrying this out take?":

LevelDefined as (abridged)Model
0a lookup or one mechanical step: a question about what is already in the conversation, a named command, a rename, a commit, a yes or nohaiku
1routine work of a known shape: an edit in one place, a fix whose cause is stated, a test for code that existssonnet
2work that needs judgment: a change across several files, a bug whose cause is not known, a review, a refactoropus
3open-ended or long work: designing something new, weighing trade-offs nobody has stated, research, many steps unattendedfable

The band after a prompt: a lookup, given to haiku

Jev answers with a probability for each level, and the one fact read from them (model_for in model.rs) is the cheapest model that is at least 75% likely to be enough: it walks up the levels adding probabilities until the sum reaches 0.75. Five rules follow, and the first is yours: with choosing off, the session's model stands and Jev is never asked. The other four give the turn to the model the fact names. Four real answers:

PromptProbabilitiesModelJev took
what branch am I on1.00, 0, 0, 0haiku185 ms
rename the variable foo to bar in main.rs0.92, 0.08, 0, 0haiku122 ms
the tests in the parser crate fail intermittently on CI and nobody knows why, find the cause and fix it0, 0, 0.99, 0.01opus196 ms
design a plugin system for this app: research how three other editors do it, weigh the trade-offs, and write the plan0, 0, 0, 1.00fable114 ms

Each cost about 540 input tokens: two thousandths of a cent.

Aside: doubt rounds up. Suppose Jev says 60% lookup, 30% routine, 10% judgment. The likeliest answer is "a lookup", and a rule that took the likeliest would send it to haiku, and be wrong four times in ten. The 75% rule sends it to sonnet. A model too large costs a few cents; a model too small costs you the turn, and then the turn again. The test doubt_about_a_prompt_rounds_up_not_down holds that still.

Aside: "yes". The shortest prompt you send is often the biggest: "yes, do that", in reply to a twelve-step plan. That is why the state carries the assistant's last message, and why the question says to judge the work asked for, not the length of the asking.

Aside: a size is not a name. Jev picks a size; what Claude Code wants in a request is a model's full id, and it is strict about it: hand it haiku and the turn ends with unrecognized_model. (The first live run of this feature did exactly that.) So the daemon keeps the id each size means, with Anthropic's own as the defaults and a [models] table in the config for anyone whose ids differ.

Run through a real session, one prompt each: "What is 2+2?" was answered by claude-haiku-4-5-20251001, "the parser tests fail now and then on CI and nobody knows why" by claude-opus-5-5, and "design a plugin system for a text editor" by claude-fable-5-1, with the session's own model set to sonnet throughout.

Two honest warnings, since this one is off unless you ask for it. The chosen model applies to every model request of the main turn, and a subagent keeps whatever model it was started with. And a conversation's prompt cache belongs to the model that read it: a turn that goes to a different model than the last one reads the conversation again at full price, so in a long session the saving on a small prompt can be less than it looks.

A turn's end: one question

When a turn is about to end, the state is the session's last five prompts (each first 3,000 characters) and the end of the assistant's final message (last 6,000 characters), and the one question is a Choice of endings: finished, waiting (needs something only the developer can give), blocked (names an obstacle), running (work still going, will report), and stopped-early (work asked for is left undone and no obstacle is named). Only at 85% or more for stopped-early is the end refused: a wrong refusal costs the user a wasted turn.

Three rules come first, on facts the daemon knows before anything is asked, and when one of them holds Jev is not asked at all: a Stop that follows this plugin's own refusal (stop_hook_active) is let through, so one refusal can never become a loop; a turn with background tasks or scheduled wakeups still pending has paused, not ended; and with no prompt or final message there is nothing to judge. That is the engine's doing, not a special case: a rule that already holds decides, and nothing is asked for a rule that can no longer matter.

Try it. cargo test -p jevhooks-rules. Every test makes up an answer the way Jev would send one (a real response body, parsed and learned from, not a shortcut past the thresholds) and asks the network what follows: ordinary_work_is_allowed, consequential_acts_are_put_to_the_user, an_unsure_answer_is_neither_allowed_nor_asked, a_command_with_no_room_waits_unless_it_is_risky_or_has_waited, stop_blocks_only_when_confident, what_is_known_already_is_never_asked_about, with_choosing_off_nothing_is_asked. Change a threshold, or swap two rules, and watch which ones notice.

And if you have a key, mise run test:live puts the questions, as they are worded today, to the real Jev: a release build is heavy and ordinary, a force push is not ordinary, a command that only writes down rm -rf ~/ is not a deletion, four prompts land on the models you would expect, and a turn that did its work is let end. It costs about a thirtieth of a cent.

For the people who maintain it

Each module has the same parts, by the same names: Fact, Value, the domain (Command, Stop, Model, which is also the jev_facts::Source of its questions), RULES, network(), state(...) (what Jev is shown), taught(...) (a response read into an Asked: the facts, and the numbers behind them for the line and the log), and Asked::made_up for tests.

PathWhat
lib.rsMODEL (jev-1.13.0, pinned beside the thresholds set against it), head and tail (clipping by characters), Never (the effect or note of a domain that has none), and how a made-up response body is written.
command.rsA command: Act and its eight definitions, the undo and load levels, WHAT_RUNS, the thresholds, Effect::Measure, needed_mb, room, patience, Asked::settle (the network run with a measured room; the verdict and the line) and Asked::is_risky.
stop.rsA turn's end: ENDINGS, before (the three facts known already), End (five ways it ends, each with its verdict and line).
model.rsA prompt: the need levels, model_for, before (whether choosing is on), End::Give(tier).

The thresholds, all reported by the daemon's status:

ConstantValueMeaning
command::ASK_AT_CONSEQUENTIAL0.60One consequential act this likely: ask.
command::ASK_AT_UNDO1.60Expected undo level this high: ask.
command::ALLOW_AT_ORDINARY0.90Ordinary acts together this likely ...
command::ALLOW_UNDO_AT_MOST1.20... and undo at most this: allow.
command::LOAD_NEEDS_MB0, 500, 3000, 6000Memory each load level should find free.
stop::BLOCK_AT_LEAST0.85stopped-early this likely: refuse the end.
model::ENOUGH_AT0.75A prompt goes to the cheapest model this likely to be enough.

← Previous: Chapter 11, jevhooks-rules/ · Up: jevhooks-rules · Next: Chapter 13, jevhooks-daemon/ →