For agents, on top of README.md, which they read first.

Notes for agents

The questions are yours to refine, without asking. Standing permission from the owner (2026-10-02). When a verdict is wrong or noisy, fix the question and say what changed. Two rules for doing it: define every term in the option or level it belongs to, as concrete acts, never a bare adjective ("destructive" stopped a harmless note in another repository); and add a question to the same request rather than overloading one, reading it into a fact a rule can test. A request mixes Choice, Score and Noul questions freely; its limit is size, not count (Jev's docs, read 2026-10-02: 64k tokens for the state plus all questions, 32k for the state plus the longest question, and at most 255 options per Choice).

Jev judges words unless told what runs. A command that only writes, prints or posts rm -rf ~/ as text was stopped as if it ran it. WHAT_RUNS in command.rs goes in the state of every command request for that reason (how_to_judge_the_command), once, where every question reads it: asked both ways on 2026-10-06, ten commands got the same verdicts with it once in the state as with it ending each of the three questions, for 212 fewer tokens. mise run test:live holds the cases that depend on it.

Only one confident act stops a command, never a sum. Summing the consequential acts' probabilities turned an unsure answer (top act 36%) into a stop. Unsure is Pass. consequential in command.rs takes the maximum for that reason.

An option's label is what Jev answers with. Act::label and ENDINGS' first fields are sent as the Choice's labels and matched back (Act::from_label, STOPPED_EARLY in stop.rs); renaming one without the other silently zeroes its probability. every_act_has_its_own_label and the_refused_ending_is_one_of_the_options guard them.

MODEL is pinned (jev-1.13.0 in lib.rs). The thresholds were set against that model's answers; moving the pin means re-reading decisions.jsonl under the new one before trusting them.

Rule order is the policy; do not reorder to tidy. In command::RULES: measuring first (so a risky command's line carries the memory reason too), the two rules on what it does before the two on room (so a risky command is never held), and "no room" asks unless patience is known to be left. In stop::RULES, "refused once already" stays first: it is what keeps one refusal of a turn's end from becoming a loop of refusals. Each is pinned by a test; a reorder that passes the tests is still a policy change to say out loud.

No catch-all rule. A rule needs at least one test, and "otherwise" was first written as Test::Known(Fact::Ordinary), which held for every answered command and so lit up "Jev is not sure" in the website's drawing of a command Jev was sure about (the owner noticed, 2026-10-06). Write the complement out as rules that are each true of what they end: not surely ordinary and not easily undone are together everything ordinary and easily undone is not. only_the_rule_that_applies_holds guards it.

A fact that is not Jev's must never reach question. asked_for is false for it, so the engine never asks for it; question returns an error for it rather than a made-up question. Keep the two in step when adding a fact.

A borderline live case is not a test of a threshold. bash cleanup.sh (a script nobody has shown Jev) came back "deletes" at 55, 56 and 59% in three runs and over 60% in a fourth on 2026-10-06, on the code before and after the move here: Jev's answers to one request vary by a few points. The live test asserts only that it is never allowed. Do not tune a threshold to make such a case pass.