For agents, on top of README.md, which they read first.
Notes for agents
The questions are yours to refine, without asking. Standing permission from the owner (2026-10-02). When a verdict is wrong or noisy, fix the question and say what changed. Two rules for doing it: define every term in the option or level it belongs to, as concrete acts, never a bare adjective ("destructive" stopped a harmless note in another repository); and add a question to the same request rather than overloading one, reading it into a fact a rule can test. A request mixes Choice, Score and Noul questions freely; its limit is size, not count (Jev's docs, read 2026-10-02: 64k tokens for the state plus all questions, 32k for the state plus the longest question, and at most 255 options per Choice).
Jev judges words unless told what runs. A command that only writes, prints or
posts rm -rf ~/ as text was stopped as if it ran it. WHAT_RUNS in command.rs
goes in the state of every command request for that reason
(how_to_judge_the_command), once, where every question reads it: asked both ways on
2026-10-06, ten commands got the same verdicts with it once in the state as with it
ending each of the three questions, for 212 fewer tokens. mise run test:live holds
the cases that depend on it.
Only one confident act stops a command, never a sum. Summing the consequential
acts' probabilities turned an unsure answer (top act 36%) into a stop. Unsure is
Pass. consequential in command.rs takes the maximum for that reason.
An option's label is what Jev answers with. Act::label and ENDINGS' first
fields are sent as the Choice's labels and matched back (Act::from_label,
STOPPED_EARLY in stop.rs); renaming one without the other silently zeroes its
probability. every_act_has_its_own_label and
the_refused_ending_is_one_of_the_options guard them.
MODEL is pinned (jev-1.13.0 in lib.rs). The thresholds were set against
that model's answers; moving the pin means re-reading decisions.jsonl under the new
one before trusting them.
Rule order is the policy; do not reorder to tidy. In command::RULES: measuring
first (so a risky command's line carries the memory reason too), the two rules on what
it does before the two on room (so a risky command is never held), and "no room" asks
unless patience is known to be left. In stop::RULES, "refused once already" stays
first: it is what keeps one refusal of a turn's end from becoming a loop of refusals.
Each is pinned by a test; a reorder that passes the tests is still a policy change to
say out loud.
No catch-all rule. A rule needs at least one test, and "otherwise" was first
written as Test::Known(Fact::Ordinary), which held for every answered command and so
lit up "Jev is not sure" in the website's drawing of a command Jev was sure about (the
owner noticed, 2026-10-06). Write the complement out as rules that are each true of
what they end: not surely ordinary and not easily undone are together everything
ordinary and easily undone is not. only_the_rule_that_applies_holds guards it.
A fact that is not Jev's must never reach question. asked_for is false for
it, so the engine never asks for it; question returns an error for it rather than a
made-up question. Keep the two in step when adding a fact.
A borderline live case is not a test of a threshold. bash cleanup.sh (a script
nobody has shown Jev) came back "deletes" at 55, 56 and 59% in three runs and over
60% in a fourth on 2026-10-06, on the code before and after the move here: Jev's
answers to one request vary by a few points. The live test asserts only that it is
never allowed. Do not tune a threshold to make such a case pass.