lmjtfy.git / apps / lmjtfy / README.md
README.mdpreviewREADME.mdsource664 lines · 36.8 KB · raw

Chapter 2: the Worker, where everything actually happens

Everything you see at https://lmjtfy.fun comes out of this folder. It is one Cloudflare Worker: a small program Cloudflare runs at its edge, started fresh for each request in milliseconds, written here in Rust and compiled to WebAssembly. There is no server to keep running. A request arrives, the Worker answers it, and that is that.

Two things have to outlive a single request: what was already asked (so it is never asked again) and how much of today's budget is left. For those the Worker has two Durable Objects, which are the closest Cloudflare comes to a tiny server with a memory: one object per name, anywhere in the world, with its own SQLite. The Worker talks to them; they remember.

flowchart LR
  V["Visitor's browser"] -->|"/, /ask, /gate"| W["Worker (src/lib.rs)"]
  V -->|"/live (socket)"| A
  G["git clone"] -->|"/lmjtfy.git"| W
  W --> A["Archive object: every call, the feed, the sockets"]
  W --> B["Budget object: today's spend, per-visitor limits"]
  A --> J["Jev (TypeSafe)"]
  A --> L["LLM (Workers AI)"]
  W --> GH["GitHub (read-only token)"]

What it serves

AddressWhat happens
/ and /?q=...The page. With ?q=, it types the question in for you and asks.
POST /askAnswers a question as a stream: the whole transcript, re-sent as each call finishes.
POST /rateA browser's 👍 or 👎 on Jev's answer, and the votes after it.
POST /commentThe words a browser adds to its vote, from the box that opens under the 👍 and 👎. Trimmed, 1,000 characters at most, five a minute a visitor, and only while the browser has a vote; answered with "Thanks" or why not, in fixed words. The text is kept (below) and is in no log line or event.
POST /seenA page's report of itself as it is left: how long it was in view, how far down, the screen; or a link followed off the site. Kept, and answered with nothing.
POST /gateRun as you type, 300 ms after each pause. First, the questions already on the feed that what is typed could be the start of, which asks nobody anything; then the facts request, so the answer starts from what Jev already said. Text that stops part way is said to be unfinished, and not judged.
/feedThe next page of "asked lately", which the feed asks for when it is scrolled to its end. It reads the archive and asks nobody anything.
/liveA WebSocket each open page holds: how many are online, toasts, live feeds, and "a new version is live".
/rulesThe rules engine with every fact clickable, asking nobody anything.
/robots.txtAsks crawlers to leave /rules?… alone: with facts set it is a page per combination, each linking to more.
/rules.svgThe rules drawn as a standalone picture, for the docs (chapter 6).
/card.pngThe picture a link unfurls with (chapter 11).
/icons/<name>.pngThe explorer's file and folder icons (chapter 15⅝).
/lmjtfy.git and the other repositoriesWith GIT_REDIRECT on, redirected with the same path and query to code.lmjtfy.fun; with it off, served here as before (below).

And at code.lmjtfy.fun, which is the same Worker answering to another name and serving only this:

AddressWhat happens
/Every repository served, with its git clone command and a copy button.
/lmjtfy.gitgit clone it, or open it in a browser and read the code (you are probably here).
/jevcrates.git, /postjevsql.git, /jevsnes.git, /jevhooks.git, /jevstrudel.git, /whiskers.gitThe client lmjtfy shares with the owner's other Jev projects, and those projects, served the same way.
/whiskers/latest/<kind>The file of that repository's latest GitHub release of that kind (below), downloaded through the Worker. GET and HEAD, Range honoured. Cached five minutes.
/whiskers/releases/<tag>/<asset>That exact file of that exact release. Cached for good.
/whiskers/latest/, /whiskers/releases/A page listing the latest release, or the recent ones, each file with its size and a link.
/icons/<name>.png, /emoji.woff2, /card.png?code=, /robots.txtWhat those pages draw themselves with.

Try it. Open https://lmjtfy.fun/rules and click answerable until it says no. Watch every rule but one go grey. That is the engine from chapter 6 running in your browser's address bar.

Live: who is here, and what just happened

Every browser with the site open keeps one socket to /live, held by the archive object. Its sockets are "hibernatable": while nothing happens, the object can sleep and nothing is billed, and the sockets stay open.

One socket per browser, not per tab: the socket belongs to a SharedWorker (src/live.js), a script the browser runs once for all of a site's tabs and keeps while any of them is open. Each tab talks to it, and it passes on everything the archive sends; a tab that opens later is handed the current count and places at once. A tab the browser freezes in the background cannot take the socket down with it, which is what goes wrong when one tab holds the socket for the rest. A browser without SharedWorker opens a socket per tab, as every page used to.

Aside. Pages of different builds use different workers (the build is in the worker's address), so after a deploy a reloaded tab gets a fresh one while a tab from before keeps its own until it reloads too.

It pushes four things:

  • How many browsers are open, and where, shown as "N online" in the top bar; hover it for a flag and a place for each, most first. The place is Cloudflare's: every request arrives with a rough country and city, and the Worker passes them to the archive when a page connects (in headers it sets over any the page sent, so a page cannot name its own). The archive keeps them on that page's socket while it is open, for the list. (What is kept for good is below, in What is kept about visitors.) Pages are told at most once a second, so a deploy, which reconnects every page at once, is one update rather than one per page. A page that reconnects while the deploy is still reaching Cloudflare's edge can come through the Worker from before, which passes no place; ten seconds on, when the deploy has settled, the archive asks it to reconnect (close code 1012), and it comes back placed.
  • A toast when someone asks a question or clones the code. A question the feed may show is shown with Jev's answer; any other is "someone asked Jev something". The page that asked does not get its own toast. The site sends pages nothing else about who is connected: no address, no browser, no history.
  • The feeds themselves. After each toast the archive sends the home page's feeds as HTML elements with ids, and a page replaces the ones it has: most asked and so far whole. A question just asked goes to the top of asked lately, taking its line from wherever it was, so the older lines a page has scrolled in stay put. No refresh.
  • The build. Every page carries the commit its Worker was built from (build.rs). A deploy restarts the archive object, every socket closes, every page reconnects, and the archive tells each one the build that is live now. A page from an older build shows a toast that stays, with a Reload button. If the deployed commit has a Release-Note: trailer, an older page also gets that line as a toast, so the people on the site hear what changed.

The code pages

git clone --recurse-submodules https://code.lmjtfy.fun/lmjtfy.git

The address is code.lmjtfy.fun. With GIT_REDIRECT on, the old https://lmjtfy.fun/lmjtfy.git (and www, and workers.dev, and every other served repository) is sent there with a 308, path and query kept: git follows the redirect of its first request, GET /lmjtfy.git/info/refs?service=git-upload-pack, and makes every request after it at the new address, so git clone https://lmjtfy.fun/lmjtfy.git still works, and so does the relative submodule URL ../jevcrates.git, which git resolves against the address the redirect gave it. A POST that reaches the old address is not git following a redirect, so it is not redirected: it is answered 405 with the address to use. A GET is a 308 because a client that must keep its method should; a person's browser follows either.

The switch, GIT_REDIRECT

The redirect is a Worker var, GIT_REDIRECT, "on" or "off" ([vars] in wrangler.toml, which says "off"), so the code host can be deployed and checked before anyone is sent to it. It is read by host::GitRedirect, which accepts exactly those two words and an unset var as off. Anything else (ON, true, an empty string) is logged as an error on every request and runs as off: off is the behaviour that cannot break a working clone, where on would send every clone to an address that may not be live.

off (default)on
lmjtfy.fun/<repo>.git…Served here, as before the code host308 to code.lmjtfy.fun, path and query kept; a POST is answered 405
www and workers.dev, a repository's path301 to lmjtfy.fun, which serves it (git follows it)308 straight to the code host
www and workers.dev, anything else301 to lmjtfy.funthe same
code.lmjtfy.funServes the repositories and the front pagethe same
The home page's clone command and "Get the code" linkhttps://lmjtfy.fun/lmjtfy.git: no page names the code hosthttps://code.lmjtfy.fun/lmjtfy.git
Code pages on the apex: the clone commandThe apexThe code host (and the apex's paths never get that far: they redirect)
lmjtfy.fun/<name>/latest…, /<name>/releases… (release files)404: not served here, in either mode308 to code.lmjtfy.fun like a clone (a POST is 405)
Staging and dev serversServe it themselves, release files toothe same: no code host there

The links follow the switch so that no page advertises an address that has not been checked. The code host answers in both modes.

Rolling out

The order matters: the code host is deployed and checked before the apex sends anyone to it.

  1. Deploy with the var off (the toml's default). This ships the code host, its front page and the custom domain; lmjtfy.fun/<repo>.git serves clones exactly as before, and the home page still names the apex. Check by hand, since tools/check-hosts expects the token to be missing: curl -sI https://code.lmjtfy.fun/ is 200, and git clone --recurse-submodules https://code.lmjtfy.fun/lmjtfy.git works, submodules included. The custom domain's certificate can take a few minutes.
  2. Flip it: lmjtfy-wrangler apps/lmjtfy deploy --var GIT_REDIRECT:on, or change the toml to "on" and deploy. Check curl -sI https://lmjtfy.fun/lmjtfy.git/info/refs?service=git-upload-pack is a 308 to the code host, and git clone --recurse-submodules https://lmjtfy.fun/lmjtfy.git (through the old address) still works.
  3. To go back: lmjtfy-wrangler apps/lmjtfy rollback (the previous version, which has no redirect), or deploy with --var GIT_REDIRECT:off. Make the toml say what is live afterwards.

The Worker gives each host its own routes (src/host.rs). The code host has the repositories, their pages and the front page, and nothing that asks Jev or reads the archive: no /ask, /feed, /live or /seen, so its pages show no online count and no toasts. Staging has no code host (a second address is a second way round Access); it serves the repositories itself, where it is.

Release files: the downloads of a private repository

src/release.rs serves the files of a repository's GitHub releases, so a release of a private repository can be downloaded by anyone (the whiskers site links to these addresses). <name> is a served repository's name without .git.

AddressWhat
/<name>/latest/<kind>The file of the latest release whose kind is <kind>. Cache-Control: public, max-age=300, because it moves.
/<name>/releases/<tag>/<asset>The exact file of the exact tag. public, max-age=86400, immutable.
/<name>/latest/, /<name>/releases/A page (the code pages' look) of the latest release, or the last twenty, with sizes and links.

Anything else under /<name>/latest or /<name>/releases is a plain 404, and so is a name that is not in Repo::ALL. A path segment may hold only letters, digits and . _ + ~ @ - (never . or ..), so there is no percent-encoding, encoded slash or dot-dot to be decoded later.

A file's kind is its name with the release's version taken out: everything after the first -<tag>- in the name (whiskers-2026.10.5-arm64.apk is arm64.apk; whiskersd-2026.10.5-x86_64-linux.tar.gz is x86_64-linux.tar.gz). A tag with a leading v is also looked for without it. A name with no -<tag>- in it, such as SHA256SUMS, is its own kind. So the address of "the arm64 build" does not change when the version does. If two files of the release have one kind, that kind is a 404 and the log names the release: it is never a guess. The pinned address is always exact.

How a file is served: the release is looked up through the GitHub API with the same read-only token (/releases/latest, /releases/tags/<tag>; kept a minute per isolate, good answers only). The file is asked for at /releases/assets/<id> with Accept: application/octet-stream; GitHub answers 302 to a short-lived signed address, which is followed by hand and fetched without the token, and its body is streamed to the visitor, never held in memory. A file is answered as the runtime's own response (Served::File, before the router): an http body is re-written through Rust as a plain stream, which is sent chunked with no length, and a download bar needs the length. The content length passes through; the type is chosen from the name (.apk is application/vnd.android.package-archive, .tar.gz is application/gzip, SHA256SUMS is text/plain, else application/octet-stream), with content-disposition: attachment and x-content-type-options: nosniff. HEAD answers from the release's own record of the size and fetches nothing. A Range header of plain bytes=… is forwarded to the signed address, which answers 206. GitHub failing is a 502 in fixed words; no GitHub text is in a response or a log.

A download from its start (no Range, or one from byte 0) is an event of what download in the archive, its detail <name>/<file name>. what is free text, so no migration was needed. The GET is not also a view (event::is_download), and nothing shows a download on the home page: no toast, no count. The two listing pages are views.

Only the code host serves these (Host::serves_downloads), and a dev server or staging, which have no code host.

To add a repository: add its name to the repositories! list in src/clone.rs and answer the compiler's match errors there. That one list is what is served, what the code host's front page lists, and what the old addresses redirect; a test holds the three together. Give the GitHub token Contents: read on it too.

The repository is private on GitHub, and is shared from here instead, so nobody needs a GitHub account and nothing is announced. The Worker speaks git's smart HTTP (src/clone.rs): the two requests a clone or fetch makes, info/refs?service=git-upload-pack and git-upload-pack, are forwarded to GitHub with a fine-grained token that can read the served repositories and nothing else. A push is refused here, and GitHub would refuse it anyway. jevcrates is served beside it at /jevcrates.git, because .gitmodules names it by the relative URL ../jevcrates.git, which git resolves against wherever lmjtfy was cloned from. The owner's three other Jev projects (postjevsql, jevsnes, jevhooks) are served too, as is whiskers (which uses Jev as its guard), and name jevcrates the same way, so each clones with its submodule from here.

Clones and pulls are counted, anonymously, per repository; "So far" shows lmjtfy's, and a clone of any project but jevcrates is toasted. jevcrates is fetched along with every project cloned with its submodules, so its toasts would double up. A pull names the commits it already has (have lines) and a clone does not; a fetch counts on the round GitHub answers with the pack, so a long negotiation counts once and a pull with nothing new is not counted.

Git only ever asks for lmjtfy.git/info/refs and lmjtfy.git/git-upload-pack, so every other path under /lmjtfy.git is free for people. The code pages (src/browse.rs) are laid out like an editor, full width. On the left, a sidebar of panels: the explorer (the whole tree, jevcrates included, opened on the way to the page you are on), the outline of the page's headings, the clone command, the latest commit, and every repository served here; jevcrates' pages add the Cargo.toml lines for depending on it, pinned to its latest commit. On the right, the page: a folder's README and CLAUDE.md as two tabs, "for people" and "for agents", between a ◀ tab for the chapter before and a ▶ tab for the chapter after (from the chapter's own last line, or, in a repository with no guide, the next folder with a README in a depth-first walk), with their diagrams drawn (click one to enlarge it); a source file with its comments rendered on the left and the code they are about on the right, coloured (chapter 12½), or rendered if it is markdown, each with a tab for its source top to bottom; ?raw gives a file as it is. /lmjtfy.git itself is the root folder, whose README is the prologue. On a phone the sidebar moves below the page.

Everything is read from GitHub's API with the same token, kept a minute per isolate: folders and files from the contents API, the explorer's tree in one request from the git trees API. What each repository says it is (the description, homepage and topics on the code host's front page, in the sidebar, the page's link preview and the picture) is GitHub's own, from GET /repos/{owner}/{name}, asked for all the repositories together and kept the same minute. Nothing about a repository is written here: change it on GitHub. If GitHub cannot be read, or a repository has no description, its card simply has none. jevcrates is browsed under third-party/jevcrates/ at the commit lmjtfy pins.

Storage

Two Durable Objects keep everything that outlasts a request. Each is one object for the whole site (id_from_name), so every visitor reads the same rows.

The archive (src/archive.rs)

A SQLite database, changed only by adding a step to MIGRATIONS. The tables share no keys, events.browser aside: an asked row's answers are what the versions behind it said, copied, and counts is a tally.

declined is the questions Jev could not take (NotAQuestion and NoQuestion), so a link to one unfurls as "Please choose an LLM instead" with its own card (Card::Declined) and not as "Let me Jev that for you". It is apart from asked on purpose: nothing was answered, so none of them is counted, listed, toasted or votable. asked wins: a question that has since been answered is shown as answered.

A question Jev declined can be voted on and commented on too, under the same thumbs and box. Its vote is keyed by answer = 'declined' (no hash of answers can be that word) and by the question, which is the declined row's key, so ratings.input = declined.input AND ratings.answer = 'declined' joins a vote to its record, and a comment to its vote as always. The thumbs mean other things there: up is "Jev was right to pass", down is "Jev could have answered". about on ratings and comments, and on both views, says which kind a row is (answered or declined); it is worked out from answer, never written. A question answered since is voted on as answered, and its earlier declined votes stay as they were.

askers is how a question counts each browser once. A browser's first page gives it a random id in a cookie (lmjtfy_browser, a year); when it asks, the Worker hashes the id with the question and the archive keeps only that. A browser asking again, or opening its own question from the feed, finds its row and is not counted, toasted or moved up the feed again. The rows of two questions from one browser have nothing in common, and the id cannot be had back from one. So the key says nothing about who asked; the browser column beside it, which the owner's backend reads, is the id itself. An ask with no cookie (a script) counts every time, as before. Counts from before 2026-10-02 stay as they were.

ratings holds the 👍 and 👎 under each answer ("Was Jev right?"): one vote per browser, keyed the same way, and pressed again to take it back. A vote is on the answer as it was kept when it was cast, by the hash of its answers, so if the question is answered differently later, that answer starts with no votes and the old one keeps its own.

comments is what a visitor adds under their vote: the box that opens when they press 👍 or 👎 ("Comments?", and a grey "Let us know what we can improve" that goes as they type). It is a table of its own because ratings is the vote now, and a vote taken back is deleted, which would take the words with it; and because events is the request log, whose columns are what an Event says. Like events it takes INSERT and nothing else (a trigger refuses the rest): a row for every Save, with the vote as it was (vote) and the same key as ratings (input, answer, who). Saving the same words twice for the same vote keeps one row. What the visitor wrote is untrusted text from the public: it reaches the page only through maud's escaping, is in no log line and no event, and the owner's backend must show it as text, never as markup.

The owner reads it through two views, which are not by the day and so have no finished copy: vote_comments is every comment (comment, at_ms, input, browser) beside the vote it was said of (vote_then) and the vote as it stands (vote_now, NULL if taken back), with latest for the last one of a vote; votes_commented is every vote now standing (vote) with its latest comment, NULL if it has none.

erDiagram
  versions {
    text sent_to PK "Jev's endpoint, or a Workers AI model id"
    text request PK "the exact body sent"
    integer version PK "1, 2, ...: each time it was sent, newest last"
    text response "the exact body that came back"
    text request_id "Jev's id for the call, if it gave one"
    integer attempts "tries the client made"
    real took_ms
    real answered_ms "when it was answered"
  }
  asked {
    text input PK "the question, cleaned"
    text answers "JSON: one Answer per question Jev answered"
    real asked_ms "last asked"
    integer times "how often it was asked"
    integer listed "1 if the feed may show it (Jev's fit fact)"
    integer llm "1 if an LLM had to be asked"
    integer moderated "the owner's say: 1 show, 0 hide, NULL leave it to listed"
  }
  declined {
    text input PK "a question Jev could not take"
    real declined_ms "last time"
    integer times "how often"
  }
  events {
    integer id PK
    real at_ms "when"
    text what "view, answer, gate, vote, comment, more, card, fetch, moved, live, left, read, out"
    text method
    text host
    text path
    text query "as it came"
    text input "the question typed, asked or voted on"
    text detail "how an ask ended, which way a vote went, clone or pull"
    real status
    real sent "calls sent for it"
    real kept "calls answered from versions"
    real llm
    real took_ms "an answer's time, or how long a page stayed"
    real first "1 on a browser's first page ever"
    real daily "1 on its first page of the UTC day"
    real session "1 on its first page in half an hour"
    text referrer "the page that linked here, whole"
    text source "utm_source or ref"
    text client "browser, git, bot, other"
    text family "Chrome, Firefox, ..."
    text os
    text device "mobile or desktop"
    text language "Accept-Language"
    text agent "User-Agent"
    text browser "the lmjtfy_browser cookie"
    text ip
    text country
    text region
    text city
    text postcode
    text timezone
    real latitude
    real longitude
    real asn "the network's number"
    text network "whose network: the ISP"
    text colo "Cloudflare's data centre"
    text protocol "HTTP and TLS versions"
    text medium "utm_medium"
    text campaign "utm_campaign"
    text screen "1920x1080, as the page says"
    text viewport "the window"
    real scroll "how far down the page was seen, 0 to 1"
    text region_code "the region within its country: CA for California"
    text continent
    text metro "Cloudflare's metro area code, where it has one"
    text bot "the kind of bot Cloudflare verified it as, if any"
    integer day "days since 1970, worked out from at_ms, and indexed"
  }
  counts {
    text name PK "sent, kept, clone:lmjtfy.git, pull:lmjtfy.git, ..."
    integer n
  }
  askers {
    text input PK "the question"
    text who PK "SHA-256 of a browser's id and the question"
    text browser "the browser's id itself, for the owner's reading"
  }
  ratings {
    text input PK "the question"
    text answer PK "SHA-256 of the answers as kept, or 'declined' for a question Jev declined"
    text who PK "as in askers"
    integer vote "1 right, -1 wrong (on declined: right to pass, could have answered)"
    text browser "as in askers"
    text about "answered or declined, worked out from answer"
  }
  comments {
    integer id PK "in the order saved"
    real at_ms
    text input "the question"
    text answer "as in ratings"
    text who "as in askers"
    text browser "as in askers"
    text about "answered or declined, as in ratings"
    integer vote "the vote it was said of, 1 right, -1 wrong"
    text comment "the visitor's own words, up to 1,000 characters: text, never markup"
  }
  migrated {
    integer version PK "MIGRATIONS steps run"
    real at_ms "when the step ran, from step 18 on"
    text build "the commit of the build that ran it"
  }
  finished {
    text name PK "a view whose days that are over are copied into a table"
    integer through "the last day copied, in days since 1970"
  }

The questions the owner's backend asks are views here, not SQL in the backend. SQLite has no functions of one's own to define, so a question that would be a function of a day is a view with a day column, asked with WHERE day = .... events.day is indexed, so asking of a day reads that day's events and no others.

ViewWhat
hitsEvery event, with what it counts as said once: viewed (a page given to a browser), asked, answered, voted, a visitor (the browser, or the address of one with no cookie), came_from (the site that linked here).
hours, days, today_by_hourThe site's numbers for each hour and each day: views, visitors, sessions, new browsers, asks, answers, votes, clones, time read.
browser_daysWhat each browser did on each day it came: viewed, typed, asked, voted, and whether it was its first.
outcomes_by_day, ask_endings_by_day, pages_by_day, questions_by_day, referrers_by_dayA label and its number n for each day: what happened, how asks ended, pages viewed, questions asked, sites that linked here.
countries_by_day, cities_by_day, networks_by_day, agents_by_dayThe same, where n is how many different visitors: by country, city, network, and kind of browser.
country_visitors, city_visitors, network_visitors, agent_browsersWho those visitors were (who), a row each. The same visitor on two days is one visitor, so a stretch of days is counted from these: COUNT(DISTINCT who).

A day that is over never changes. So each view by the day (all of those above but hits, days and today_by_hour, and the trees' below) is three things: <name>_live works it out from events, <name>_finished is a table the days that are over are copied into, once, and <name> is the two together, which is the one to ask. finished says how far each has been copied. The first event of a new day is what copies the day before (Shelf::finish).

Aside. Other databases call this a materialized view and keep it up to date themselves. SQLite has none. But a view of a finished day has nothing left to keep up with, so copying it once is the whole job.

Two trees are for drilling into: the site's pages, and where visitors came from.

ViewWhat
pagesA row per page with its views by people, browsers, views by bots and not-found answers, for all time. /lmjtfy.git/apps/x counts for /lmjtfy.git, for /lmjtfy.git/apps and for itself. WHERE parent IS NULL is the top of the site; WHERE parent = '/lmjtfy.git' is what is under it.
placesThe same for country, region, city and network: events, views, browsers and addresses for each place, with a region's code and the latitude and longitude of the middle of where its visitors were. WHERE parent IS NULL is the countries.
page_days, place_daysA tree's rows for each day, kept as the views above are, for asking of a stretch of days.
page_visitors, place_visitors, place_addressesWho was on each page and from each place each day (who), for counting once over many days. pages and places are these added up.
vote_comments, votes_commentedThe comments under votes: every comment beside the vote it was said of and the vote now; and every standing vote beside its latest comment; both with about (answered or declined: up and down mean different things). Not by the day.
page_hits, place_hitsEvery event once for each level it belongs to, with its at_ms: for a question the others cannot answer, such as who read one page. They read events, so ask them of a stretch of time.

Aside. SQLite has no function to split a path, and a Durable Object's SQLite takes no functions of our own. So page_hits turns a path into a JSON array (/a/b becomes ["","a","b"]) and reads it back as rows with json_each, which SQLite does have: a row per folder.

versions, events and comments can be added to and nothing else: a trigger on each refuses a change or a deletion. What was kept stays as it was kept.

It also holds, in memory only, the calls on the wire (so identical asks wait on one) and the /live sockets of every open page.

The budgets (src/meter.rs)

Key-value storage, a JSON value per key.

erDiagram
  neurons {
    integer day "days since 1970, UTC"
    real used "Workers AI neurons spent that day"
  }
  jev_dollars {
    integer day
    real used "dollars of Jev spent that day"
  }
  account {
    integer day
    real at_ms "when Cloudflare's figure was read"
    real neurons "the whole account's neurons that day"
  }

Each visitor's count for the minute (budget::Visits) is in memory and written nowhere by the site.

What is kept about visitors

Everything a request says, for the owner alone. Until 2026-10-03 the site kept nothing about who asked; the owner then ruled the other way ("anywhere in the app where we are dropping data we should plug"), so this is the true account now.

Every request worth keeping is a row of events in the archive (packages/archive/src/event.rs): a page viewed, an ask and how it ended, each as-you-type request, a vote, a clone, a redirect from an old address, a page connecting and leaving, and what each page reports of itself as it is left (how long it was in view, how far down it was read, the screen) or when a link off the site is followed. A row has the question, the visitor's address, their browser's id (the lmjtfy_browser cookie, which ties one browser's rows together), the user agent, the referring page, and what Cloudflare says of where the request came from: the network (the ISP), country, region, city, postcode, timezone and coordinates. The Worker writes the row after the response has gone (wait_until), so keeping it delays nobody.

None of it is shown on the site. A question Jev judged unfit for the feed was always kept (asked.listed); now so is every question that ended any other way, in events.

Visits are counted from a second cookie, lmjtfy_visit, which holds the time the browser was last counted: a page view compares it with now and marks itself the browser's first ever, first today, or first in half an hour.

The admin door

The owner reads all of this from a separate, private Worker, which binds this Worker's archive object (script_name in its wrangler.toml) and posts archive::Admin messages to the object's /admin path:

  • Select: one SQL statement that only reads (archive::reads_only), and its rows back, with how many rows SQLite read to find them. Cloudflare meters rows read by the day, and a small answer can cost a whole table.

  • Moderate: the owner's say on whether a question is on the feed, kept in asked.moderated beside Jev's own verdict in listed. The feed shows COALESCE(moderated, listed).

  • Bookmark and Restore: Cloudflare keeps thirty days of the archive's history. Bookmark names a moment in it; Restore has the archive start again as it was then, everything since lost, and answers with the bookmark that undoes it. They are for a migration that went wrong, and they work when nothing else does: an archive whose migrations did not run refuses everything but these two. migrated.at_ms is when each step ran, which is the moment to go back to just before.

This Worker has no route that reaches that path. A visitor's request is passed to the object only as a socket upgrade on /live.

What Cloudflare keeps

Cloudflare, which runs the site, keeps its own record, in the owner's account and nowhere public:

  • Workers Logs, for 3 days: one entry per request the Worker or its objects handle, with the request as it arrived, address included.
  • Traces, for 7 days: each request's timing, through the fetches and the Durable Object calls it made.
  • Metrics, the request counts and errors every Worker has.
  • Web Analytics: on lmjtfy.fun, Cloudflare adds its beacon script (static.cloudflareinsights.com/beacon.min.js) to each page it serves to a browser, and the browser reports page views and load times to lmjtfy.fun/cdn-cgi/rum. It sets no cookie. The Worker's own HTML does not contain the script: Cloudflare adds it on the way out, and only for browsers, so curl does not see it.

The owner turned on the logs and traces (2026-10-02) and Web Analytics with the domain (2026-10-03). They are configured in wrangler.toml under [observability].

Credentials

All from 1Password, none ever written to a file here. TypeSafe's API refuses browser origins anyway, so every Jev call goes through the Worker and the key never reaches a page.

  • LMJTFY_TYPESAFE_API_KEY, lmjtfy's own Jev key, is a Worker secret. Under wrangler dev it comes from the process environment, which op-env-run fills from the host's 1Password Environment. Deployed, it is set with lmjtfy-secret. Without it the Worker still runs, and the page says Jev is offline.
  • The Cloudflare API token. Workers AI has no local emulation, so even wrangler dev calls the real models on the real account. lmjtfy-wrangler (from the owner's devshell) reads the token from 1Password per run and execs wrangler with it.
  • LMJTFY_GITHUB_TOKEN, a Worker secret: the clone proxy's GitHub fine-grained token, Contents: read on every repository Repo::ALL serves (src/clone.rs) and no other permission (Metadata: read, which the repository's description and topics need, is implicit for a fine-grained token). GitHub's API cannot mint a fine-grained token, so it is made on GitHub and kept in the host's 1Password Environment, which op-env-run hands to wrangler dev; deployed, it is set with op-env-run -- lmjtfy-secret github. Without it a clone is answered with 503. Release files (src/release.rs) need the same token to read the repository's releases and their assets, which Contents: read covers.
  • CLOUDFLARE_ANALYTICS_TOKEN and CLOUDFLARE_ACCOUNT_ID, Worker secrets, are how the budget object reads the account's usage. The token can read analytics and nothing else; nixos-config's infra declares it (cloudflare_account_token.lmjtfy-analytics). The dev server runs without them and counts only itself.

In this folder

PathWhat
src/The Worker's code. Chapter 3 walks through it.
wrangler.tomlThe Worker's name, the pinned Jev model, the chosen LLM, the Jev budget, the bindings and the declared secrets; and the same again for staging, a second Worker with its own archive where a change is tried first.
build.rsStamps the build with the commit it came from (LMJTFY_BUILD).
Cargo.tomlThe crate: a cdylib for the Worker, and an rlib so its tests run natively.

build/ and .wrangler/ are the build's and the dev server's and are not committed.

← Previous: Chapter 1, apps/ · Up: apps · Next: Chapter 3, the source →