1--- 2title: SNEK HISS 3theme: jev 4status: made 5song: song.js 6critic: 7 date: 2026-10-03 8 model: jev-1.13.0 9 method: rubric-1 10 art: 0.68 11 artRuns: [0.68, 0.68, 0.68] 12 criteria: 13 hook: 2.8 14 development: 3.7 15 structure: 3.4 16 concept: 4.1 17 originality: 3.7 18 craft: 2.8 19measured: {"date":"2026-10-03","song":"a77b95f7e1eb","total":{"loudnessDbfs":-21.8,"peakDbfs":-1.4},"parts":{"bass":{"loudnessDbfs":-30.6,"peakDbfs":-17.7},"drums":{"loudnessDbfs":-26.4,"peakDbfs":-4.9},"ending":{"loudnessDbfs":-43.1,"peakDbfs":-6.7},"keys":{"loudnessDbfs":-32.3,"peakDbfs":-11.9},"lead":{"loudnessDbfs":-31.4,"peakDbfs":-10.8},"other":{"loudnessDbfs":-45.8,"peakDbfs":-14.2},"snake":{"loudnessDbfs":-33,"peakDbfs":-10},"vocals":{"loudnessDbfs":-27.7,"peakDbfs":-2.2}},"sections":{"1-8":{"total":{"loudnessDbfs":-27.1,"peakDbfs":-9.4},"parts":{"drums":{"loudnessDbfs":-29,"peakDbfs":-12.9},"keys":{"loudnessDbfs":-33.4,"peakDbfs":-20.9},"snake":{"loudnessDbfs":-36.3,"peakDbfs":-18.2}}},"9-16":{"total":{"loudnessDbfs":-21.2,"peakDbfs":-2.5},"parts":{"bass":{"loudnessDbfs":-29.7,"peakDbfs":-20.2},"drums":{"loudnessDbfs":-25.9,"peakDbfs":-8.2},"keys":{"loudnessDbfs":-34.3,"peakDbfs":-21.3},"snake":{"loudnessDbfs":-32.1,"peakDbfs":-18},"vocals":{"loudnessDbfs":-25.1,"peakDbfs":-2.9}}},"17-24":{"total":{"loudnessDbfs":-23.2,"peakDbfs":-5.3},"parts":{"bass":{"loudnessDbfs":-46.5,"peakDbfs":-20.2},"drums":{"loudnessDbfs":-29.2,"peakDbfs":-12.3},"keys":{"loudnessDbfs":-42.1,"peakDbfs":-21.8},"other":{"loudnessDbfs":-36.7,"peakDbfs":-14.2},"snake":{"loudnessDbfs":-27.7,"peakDbfs":-11.3},"vocals":{"loudnessDbfs":-28.1,"peakDbfs":-6}}},"25-32":{"total":{"loudnessDbfs":-18.4,"peakDbfs":-1.4},"parts":{"bass":{"loudnessDbfs":-27.5,"peakDbfs":-17.7},"drums":{"loudnessDbfs":-23.7,"peakDbfs":-5.1},"keys":{"loudnessDbfs":-30.5,"peakDbfs":-17.6},"lead":{"loudnessDbfs":-25.4,"peakDbfs":-10.8},"other":{"loudnessDbfs":-45.6,"peakDbfs":-15.8},"snake":{"loudnessDbfs":-27.3,"peakDbfs":-10},"vocals":{"loudnessDbfs":-25.5,"peakDbfs":-2.2}}},"33-40":{"total":{"loudnessDbfs":-22.1,"peakDbfs":-5.3},"parts":{"bass":{"loudnessDbfs":-29.6,"peakDbfs":-19.9},"drums":{"loudnessDbfs":-26,"peakDbfs":-7.7},"keys":{"loudnessDbfs":-35.1,"peakDbfs":-18.2},"lead":{"loudnessDbfs":-146.7,"peakDbfs":-120.1},"snake":{"loudnessDbfs":-37.9,"peakDbfs":-15.4},"vocals":{"loudnessDbfs":-26.5,"peakDbfs":-7.4}}},"41-48":{"total":{"loudnessDbfs":-23.4,"peakDbfs":-3.9},"parts":{"bass":{"loudnessDbfs":-32.1,"peakDbfs":-20},"drums":{"loudnessDbfs":-29.3,"peakDbfs":-12.3},"keys":{"loudnessDbfs":-31.5,"peakDbfs":-16.4},"lead":{"loudnessDbfs":-532,"peakDbfs":-506.9},"snake":{"loudnessDbfs":-81.1,"peakDbfs":-47.6},"vocals":{"loudnessDbfs":-26.7,"peakDbfs":-4.5}}},"49-56":{"total":{"loudnessDbfs":-22.6,"peakDbfs":-4.5},"parts":{"bass":{"loudnessDbfs":-30.3,"peakDbfs":-20},"drums":{"loudnessDbfs":-26.6,"peakDbfs":-7.7},"keys":{"loudnessDbfs":-31,"peakDbfs":-13.5},"lead":{"loudnessDbfs":null,"peakDbfs":null},"vocals":{"loudnessDbfs":-27.9,"peakDbfs":-4.5}}},"57-64":{"total":{"loudnessDbfs":-18.7,"peakDbfs":-2},"parts":{"bass":{"loudnessDbfs":-27.7,"peakDbfs":-17.8},"drums":{"loudnessDbfs":-23.7,"peakDbfs":-4.9},"keys":{"loudnessDbfs":-28,"peakDbfs":-11.9},"lead":{"loudnessDbfs":-24.3,"peakDbfs":-12.1},"vocals":{"loudnessDbfs":-26.5,"peakDbfs":-3.8}}},"65-72":{"total":{"loudnessDbfs":-24.9,"peakDbfs":-6.9},"parts":{"bass":{"loudnessDbfs":-31.1,"peakDbfs":-20},"drums":{"loudnessDbfs":-27.6,"peakDbfs":-12.2},"ending":{"loudnessDbfs":-33.4,"peakDbfs":-6.7},"keys":{"loudnessDbfs":-36.3,"peakDbfs":-16.4},"lead":{"loudnessDbfs":-55.7,"peakDbfs":-29.8}}}}} 20ideas: 21 - date: 2026-10-02 22 model: claude-opus-5-5 23 harness: claude-code 24 prompt: "Make a new song." 25 brief: "WanSong v1.0 (2026), \"WanSong v1.0 Technical Report\": argues that discrete autoregressive models inevitably create unnatural vocal timbre and \"mush\". Instead, it uses an end-to-end hybrid Multimodal Diffusion Transformer (MMDiT) operating on continuous latent tokens. By using flow matching and directly outputting separated vocal and backing track stems simultaneously, it avoids inter-track phase modulation and vocoder sizzle. Solving neural audio codec artifacts in the 3 kHz to 8 kHz band: Towards Neural Audio Codec for High-Fidelity Music Streaming (He et al., 2026) analyzed reconstruction errors across standard neural codecs (like DAC and EnCodec) and found that Residual Vector Quantization severely struggles between 3,000 Hz and 8,000 Hz, the exact presence and sibilance frequency band where human ears detect harsh, unnatural \"robotic\" hiss. The paper proposes TQCodec, introducing residual linear projection quantization (SimVQ) and perception-driven band-wise bit allocation to clear up the mid-to-high frequency band. It's almost metallic, like a snek hiss." 26revisions: 27 - rev: 1 28 date: 2026-10-02 29 model: claude-opus-5-5 30 harness: claude-code 31 prompt: "WanSong v1.0 (2026), \"WanSong v1.0 Technical Report\": argues that discrete autoregressive models inevitably create unnatural vocal timbre and \"mush\". Instead, it uses an end-to-end hybrid Multimodal Diffusion Transformer (MMDiT) operating on continuous latent tokens. By using flow matching and directly outputting separated vocal and backing track stems simultaneously, it avoids inter-track phase modulation and vocoder sizzle. Solving neural audio codec artifacts in the 3 kHz to 8 kHz band: Towards Neural Audio Codec for High-Fidelity Music Streaming (He et al., 2026) analyzed reconstruction errors across standard neural codecs (like DAC and EnCodec) and found that Residual Vector Quantization severely struggles between 3,000 Hz and 8,000 Hz, the exact presence and sibilance frequency band where human ears detect harsh, unnatural \"robotic\" hiss. The paper proposes TQCodec, introducing residual linear projection quantization (SimVQ) and perception-driven band-wise bit allocation to clear up the mid-to-high frequency band. It's almost metallic, like a snek hiss." 32 output: null # overwritten by rev 2 33 art: 0.65 34 artRuns: [0.65, 0.64, 0.65] 35 - rev: 2 36 date: 2026-10-03 37 model: claude-opus-5-5 38 harness: claude-code 39 prompt: "Keep Jev from removing what a section is made of: the sung lines are never out (back or full), and the intro's first four bars of arpeggio, the hunt's solo snake, the snake's answer in chorus 1 (one swell across both bars) and \"Two stems, one run\" always play. Make the hats what the style says, offbeat eighths with sixteenths added in the choruses. Bring the backing in on bar 50, straight after the line sung alone. Make the choruses lift: the lead level with the vocals, a louder riser into chorus 1 only, the snake pulling back in the hunt's last bar, and the energy asked about the section being entered. A dropout stops held sounds too. From the flow section on no transition is made of noise: a pad swell leads into the clean chorus, and the opening has no transition. Echoes follow the tempo. Level every sung line to the same loudness, within 1 dB of -19 LUFS. Move the sine sub's F up an octave, and let the last chord end inside its bar. Say only what the papers say about themselves: the hiss is what a listener hears, TQCodec is a streaming codec that found its own quantizer noisy at 3000 to 8000 Hz, and WanSong, which names no band and no hiss, is the other road." 40 output: null # overwritten by rev 3 41 art: 0.68 42 artRuns: [0.68, 0.68, 0.68] 43 - rev: 3 44 date: 2026-10-03 45 model: claude-opus-5-5 46 harness: claude-code 47 prompt: "Make the lead a melody, not the chord again: stepwise phrases answering each sung line, and a last phrase of its own in the clean chorus, ending on a held note. Replace \"More bits, down low\", which is heard as \"download\" and is about the low end, with a line about the snake's own band: \"Keep the mid band detail\". Take out the way from chorus 1 back to the hunt: the snake is found once." 48 output: song.js 49 art: 0.68 50 artRuns: [0.68, 0.68, 0.68] 51--- 52 53# SNEK HISS — spec 54 55`song.js` is generated from this spec. Change the spec first, then the song. 56 57## Intent 58 59A listener hears a tell in AI-made music: a thin, almost metallic hiss in 60the presence band, like a snake. That is the listener's ear, not a finding of 61either paper. The song makes the hiss a character. It creeps in while a codec 62picks tokens one at a time, the song hunts it band by band and finds it 63between 3 and 8 kHz, where one 2026 paper found its own quantizer noisy, and 64fixes it as that paper does. Then it takes the other road, a second 2026 65paper's: no codebook at all. The last chorus plays with nothing metallic in 66it, and the last sound of the song is one small hiss. 67 68The two papers, read at the source on 2026-10-02 and again on 2026-10-03 69(arXiv's HTML of each): 70 71- **TQCodec**: "TQCodec: Towards neural audio codec for high-fidelity music 72 streaming", arXiv 2603.01592 (Tencent Music Entertainment and others, 73 submitted 2026-03-02), https://arxiv.org/abs/2603.01592. 74- **WanSong**: "WanSong v1.0 Technical Report", arXiv 2607.14749 (Wan Team, 75 Alibaba Group, July 2026), https://arxiv.org/abs/2607.14749. 76 77**What the brief says that the papers do not.** The brief is a summary of 78the papers, and some of it is the summary's, not theirs. The lyrics sing 79only what the papers say: 80 81| The brief | The paper | 82|---|---| 83| The 3 to 8 kHz band is why AI-made music sounds synthetic | Neither paper says so; the link is the brief's. TQCodec is a codec for music streaming (44.1 kHz, 32 to 128 kbps) that found its own RVQ noisy in that band when reconstructing real music. Its text gives 3000–8000 Hz; its Fig. 2 shows the RVQ failing "above 4000Hz", "a blurred spectrogram between 5000–10000 Hz" | 84| WanSong is about the same artifact | WanSong never mentions a hiss, a codec artifact or a frequency band. It is a song generator that works on continuous tokens, so it has no codebook to pick from: the song sings it as the other road, not as a fix for the band | 85| RVQ "severely struggles between 3,000 Hz and 8,000 Hz" | Yes: "the RVQ struggles to accurately model the mid-frequency range (3000–8000 Hz)", with decoded audio that "still contains perceptible noise" (TQCodec, 3.1.2) | 86| That band is where ears hear a harsh "robotic" hiss; it is the sibilance band | Not in the paper, which never says robotic, hiss or sibilance. "Almost metallic, like a snake hiss" is the listener's ear, and the song sings it as that | 87| "Residual linear projection quantization (SimVQ)" | The paper's words: SimVQ, "whose codebook is frozen and the codes are implicitly generated through a linear projection", extended to "the residual SimVQ" | 88| Band-wise bit allocation "to clear up the mid-to-high frequency band" | The reverse: the allocation is "to prioritize perceptually critical lower frequencies". SimVQ is what the paper credits for the mid frequencies | 89| WanSong argues autoregressive models "inevitably create unnatural vocal timbre and mush" | Not in the report. It says AR models "can lead to challenges in generation efficiency and in maintaining consistent perceptual quality for long-form audio" | 90| Dual stems avoid "inter-track phase modulation and vocoder sizzle" | Not in the report. Its reason: one token for vocals and backing makes the model over-focus on the vocals and suppress the backing; modelling them apart prevents "cross-interference" | 91| Continuous tokens, a hybrid MMDiT, flow matching, vocal and backing stems at once | Yes: "audios are treated as continuous tokens", a "hybrid-MMDit backbone", flow matching, "dual stems (vocals and background music) in a single run", songs "up to 5 minutes" | 92 93**Jev's place in it.** Jev is a typed-decision model. Asked which section 94comes next, or how much of a part a section should have, it returns a 95probability for every option the song wrote down. The song does the rest: it 96draws the section from those probabilities and rounds each level to a whole 97step. One of the levels Jev sets is the snake's. 98 99## Style 100 101- House, 124 bpm, A minor: Am – F – C – G, two bars per chord, so each 102 vocoded line sits on one chord. Four-on-the-floor kick, clap on 2 and 4, 103 offbeat hats (one on each offbeat eighth; the choruses add sixteenths 104 around them), a plucked saw arpeggio, a saw-and-sub bass whose sine sub stays 105 above 45 Hz (its F is an octave over the saw's, which would sit at 43.7 Hz), a square lead in the choruses. 106- The lead is a melody, not the chord again: a two-note pickup under the end 107 of each sung line's bar, then a stepwise phrase in the next. Chorus 2 108 plays it an octave up and ends on a phrase of its own, a held A. 109- **The snake** is the hiss made audible on purpose: white noise held 110 between a 3 kHz high-pass and an 8 kHz low-pass, both resonant so the band's 111 edges ring (the metal), with a quick run of inharmonic sine tones inside the 112 band (the warble a codec makes). It breathes in long swells. 113- **The codec.** Until the fix, every part goes through a bit-crusher and a 114 sample-rate reducer (`crush`, `coarse`): the first four sections are 115 rough, verse 2 is half fixed, and from the flow section on nothing is 116 crushed. The snake's level follows the same arc. 117- **Mix.** Peaks under −1 dBFS, measured, set by one master level. The 118 vocals sit over the band; the snake is never louder than the vocals except 119 where it plays alone. Each chorus measures at least 2 dB over the verse 120 before it: the choruses are played at full velocity as written and every 121 other section at 78% of it. In the choruses the lead sits level with the 122 vocals, within 1.5 dB. Measured 2026-10-03: verse 1 −21.1, chorus 1 −18.9; 123 verse 2 −22.1, chorus 2 −18.9; the peak −1.7 dBFS. 124- **Echoes** follow the tempo (a dotted eighth), never a time in seconds. 125 126## Structure 127 12872 bars (~2:19), nine 8-bar sections. 129 130| Bars | Section | Contents | 131|---|---|---| 132| 1–8 | Intro | The crushed arpeggio alone; the snake creeps in from bar 3; kick from bar 5 | 133| 9–16 | Verse 1 | Lines 1–4 over the crushed band; the snake under everything | 134| 17–24 | Hunt | Two halves. First (17–20), lines 5–6 while a band-pass steps up through the band, one band a bar (250 Hz, 700 Hz, 1.6 kHz, 3.2 kHz). Second (21–24), everything stops but the kick and the snake, alone and loud: "That's a snake"; the snake pulls back in bar 24, so the chorus lands over it | 135| 25–32 | Chorus 1 | "Three to eight kilohertz / Almost metallic / Like a snake's hiss", the lead, sixteenth hats; in the last two bars the snake answers with one long swell, a single note across both bars | 136| 33–40 | Verse 2 | TQCodec's fix, lines 10–13: the crushing halves and the snake fades bar by bar | 137| 41–48 | Flow | The other road (WanSong), lines 14–17: full-band noise that fades away under a pad that fades in (noise to song); nothing is crushed from here; kick from bar 45 | 138| 49–56 | Stems | "Two stems, one run" sung alone (bar 49); the backing alone (bars 50–52); "The vocals, and the backing" with both (53–54), then the band | 139| 57–64 | Chorus 2 | "Three to eight kilohertz / Nothing metallic / No snake": clean, the lead up an octave, carrying the two bars where the snake used to answer | 140| 65–72 | Outro | Arpeggio and bass closing down, resolving to A minor on bar 71, a chord that has died away by the end of its bar; on bar 72 one small hiss, alone | 141 142## Vocals 143 144Windows SAPI (Microsoft David) through the channel vocoder, each line on the 145chord it lands on, levelled by loudness, not by peak: every clip within 1 dB 146of −19 LUFS (`samples/README.md` has each clip's figure). The levelling 147changes each mp3's own gain, in 1.5 dB steps, without re-encoding: a second 148pass through a 64 kb/s encoder cost `small.en` two lines ("and get code from 149the book", "or is H.O.L.S.M.VQ?"). One line per 2 bars: 150 151| # | Section | Line | Chord | Source | 152|---|---|---|---|---| 153| 1 | Verse 1 | "One token at a time" | Am | WanSong: most systems are autoregressive, on discrete tokens | 154| 2 | | "Pick a code from the book" | F | vector quantization: each frame becomes a codebook entry | 155| 3 | | "Every pick, a little off" | C | the quantizer's error, which RVQ's later stages chase | 156| 4 | | "And something starts to hiss" | G | TQCodec: decoded audio "still contains perceptible noise" | 157| 5 | Hunt | "Check it, band by band" | Am | TQCodec: "analysis of the band-wise reconstruction error" | 158| 6 | | "Three to eight kilohertz" | F | TQCodec: "(3000–8000 Hz)" | 159| 7 | | "That's a snake" | C | the listener's ear | 160| 8 | Chorus 1 | "Three to eight kilohertz" | Am | | 161| 9 | | "Almost metallic" / "Like a snake's hiss" | F / C | the listener's ear | 162| 10 | Verse 2 | "Freeze the codebook" | Am | SimVQ: "whose codebook is frozen" | 163| 11 | | "A linear projection" | F | "the codes are implicitly generated through a linear projection" | 164| 12 | | "Keep the mid band detail" | C | "SimVQ for mid-frequency detail preservation" | 165| 13 | | "Residual sim V Q" | G | "the residual SimVQ" | 166| 14 | Flow | "Continuous tokens" | Am | WanSong: "audios are treated as continuous tokens" | 167| 15 | | "Flow matching" | F | "The flow matching framework … is used" | 168| 16 | | "From noise, to a song" | C | flow matching: from "a random noise" to the audio latent | 169| 17 | | "Five minutes long" | G | "songs up to 5 minutes" | 170| 18 | Stems | "Two stems, one run" | Am | "dual stems (vocals and background music) in a single run" | 171| 19 | | "The vocals, and the backing" | C | the same | 172| 20 | Chorus 2 | "Three to eight kilohertz" / "Nothing metallic" / "No snake" | Am / F / C | | 173 174A line is kept only when whisper `small.en` and `medium.en` read it back 175(`samples/README.md` has each reading). Line 12 was "More bits, down low" 176until rev 3: both models heard "download", and the paper's extra bits go to 177the low end, not the snake's band. "That's a snek" was tried for line 7 and 178both models heard "snack", so the snake keeps its vowel. 179 180## Jev in the booth 181 182Two jevs, declared at the top in TypeSafe's own shape (`jev({ state, 183questions })`), both given the song as state, asked every 8 bars: 184 185- **`formJev`** asks one choice: which **section** plays next. 186- **`djJev`** asks, once that is decided (`after: { section }`), a **level** 187 for each part (keys, lead, bass, drums and the snake: out, back or full; 188 the vocals: back or full), how the song moves **into** the section, and 189 its **energy**. Every level question says that bars a part already sits 190 out as written are not Jev's to decide: the level is for the bars where 191 the part plays. The energy is asked about the section being entered, and 192 each of its four levels says what it is for: held back (55% velocity) for 193 a section that waits, steady (70%) for a verse, driving (85%) for a 194 section that builds, all out (as written) for a payoff, the choruses. 195 196**The form.** The story runs one way (hiss, hunt, fix, clean), so Jev may 197linger but not go back across the fix: 198 199| Section | Bars as written | May be followed by | Plays at most | 200|---|---|---|---| 201| `intro` | 1–8 | verse1 | | 202| `verse1` | 9–16 | hunt | | 203| `hunt` | 17–24 | chorus1 | | 204| `chorus1` | 25–32 | verse2 | | 205| `verse2` | 33–40 | flow | | 206| `flow` | 41–48 | stems | | 207| `stems` | 49–56 | chorus2 | | 208| `chorus2` | 57–64 | outro, chorus2 | twice | 209| `outro` | 65–72 | (the end) | | 210 211The cap is 10 sections. The snake is found once: the only choice in the 212form is whether the clean chorus plays again before the outro (until rev 3 213chorus 1 could send the hunt out again, a path Jev took about once in 260 214plays). Nothing after the fix leads back before it, so the snake never returns once it is gone, except 215in the outro's last bar. 216 217**The parts.** Each plays on its own orbit, named for its question, at 218Jev's level: out is silent, back is half velocity, full is as written. The 219snake is a part like the others, asked about only where it plays as written 220(intro to verse 2). 221 222A part that is all a section has is not a level. These always play, outside 223the levels: 224 225- "That's a snake", the hunt's punchline; 226- "Two stems, one run", the line sung alone (so it is at full level, never 227 back); 228- the hunt's solo snake (bars 21–24); 229- the snake's answer in chorus 1 (bars 31–32); 230- the intro's arpeggio in bars 1–4, before anything else has come in. 231 232The sung lines are never out: the vocals' question has two levels, back and 233full. The snake's level is for the bars where it plays under the band 234(3–20, 25–30, 33–40). The outro's last hiss is the ending, outside the 235levels and the energy: it always lands. 236 237**The transitions** play over the last bar before a section: as written, a 238tom fill, a noise riser, or a dropout of the last beat. As written is a 239noise riser into chorus 1, a pad swell into the clean chorus, and straight 240in elsewhere. 241 242- From the flow section on, no transition is made of noise: the song has 243 said the noise is gone. Into the flow, the stems, the clean chorus and the 244 outro Jev picks from as written, a fill or a dropout. 245- Into chorus 1 the choices are the same three: as written already is the 246 riser. 247- The opening has no transition: nothing is playing before it. 248- A dropout stops held sounds too. The snake's swell and the pad are cut 249 short at the last beat with everything else, so the gap is a gap; only 250 reverb and echo tails ring into it. 251 252Without Jev the song plays as written, every part in full.