SPEC.mdpreviewSPEC.mdsource252 lines · 21.0 KB · raw

title: SNEK HISS theme: jev status: made song: song.js critic: date: 2026-10-03 model: jev-1.13.0 method: rubric-1 art: 0.68 artRuns: [0.68, 0.68, 0.68] criteria: hook: 2.8 development: 3.7 structure: 3.4 concept: 4.1 originality: 3.7 craft: 2.8 measured: {"date":"2026-10-03","song":"a77b95f7e1eb","total":{"loudnessDbfs":-21.8,"peakDbfs":-1.4},"parts":{"bass":{"loudnessDbfs":-30.6,"peakDbfs":-17.7},"drums":{"loudnessDbfs":-26.4,"peakDbfs":-4.9},"ending":{"loudnessDbfs":-43.1,"peakDbfs":-6.7},"keys":{"loudnessDbfs":-32.3,"peakDbfs":-11.9},"lead":{"loudnessDbfs":-31.4,"peakDbfs":-10.8},"other":{"loudnessDbfs":-45.8,"peakDbfs":-14.2},"snake":{"loudnessDbfs":-33,"peakDbfs":-10},"vocals":{"loudnessDbfs":-27.7,"peakDbfs":-2.2}},"sections":{"1-8":{"total":{"loudnessDbfs":-27.1,"peakDbfs":-9.4},"parts":{"drums":{"loudnessDbfs":-29,"peakDbfs":-12.9},"keys":{"loudnessDbfs":-33.4,"peakDbfs":-20.9},"snake":{"loudnessDbfs":-36.3,"peakDbfs":-18.2}}},"9-16":{"total":{"loudnessDbfs":-21.2,"peakDbfs":-2.5},"parts":{"bass":{"loudnessDbfs":-29.7,"peakDbfs":-20.2},"drums":{"loudnessDbfs":-25.9,"peakDbfs":-8.2},"keys":{"loudnessDbfs":-34.3,"peakDbfs":-21.3},"snake":{"loudnessDbfs":-32.1,"peakDbfs":-18},"vocals":{"loudnessDbfs":-25.1,"peakDbfs":-2.9}}},"17-24":{"total":{"loudnessDbfs":-23.2,"peakDbfs":-5.3},"parts":{"bass":{"loudnessDbfs":-46.5,"peakDbfs":-20.2},"drums":{"loudnessDbfs":-29.2,"peakDbfs":-12.3},"keys":{"loudnessDbfs":-42.1,"peakDbfs":-21.8},"other":{"loudnessDbfs":-36.7,"peakDbfs":-14.2},"snake":{"loudnessDbfs":-27.7,"peakDbfs":-11.3},"vocals":{"loudnessDbfs":-28.1,"peakDbfs":-6}}},"25-32":{"total":{"loudnessDbfs":-18.4,"peakDbfs":-1.4},"parts":{"bass":{"loudnessDbfs":-27.5,"peakDbfs":-17.7},"drums":{"loudnessDbfs":-23.7,"peakDbfs":-5.1},"keys":{"loudnessDbfs":-30.5,"peakDbfs":-17.6},"lead":{"loudnessDbfs":-25.4,"peakDbfs":-10.8},"other":{"loudnessDbfs":-45.6,"peakDbfs":-15.8},"snake":{"loudnessDbfs":-27.3,"peakDbfs":-10},"vocals":{"loudnessDbfs":-25.5,"peakDbfs":-2.2}}},"33-40":{"total":{"loudnessDbfs":-22.1,"peakDbfs":-5.3},"parts":{"bass":{"loudnessDbfs":-29.6,"peakDbfs":-19.9},"drums":{"loudnessDbfs":-26,"peakDbfs":-7.7},"keys":{"loudnessDbfs":-35.1,"peakDbfs":-18.2},"lead":{"loudnessDbfs":-146.7,"peakDbfs":-120.1},"snake":{"loudnessDbfs":-37.9,"peakDbfs":-15.4},"vocals":{"loudnessDbfs":-26.5,"peakDbfs":-7.4}}},"41-48":{"total":{"loudnessDbfs":-23.4,"peakDbfs":-3.9},"parts":{"bass":{"loudnessDbfs":-32.1,"peakDbfs":-20},"drums":{"loudnessDbfs":-29.3,"peakDbfs":-12.3},"keys":{"loudnessDbfs":-31.5,"peakDbfs":-16.4},"lead":{"loudnessDbfs":-532,"peakDbfs":-506.9},"snake":{"loudnessDbfs":-81.1,"peakDbfs":-47.6},"vocals":{"loudnessDbfs":-26.7,"peakDbfs":-4.5}}},"49-56":{"total":{"loudnessDbfs":-22.6,"peakDbfs":-4.5},"parts":{"bass":{"loudnessDbfs":-30.3,"peakDbfs":-20},"drums":{"loudnessDbfs":-26.6,"peakDbfs":-7.7},"keys":{"loudnessDbfs":-31,"peakDbfs":-13.5},"lead":{"loudnessDbfs":null,"peakDbfs":null},"vocals":{"loudnessDbfs":-27.9,"peakDbfs":-4.5}}},"57-64":{"total":{"loudnessDbfs":-18.7,"peakDbfs":-2},"parts":{"bass":{"loudnessDbfs":-27.7,"peakDbfs":-17.8},"drums":{"loudnessDbfs":-23.7,"peakDbfs":-4.9},"keys":{"loudnessDbfs":-28,"peakDbfs":-11.9},"lead":{"loudnessDbfs":-24.3,"peakDbfs":-12.1},"vocals":{"loudnessDbfs":-26.5,"peakDbfs":-3.8}}},"65-72":{"total":{"loudnessDbfs":-24.9,"peakDbfs":-6.9},"parts":{"bass":{"loudnessDbfs":-31.1,"peakDbfs":-20},"drums":{"loudnessDbfs":-27.6,"peakDbfs":-12.2},"ending":{"loudnessDbfs":-33.4,"peakDbfs":-6.7},"keys":{"loudnessDbfs":-36.3,"peakDbfs":-16.4},"lead":{"loudnessDbfs":-55.7,"peakDbfs":-29.8}}}}} ideas:

  • date: 2026-10-02 model: claude-opus-5-5 harness: claude-code prompt: "Make a new song." brief: "WanSong v1.0 (2026), "WanSong v1.0 Technical Report": argues that discrete autoregressive models inevitably create unnatural vocal timbre and "mush". Instead, it uses an end-to-end hybrid Multimodal Diffusion Transformer (MMDiT) operating on continuous latent tokens. By using flow matching and directly outputting separated vocal and backing track stems simultaneously, it avoids inter-track phase modulation and vocoder sizzle. Solving neural audio codec artifacts in the 3 kHz to 8 kHz band: Towards Neural Audio Codec for High-Fidelity Music Streaming (He et al., 2026) analyzed reconstruction errors across standard neural codecs (like DAC and EnCodec) and found that Residual Vector Quantization severely struggles between 3,000 Hz and 8,000 Hz, the exact presence and sibilance frequency band where human ears detect harsh, unnatural "robotic" hiss. The paper proposes TQCodec, introducing residual linear projection quantization (SimVQ) and perception-driven band-wise bit allocation to clear up the mid-to-high frequency band. It's almost metallic, like a snek hiss." revisions:
  • rev: 1 date: 2026-10-02 model: claude-opus-5-5 harness: claude-code prompt: "WanSong v1.0 (2026), "WanSong v1.0 Technical Report": argues that discrete autoregressive models inevitably create unnatural vocal timbre and "mush". Instead, it uses an end-to-end hybrid Multimodal Diffusion Transformer (MMDiT) operating on continuous latent tokens. By using flow matching and directly outputting separated vocal and backing track stems simultaneously, it avoids inter-track phase modulation and vocoder sizzle. Solving neural audio codec artifacts in the 3 kHz to 8 kHz band: Towards Neural Audio Codec for High-Fidelity Music Streaming (He et al., 2026) analyzed reconstruction errors across standard neural codecs (like DAC and EnCodec) and found that Residual Vector Quantization severely struggles between 3,000 Hz and 8,000 Hz, the exact presence and sibilance frequency band where human ears detect harsh, unnatural "robotic" hiss. The paper proposes TQCodec, introducing residual linear projection quantization (SimVQ) and perception-driven band-wise bit allocation to clear up the mid-to-high frequency band. It's almost metallic, like a snek hiss." output: null # overwritten by rev 2 art: 0.65 artRuns: [0.65, 0.64, 0.65]
  • rev: 2 date: 2026-10-03 model: claude-opus-5-5 harness: claude-code prompt: "Keep Jev from removing what a section is made of: the sung lines are never out (back or full), and the intro's first four bars of arpeggio, the hunt's solo snake, the snake's answer in chorus 1 (one swell across both bars) and "Two stems, one run" always play. Make the hats what the style says, offbeat eighths with sixteenths added in the choruses. Bring the backing in on bar 50, straight after the line sung alone. Make the choruses lift: the lead level with the vocals, a louder riser into chorus 1 only, the snake pulling back in the hunt's last bar, and the energy asked about the section being entered. A dropout stops held sounds too. From the flow section on no transition is made of noise: a pad swell leads into the clean chorus, and the opening has no transition. Echoes follow the tempo. Level every sung line to the same loudness, within 1 dB of -19 LUFS. Move the sine sub's F up an octave, and let the last chord end inside its bar. Say only what the papers say about themselves: the hiss is what a listener hears, TQCodec is a streaming codec that found its own quantizer noisy at 3000 to 8000 Hz, and WanSong, which names no band and no hiss, is the other road." output: null # overwritten by rev 3 art: 0.68 artRuns: [0.68, 0.68, 0.68]
  • rev: 3 date: 2026-10-03 model: claude-opus-5-5 harness: claude-code prompt: "Make the lead a melody, not the chord again: stepwise phrases answering each sung line, and a last phrase of its own in the clean chorus, ending on a held note. Replace "More bits, down low", which is heard as "download" and is about the low end, with a line about the snake's own band: "Keep the mid band detail". Take out the way from chorus 1 back to the hunt: the snake is found once." output: song.js art: 0.68 artRuns: [0.68, 0.68, 0.68]

SNEK HISS — spec

song.js is generated from this spec. Change the spec first, then the song.

Intent

A listener hears a tell in AI-made music: a thin, almost metallic hiss in the presence band, like a snake. That is the listener's ear, not a finding of either paper. The song makes the hiss a character. It creeps in while a codec picks tokens one at a time, the song hunts it band by band and finds it between 3 and 8 kHz, where one 2026 paper found its own quantizer noisy, and fixes it as that paper does. Then it takes the other road, a second 2026 paper's: no codebook at all. The last chorus plays with nothing metallic in it, and the last sound of the song is one small hiss.

The two papers, read at the source on 2026-10-02 and again on 2026-10-03 (arXiv's HTML of each):

  • TQCodec: "TQCodec: Towards neural audio codec for high-fidelity music streaming", arXiv 2603.01592 (Tencent Music Entertainment and others, submitted 2026-03-02), https://arxiv.org/abs/2603.01592.
  • WanSong: "WanSong v1.0 Technical Report", arXiv 2607.14749 (Wan Team, Alibaba Group, July 2026), https://arxiv.org/abs/2607.14749.

What the brief says that the papers do not. The brief is a summary of the papers, and some of it is the summary's, not theirs. The lyrics sing only what the papers say:

The briefThe paper
The 3 to 8 kHz band is why AI-made music sounds syntheticNeither paper says so; the link is the brief's. TQCodec is a codec for music streaming (44.1 kHz, 32 to 128 kbps) that found its own RVQ noisy in that band when reconstructing real music. Its text gives 3000–8000 Hz; its Fig. 2 shows the RVQ failing "above 4000Hz", "a blurred spectrogram between 5000–10000 Hz"
WanSong is about the same artifactWanSong never mentions a hiss, a codec artifact or a frequency band. It is a song generator that works on continuous tokens, so it has no codebook to pick from: the song sings it as the other road, not as a fix for the band
RVQ "severely struggles between 3,000 Hz and 8,000 Hz"Yes: "the RVQ struggles to accurately model the mid-frequency range (3000–8000 Hz)", with decoded audio that "still contains perceptible noise" (TQCodec, 3.1.2)
That band is where ears hear a harsh "robotic" hiss; it is the sibilance bandNot in the paper, which never says robotic, hiss or sibilance. "Almost metallic, like a snake hiss" is the listener's ear, and the song sings it as that
"Residual linear projection quantization (SimVQ)"The paper's words: SimVQ, "whose codebook is frozen and the codes are implicitly generated through a linear projection", extended to "the residual SimVQ"
Band-wise bit allocation "to clear up the mid-to-high frequency band"The reverse: the allocation is "to prioritize perceptually critical lower frequencies". SimVQ is what the paper credits for the mid frequencies
WanSong argues autoregressive models "inevitably create unnatural vocal timbre and mush"Not in the report. It says AR models "can lead to challenges in generation efficiency and in maintaining consistent perceptual quality for long-form audio"
Dual stems avoid "inter-track phase modulation and vocoder sizzle"Not in the report. Its reason: one token for vocals and backing makes the model over-focus on the vocals and suppress the backing; modelling them apart prevents "cross-interference"
Continuous tokens, a hybrid MMDiT, flow matching, vocal and backing stems at onceYes: "audios are treated as continuous tokens", a "hybrid-MMDit backbone", flow matching, "dual stems (vocals and background music) in a single run", songs "up to 5 minutes"

Jev's place in it. Jev is a typed-decision model. Asked which section comes next, or how much of a part a section should have, it returns a probability for every option the song wrote down. The song does the rest: it draws the section from those probabilities and rounds each level to a whole step. One of the levels Jev sets is the snake's.

Style

  • House, 124 bpm, A minor: Am – F – C – G, two bars per chord, so each vocoded line sits on one chord. Four-on-the-floor kick, clap on 2 and 4, offbeat hats (one on each offbeat eighth; the choruses add sixteenths around them), a plucked saw arpeggio, a saw-and-sub bass whose sine sub stays above 45 Hz (its F is an octave over the saw's, which would sit at 43.7 Hz), a square lead in the choruses.
  • The lead is a melody, not the chord again: a two-note pickup under the end of each sung line's bar, then a stepwise phrase in the next. Chorus 2 plays it an octave up and ends on a phrase of its own, a held A.
  • The snake is the hiss made audible on purpose: white noise held between a 3 kHz high-pass and an 8 kHz low-pass, both resonant so the band's edges ring (the metal), with a quick run of inharmonic sine tones inside the band (the warble a codec makes). It breathes in long swells.
  • The codec. Until the fix, every part goes through a bit-crusher and a sample-rate reducer (crush, coarse): the first four sections are rough, verse 2 is half fixed, and from the flow section on nothing is crushed. The snake's level follows the same arc.
  • Mix. Peaks under −1 dBFS, measured, set by one master level. The vocals sit over the band; the snake is never louder than the vocals except where it plays alone. Each chorus measures at least 2 dB over the verse before it: the choruses are played at full velocity as written and every other section at 78% of it. In the choruses the lead sits level with the vocals, within 1.5 dB. Measured 2026-10-03: verse 1 −21.1, chorus 1 −18.9; verse 2 −22.1, chorus 2 −18.9; the peak −1.7 dBFS.
  • Echoes follow the tempo (a dotted eighth), never a time in seconds.

Structure

72 bars (~2:19), nine 8-bar sections.

BarsSectionContents
1–8IntroThe crushed arpeggio alone; the snake creeps in from bar 3; kick from bar 5
9–16Verse 1Lines 1–4 over the crushed band; the snake under everything
17–24HuntTwo halves. First (17–20), lines 5–6 while a band-pass steps up through the band, one band a bar (250 Hz, 700 Hz, 1.6 kHz, 3.2 kHz). Second (21–24), everything stops but the kick and the snake, alone and loud: "That's a snake"; the snake pulls back in bar 24, so the chorus lands over it
25–32Chorus 1"Three to eight kilohertz / Almost metallic / Like a snake's hiss", the lead, sixteenth hats; in the last two bars the snake answers with one long swell, a single note across both bars
33–40Verse 2TQCodec's fix, lines 10–13: the crushing halves and the snake fades bar by bar
41–48FlowThe other road (WanSong), lines 14–17: full-band noise that fades away under a pad that fades in (noise to song); nothing is crushed from here; kick from bar 45
49–56Stems"Two stems, one run" sung alone (bar 49); the backing alone (bars 50–52); "The vocals, and the backing" with both (53–54), then the band
57–64Chorus 2"Three to eight kilohertz / Nothing metallic / No snake": clean, the lead up an octave, carrying the two bars where the snake used to answer
65–72OutroArpeggio and bass closing down, resolving to A minor on bar 71, a chord that has died away by the end of its bar; on bar 72 one small hiss, alone

Vocals

Windows SAPI (Microsoft David) through the channel vocoder, each line on the chord it lands on, levelled by loudness, not by peak: every clip within 1 dB of −19 LUFS (samples/README.md has each clip's figure). The levelling changes each mp3's own gain, in 1.5 dB steps, without re-encoding: a second pass through a 64 kb/s encoder cost small.en two lines ("and get code from the book", "or is H.O.L.S.M.VQ?"). One line per 2 bars:

#SectionLineChordSource
1Verse 1"One token at a time"AmWanSong: most systems are autoregressive, on discrete tokens
2"Pick a code from the book"Fvector quantization: each frame becomes a codebook entry
3"Every pick, a little off"Cthe quantizer's error, which RVQ's later stages chase
4"And something starts to hiss"GTQCodec: decoded audio "still contains perceptible noise"
5Hunt"Check it, band by band"AmTQCodec: "analysis of the band-wise reconstruction error"
6"Three to eight kilohertz"FTQCodec: "(3000–8000 Hz)"
7"That's a snake"Cthe listener's ear
8Chorus 1"Three to eight kilohertz"Am
9"Almost metallic" / "Like a snake's hiss"F / Cthe listener's ear
10Verse 2"Freeze the codebook"AmSimVQ: "whose codebook is frozen"
11"A linear projection"F"the codes are implicitly generated through a linear projection"
12"Keep the mid band detail"C"SimVQ for mid-frequency detail preservation"
13"Residual sim V Q"G"the residual SimVQ"
14Flow"Continuous tokens"AmWanSong: "audios are treated as continuous tokens"
15"Flow matching"F"The flow matching framework … is used"
16"From noise, to a song"Cflow matching: from "a random noise" to the audio latent
17"Five minutes long"G"songs up to 5 minutes"
18Stems"Two stems, one run"Am"dual stems (vocals and background music) in a single run"
19"The vocals, and the backing"Cthe same
20Chorus 2"Three to eight kilohertz" / "Nothing metallic" / "No snake"Am / F / C

A line is kept only when whisper small.en and medium.en read it back (samples/README.md has each reading). Line 12 was "More bits, down low" until rev 3: both models heard "download", and the paper's extra bits go to the low end, not the snake's band. "That's a snek" was tried for line 7 and both models heard "snack", so the snake keeps its vowel.

Jev in the booth

Two jevs, declared at the top in TypeSafe's own shape (jev({ state, questions })), both given the song as state, asked every 8 bars:

  • formJev asks one choice: which section plays next.
  • djJev asks, once that is decided (after: { section }), a level for each part (keys, lead, bass, drums and the snake: out, back or full; the vocals: back or full), how the song moves into the section, and its energy. Every level question says that bars a part already sits out as written are not Jev's to decide: the level is for the bars where the part plays. The energy is asked about the section being entered, and each of its four levels says what it is for: held back (55% velocity) for a section that waits, steady (70%) for a verse, driving (85%) for a section that builds, all out (as written) for a payoff, the choruses.

The form. The story runs one way (hiss, hunt, fix, clean), so Jev may linger but not go back across the fix:

SectionBars as writtenMay be followed byPlays at most
intro1–8verse1
verse19–16hunt
hunt17–24chorus1
chorus125–32verse2
verse233–40flow
flow41–48stems
stems49–56chorus2
chorus257–64outro, chorus2twice
outro65–72(the end)

The cap is 10 sections. The snake is found once: the only choice in the form is whether the clean chorus plays again before the outro (until rev 3 chorus 1 could send the hunt out again, a path Jev took about once in 260 plays). Nothing after the fix leads back before it, so the snake never returns once it is gone, except in the outro's last bar.

The parts. Each plays on its own orbit, named for its question, at Jev's level: out is silent, back is half velocity, full is as written. The snake is a part like the others, asked about only where it plays as written (intro to verse 2).

A part that is all a section has is not a level. These always play, outside the levels:

  • "That's a snake", the hunt's punchline;
  • "Two stems, one run", the line sung alone (so it is at full level, never back);
  • the hunt's solo snake (bars 21–24);
  • the snake's answer in chorus 1 (bars 31–32);
  • the intro's arpeggio in bars 1–4, before anything else has come in.

The sung lines are never out: the vocals' question has two levels, back and full. The snake's level is for the bars where it plays under the band (3–20, 25–30, 33–40). The outro's last hiss is the ending, outside the levels and the energy: it always lands.

The transitions play over the last bar before a section: as written, a tom fill, a noise riser, or a dropout of the last beat. As written is a noise riser into chorus 1, a pad swell into the clean chorus, and straight in elsewhere.

  • From the flow section on, no transition is made of noise: the song has said the noise is gone. Into the flow, the stems, the clean chorus and the outro Jev picks from as written, a fill or a dropout.
  • Into chorus 1 the choices are the same three: as written already is the riser.
  • The opening has no transition: nothing is playing before it.
  • A dropout stops held sounds too. The snake's swell and the pad are cut short at the last beat with everything else, so the gap is a gap; only reverb and echo tails ring into it.

Without Jev the song plays as written, every part in full.