SPEC.mdpreviewSPEC.mdsource252 lines · 21.0 KB · raw
1---
2title: SNEK HISS
3theme: jev
4status: made
5song: song.js
6critic:
7  date: 2026-10-03
8  model: jev-1.13.0
9  method: rubric-1
10  art: 0.68
11  artRuns: [0.68, 0.68, 0.68]
12  criteria:
13    hook: 2.8
14    development: 3.7
15    structure: 3.4
16    concept: 4.1
17    originality: 3.7
18    craft: 2.8
19measured: {"date":"2026-10-03","song":"a77b95f7e1eb","total":{"loudnessDbfs":-21.8,"peakDbfs":-1.4},"parts":{"bass":{"loudnessDbfs":-30.6,"peakDbfs":-17.7},"drums":{"loudnessDbfs":-26.4,"peakDbfs":-4.9},"ending":{"loudnessDbfs":-43.1,"peakDbfs":-6.7},"keys":{"loudnessDbfs":-32.3,"peakDbfs":-11.9},"lead":{"loudnessDbfs":-31.4,"peakDbfs":-10.8},"other":{"loudnessDbfs":-45.8,"peakDbfs":-14.2},"snake":{"loudnessDbfs":-33,"peakDbfs":-10},"vocals":{"loudnessDbfs":-27.7,"peakDbfs":-2.2}},"sections":{"1-8":{"total":{"loudnessDbfs":-27.1,"peakDbfs":-9.4},"parts":{"drums":{"loudnessDbfs":-29,"peakDbfs":-12.9},"keys":{"loudnessDbfs":-33.4,"peakDbfs":-20.9},"snake":{"loudnessDbfs":-36.3,"peakDbfs":-18.2}}},"9-16":{"total":{"loudnessDbfs":-21.2,"peakDbfs":-2.5},"parts":{"bass":{"loudnessDbfs":-29.7,"peakDbfs":-20.2},"drums":{"loudnessDbfs":-25.9,"peakDbfs":-8.2},"keys":{"loudnessDbfs":-34.3,"peakDbfs":-21.3},"snake":{"loudnessDbfs":-32.1,"peakDbfs":-18},"vocals":{"loudnessDbfs":-25.1,"peakDbfs":-2.9}}},"17-24":{"total":{"loudnessDbfs":-23.2,"peakDbfs":-5.3},"parts":{"bass":{"loudnessDbfs":-46.5,"peakDbfs":-20.2},"drums":{"loudnessDbfs":-29.2,"peakDbfs":-12.3},"keys":{"loudnessDbfs":-42.1,"peakDbfs":-21.8},"other":{"loudnessDbfs":-36.7,"peakDbfs":-14.2},"snake":{"loudnessDbfs":-27.7,"peakDbfs":-11.3},"vocals":{"loudnessDbfs":-28.1,"peakDbfs":-6}}},"25-32":{"total":{"loudnessDbfs":-18.4,"peakDbfs":-1.4},"parts":{"bass":{"loudnessDbfs":-27.5,"peakDbfs":-17.7},"drums":{"loudnessDbfs":-23.7,"peakDbfs":-5.1},"keys":{"loudnessDbfs":-30.5,"peakDbfs":-17.6},"lead":{"loudnessDbfs":-25.4,"peakDbfs":-10.8},"other":{"loudnessDbfs":-45.6,"peakDbfs":-15.8},"snake":{"loudnessDbfs":-27.3,"peakDbfs":-10},"vocals":{"loudnessDbfs":-25.5,"peakDbfs":-2.2}}},"33-40":{"total":{"loudnessDbfs":-22.1,"peakDbfs":-5.3},"parts":{"bass":{"loudnessDbfs":-29.6,"peakDbfs":-19.9},"drums":{"loudnessDbfs":-26,"peakDbfs":-7.7},"keys":{"loudnessDbfs":-35.1,"peakDbfs":-18.2},"lead":{"loudnessDbfs":-146.7,"peakDbfs":-120.1},"snake":{"loudnessDbfs":-37.9,"peakDbfs":-15.4},"vocals":{"loudnessDbfs":-26.5,"peakDbfs":-7.4}}},"41-48":{"total":{"loudnessDbfs":-23.4,"peakDbfs":-3.9},"parts":{"bass":{"loudnessDbfs":-32.1,"peakDbfs":-20},"drums":{"loudnessDbfs":-29.3,"peakDbfs":-12.3},"keys":{"loudnessDbfs":-31.5,"peakDbfs":-16.4},"lead":{"loudnessDbfs":-532,"peakDbfs":-506.9},"snake":{"loudnessDbfs":-81.1,"peakDbfs":-47.6},"vocals":{"loudnessDbfs":-26.7,"peakDbfs":-4.5}}},"49-56":{"total":{"loudnessDbfs":-22.6,"peakDbfs":-4.5},"parts":{"bass":{"loudnessDbfs":-30.3,"peakDbfs":-20},"drums":{"loudnessDbfs":-26.6,"peakDbfs":-7.7},"keys":{"loudnessDbfs":-31,"peakDbfs":-13.5},"lead":{"loudnessDbfs":null,"peakDbfs":null},"vocals":{"loudnessDbfs":-27.9,"peakDbfs":-4.5}}},"57-64":{"total":{"loudnessDbfs":-18.7,"peakDbfs":-2},"parts":{"bass":{"loudnessDbfs":-27.7,"peakDbfs":-17.8},"drums":{"loudnessDbfs":-23.7,"peakDbfs":-4.9},"keys":{"loudnessDbfs":-28,"peakDbfs":-11.9},"lead":{"loudnessDbfs":-24.3,"peakDbfs":-12.1},"vocals":{"loudnessDbfs":-26.5,"peakDbfs":-3.8}}},"65-72":{"total":{"loudnessDbfs":-24.9,"peakDbfs":-6.9},"parts":{"bass":{"loudnessDbfs":-31.1,"peakDbfs":-20},"drums":{"loudnessDbfs":-27.6,"peakDbfs":-12.2},"ending":{"loudnessDbfs":-33.4,"peakDbfs":-6.7},"keys":{"loudnessDbfs":-36.3,"peakDbfs":-16.4},"lead":{"loudnessDbfs":-55.7,"peakDbfs":-29.8}}}}}
20ideas:
21  - date: 2026-10-02
22    model: claude-opus-5-5
23    harness: claude-code
24    prompt: "Make a new song."
25    brief: "WanSong v1.0 (2026), \"WanSong v1.0 Technical Report\": argues that discrete autoregressive models inevitably create unnatural vocal timbre and \"mush\". Instead, it uses an end-to-end hybrid Multimodal Diffusion Transformer (MMDiT) operating on continuous latent tokens. By using flow matching and directly outputting separated vocal and backing track stems simultaneously, it avoids inter-track phase modulation and vocoder sizzle. Solving neural audio codec artifacts in the 3 kHz to 8 kHz band: Towards Neural Audio Codec for High-Fidelity Music Streaming (He et al., 2026) analyzed reconstruction errors across standard neural codecs (like DAC and EnCodec) and found that Residual Vector Quantization severely struggles between 3,000 Hz and 8,000 Hz, the exact presence and sibilance frequency band where human ears detect harsh, unnatural \"robotic\" hiss. The paper proposes TQCodec, introducing residual linear projection quantization (SimVQ) and perception-driven band-wise bit allocation to clear up the mid-to-high frequency band. It's almost metallic, like a snek hiss."
26revisions:
27  - rev: 1
28    date: 2026-10-02
29    model: claude-opus-5-5
30    harness: claude-code
31    prompt: "WanSong v1.0 (2026), \"WanSong v1.0 Technical Report\": argues that discrete autoregressive models inevitably create unnatural vocal timbre and \"mush\". Instead, it uses an end-to-end hybrid Multimodal Diffusion Transformer (MMDiT) operating on continuous latent tokens. By using flow matching and directly outputting separated vocal and backing track stems simultaneously, it avoids inter-track phase modulation and vocoder sizzle. Solving neural audio codec artifacts in the 3 kHz to 8 kHz band: Towards Neural Audio Codec for High-Fidelity Music Streaming (He et al., 2026) analyzed reconstruction errors across standard neural codecs (like DAC and EnCodec) and found that Residual Vector Quantization severely struggles between 3,000 Hz and 8,000 Hz, the exact presence and sibilance frequency band where human ears detect harsh, unnatural \"robotic\" hiss. The paper proposes TQCodec, introducing residual linear projection quantization (SimVQ) and perception-driven band-wise bit allocation to clear up the mid-to-high frequency band. It's almost metallic, like a snek hiss."
32    output: null  # overwritten by rev 2
33    art: 0.65
34    artRuns: [0.65, 0.64, 0.65]
35  - rev: 2
36    date: 2026-10-03
37    model: claude-opus-5-5
38    harness: claude-code
39    prompt: "Keep Jev from removing what a section is made of: the sung lines are never out (back or full), and the intro's first four bars of arpeggio, the hunt's solo snake, the snake's answer in chorus 1 (one swell across both bars) and \"Two stems, one run\" always play. Make the hats what the style says, offbeat eighths with sixteenths added in the choruses. Bring the backing in on bar 50, straight after the line sung alone. Make the choruses lift: the lead level with the vocals, a louder riser into chorus 1 only, the snake pulling back in the hunt's last bar, and the energy asked about the section being entered. A dropout stops held sounds too. From the flow section on no transition is made of noise: a pad swell leads into the clean chorus, and the opening has no transition. Echoes follow the tempo. Level every sung line to the same loudness, within 1 dB of -19 LUFS. Move the sine sub's F up an octave, and let the last chord end inside its bar. Say only what the papers say about themselves: the hiss is what a listener hears, TQCodec is a streaming codec that found its own quantizer noisy at 3000 to 8000 Hz, and WanSong, which names no band and no hiss, is the other road."
40    output: null  # overwritten by rev 3
41    art: 0.68
42    artRuns: [0.68, 0.68, 0.68]
43  - rev: 3
44    date: 2026-10-03
45    model: claude-opus-5-5
46    harness: claude-code
47    prompt: "Make the lead a melody, not the chord again: stepwise phrases answering each sung line, and a last phrase of its own in the clean chorus, ending on a held note. Replace \"More bits, down low\", which is heard as \"download\" and is about the low end, with a line about the snake's own band: \"Keep the mid band detail\". Take out the way from chorus 1 back to the hunt: the snake is found once."
48    output: song.js
49    art: 0.68
50    artRuns: [0.68, 0.68, 0.68]
51---
52
53# SNEK HISS — spec
54
55`song.js` is generated from this spec. Change the spec first, then the song.
56
57## Intent
58
59A listener hears a tell in AI-made music: a thin, almost metallic hiss in
60the presence band, like a snake. That is the listener's ear, not a finding of
61either paper. The song makes the hiss a character. It creeps in while a codec
62picks tokens one at a time, the song hunts it band by band and finds it
63between 3 and 8 kHz, where one 2026 paper found its own quantizer noisy, and
64fixes it as that paper does. Then it takes the other road, a second 2026
65paper's: no codebook at all. The last chorus plays with nothing metallic in
66it, and the last sound of the song is one small hiss.
67
68The two papers, read at the source on 2026-10-02 and again on 2026-10-03
69(arXiv's HTML of each):
70
71- **TQCodec**: "TQCodec: Towards neural audio codec for high-fidelity music
72  streaming", arXiv 2603.01592 (Tencent Music Entertainment and others,
73  submitted 2026-03-02), https://arxiv.org/abs/2603.01592.
74- **WanSong**: "WanSong v1.0 Technical Report", arXiv 2607.14749 (Wan Team,
75  Alibaba Group, July 2026), https://arxiv.org/abs/2607.14749.
76
77**What the brief says that the papers do not.** The brief is a summary of
78the papers, and some of it is the summary's, not theirs. The lyrics sing
79only what the papers say:
80
81| The brief | The paper |
82|---|---|
83| The 3 to 8 kHz band is why AI-made music sounds synthetic | Neither paper says so; the link is the brief's. TQCodec is a codec for music streaming (44.1 kHz, 32 to 128 kbps) that found its own RVQ noisy in that band when reconstructing real music. Its text gives 3000–8000 Hz; its Fig. 2 shows the RVQ failing "above 4000Hz", "a blurred spectrogram between 5000–10000 Hz" |
84| WanSong is about the same artifact | WanSong never mentions a hiss, a codec artifact or a frequency band. It is a song generator that works on continuous tokens, so it has no codebook to pick from: the song sings it as the other road, not as a fix for the band |
85| RVQ "severely struggles between 3,000 Hz and 8,000 Hz" | Yes: "the RVQ struggles to accurately model the mid-frequency range (3000–8000 Hz)", with decoded audio that "still contains perceptible noise" (TQCodec, 3.1.2) |
86| That band is where ears hear a harsh "robotic" hiss; it is the sibilance band | Not in the paper, which never says robotic, hiss or sibilance. "Almost metallic, like a snake hiss" is the listener's ear, and the song sings it as that |
87| "Residual linear projection quantization (SimVQ)" | The paper's words: SimVQ, "whose codebook is frozen and the codes are implicitly generated through a linear projection", extended to "the residual SimVQ" |
88| Band-wise bit allocation "to clear up the mid-to-high frequency band" | The reverse: the allocation is "to prioritize perceptually critical lower frequencies". SimVQ is what the paper credits for the mid frequencies |
89| WanSong argues autoregressive models "inevitably create unnatural vocal timbre and mush" | Not in the report. It says AR models "can lead to challenges in generation efficiency and in maintaining consistent perceptual quality for long-form audio" |
90| Dual stems avoid "inter-track phase modulation and vocoder sizzle" | Not in the report. Its reason: one token for vocals and backing makes the model over-focus on the vocals and suppress the backing; modelling them apart prevents "cross-interference" |
91| Continuous tokens, a hybrid MMDiT, flow matching, vocal and backing stems at once | Yes: "audios are treated as continuous tokens", a "hybrid-MMDit backbone", flow matching, "dual stems (vocals and background music) in a single run", songs "up to 5 minutes" |
92
93**Jev's place in it.** Jev is a typed-decision model. Asked which section
94comes next, or how much of a part a section should have, it returns a
95probability for every option the song wrote down. The song does the rest: it
96draws the section from those probabilities and rounds each level to a whole
97step. One of the levels Jev sets is the snake's.
98
99## Style
100
101- House, 124 bpm, A minor: Am – F – C – G, two bars per chord, so each
102  vocoded line sits on one chord. Four-on-the-floor kick, clap on 2 and 4,
103  offbeat hats (one on each offbeat eighth; the choruses add sixteenths
104  around them), a plucked saw arpeggio, a saw-and-sub bass whose sine sub stays
105  above 45 Hz (its F is an octave over the saw's, which would sit at 43.7 Hz), a square lead in the choruses.
106- The lead is a melody, not the chord again: a two-note pickup under the end
107  of each sung line's bar, then a stepwise phrase in the next. Chorus 2
108  plays it an octave up and ends on a phrase of its own, a held A.
109- **The snake** is the hiss made audible on purpose: white noise held
110  between a 3 kHz high-pass and an 8 kHz low-pass, both resonant so the band's
111  edges ring (the metal), with a quick run of inharmonic sine tones inside the
112  band (the warble a codec makes). It breathes in long swells.
113- **The codec.** Until the fix, every part goes through a bit-crusher and a
114  sample-rate reducer (`crush`, `coarse`): the first four sections are
115  rough, verse 2 is half fixed, and from the flow section on nothing is
116  crushed. The snake's level follows the same arc.
117- **Mix.** Peaks under −1 dBFS, measured, set by one master level. The
118  vocals sit over the band; the snake is never louder than the vocals except
119  where it plays alone. Each chorus measures at least 2 dB over the verse
120  before it: the choruses are played at full velocity as written and every
121  other section at 78% of it. In the choruses the lead sits level with the
122  vocals, within 1.5 dB. Measured 2026-10-03: verse 1 −21.1, chorus 1 −18.9;
123  verse 2 −22.1, chorus 2 −18.9; the peak −1.7 dBFS.
124- **Echoes** follow the tempo (a dotted eighth), never a time in seconds.
125
126## Structure
127
12872 bars (~2:19), nine 8-bar sections.
129
130| Bars | Section | Contents |
131|---|---|---|
132| 1–8 | Intro | The crushed arpeggio alone; the snake creeps in from bar 3; kick from bar 5 |
133| 9–16 | Verse 1 | Lines 1–4 over the crushed band; the snake under everything |
134| 17–24 | Hunt | Two halves. First (17–20), lines 5–6 while a band-pass steps up through the band, one band a bar (250 Hz, 700 Hz, 1.6 kHz, 3.2 kHz). Second (21–24), everything stops but the kick and the snake, alone and loud: "That's a snake"; the snake pulls back in bar 24, so the chorus lands over it |
135| 25–32 | Chorus 1 | "Three to eight kilohertz / Almost metallic / Like a snake's hiss", the lead, sixteenth hats; in the last two bars the snake answers with one long swell, a single note across both bars |
136| 33–40 | Verse 2 | TQCodec's fix, lines 10–13: the crushing halves and the snake fades bar by bar |
137| 41–48 | Flow | The other road (WanSong), lines 14–17: full-band noise that fades away under a pad that fades in (noise to song); nothing is crushed from here; kick from bar 45 |
138| 49–56 | Stems | "Two stems, one run" sung alone (bar 49); the backing alone (bars 50–52); "The vocals, and the backing" with both (53–54), then the band |
139| 57–64 | Chorus 2 | "Three to eight kilohertz / Nothing metallic / No snake": clean, the lead up an octave, carrying the two bars where the snake used to answer |
140| 65–72 | Outro | Arpeggio and bass closing down, resolving to A minor on bar 71, a chord that has died away by the end of its bar; on bar 72 one small hiss, alone |
141
142## Vocals
143
144Windows SAPI (Microsoft David) through the channel vocoder, each line on the
145chord it lands on, levelled by loudness, not by peak: every clip within 1 dB
146of −19 LUFS (`samples/README.md` has each clip's figure). The levelling
147changes each mp3's own gain, in 1.5 dB steps, without re-encoding: a second
148pass through a 64 kb/s encoder cost `small.en` two lines ("and get code from
149the book", "or is H.O.L.S.M.VQ?"). One line per 2 bars:
150
151| # | Section | Line | Chord | Source |
152|---|---|---|---|---|
153| 1 | Verse 1 | "One token at a time" | Am | WanSong: most systems are autoregressive, on discrete tokens |
154| 2 | | "Pick a code from the book" | F | vector quantization: each frame becomes a codebook entry |
155| 3 | | "Every pick, a little off" | C | the quantizer's error, which RVQ's later stages chase |
156| 4 | | "And something starts to hiss" | G | TQCodec: decoded audio "still contains perceptible noise" |
157| 5 | Hunt | "Check it, band by band" | Am | TQCodec: "analysis of the band-wise reconstruction error" |
158| 6 | | "Three to eight kilohertz" | F | TQCodec: "(3000–8000 Hz)" |
159| 7 | | "That's a snake" | C | the listener's ear |
160| 8 | Chorus 1 | "Three to eight kilohertz" | Am | |
161| 9 | | "Almost metallic" / "Like a snake's hiss" | F / C | the listener's ear |
162| 10 | Verse 2 | "Freeze the codebook" | Am | SimVQ: "whose codebook is frozen" |
163| 11 | | "A linear projection" | F | "the codes are implicitly generated through a linear projection" |
164| 12 | | "Keep the mid band detail" | C | "SimVQ for mid-frequency detail preservation" |
165| 13 | | "Residual sim V Q" | G | "the residual SimVQ" |
166| 14 | Flow | "Continuous tokens" | Am | WanSong: "audios are treated as continuous tokens" |
167| 15 | | "Flow matching" | F | "The flow matching framework … is used" |
168| 16 | | "From noise, to a song" | C | flow matching: from "a random noise" to the audio latent |
169| 17 | | "Five minutes long" | G | "songs up to 5 minutes" |
170| 18 | Stems | "Two stems, one run" | Am | "dual stems (vocals and background music) in a single run" |
171| 19 | | "The vocals, and the backing" | C | the same |
172| 20 | Chorus 2 | "Three to eight kilohertz" / "Nothing metallic" / "No snake" | Am / F / C | |
173
174A line is kept only when whisper `small.en` and `medium.en` read it back
175(`samples/README.md` has each reading). Line 12 was "More bits, down low"
176until rev 3: both models heard "download", and the paper's extra bits go to
177the low end, not the snake's band. "That's a snek" was tried for line 7 and
178both models heard "snack", so the snake keeps its vowel.
179
180## Jev in the booth
181
182Two jevs, declared at the top in TypeSafe's own shape (`jev({ state,
183questions })`), both given the song as state, asked every 8 bars:
184
185- **`formJev`** asks one choice: which **section** plays next.
186- **`djJev`** asks, once that is decided (`after: { section }`), a **level**
187  for each part (keys, lead, bass, drums and the snake: out, back or full;
188  the vocals: back or full), how the song moves **into** the section, and
189  its **energy**. Every level question says that bars a part already sits
190  out as written are not Jev's to decide: the level is for the bars where
191  the part plays. The energy is asked about the section being entered, and
192  each of its four levels says what it is for: held back (55% velocity) for
193  a section that waits, steady (70%) for a verse, driving (85%) for a
194  section that builds, all out (as written) for a payoff, the choruses.
195
196**The form.** The story runs one way (hiss, hunt, fix, clean), so Jev may
197linger but not go back across the fix:
198
199| Section | Bars as written | May be followed by | Plays at most |
200|---|---|---|---|
201| `intro` | 1–8 | verse1 | |
202| `verse1` | 9–16 | hunt | |
203| `hunt` | 17–24 | chorus1 | |
204| `chorus1` | 25–32 | verse2 | |
205| `verse2` | 33–40 | flow | |
206| `flow` | 41–48 | stems | |
207| `stems` | 49–56 | chorus2 | |
208| `chorus2` | 57–64 | outro, chorus2 | twice |
209| `outro` | 65–72 | (the end) | |
210
211The cap is 10 sections. The snake is found once: the only choice in the
212form is whether the clean chorus plays again before the outro (until rev 3
213chorus 1 could send the hunt out again, a path Jev took about once in 260
214plays). Nothing after the fix leads back before it, so the snake never returns once it is gone, except
215in the outro's last bar.
216
217**The parts.** Each plays on its own orbit, named for its question, at
218Jev's level: out is silent, back is half velocity, full is as written. The
219snake is a part like the others, asked about only where it plays as written
220(intro to verse 2).
221
222A part that is all a section has is not a level. These always play, outside
223the levels:
224
225- "That's a snake", the hunt's punchline;
226- "Two stems, one run", the line sung alone (so it is at full level, never
227  back);
228- the hunt's solo snake (bars 21–24);
229- the snake's answer in chorus 1 (bars 31–32);
230- the intro's arpeggio in bars 1–4, before anything else has come in.
231
232The sung lines are never out: the vocals' question has two levels, back and
233full. The snake's level is for the bars where it plays under the band
234(3–20, 25–30, 33–40). The outro's last hiss is the ending, outside the
235levels and the energy: it always lands.
236
237**The transitions** play over the last bar before a section: as written, a
238tom fill, a noise riser, or a dropout of the last beat. As written is a
239noise riser into chorus 1, a pad swell into the clean chorus, and straight
240in elsewhere.
241
242- From the flow section on, no transition is made of noise: the song has
243  said the noise is gone. Into the flow, the stems, the clean chorus and the
244  outro Jev picks from as written, a fill or a dropout.
245- Into chorus 1 the choices are the same three: as written already is the
246  riser.
247- The opening has no transition: nothing is playing before it.
248- A dropout stops held sounds too. The snake's swell and the pad are cut
249  short at the last beat with everything else, so the gap is a gap; only
250  reverb and echo tails ring into it.
251
252Without Jev the song plays as written, every part in full.