The Voice-Acting Actor: Instrument, Tools, Ears, and the Road There

A self-contained technical report on the MOSS voice-acting line, from the first prototype on MOSS-TTS-Local-Transformer v1.5 to the directing agent of September 2026 — told along the project leader's actor metaphor: the Instrument (the speech model), the Tools (adapters, guidance, prompt forms, steering), the Ears (learned scorers) and the Mind (the directing LLM) — with every failure, retraction and measurement error kept.
LAION · Christoph Schuhmann and collaborators · rendered 2026-09-09 19:20 by $SC/code/vareport/en_build.py. No external resources.

Contents
  1. How to read this report
  2. Part I — The instrument and how it was built
  3. I.1 What the system is
  4. I.2 Terms, defined once
  5. I.3 The base model and the two full fine-tunes: v1.5 → v1 → v2
  6. I.4 The adapter ecosystem built on v1 and v2 — the first Tools
  7. I.5 Five hundred synthetic voice profiles
  8. I.6 The measurement instruments, and how they were trained — the Ears before the actor
  9. I.7 The prompting model, the manual, and the v2 casting agent — the actor's first draft
  10. I.8 The timed script
  11. I.9 Supervised rounds: timing solved, emotion lost, format regained
  12. I.10 Preference tuning: seven runs, four corpora, and one metric that lies
  13. I.11 GRPO: five runs, and one structural negative result
  14. I.12 Specialist adapters: the only thing that moved intensity before the agent
  15. I.13 Two new contrastive pair families — and the model they produced
  16. I.14 Three inference-time levers, measured end to end (28–29 August)
  17. I.15 The trajectory corpus
  18. I.16 Prehistory kept for the record: the DramaBox/LaionBox audio DiT, the search agent, and the first evaluation stack
  19. Part II — The actor: Mind, Instrument, Tools, Ears, Selection
  20. II.0 The metaphor and its technical counterpart
  21. II.1 Actor's Mind — the directing model
  22. II.2 Actor's Instrument — SFT3
  23. II.3 Actor's Tools — adapters, guidance, prompt forms, steering
  24. II.4 Layer forensics — where in the network each attribute lives
  25. II.5 Actor's Ears — the perception module
  26. II.6 Selection — Best-of-N
  27. II.7 The babble bug — the fault that named the shipped configuration
  28. II.8 The tools programme after 29 August, part 1 — conditioning datasets, transitions, tempo, crossfade, the quality adapters (protocol §32–§42)
  29. II.9 The tools programme after 29 August, part 2 — the vocal-burst adapters: dose, labels, manufactured data, guidance on the burst (protocol §43–§56)
  30. II.10 The tools and ears programme after 29 August, part 3 — real burst data, the 17-class detector, the scream gain, consolidation, and two predictors put to a human vote (protocol §57–§70)
  31. Part III — The journey: failures, retractions and measurement errors, in order
  32. III.1 When a training metric lies — six cases before the agent
  33. III.2 Engineering failures worth publishing
  34. III.3 Every claim that was made and later withdrawn, in order
  35. III.4 The agent era, September 2026 — what went wrong and what followed
  36. III.5 Contradictions between sources, recorded rather than smoothed over
  37. III.6 Chronology
  38. Part IV — What is published, what is open
  39. IV.1 Open questions of the agent era — the human evaluation first
  40. IV.2 What was open on 29 August and is still open
  41. IV.3 Everything published, with URLs
  42. IV.4 The application that is ready now: dubbing under a duration budget
  43. IV.5 Limitations of everything above
  44. IV.6 The one-paragraph summary, for anyone who read only this section
  45. Appendix — Full tables, computed at render time
  46. A.1 The 18-run ranking of 29 August (corrected percentile scale, 80 prompts × 4 completions, plateau yardstick)
  47. A.2 The shipped configuration, every constant
  48. A.3 Best-of-N optimisation ($SC/out/vb_opt/*.json)
  49. A.4 The 17-class detector's recall floor (per_class_recall.json)
  50. A.5 Encoder comparison (encoder_compare.json)
  51. A.6 The guidance-contrast preference pairs ($SC/cfg_rows)
  52. A.7 The combination study (~/combination_study/stats/analysis.json)
  53. A.8 The burst recipes as served (~/wikiskills/VOCAL_BURSTS.md)
  54. A.9 The coefficient table shipped with the three levers (coefficients.json)
  55. A.10 Index of the protocol
  56. Provenance of the numbers

How to read this report

Every number below was measured, and the measurement is named — the source file, the protocol section, or the earlier report it is transcribed from, with its n, its effect size and its uncertainty wherever the source recorded them; where a source did not record one, the text says so rather than inventing it. Numbers in tables that come from files on disk are computed when this page is rendered. Where a result is negative it is reported as a result. Where two sources disagree, the disagreement is recorded (Part III.5), not smoothed over. Numbers use English notation: a point for decimals, a comma for thousands. German protocol quotations are translated.

The metaphor. The project leader describes the system as an actor. The Actor's Mind is the directing language model that reads the scene and decides what to say and how. The Instrument is the speech model it plays. The Tools are what it can do to the instrument at the moment of speaking — adapters, guidance, prompt forms, steering vectors. The Ears are the learned scorers it listens with, and Selection is how it picks a take. Part I is the biography of the instrument, the tools and the ears before the actor existed. Part II is the actor. Part III is the journey — everything that went wrong, in order, and what it changed. Part IV is what is public and what is open, with the human evaluation first, because it is the one measurement that has never been made.

One caveat applies to the whole document and is repeated where it bites: no human listening study has been completed. As of 9 September 2026, 2,000 stimulus pairs are published and a listening Space exists, but no votes have been collected. The one occasion on which a human checked a lever — the project leader's listening test of the steering vectors — left no quantitative record (Part II.3.3). Every other quality figure in this report is the output of a learned scorer.

Part I — The instrument and how it was built

In the actor metaphor that organises this report, the Instrument is the speech model itself — the thing that produces sound when the actor's mind decides what to say and how. Part I is the biography of that instrument: how a general-purpose text-to-speech base model became, over two full fine-tunes, three supervised rounds on a timed-script format, seven preference runs, five reinforcement-learning runs and roughly a thousand adapters, the model the actor of Part II plays. Part I also builds the Tools (adapters, doses, prompt forms) and the Ears (the learned scorers), because both were made before the actor existed and both were shaped by the failures recorded here. Sources: [TR-0829] §1–§17 and §24 are the backbone; [TR-0826] and [REF-V1] supply what later versions dropped; the protocol supplies the details no report carried.

I.1 What the system is

MOSS Voice-Acting is a text-to-speech model that is directed rather than configured. Instead of exposing a handful of knobs — speed, pitch, an “emotion” enum — it takes a written stage direction in natural language, of the kind a director gives a voice actor, and performs the line accordingly. It speaks English and German, outputs 48 kHz audio, can clone a voice from a reference recording, and can produce non-verbal vocal bursts — sighs, gasps, chuckles, groans — inside a spoken line. Source: [TR-0829] §1.

The production checkpoint at the time of the 29 August report was laion/moss-tts-local-transformer-4.55b-voice-acting-v2, 4.55 billion parameters; the checkpoint the actor plays today is its third supervised successor, SFT3 (…-voice-acting-v2-sft3, Part II). Neither appeared in one step. Section I.3 reconstructs the road: an off-the-shelf OpenMOSS base model, a full fine-tune on distilled voice-acting data plus a rank-256 adapter merged into it (“v1”), then a second six-dataset full fine-tune selected on generated audio rather than on loss (“v2”), then the timed-script rounds of I.9.

Around that model sits an ecosystem that this report describes end to end:

The short version of the result, as of 29 August ([TR-0829] §1). Timing control works: ask for a sentence that takes 3.4 seconds followed by a 0.6-second pause and a 0.3-second sigh, and you get it — the median error on the total clip length is 0.08 s and essentially every clip lands within 0.5 s. Intensity control does not work yet: ask for an emotion in the 90th-to-98th percentile of how strongly that emotion appears in the corpus, and the model delivers something around the 34th percentile. Four training objectives were aimed at that gap. Three things moved it, all by less than a hundredth: data — a small adapter trained on the most intense and most genuine 8 % of the corpus (I.12.3); overdriving that adapter past its trained strength (I.12.4); and a better-constructed contrast in the preference corpus, which as of 28 August produced the best model in the project and was the first preference tuning here that raised intensity without costing intelligibility (I.13.4). What happened after 29 August — the agent, the shipped adapter stack, Best-of-N with new ears, the low-α steering result that superseded the negative one — is Part II.

I.2 Terms, defined once

Terms used throughout ([TR-0829] §2).
TermMeaning
TTSText-to-speech.
LoRALow-rank adaptation. Instead of retraining billions of weights, a small pair of low-rank matrices is trained whose product is added to selected weight matrices. A “rank-16” adapter here is about 34 M trainable parameters against 4.13 B frozen ones. The added delta is scaled by alpha / r; multiplying that by a merge weight or dose λ dials the adapter up or down at inference.
SFTSupervised fine-tuning: ordinary next-token training on (prompt, target) pairs.
DPODirect preference optimisation: training on pairs of outputs labelled better/worse, pushing up the likelihood of the preferred one relative to the other.
GRPOGroup-relative policy optimisation: generate several outputs for the same prompt, score them, push the policy toward the ones that scored above the group average.
WERWord error rate: run a speech recogniser on the generated audio and compare its transcript to the words requested. Used here as an intelligibility and hallucination detector, not as an ASR benchmark.
RVQ codec / audio tokenResidual vector quantisation. The audio is generated as discrete symbols from a learned codec and decoded back to sound.
Vocal burstA non-verbal vocalisation inside speech: a laugh, sob, gasp, sigh, groan, scoff.
Best-of-NGenerate N candidate takes for one prompt, score them all, keep the best. Used as an inference strategy and as a data-filtering strategy.
ECDF percentileEmpirical cumulative distribution function. A raw score is replaced by its rank within the whole corpus so that scores from incomparable heads become comparable. I.6.7 explains why this was necessary and how it was still wrong at first.
Actor / Mind / Instrument / Tools / EarsThe metaphor of Part II: the directing LLM (Mind) plays the speech model (Instrument) using adapters, guidance and prompt forms (Tools), and judges the result with learned scorers (Ears).

I.3 The base model and the two full fine-tunes: v1.5 → v1 → v2

The production model was not trained from scratch and did not arrive in one step. It is the end of a four-stage chain, and every stage constrains what the later ones could do. The emotional-intensity ceiling that the later sections keep running into is partly inherited from it. Source: [TR-0829] §3.

  1. OpenMOSS-Team/MOSS-TTS-Local-Transformer-v1.5 — a general-purpose autoregressive TTS base model, 12-codebook RVQ at 48 kHz.
  2. Voice-acting v1 — a full-parameter fine-tune on distilled voice-acting data, plus a rank-256 LoRA trained on 1 M samples and merged in. Public as laion/moss-tts-local-transformer-4.55b-voice-acting.
  3. Voice-acting v2 — a second full fine-tune over six datasets (185,503 rows), the checkpoint selected on generated audio quality, not on loss. Public as …-voice-acting-v2. This is the model every adapter in I.4 is attached to.
  4. The timed-script round (I.8–I.13) — three more supervised rounds on a new prompt format, then preference tuning and RL on top of that; SFT3 is the instrument of Part II.

I.3.1 The base checkpoint and its representation

MOSS-TTS-Local-Transformer-v1.5 is an autoregressive local-transformer TTS model: a large semantic backbone predicts one hidden state per audio frame, and a small second transformer expands that hidden state into the codec tokens for the frame.

Components of the v1.5 architecture ([TR-0829] §3.1).
ComponentSizeRole
Semantic transformer (Qwen3-style, 36 layers)≈4 BReads the prompt — instruction, reference audio, language, length budget, text — and produces one hidden state per audio frame.
Local “talker” transformer (GPT-2-style decoder)≈550 MExpands each frame's hidden state into the 12 codec tokens that make up that frame.
12 audio LM heads + 1 text headOne output head per codebook. These are weight-tied to the corresponding audio embedding tables — a detail that later matters a great deal (III.2.2).

4.13 B of the 4.55 B parameters are trainable; the remainder is the frozen codec-side embedding. The audio tokenizer is MOSS-Audio-Tokenizer-v2: a residual vector-quantised codec with 12 codebooks of 1,024 entries at 12.5 frames per second, decoding to native 48 kHz audio. Two consequences run through everything below:

The base model's prompt surface, which v1 and v2 both inherit, is two fields: instruction carrying a GENERAL: voice description and a SCRIPT: block with inline parenthesised delivery cues and [pause] markers, and text carrying the same spoken words plain. Recommended sampling with a reference clip: temperature 0.8–0.9, top-p 0.9, top-k 40–50.

I.3.2 Stage one — voice-acting v1

Two things happened here, and they are often conflated because the public artefact is their merge.

(a) A full-parameter fine-tune of v1.5 on DramaBox-style voice-acting data. “DramaBox-style” means the data is formatted as a structured director's instruction plus the spoken line, and it was filtered for vocal bursts and strong emotion. The checkpoint, with its checkpoint-cont-* continuations, lives in the private repository TTS-AGI/moss-dramabox-ft.

(b) A rank-256 voice-acting LoRA on top of that, α = 256, over all attention and MLP projections, trained on 1 million samples (the adapter repository's root is the 100 k checkpoint; checkpoint-1m/ is the full run), then merged into the weights. Training: 8×GPU DDP, bf16, gradient checkpointing, global batch 128, warmup 0.03 with linear decay.

The public release laion/moss-tts-local-transformer-4.55b-voice-acting is that merged model — “the voice-acting LoRA is already folded into the base fine-tune”, in the card's own words. Its card names four data sources, and this mix is the first place where the balance between synthetic and real material is set:

Data sources of voice-acting v1 ([TR-0829] §3.2, from the model card).
SourceWhat it is
DramaBox distillationVoice-acting performances distilled from DramaBox-style prompts — structured director's instructions plus spoken lines — filtered for vocal bursts and strong emotion.
Gemini TTS distillationAudio-matched fixed prompts paired with MOSS RVQ tokens. This is the synthetic arm: material generated by a commercial TTS system and re-tokenised.
EmoliaGerman emotional speech — real recorded material.
Podcast snippetsPublicly available podcast audio — the “audio snippets” corpus, also real.

(c) A vocal-burst LoRA. Rank 256 again, trained on a 100 k inline-burst dataset and resumed from the 1 M voice-acting adapter rather than started fresh. Its training data (TTS-AGI/moss-inline-vocal-bursts, private) writes the bursts into both the instruction and the text, which is what teaches the model that a burst tag corresponds to an event in the audio. The merged result was called merged_inline, and a further all-parameter fine-tune was run from it. A separate library of per-class burst adapters — chuckle, breathy giggle, growl, fearful gasp, contented sigh, and so on — is published as laion/vocal-burst-lora-adapters.

The measurement that justified a burst adapter at all ([TR-0829] §3.2). Burst hit rate — how often a requested burst actually lands in the audio — was 23.6 % from the prompt alone, 16.7 % with only an emotion adapter, and 71.9 % with the burst adapter at full dose. The dose knee is λ = 0.75: it delivers 91 % of the full-merge burst presence at a much better blend. Across 2,304 generations the single best cell was burst @ 1.0 + emotion @ 0.5; half-burst plus half-emotion was the worst combination of the grid, because the two adapters interfere. Placement rules from the same study: put the tag inline in the SCRIPT: block, never open or close a part with a burst, and never place silence directly after one. The n of this grid is the 2,304 generations; no confidence intervals were recorded for the hit rates.

I.3.3 Stage two — voice-acting v2, the six-dataset full fine-tune

This is the production model of the pre-agent era. It started on 2026-07-22 and ran roughly 14–15 hours on 8×A100-80GB — not on JUPITER. Its training scripts lived on an ephemeral RAM disk that has since been wiped; the surviving record is the run notes, the published repositories and the live results pages, which is why this section is reconstructed rather than quoted from a script ([TR-0829] §3.3).

Training configuration of voice-acting v2 ([TR-0829] §3.3).
SettingValue
Data185,503 rows across six datasets — prep_top3_neutral, prep_emolia_elise, prep_a1plus, our_top3, our_emolia, our_adult — materialised to va_data/*.parquet and sampled with equal probability per dataset, not proportionally to size.
Schedule6 epochs · learning rate 1e-5 linear · warmup = the first half of epoch 1. (The architecture's official default is 2e-5; 1e-5 was chosen to match the earlier runs so the two would be comparable.)
BatchPer-device batch 1 · gradient accumulation 8 → effective batch 64 · channel-wise loss weight (1, 32) — the text channel counts 1, the twelve audio channels count 32.
AugmentationPer-sample caption dropout: 15 % drop the GENERAL: block, 15 % drop the inline cues, 70 % drop both (and half of those also drop a third of the remaining inline cues and pauses). Reference-audio dropout keeps 66 %.
EvaluationEvery half epoch: per-dataset validation loss plus five generated clips per dataset, each paired with its ground truth.
Outputslaion/moss-tts-local-transformer-4.55b-voice-acting-v2; rolling checkpoint backup in the private TTS-AGI/moss-voiceacting-ft-checkpoints.

Where the six datasets come from. The prep_* half is prepared external material: prep_top3_neutral is the top-3 reward-selected takes of a neutral-delivery set, prep_emolia_elise is a cut of the Emolia German emotional-speech corpus, and prep_a1plus is an additional voice-acting collection (published tokenised, without per-clip scores, as TTS-AGI/additional-data-a1-plus, 31,306 rows). The our_* half is the model family's own generations, kept by a best-of-N reward: our_top3 the top-3 takes per prompt, our_emolia the Emolia-conditioned generations, our_adult the adult-register set. The Gemini-generated named-character material (TTS-AGI/gemini-adult-voices, 30 named characters × 1,000 clips) sits in the same family of synthetic sources.

How v2 was selected. Not by validation loss. The released weights are the best checkpoint by cumulative genuineness + vocal-burst blend on a held-out set — that is, by scoring generated audio with the two learned proxies of I.6.2–I.6.3. Every half-epoch evaluation therefore produced audio as well as a loss, and the audio decided. This is the rule the whole project follows afterwards, and Part III.1 is six separate demonstrations of why it is not optional: a training-time number has been clear, consistent, monotone and wrong on at least four occasions.

I.3.4 Determinism and serving

Randomness lives entirely in the autoregressive token sampling (the torch RNG, fixed by a seed); the codec decoder is deterministic. The same seed therefore gives byte-identical audio despite a non-zero sampling temperature — verified at MAE 0.0 across two seeded generations. That is what makes the paired sweeps of I.12.4 and the paired Best-of-N experiments of Part II possible at all.

Serving throughput was characterised with SGLang-Omni on a JUPITER GH200 (aarch64): 27.7× realtime at concurrency 32 and 31.2× at concurrency 128 on a 5-second sweep, against 18.8× for the plain transformers baseline — a 1.47–1.66× speedup, saturating around concurrency 96–128 at roughly 6.4 clips/s. On 30-second clips it peaks at 85.3× realtime at concurrency 64. The lesson in that pair of numbers is to always quote a realtime factor together with the clip length it was measured on; the advertised ~35×-per-A100 figure was not reproduced locally. Source: [TR-0829] §3.4.

I.3.5 The numerics that cost hours, recorded because they generalise

Two of these — bf16 eps and the checkpointing path — are the same class of failure: a numerical default that is correct in fp32 and silently wrong in bf16. The sibling diffusion-model project in the same group (I.16) hit the identical class from the other side, where bf16's unit in the last place exceeded a typical Adam update and training silently froze; the fix there was fp32 master weights. The transferable rule is that bf16 needs its optimizer numerics checked explicitly, not inherited. Source: [TR-0829] §3.5; the diffusion case is [REF-V1] Part II.

I.4 The adapter ecosystem built on v1 and v2 — the first Tools

Before the timed-script round and long before the agent, the project's main mechanism for control was adapters. In the metaphor these are the actor's first tools: things you hand the instrument to change what it does without rebuilding it. Every one is a PEFT LoRA attached to the same frozen base model, and — this is what makes them stackable — they all declare the same 23 target-module patterns: the global Qwen3 stack's q, k, v, o, gate, up, down_proj; the local GPT-2 decoder's c_attn, c_proj, fc_in, fc_out; and all twelve audio_lm_heads.0…11. They differ only in rank and in what they were trained on. Source: [TR-0829] §4.

Adapter families before the current round ([TR-0829] §4).
FamilyAdaptersWhat it movesTypical dose λ
voice profiles (pilot)10 (+60 ablation)Who is speaking — one identity across the whole expressive range1.0
voice profiles (production)500The same, for the full synthetic voice corpus; rank 41.0
emotion LoRAs v340One adapter per emotion, rank 32 / α 640.35–0.75
VoiceNet dimension LoRAs11457 timbre/prosody dimensions × {high, low}0.25–1.25, per dimension and direction
vocal-burst LoRAs64One per burst class — sobs, sharp inhale, chuckle, scream …0.5–0.75
character LoRAs, genuine / refined / 12-cluster120 / 120 / 12Character archetypes (orc, dragon, fairy, …)1.0
explicitness LoRAs (gated)6Adult / unguarded register; access-restricted behind an age check0.2–0.8
German broadcast · sports commentary9 · 6Domain registers0.25 · 1.0

I.4.1 Emotion LoRAs: three generations, and what each one taught

v1 — the rescue result (2026-07-26). Evolutionary search over natural-language prompts had failed outright to elicit Anger and Amusement from the base model. Training a per-emotion rank-64 LoRA on emotion-heavy audio rescued them. The ablation used the best checkpoint and the same emotional prompt in both arms, so the LoRA is the only variable (EmoNet z; higher = more of the emotion). No n per cell is recorded in the source for this first ablation.

Emotion LoRA v1 rescue ablation ([TR-0829] §4.1).
EmotionBaseline v2, prompt only+ rank-64 LoRANote
Amusement−0.51+2.50the prompt failed completely; rank 64 ≫ rank 128 (2.50 vs 0.99)
Anger−0.37+0.62ranks roughly tie; rank 128 keeps speaker similarity higher (0.53)
Fear+0.34+1.20rank 64, best at an intermediate checkpoint (≈6 epochs)
Sadness+0.80+0.57the LoRA does not help — the prompt already sufficed

Rank 64 became the default (rank 128 overfits the ≈1,000-clip buckets), and the best adapter was consistently an intermediate checkpoint — 20 epochs drives the loss to 0.03–0.4 and overfits. Speaker similarity was preserved at 0.41–0.58.

Overnight scale-up and the mixture study (2026-07-27). 34 emotions were trained on natural EmoLia top-1000 subsets. Two things came out. First, there is no single best rank: Anger and Fear peak at rank 128, Sadness at 64, Amusement only moves at all at rank 32. Second, and more consequential, a seven-way mixture study over three sources — D cross-folder DACVAE, E natural EmoLia, G LAION's Got Talent (high-emotion synthetic voice acting) — found that Got-Talent wins: pure Got-Talent is best for Fear (z = +2.62), Distress, Sadness and Anger; only Amusement preferred a three-way mix. High-emotion synthetic data became the primary emotion-LoRA training source. Clear intensity lifts at this stage: Amusement +1.29, Sadness +0.59, Anger +0.52, Fear +0.47 — with many emotions near zero, because natural speech is subtler.

v3 — 40 of 40, and the synthetic character dialled back. The shipped set, TTS-AGI/moss-emotion-loras-v3, is 40 adapters at rank 32, α 64, single-phase, trained against v2. Its distinguishing change is a cap: Got-Talent material is limited to ≤ 25 % per bucket, so at least ≈60 % of every mix is natural or expressive non-Got-Talent material — EmoLia, DACVAE, the Gemini-generated adult-voice set, and the mitermix collection. The card states the reason plainly: this “removes the synthetic Got-Talent character that v2 had”. So the trajectory of this line is a full round trip — synthetic data was found to be the strongest single source, and then had to be capped because it stamped its own character on the output.

I.4.2 The full 40-emotion sweep: 11 responders and 29 that barely move

One adapter per emotion, best of ranks {16, 32, 64}, about 1,000–1,400 clips per emotion (natural EmoLia topped up with Got-Talent, and about 25 % of rows carrying a same-voice contrasting-emotion reference). The best adapter per emotion was chosen by a gated composite, z + 0.15·blend + 0.5·spk + 0.1·genu — the emotion term is gated and the quality terms are already there, which is the ancestor of every reward used later in this report, including the Best-of-N reward of the agent.

Strong responders — z-lift ≥ 1.0 (11 of 40): Amusement +1.99 · Helplessness +1.79 · Intoxication +1.58 · Fatigue +1.47 · Sexual Lust +1.38 · Distress +1.35 · Anger +1.34 · Fear +1.26 · Sadness and Teasing +1.16 · Malevolence +1.15.

Weak or inverted (29 of 40): high-valence and subtle states — Awe, Elation, Pride, Contentment, Pleasure, Hope, Thankfulness, Astonishment, Infatuation, Contempt … Two are self-reducers and are usable as minimising directions: Contemplation (+0.06 → −0.56) and Concentration (−0.30 → −0.40).

The central emotion-LoRA finding. An emotion LoRA raises the target emotion and lowers genuineness, vocal-burst blend, audio quality and intelligibility. The best total reward is therefore frequently prompt-only, or the LoRA merged at about 50 %, rather than the LoRA at full strength. This trade-off — not raw emotion maximum — drives every recipe in the public manual, recurs unchanged in the timed-script round (I.12), and is the reason the agent's shipped stack of Part II keeps the emotion adapter at 1.0 only because the surrounding quality adapters pull the other way.

A reference-audio ablation on the same sweep is worth one line because it is so emotion-specific: attaching a same-voice, contrasting-emotion reference to ≈50 % of rows helps Amusement enormously (+2.93 z, beating even the best reference-free mixture) and is mixed or small for everything else, with no improvement in speaker similarity. It is a per-emotion option at around 25 % of rows, not a default.

I.4.3 Character LoRAs, and the geometry of voice space

Character adapter sets ([TR-0829] §4.3).
SetCountMethod
Genuine120Mined character voices from v2 generations plus Gemini-written casting profiles; rank-32, from scratch, ≥ 50 DNSMOS-filtered clips per class (a further 46 warm-started adapters cover the smallest classes).
Refined120Self-distillation: 640 candidates per character → filter → SIDON restoration → keep the top 20 % by profile reward → warm-start fine-tune for 3 epochs. Ships three epoch checkpoints each.
12-cluster representatives12The nearest genuine adapter to each KMeans centroid over the 57 VoiceNet voice-quality dimensions.

Voice space is a continuum, not a set of types. The 567 character reference voices, z-scored over their 57 VoiceNet dimensions and clustered at k = 12 and k = 18, give a silhouette score below 0.2 everywhere; k = 18 adds only about 5 % of variance and is less separated than k = 12. Twelve is a workable practical taxonomy; eighteen just subdivides the same families. The poles of the space are bright / high / feminine at one end and dark / low / gravelly / masculine at the other. A caveat travels with it: the 120 creature adapters skew masculine and monstrous, so the feminine and clean clusters get loose matches — the “High-Bright-Feminine” centroid maps to a guttural imp, which is a conceptual mismatch. Good matches: Playful-Feminine → elf, Full-Chesty → ogre, Silken → titan, Gravel → skeleton.

A cross-lingual reference does not clone a voice; it invents one. Over 1,280 generations conditioning German and English output on Japanese anime references, mean speaker similarity to the reference was 0.188 against an unrelated-speaker floor of 0.105 and a self-similarity of 1.000. Pitch tracks (r = 0.591) but compresses toward the model's own register (median F0 182 → 157 Hz). English transfers better than German (0.217 vs 0.158), and best-of-8 lifts it to 0.294. The recommendation is to use cross-lingual references to mine new character voices, never to clone — a limit that returns in III.4 as a constraint on dubbing.

The corpora behind the character work ([TR-0829] §4.3).
DatasetSize (rows)What
scientifi-papers/chavo-annotated750,242Game character clips with the full 57 VoiceNet dimensions, 40 emotions and genuineness/blend, tokenised for MOSS v2. Heavily skewed: 668 k of the rows are creature, across 807 races.
laion/moss-character-voices-top3-captioned≈38,400Designed takes, 128 shards, ≈5,500 per archetype.
TTS-AGI/gemini-adult-voices (private)30,00030 named characters × 1,000 clips — the Gemini-generated synthetic arm; tokenised, no per-clip scores.
TTS-AGI/additional-data-a1-plus (private)31,306Voice-acting clips; tokenised, no per-clip scores. This is prep_a1plus in I.3.3.

I.4.4 VoiceNet-dimension LoRAs

One rank-32 adapter per pole of each VoiceNet dimension, trained on v2. The published evolution study covers 86 dimensions across high and low poles. Δ100 is the measured shift at full dose. Most responsive: Ranting style S_RANT high +0.78 · Tension TENS high +0.67 · Dynamic arc DARC high +0.65 · Respiration RESP high +0.65. Least responsive: Mixed resonance R_MIXD low +0.04 · Roughness ROUG low +0.06 · Mixed resonance high +0.06.

Averaged over all of them, λ = 0.5 delivers about 36 % of the full-strength shift (mean in-direction shift +0.350 at λ = 1.0 against +0.157 at λ = 0.5) — so a half-strength merge buys a large fraction of the effect at much lower risk to speaker identity. That matches the search agent's independent canary finding ([REF-V1] Part III) that tempered doses of 0.4–0.5 preserve identity where 1.0 and above maximise the target and destroy it.

A practical trick from the manual that has no obvious analogue elsewhere: use VoiceNet adapters as stabilisers. Adding vn_CLRT_high @ 0.4 and vn_STRU_high @ 0.3 (plus vn_S_NARR_high for storytelling) protects intelligibility when a strong emotion or style adapter would otherwise push word error rate up; adding +0.5·ESTH to the reward biases toward “nice to listen to”. This is the direct ancestor of the esthetics_high adapter in the agent's shipped stack (Part II).

Bucket feasibility, measured on game-voice data. Per-dimension buckets at three poles — high (≥ 5), mid (3–4), low (≤ 1), capped at 1,500 each — give 154 trainable buckets (≥ 200 clips) across 57 dimensions. Sixteen dimensions have no high bucket at all in game voices (S_TECH, VALN, S_NEWS, STRU, SMTH, R_ORAL, R_MASK, R_MIXD, R_HEAD, EXPL, DFLU, ESTH, COGL, CHNK, BRGT, BKGN) and need podcast or empathic-speech fill; EXPL has essentially no variation in that data at all.

I.4.5 Explicitness LoRAs

Six adapters, published both publicly and in a gated form requiring manual approval, an 18+ confirmation and name/email collection: TTS-AGI/moss-explicitness-loras. The set is aesthetic_mix_r32 (the recommended default), adult_r32, a1_r32, and raw_r16 / raw_r32 / raw_r64, of which raw_r32 gives the largest Sexual-Lust lift. The recommended recipe is the aesthetic adapter at 0.2–0.8 with the emotion adapter held at 0.35–0.75. Source: [TR-0829] §4.5.

I.4.6 A negative result worth keeping: conditioning on 99 numeric scores

An attempt to bypass natural language entirely. Ninety-nine numbers — the 40 EmoNet emotions, the 57 VoiceNet dimensions, genuineness and vocal-burst blend — were each projected into the model's 2,560-dimensional latent space by a tiny per-score projector, and the 99 resulting rows were spliced into the prompt between the backbone's <|vision_start|> and <|vision_end|> tokens, using the otherwise-unused <|image_pad|> token as a placeholder. A rank-128 LoRA and the 99 projectors were trained jointly, warm-started from an existing adapter. [REF-V1] adds the slot layout: slots 0–39 the emotions, 40–96 the VoiceNet dimensions, 97 genuineness, 98 blend.

The first thing measured was the ceiling: how much of a conditioning signal survives a round trip through the codec at all. Encode a clip, decode it, re-score it, and correlate:

Codec round-trip correlation of the 99 conditioning signals ([TR-0829] §4.6).
SignalRound-trip correlationReading
Overall0.993
40 emotions0.993Emotion survives the codec almost perfectly.
57 VoiceNet dimensions0.995So does timbre.
Genuineness0.951Good.
Vocal-burst blend0.823The hardest of the 99. Bursts are fine, transient events that a 12.5 Hz codec partly smears — so the achievable control on blend is intrinsically lower than on anything else.

The outcome of the experiment itself was negative: the projected-score conditioning path came out slightly below plain v2 on the two proxies in a head-to-head evaluation (genuineness −0.12, blend −0.44; n not recorded in the source). Numeric conditioning is promising and did not beat text in this iteration. The correlation ceilings, though, are permanently useful, and the 0.823 for blend explains a recurring difficulty later — including why the blend term of the agent's Best-of-N reward correlates negatively with the human-proxy target (Part II).

I.4.7 Merging: what a dose actually does

Merge weight is an intensity dial, and it is roughly linear. Averaged over adapters and multiple seeds, λ = 0.5 recovers about 40 % of the full-strength gain, with no snap-on, no collapse and naturalness held. (With a fixed seed the same curve looks like a staircase rather than a smooth line — an artefact of sampling, not of the adapter.)

And it destroys speaker identity, monotonically. Measured as ECAPA speaker-embedding similarity to the reference clip, on cross-lingual Japanese → German/English generations:

Dose against speaker similarity ([TR-0829] §4.7).
Dose λSimilarity to referenceReading
0.00.62baseline cloning fidelity
0.50.57usable
1.00.50degrading
1.5−0.03below the 0.105 unrelated-speaker floor — the output voice has no relationship to the reference at all

Anchors from the same encoder: a reference against itself scores 1.000, two different speakers score 0.105. Published guidance: keep λ ≤ 0.5 if reference fidelity matters, 0.75–1.0 if only the emotion matters. Overdriving also truncates: in a controlled stacking run, one sentence and one seed produced 3.84 s with the voice adapter alone and 1.68 s — under half — with an emotion adapter at dose 1.9 stacked on top. An overdriven stack does not merely sound wrong; it stops early. This is the earliest observation of the mechanism behind the agent's babble bug (Part II.7).

Difference vectors. Taking Effect = mean z at λ = 1.0 minus mean z at λ = 0, and recording for each measured dimension which adapter raises it most and which lowers it most, turns the adapter library into a set of directions rather than a set of labels. Highlights: the Amusement adapter raises Intoxication by +1.19 as well as Amusement by +1.80; the Sadness adapter is a broad negative lever, raising Embarrassment +1.08, Disappointment +0.85 and Fear +0.66; and subtracting Teasing (−1.70 on Concentration) is a “more concentrated” direction. Hope at −0.83 on Contentment behaves the same way.

I.4.8 Vocal-burst merging: the dose knee and a blind evaluator

An in-sentence merge-dose curve over 1,536 generations, 8 families × 64 classes:

Burst dose curve, n = 1,536 ([TR-0829] §4.8).
Burst dose λBurst presentBlendGenuinenessComposite
0.2527.3 %5.731.480.080
0.5050.5 %4.761.890.146
0.75 — the knee64.8 %4.171.970.199
1.0071.4 %3.971.940.226

λ = 0.75 gives 91 % of the full-dose burst presence at the peak of genuineness and a better blend; presence is bought against blend roughly one for one. Two exceptions: Sexual Lust peaks at 0.75 and then drops, Fear is best at 1.0. Five classes never fire inside a sentence at all — hiss, kissing, lip-smack, playful whistle, slurping. Rank barely matters (composite 0.416 / 0.424 / 0.427 for ranks 16 / 32 / 64). The bare-burst prior is the opposite: in isolated-burst sweeps λ = 1.0 won every class. Do not carry that setting into sentence generation.

The evaluator was blind to a third of the taxonomy. The locator-plus-classifier pipeline could not recognise 26 of 77 burst classes even in real human recordings. On the 45 measurable classes, 84 % improved (mean +0.132); pooled over all 77 that dilutes to +0.060. A separate Gemini listening judge confirmed 64 of 77 classes at mean ≥ 1.5 — including 19 of the 26 the metric evaluator cannot see (controls: real human audio 1.874, base model with no adapter 0.411, adapter 1.755). The same label filter discarded 83 % of the available data, 97 k rows down to 16.7 k usable. The transferable point: when a metric reports a small effect, check whether the metric can see the effect at all before concluding the effect is small. The agent-era burst detector (Part II, Ears) was built because of this.

I.4.9 Contained emotion — a mechanism that works, and where it reverses

“Contained” means the emotion is present but held in, which is what most adult dramatic performance actually is. The mechanism that works is specific: keep the emotion and drop vulnerability — a light vn_VULN_low at 0.06–0.18 plus a parenthetical restraint cue, over a character-adapter scaffold with the emotion at about 0.4. What does not work: turning the emotion down, adding tension, or applying heavy Emotional-Numbness or vn_TENS_high — all of those armour the delivery into a news-anchor flatness.

A deep paired re-run over 1,878 clips measured ΔVULN as the suppressed cohort's mean minus the free cohort's; negative means the masking worked:

Contained-emotion re-run, n = 1,878 clips ([TR-0829] §4.9).
VerdictEmotions, with ΔVULN at moderate / intense
Masks cleanlyFatigue −0.17 / −0.27 · Embarrassment −0.11 / −0.25 · Pain −0.18 / −0.01
Masks at intense onlyEmotional Numbness −0.01 / −0.55 · Intoxication +0.02 / −1.05 · Helplessness +0.18 / −0.23
ReversesSadness +1.29 / +0.45 · Pride +1.09 / +1.14 · Distress +0.43 / +0.82 · Affection +0.18 / +0.59
Erasure, not maskingFear (intense −0.28, and the emotion itself −0.32) · Doubt (−0.75, −0.56)

The rule that falls out: masking works when the audible signal is effort, arousal or social surface and vulnerability is incidental; it fails or reverses when vulnerability is the emotion's core — for Sadness, Pride, Distress and Affection, ship the free take. This result overturned an earlier “moderate contains, intense flattens” story that turned out to be an artefact of four anchors and sparse unpaired data — the first of the retracted claims collected in Part III.

I.4.10 “Text carries the condition”

A large ablation — 106 growing to 124 arms, up to 3,968 candidates — found that several apparently “vocal” dimensions are in fact properties of the written line, and no adapter moves them. Source: [TR-0829] §4.10.

I.4.11 Notation, ranking and the reward form

A notation bug that inverted a whole experimental round. Inline cues and bursts are written in round brackets, pauses in square brackets, tags in lower case. One pipeline emitted <sobs> instead. The angle-bracket form was not matched by the WER stripper, so the word “sobs” stayed in the reference string and every burst-carrying take was charged with failing to pronounce it. The ranker then selected against tagged takes, and round 2's conclusion that “tags lose” was an artefact. With the corrected stripper (\([^)]*\)|\[[^\]]*\]|<[^>]{1,40}>), round brackets beat angle brackets on emotion, blend and genuineness alike.

Three further ranking decisions were each measured on 28,698 stored candidates:

The production matrix at the end of this line was 169 distinct adapter keys — 13 character, 40 emotion, 114 VoiceNet, plus explicitness and sports.

I.4.12 Best-of-N economics, and why ASR is the swappable cost

Best-of-N winner reward against N ([TR-0829] §4.12).
Candidates N per group18163264
Mean reward of the winner0.3830.5200.5500.5740.596

The 16 → 32 step is worth +4.46 % reward, −8.95 % duration error and +0.254 blend; 32 → 64 costs 3.4× for the same increment. Crucially, generation time is flat from batch 16 to 32 (20.34 → 20.54 s on a GH200) because the decode loop is latency-bound, so a batch of 16 wastes half the GPU. N = 32 is the operating point of the corpus-building era, and peak VRAM there is about 23 GB. (The agent of Part II runs BON_N = 8 because it is interactive and must answer within one turn; the cost argument is the same, the budget differs.)

The dominant cost is then the speech recogniser, and it is swappable: Parakeet-v3 runs at 52× realtime and is 12 % of a 32-candidate group's cost; Whisper-turbo 12.7× and 32 %; CrisperWhisper 3.1× and 66 %. Rank on Parakeet and re-transcribe only the top three with CrisperWhisper — on a 40 × 3,000 production run that took the bill from 2,952 to 1,148 GPU-hours. Parakeet-TDT 0.6b v3 is the ear the agent uses for WER today.

Two more rules from the same body of work, both of which recur later: rank whole assemblies, not parts — across 108 rounds the top-ranked assembly used the best candidate of every part in only 11 rounds (10.2 %) — and n = 8 per cell is noise; use n ≥ 30 with the base model in the same batch.

I.4.13 What building all of them taught about rank, dose and stacking

Rank is not where the quality is. The voice-profile pilot trained each of ten voices at ranks 16, 8 and 4 in a single process against one shared frozen base, stepping on the same micro-batches with the same seed — so rank is the only difference between the three adapters. On 1,920 held-out clips per arm:

Voice-profile rank ablation, n = 1,920 clips per arm ([TR-0829] §4.13).
ComparisonΔ speaker similarity95 % CIp
base → rank 4 (stage 2)+0.2102[+0.2030, +0.2175]< 1e−300
rank 16 → rank 4+0.0023[−0.0014, +0.0060]0.22
rank 16 → rank 8+0.0017[−0.0018, +0.0052]0.35

The base → rank-4 effect is about ninety times the rank-16 → rank-4 gap, and the inter-rank gaps are indistinguishable from zero on every axis measured. Nine of the ten voices ship at rank 4 — 8.6 M trainable parameters. (Speaker similarity here is cosine similarity between ECAPA embeddings from speechbrain/spkrec-ecapa-voxceleb.)

Dose is not what carries the result. Across the acting corpus, emotion-adapter dose correlates −0.226 with the measured emotion peak and +0.084 with genuineness. In the manual's words: “you cannot buy a turn by raising λ.” 98.6 % of the 492 doses actually used across the corpus sit in the 0.35–0.75 band. What carries the result is the written line (I.4.10).

Only 22 of 40 emotion adapters beat a good prompt. The manual's own emotion index reports that only 22 of the 40 emotion LoRAs produce a positive lift over a well-written prompt with no adapter. Biggest wins: Sexual Lust +0.25, Fatigue +0.21, Helplessness +0.18, Jealousy and Envy +0.17. Biggest losses: Distress −0.18, Concentration −0.12, Anger −0.12, Relief −0.11, Disappointment −0.10. Eight emotions — Interest, Concentration, Impatience and Irritability, Amusement, Hope, Doubt, Disgust, Pride — are best served by a neutral prompt and no adapter at all.

Stacking is a real engineering problem. Because every family targets the same 23 modules, a voice adapter and an emotion adapter are always fighting over the same weights, and at equal λ the larger-rank adapter dominates. Voice adapters are rank 4; emotion, burst, VoiceNet and character adapters are rank 32. That asymmetry — not any target-module difference — is why every non-voice λ is below 1 unless that adapter is the condition. Two library-level traps are documented with their symptoms, and both killed multi-hour runs: PeftModel.active_adapter is a plain attribute that base_model.set_adapter() does not update, so it stays pinned to the first adapter ever loaded and raises KeyError forever once that adapter is evicted; and a snapshot of the trained scalings goes stale the moment another adapter is loaded, so a dose computation must consult both dictionaries and record the trained value on first sight. A third surfaced on 28 August and is recorded with the published cards (IV.1).

I.5 Five hundred synthetic voice profiles

A large part of the training material for the timed-script round is not recorded speech but voice profiles: synthetic speakers, each generated by the v2 model with the adapters of I.4, each exercised across the whole expressive range — 40 emotions, 57 VoiceNet dimensions, acting edge cases, character clusters, vocal bursts, English and German — and each filtered by the scoring models before anything was kept. Source: [TR-0829] §5.

Scale, measured on 12 verified voices: 40,143 non-empty clips per voice · 108.4 h of playtime per voice · 842 condition groups · 30,869 unique utterances · mean clip 9.72 s. Projected to 500 voices: ≈20.1 M clips, ≈54,200 h — about 6.2 years of audio. Cost: 31.2 GPU-hours per voice billed → ≈16,150 GPU-h ≈ 1.16 M core-hours for 500 voices. Throughput held at scale: 1,354–1,451 generations per GPU-hour against a pilot mean of 1,405.

The selection rule that produced them is the direct ancestor of the reward used in the timed-script round and of the agent's Best-of-N reward:

reward = (genuineness + w_blend·blend + 1.25·target) × (1 − WER) · score = reward × durationFloor × lengthCeiling × identityRank × floorPenalty

Three of its design choices were each measured on 28,698 stored candidates, and each is a small lesson (the first two are I.4.11 seen from the data side):

The honest headline of this stage. 43.4 % of the voice-profile corpus falls below the project's own 0.40 speaker-similarity floor — and 57.1 % within the intense-emotion block. Identity was ranked, not gated, so the adapters were trained on takes that partly drift off the reference voice. Caveat: re-scored with an independent WavLM timbral embedder that played no role in selection, 75.4 % of the takes classified as failures sit above that model's own published same-speaker threshold, and the two embedders correlate at only r = 0.66. So “43.4 % is not recognisably the same speaker” is a claim about one similarity scale, not an established perceptual fact. Consequence: strong emotion and stable identity are, on this system, in direct tension — which is I.7 and, in the agent era, the has_voice split of Part II.

The largest published artefact from this stage is laion/laion-voice-profiles-annotated — 28,212,933 rows, 5.8 TB.

[REF-V1] adds two details that later versions dropped: the voice-profile pipeline used 108 rounds of assembly ranking (the 10.2 % figure of I.4.12 comes from there), and the per-voice adapters of the pilot were trained in a single process against one shared base so that rank comparisons would be paired — the design that made the p = 0.22 / 0.35 result of I.4.13 possible.

I.6 The measurement instruments, and how they were trained — the Ears before the actor

Nothing in this project can be evaluated by listening at the scale at which it is generated — a single GRPO run produces 1,400–1,700 fully scored clips per hour. So the project runs on a set of learned scoring models. They do three different jobs, and it is worth keeping those apart: selection (which of 32 or 64 candidate takes to keep), reward (what an RL objective maximises) and evaluation (the numbers in this report). The same models do all three, which is a known circularity and is discussed in I.6.8. In the metaphor these are the actor's Ears; Part II.5 describes how the agent uses them, this section how they were built. Source: [TR-0829] §6.

I.6.1 VoiceCLAP — the shared audio embedding

VoiceCLAP is a CLAP-style (Contrastive Language–Audio Pretraining) dual-tower model: an audio encoder and a text encoder trained so that a clip and a sentence describing how it sounds land close together in one shared vector space. The workhorse is laion/voiceclap-commercial:

VoiceCLAP-commercial ([TR-0829] §6.1, from the card).
PartWhat it is
Audio towerBUD-E-Whisper-Small — 12 layers × 768 dim × 12 heads, 80-mel at 16 kHz
Text towerall-MiniLM-L6-v2, 6 layers × 384 dim, mean-pooled
Joint space768-d, L2-normalised; ≈110 M parameters total
LossSigLIP sigmoid contrastive plus a prototypical-contrastive auxiliary term — 39 learned emotion prototypes, weight 0.2, cross-entropy against z-scored pseudo-labels derived from Emolia's emotion scalars
Training dataOnly commercially licensable sources: Emilia-YODAS (the CC-BY-4.0 subset of Emilia), LAION's Got Talent, and the in-house Majestrino corpus. The CC-BY-NC clips are filtered out at training time by clip id; Expresso and EARS are not used.

A controlled ablation reported on the card shows that removing the non-commercial data costs nothing — the commercial-only model wins 4 of 5 benchmarks against the full-mix model, trailing only on fine-grained intensity ranking of synthetic audio. Two recovery experiments were tried and discarded: adding VoxCeleb1/2 hurt, and upweighting the safe corpora was flat.

A larger sibling exists, laion/voiceclap-large-v2 — a rank-16 LoRA (α 32) on the 7 B LCO-Embedding-Omni-7B (Qwen2.5-Omni thinker), producing a 3,584-d embedding and setting the better numbers on the human-annotated VoiceNet benchmark (emotion balanced accuracy 0.7069 against voiceclap-small's 0.6754). It is not what the pipeline runs on, because the 768-d commercial model comes within a hair of it on the two heads that matter at 1/4.7 the embedding width. In the agent era the two were compared directly as burst-detector backbones (Part II.5); the large model wins the ranked-vs-measured comparison there.

The architectural fact that matters most: one VoiceCLAP encode feeds three separate instruments. In the production scorer (pp_scores_fast.FastScorer) a clip is encoded exactly twice in total — once by VoiceCLAP-commercial and once by BUD-E-Whisper — and every downstream head reads one of those two vectors.

I.6.2 Genuineness: does it sound like a person or like a reading?

There are two generations of this predictor, and conflating them is easy, so both are described.

Generation 1 — the Empathic-Insight “Authenticity” head. The original genuineness signal is one of the expert heads in the Empathic-Insight-Voice suite: the repository ships 55 checkpoints, of which 40 are emotions and 15 are demographic/vocal/psychological attributes — and Authenticity and Arousal are two of those 15. It is trained exactly like an emotion expert: a BUD-E-Whisper encoder is fine-tuned once and then frozen, and a single MLP is trained on its flattened sequence embedding. This is the head occupying slot 97 of the project's 99-dimensional scorer (va_rescore.NinetyNineScorer), and it is what filtered the data for the v1/v2-era emotion buckets. Its published private siblings are laion/genuineness and laion/vocal-burst-blend. One number from that generation is worth carrying forward: inter-rater agreement on authenticity is very low — Cronbach's α ≈ 0.185 on n = 300, far weaker than the emotion labels. The head is modelling a genuinely subjective target. That is precisely why genuineness is down-weighted (≈0.1 in the selection score of I.4.2) and used as a gate rather than as a primary objective.

Generation 2 — the VoiceCLAP head the pipeline uses. laion/voiceclap-commercial-genuineness predicts a single continuous score in [0, 6] for how much a clip sounds like a real, lived-in spoken moment rather than a rehearsed or synthetic read. Anchors from its own card: 0–2 reads as rehearsed, flat or synthetic; 3–4 plausibly natural; 5–6 sounds like a genuine, unscripted moment. It explicitly is not about audio fidelity — it is about natural timing, breath, micro-imperfections and emotional grounding.

Genuineness head, generation 2 ([TR-0829] §6.2; verified against ~/scorer_facts.md).
AspectValue
ArchitectureFrozen VoiceCLAP-commercial 768-d embedding → L2-normalise → standardise with stored μ/σ → Linear(768,50) → GELU → Dropout(0.2) → Linear(50,1), about 38 k parameters. The backbone is never fine-tuned.
Training data≈10,000 speech clips from a mix of various TTS systems and the Emolia expressive-speech collection, spanning a wide range of naturalness.
LabelsEach clip scored 0–6 by Gemini-3.1-Pro.
LossHuber (δ 1.5), Adam, standardised inputs.
Heads shippedTwo: full (default, trained on the whole label distribution) and balanced (retrained on a class-balanced subset — flatter per-bucket error, similar overall MAE).
Validation140 held-out clips, 20 per integer score 0–6, stratified, seed 1234. Full head: MAE 1.00 · Pearson r 0.77 · RMSE 1.33.
License / packagingCC-BY-4.0; fully standalone (≈450 MB) — the frozen embedder is bundled, so nothing is fetched at inference.

I.6.3 Vocal-burst blend: does the sigh belong to the sentence?

laion/voiceclap-commercial-vocalburst-blend predicts, on a 0–10 scale, how naturally a non-verbal burst blends into the speech around it. Anchors from its card: 0 = disconnected — spliced or pasted in, robotic, or the wrong emotion for the context, and also assigned to clips with no genuine burst at all; 5 = fits the context but sounds performed or stagey; 10 = totally organic, indistinguishable from a natural human reaction. Intermediate values interpolate; output is clamped to [0, 10]. Architecturally it is the same shape as the genuineness head — frozen VoiceCLAP-commercial 768-d → L2-normalise → standardise → Linear(768,50) → GELU → Dropout(0.2) → Linear(50,1), about 38 k parameters — which is why one encode yields both scores. Its card publishes a comparison at equal head architecture across three embedders:

Blend head across embedders ([TR-0829] §6.3, from the card).
EmbedderWidthVal MAE ↓Val corr ↑
VoiceCLAP-commercial (shipped)7682.0570.625
VoiceCLAP-small7682.3600.418
VoiceCLAP-large-v23,5841.8950.645
What the blend card does not say. It gives no training-set size, no label source and no train/val split description — only the comparison above. Nothing here should be read as saying the blend head was labelled by Gemini or trained on Emolia; that is stated by the genuineness card about itself, and does not transfer. Neither card publishes a human-agreement figure. The blend card's own stated limits: the frozen encoder caps accuracy; mid-range scores (4–6) carry the most uncertainty and the model is most reliable at the extremes; it is a fast proxy for ranking and reward shaping rather than a replacement for a strong multimodal judge or a human listener; its behaviour on music, noise-only or non-speech audio is undefined. An earlier draft of the 26 August report asserted training details for this head that its card does not contain; they were removed (III.3).

These two heads are why the project's data selection is not naive. As I.12.1 shows, the most intense recordings in the corpus are on average the least genuine, so selecting on intensity alone selects for overacting. Genuineness and blend are the gate that stops that — and they are two of the four terms of the agent's Best-of-N reward.

I.6.4 Empathic-Insight-Voice-Plus: 40 emotion heads and 4 quality heads

laion/Empathic-Insight-Voice-Plus is not one model but a suite of independent single-output MLP regression experts on top of a frozen BUD-E-Whisper encoder (a fine-tuned Whisper-Small producing a 1,500 × 768 sequence).

The production scorer loads 44 heads — the 40 emotions and the 4 quality experts — and leaves the attribute heads unloaded, because the VoiceNet regressor (I.6.5) covers those dimensions. Emotion outputs are unipolar, nominally 0 → 4 (0 = the emotion is absent, 4+ = extreme manifestation, card maximum 4.5), but the heads are unbounded regressors and real corpus values overshoot — the project's own normalisation statistics record a maximum of 8.20 on the Anger head.

The 40 emotions, verified against the repository's file listing, are the axis along which everything in this report is requested and measured: Affection · Amusement · Anger · Astonishment/Surprise · Awe · Bitterness · Concentration · Confusion · Contemplation · Contempt · Contentment · Disappointment · Disgust · Distress · Doubt · Elation · Embarrassment · Emotional Numbness · Fatigue/Exhaustion · Fear · Helplessness · Hope/Enthusiasm/Optimism · Impatience and Irritability · Infatuation · Interest · Intoxication/Altered States of Consciousness · Jealousy & Envy · Longing · Malevolence/Malice · Pain · Pleasure/Ecstasy · Pride · Relief · Sadness · Sexual Lust · Shame · Sourness · Teasing · Thankfulness/Gratitude · Triumph.

A naming bug found while writing the 26 August report. The heads are keyed by name, taken from the checkpoint filenames. In the predictor repository the file is model_Jealousy_&_Envy_best.pth, so the runtime scorer's emotion dictionary is keyed Jealousy_&_Envy — with an ampersand. Everything else in the project — the corpus columns, the ECDF tables, the bucket names, the adapter directories — uses Jealousy_and_Envy. Every evaluation therefore executes scores["emonet"].get("Jealousy_and_Envy", 0.0), misses, and records a hard 0.0. The effect is visible in the data: in the 17-adapter grid (I.12.5) and in the merge-weight sweep (I.12.4), Jealousy_and_Envy reads exactly 0.000 on every single clip in every arm. Two of the 80 evaluation prompts — the English and German Jealousy prompts, so 8 of 320 clips — are pinned at zero in every model's emotion average. Comparisons between models are unaffected (a constant offset); absolute emotion percentiles are biased slightly low, by roughly one fortieth of whatever that emotion would have scored. An earlier draft attributed the zeros to the percentile correction of I.6.7 behaving as designed — the 26 August report still says so in its §12.5 ([TR-0826]). That explanation was wrong; the cause is a key mismatch, and it is a recurrence of the defect that once left this emotion out of the head list entirely: the emotion list originally had 39 entries — Jealousy_&_Envy was simply missing, so its head was never loaded. Because heads are keyed by name and not by index, the other 39 were still correctly labelled. It was a missing emotion, not a misaligned one, and only a count caught it.

I.6.5 VoiceNet: 57 dimensions of how a voice sounds

laion/voicenet-dimension-predictors-commercial is a bank of 114 small MLP heads — 57 dimensions × 2 (a regression head and a classification head) — on the same frozen 768-d VoiceCLAP-commercial embedding. One encode yields all 57 dimensions. Outputs are ordinal levels, 0–6 for most dimensions, 0–4 for background noise and 0–2 for content appropriateness.

How it was trained. Labels are laion/emolia-voicenet-gemini-annotations — Emolia clips each re-scored on a single VoiceNet dimension by Gemini-3.5-flash, combining three annotation rounds (round 3 topped up the most data-starved score levels of the hardest dimensions). Each unique clip is encoded once. For every dimension three class-balancing variants are trained (buckets capped at 3× / 4× / 5× the second-smallest usable bucket), and the shipped default is chosen per dimension on a frozen validation split — lowest MAE for the regression head, highest accuracy for the classification head, chosen independently. Regression uses Huber (δ 1.5), H = 64, dropout 0.33; classification uses cross-entropy, H = 48, dropout 0.33; Adam, 50 epochs, best-val checkpoint. Best-per-dimension results: mean regression MAE 0.744, mean Pearson 0.79, mean within-one-level accuracy 0.870, mean classification accuracy 0.617. Per-dimension quality varies a great deal — head resonance reaches MAE 0.478 with r = 0.88, while cognitive load sits at 1.052.

The 57 VoiceNet dimensions ([TR-0829] §6.5).
GroupCodes
Core affect & dynamics (10)AROU arousal · ARSH arousal shift · VALN valence · VALS valence shift · VOLT volatility · TENS tension · VULN vulnerability · DARC dynamic arc · FOCS focus · COGL cognitive load
Speaker identity (2)AGEV voice age · GEND perceived gender
Prosody & delivery (9)TEMP · VFLX velocity flux · RANG pitch range · REGS register · EMPH emphasis · ATCK attack · SMTH smoothness · CHNK chunking · RESP respiration
Articulation & discourse (3)CLRT clarity · DFLU disfluency · STRU structure
Timbre & voice quality (8)BRGT · WARM · FULL · ROUG · HARM · METL · ESTH esthetics · STNC stance
Resonance placement (7)R_CHST · R_HEAD · R_MASK · R_NASL · R_ORAL · R_THRT · R_MIXD
Speaking styles (15)S_ASMR · S_AUTH · S_CART · S_CASU · S_CONV · S_DRAM · S_FORM · S_MONO · S_NARR · S_NEWS · S_PLAY · S_RANT · S_STRY · S_TECH · S_WHIS
Recording & content (3)RCQL recording quality · BKGN background noise · EXPL content appropriateness
Two polarity traps, and a bug they caused. GEND runs feminine at 0 to masculine at 6, and BKGN runs noisy at 0 to clean at 4 — higher means less noise. Both are counter-intuitive, and both were inverted in an earlier version of the procedural caption generator, so that corpus-wide the prose descriptions of gender and background noise ran backwards while the numeric columns were correct throughout. The fix was verified end to end on the shipped strings rather than on the ladder tables: across 40 sampled voices, zero captions contradict the voice's measured gender class. Anyone who trained on the caption text from that period learned the opposite of the truth for those two attributes.

I.6.6 The burst locator and the burst detector

Bursts are handled by two models in sequence, because “where is it” and “what is it” are different problems. laion/vocalburst-locator does frame-level binary segmentation at 20 ms resolution: a Whisper-small encoder adapted with a rank-8 (α 16) LoRA on q_proj/v_proj and then merged, followed by Linear(768→384) → GELU → Conv1d(384, kernel 7) → Linear(384→1) → sigmoid, giving 1,500 per-frame probabilities over a 30-second window, post-processed into (start, end, confidence) events. Event F1 at IoU 0.5 on 992 real in-the-wild expressive-speech clips: 0.607 for v2 against 0.152 for v1 — and v1 fails in a specific way, reaching usable precision only at threshold 0.80 where recall collapses to 0.21.

Post-processing dominated the weights. Ground-truth bursts have a median duration of about 180 ms, so the v1-era default minimum-duration filter of 0.5 s discards roughly 96 % of real bursts before matching ever happens. On one identical checkpoint, changing only the post-processing moved event F1 from 0.243 to 0.598 — a larger effect than any training change made for v2. Three subsequent attempts to beat 0.607 by training all failed, one of them (1,044,713 edge-case windows that were 100 % positive) dropping F1 to 0.482 because the extractor had thrown away the negatives.

laion/vocal-burst-detector-v2 then classifies each located span: an MLP (Linear(768,256) → BatchNorm → GELU → Dropout(0.3) → Linear(256,83), ≈219 k parameters) on the VoiceCLAP-commercial embedding of the cut, over an 83-class taxonomy — 82 vocalisations plus no_burst at index 82. Argmax accuracy on the official 1,940-clip validation set is 58.1 % overall: 70.4 % on clips with a single pure burst, 39.4 % on composites. Five classes that are hand or body impacts rather than vocalisations — blowing a kiss, finger snaps, hand scratching head, hand slaps, slap face — are folded into no_burst and masked at inference, so the recognised label space is 77 burst classes plus no_burst. A burst is additionally vetoed when P(no_burst) ≥ 0.50. The agent era replaced this detector with a 17-class successor, laion/vocal-burst-detector-x2, because 83 classes at 58.1 % could not be trusted for a reward term; Part II.5 has its numbers.

The pre-agent reward pipeline ran the locator at threshold 0.5, merge gap 0.2 s, minimum duration 0.08 s, padded each span by 0.05 s before classification, and dropped anything under 0.10 s (LOC_THR, LOC_GAP, LOC_MINDUR = 0.5, 0.2, 0.08; BURST_PAD, MIN_BURST, NOBURST_GATE = 0.05, 0.10, 0.50). The card's own recommendation is slightly tighter — merge gap 0.10 and minimum duration 0.10 — and the difference has not been ablated.

I.6.7 Turning raw scores into a scale you can ask for

Everything above produces numbers on incomparable scales. Measured on the project corpus, the Interest head has a median of 2.082 and never outputs zero, while the Infatuation head has a median of −0.017 and outputs zero on 87.7 % of clips. A reward built on raw values would have been an Interest detector. The answer is the corpus ECDF: every raw score is mapped to its percentile within the full 3,147,802-row corpus, per head. Four intensity bands are defined on that percentile scale, and the same four cutoffs are used by the prompt builder, by the reward function and by the evaluation harness, so what the prompt asks for is exactly what gets measured.

Intensity bands on the percentile scale ([TR-0829] §6.7).
BandPercentileAdverbs used in the prompt
faint0.40 – 0.70barely, faintly, only slightly, just a little
moderate0.70 – 0.90clearly, plainly, noticeably, unmistakably
intense0.90 – 0.98strongly, intensely, very, deeply
extreme0.98 – 1.00overwhelmingly, extremely, utterly, completely

The zero-point defect, measured 25 August 2026. A plain ECDF percentile turned out to be still not comparable across the 40 heads, for a reason that is obvious in retrospect. A clip with a raw score of exactly zero — no measurable emotion at all — lands at a percentile that depends entirely on how often that head outputs zero:

Percentile of a raw score of zero, per head ([TR-0829] §6.7).
HeadPercentile of raw 0HeadPercentile of raw 0
Awe0.822Contempt0.510
Infatuation0.766Amusement0.487
Pain0.703Jealousy and Envy0.004
Distress0.628Affection0.002

Asking for band 0.90–0.98 was therefore a different task for every emotion: on Awe the model starts 0.08 below the target having produced nothing, while on Affection it starts at zero. Averaged over 40 heads, the reported figure was measuring where each head's zero point happens to fall, not how well the emotion was performed. The fix is a per-head affine rescale of the tail above the zero point, p′ = max(0, (p − p₀) / (1 − p₀)) where p₀ is the percentile of a raw score of zero. It was verified to map raw 0 to exactly 0 on all 40 heads while remaining monotone within each head — so every comparison between two clips of the same emotion, which is what DPO pairs and GRPO group rankings consist of, is unchanged. Only absolute targets and cross-head averages move.

Consequence for reading older numbers. Evaluations run before this correction and after it are on two different scales. The same audio, scored both ways, differs by exactly this much: emotion percentile 0.5095 → 0.3494, composite reward 0.4900 → 0.4584; word error rate, blend, genuineness, burst metrics and every duration field are byte-identical between the two files. Sorting all 22 stored evaluations by timestamp gives a clean cut at 2026-08-25 12:00 UTC: every run before it reports an emotion percentile ≥ 0.5095, every run after ≤ 0.3572. Six runs are on the old scale and sixteen on the new one, and the two sets must not share a ranking column. The superseded SFT-3 directory is kept under an explicit _UNCORRECTED_ECDF name so it cannot be quoted by accident. No cross-model number from before that point appears in this report.

I.6.8 How far these instruments can be trusted

Two mitigations have actually worked and are recommended rather than merely stated: always run a control arm that isolates the claimed cause, and verify identity from an independent namespace — never infer “this file is that record” from a shared key. The agent-era Ears add a third: a human-proxy target (the burst detector's own AUC against a held-out human-labelled set) that the reward terms are correlated against, so that a term's usefulness is a measured number rather than an assumption (Part II.6).

I.7 The prompting model, the manual, and the v2 casting agent — the actor's first draft

The shipped v2 model is documented in a public manual: *MOSS Voice-Acting — Manual & Studies Hub*. It is 25 chapters: a hub, an emotion-conditioning manual with one page per emotion, twelve recipe chapters, and an eight-page tree recording an autonomous casting agent's four generations of runs. It is unusually forthright — a large fraction of it is a record of things that did not work, with the measurements attached. Source: [TR-0829] §7–§8.

I.7.1 The prompt surface the model ships with

Two fields, exactly:

instruction = GENERAL: <the standing description of the voice and the situation> / SCRIPT: (delivery cue) "the spoken line, (burst) with tags inline, [pause] and pauses in squares" — and text = <exactly the same spoken words, plain — no cues, no brackets>.

The words appear in both fields — plain in text, cue-annotated inside SCRIPT: — and that redundancy is deliberate and matches the training data. The manual's notation rules, verbatim:

A worked example, copied verbatim from the manual's casting-agent prompt archive (adapters Scream @ 0.9, Anger @ 0.4, Distress @ 0.35): GENERAL: Start wounded and breathless, then rebuild into nearly screaming but controlled rage. The final words should land cold and absolute, with the voice hardened by pain. Place the vocal burst inline and continue speaking immediately afterward. SCRIPT: (a broken breath turning into a controlled eruption) "Don't you dare touch me." (Scream) "You don't get to cry now. You don't get to make this about your guilt. Pack your things, and get out before I say something neither of us can take back."

There are no duration tags in the shipped prompt format. Length is controlled only indirectly, through the tokens field (target frames, ≈ words × 6) and max_new_frames. The manual gives a specific warning for non-verbal parts: tokens defaults to the number of words, and a scream's text has almost none, so the model is told to make ≈5 tokens and stops after 0.3 s — for non-verbal material you must set about 12.5 tokens per intended second by hand. This is exactly the gap the timed-script round (I.8) closes.

I.7.2 How intensity is requested, and why the manual says the prose does the work

Intensity in the shipped format is not a tag and not a number. It is requested three ways, and the manual is emphatic about their ranking:

One finding here anticipates the whole of the GRPO story (I.11): a maximise objective cannot produce a moderate emotion. The manual states it directly — “a maximize objective always drifts to intense; a believable moderate/subtle emotion scores low against ‘maximize’ and dies out of the population.” The fix used there is a target-band fitness function, which is precisely the shape the timed-script reward uses. The manual also independently reports the cross-head incomparability of I.6.7, in its own units: “Anger/Sadness/Fear reach ≈2.4–3.7 on the detector; Amusement tops out ≈1.7 even when strong — do not compare raw emotion scores across emotions; compare each against its own free take.”

I.7.3 Best-of-N is not optional, and it is nearly free

Because emotion strength, burst placement and prosody are seed-dependent, the recommended way to use this model is to generate many candidates and rank them. The economics are in I.4.12; the manual adds: for a mid-sentence burst specifically, a burst appears in roughly 50 % of takes, so five candidates give about a 97 % chance that at least one has it; rank whole assemblies, not parts (10.2 % agreement between the best assembly and the part-wise best); at scale, transcribe for ranking with a fast ASR model and re-transcribe only the top three with the accurate one. Best-of-N is the one element of the v2-era practice that survives unchanged into the agent (Part II.6).

I.7.4 The v2 casting agent and its reliability limits, stated plainly

The manual documents an agent built on the v2 model: it writes a ≈30-second performance as 2–4 parts, generates each best-of-16, assembles the best combination across parts, ranks with the local scorer stack, and has a supervisor listen afterwards and give feedback for the next round. Seven generations of it were run — four complete, three aborted — over nine challenges, roughly 2,000 generations. A separate autonomous LoRA-plus-prompt search agent (LAION-AI/voice-acting-search-agent) ran an overnight programme of 163 agent generations and ≈10 k scored takes on a single A100; its notable finding is about the brain, not the speech model — a Gemma-4-31B QAT policy won decisively on single-dimension missions using only 22 tool calls, while a Gemini supervisor was an effective active director (tracking the objective at r = 0.72) and a local audio judge was usable only as a lenient score-only gate, because active steering by it was harmful. That finding — a capable local Mind, a supervisor that directs, an audio judge that only gates — is the design of the Part II agent in embryo.

[REF-V1] Part III preserves more of the search agent than any later version: the policy comparison ran Gemma-4-31B against Gemma-4-12B and a Gemini policy; the winning runs used 22 tool calls against a budget of 40; the canary that caught dose overdrive was a fixed neutral line regenerated after every merge; and the “gotchas” list of that report (adapter eviction, the stale-scaling snapshot, the active_adapter attribute, capitalised tags, empty decodes) is the same list that I.4.13 and I.7.5 carry. The autonomous runs cost about 12 GPU-hours in total on one A100.

That agent can merge the LoRAs, and the audio it produces is more emotional than what the SFT/DPO line of I.9–I.10 produces. It is also, on the project's own measurements, not reliable on character consistency.

I.7.5 The voice drifts

And the thresholds themselves are wrong when set high. The casting playbook originally used similarity gates of 0.82 / 0.75 / 0.68. The voice-consistency chapter overrides them: regenerate below 0.58, repair with voice conversion below 0.45, reject below 0.40 — and calls that correction “the single most important thing on this page”. The evidence: among 170 human/model-confirmed same-speaker pairs the median ECAPA similarity is 0.632, 55 % fall below 0.68 and 84 % below 0.82, while the highest score reached by a genuinely different speaker was 0.816. Four embedding models were tried (ECAPA AUC 0.847, VoiceNet 0.814, WavLM timbre 0.794, VoiceCLAP 0.770) and none gives a clean threshold. The conclusion: ranking on identity is fine, gating on it is not — and never raise the floor “to be safe”, because the failure that creates (flat, over-clean delivery) is less visible than a voice change but affects every take rather than a few. The agent's speaker-similarity ear (Part II.5) inherits the 0.40 floor and the rank-don't-gate rule.

I.7.6 Under strong emotion, output goes strange

I.7.7 And the local reward does not predict a listener

“Short answer: on spoken scenes, a little — ρ ≈ 0.21, about 4 % of the variance. On non-verbal scenes, nothing, and the sign is wrong. There is currently no configuration in which the local stack can replace the listener.” Two of the reward's own components point the wrong way against listener preference — genuineness at −0.143 and blend at −0.148; only emotion peak (+0.179) is useful. An optimally weighted linear combination of everything the local stack measures explains under 3 % of a listener's within-round preference. The manual's conclusion is that re-weighting cannot fix this, because the information is not in the feature set. The n behind ρ ≈ 0.21 is the listener's within-round rankings across the casting agent's generations; the source does not state it as a single number, and neither does this report.

Two further findings sharpen it. The casting score in use was measured to correlate +0.91 to +0.98 with its own (1 − WER) multiplier and +0.00 to +0.04 with genuineness — it ranks assemblies by intelligibility and speaker continuity and is close to blind to everything else it claims to measure. And three of its eight terms were dead: the emotion-peak term was saturated at its cap in 92.4 % of assemblies, and the arc term sat at exactly 1.0 in 84 % and took three distinct values ever. Compare the agent-era Best-of-N reward of Part II, where the same question — which term carries the signal — was answered by correlating each term against a human-proxy target on 15,517 candidates: genuineness −0.0085, blend −0.0839, CLAP +0.1482, 1−WER +0.0388, burst +0.3053. The two reward generations agree on the sign of genuineness and blend.

What this adds up to, for anyone who wants to use the v2 agent. It produces more emotional performances than the supervised line does. It is not production-reliable on character consistency: it drifts, and it drifts hardest exactly where the performance is strongest and on non-verbal material. Under strong conditioning it can garble, truncate, or quietly stop acting. The practical way to use it is the way the manual recommends — large candidate groups (32 is the economic sweet spot), reward-ranking with hard exclusion rather than penalty on identity, anchor-plus-tail continuation to hold a voice, byte-identical voice descriptions across parts, and a human in the loop for the final pick, because the local reward will reliably reject the worst takes and will not reliably find the best. Every one of these rules is a design input of the Part II agent.

I.8 The timed script

Listening to the round-1 supervised-plus-preference model, two faults dominated: uncontrolled vocal bursts — a sigh that should last 0.5 s runs 5 s — and hallucinated content — text appears that was never prompted. Both are the same missing signal: the prompt never stated how long anything may take. The timed-script round exists to supply it. All of it runs on JUPITER: GH200, 4 GPUs per node, 102 GB each, SLURM account reformo, partition booster. Source: [TR-0829] §9.

I.8.1 The format

The instruction block is unchanged in shape from the shipped format — the same Reference(s) / Instruction / Tokens / Quality / Sound Event / Ambient Sound / Language / Text template — but the SCRIPT: block is now a rendered timed script: [0.9 seconds pause] [3.4 seconds duration] Erster Satz. [0.6 seconds pause] (contented sigh, 0.4 seconds) [4.1 seconds duration] Zweiter Satz. There are exactly four kinds of tag, told apart by their brackets:

The four tags of the timed script ([TR-0829] §9.1).
TagMeansRule
[3.9 seconds duration]The next sentence must take this longSquare brackets. One per speech segment, standing before the text it applies to.
[0.8 seconds pause]Silence of this lengthSquare brackets. Every gap of 0.2 s or more, including before the first word and after the last.
(contented sigh, 0.2 seconds)A non-speech vocalisation of this lengthRound brackets with a duration. Label first, then the seconds.
(clearly amused, warm and open, unguarded)How to perform the next sentenceRound brackets without a duration. Stands before the duration tag.

The disambiguation rule in one line: square bracket = a number of seconds; round bracket with a number = a vocal burst; round bracket without a number = a delivery direction. That is the only thing separating a burst from a direction, which is why directions are always emitted with no number at all. There is one soft spot: in the untimed 30 % of samples nothing carries a number, so a burst renders as a bare (chuckle) — the same shape as a direction. What still separates them there is length: a burst label is one or two words, a direction is a phrase.

Segmentation rules, all measured against the corpus: split at sentence ends; split again at a vocal burst, at the nearest word gap; split any remaining segment over 12 s at its largest internal gap provided both halves stay above 2 s. A duration is measured from the first word onset to the last word offset of that segment, so two sentences of 12 s and 8 s produce [12.0 seconds duration] and [8.0 seconds duration], never a single 20-second tag. Gaps under 0.20 s are folded into the neighbouring speech rather than printed, so the printed numbers still sum to the clip length. Bursts under 0.05 s are dropped, and a burst that overlaps speech prints only the part that does not overlap.

Traps in building the timed script ([TR-0829] §9.1).
TrapMeasured cost
Rebuilding the transcript by joining word tokens — the forced aligner drops numerals (“154 Euro” → “Euro”)would lose text on 9.9 % of real-speech rows
Taking bursts from the parenthetical cues in the source text instead of from detections39.4 % of those cues were never confirmed by the detector
Printing a burst's full length when it overlaps speech49.0 % of detected bursts overlap a word
Printing sub-threshold gapswould break the “numbers sum to the clip length” invariant

Verification, done the right way. The emitted string is re-parsed by a separate function with no access to the source data, and the recovered numbers compared against the true clip length. Residual median 0.010 s on 8,614 real-speech rows and 0.020 s on 14,359 voice-profile rows; within rounding tolerance on 99.9 % / 98.7 %; the transcript comes back byte-identical on 100 % of rows. This is the pattern I.6.8 recommends: make at least one leg of a check independent of the pipeline that produced the thing being checked.

I.8.2 The inline delivery direction

The round-2 format dropped the parenthesised delivery direction that round 1 had carried — by accident, and the accident became the round's central negative result (I.9.2). Round 3 puts it back, inside the script, before each duration tag. A direction is composed of five pieces:

Pieces of a delivery direction ([TR-0829] §9.2).
PieceSourceExample
Intensity adverbThe requested percentile bandoverwhelmingly
Emotion nameThe clip's own label — from its identifier or from the caption's *reads as …* clause. Never guessed.amused
Optional second emotionSame source, as a shadingwith a hint of bitterness
Performance anchorHarvested from the corpus, not written. 187,528 voice-profile rows were scanned and the leading direction of each grouped by emotion.letting it out / not hiding it, warm and open, emotionally reachable, unguarded
Manner wordsThe VoiceNet ladders the row's own caption already statesslightly bright, relaxed

Only the first segment gets the full direction; later segments get a short reminder ((still clearly amused)), and segments under 1.6 s often get none — repeating a thirty-word note in front of a 0.6-second line buries the line it is meant to shape. When a row has no emotion label at all — a VoiceNet-axis row, for instance — the opening segment gets the corpus's own most common direction instead.

A limitation kept deliberately rather than silently repaired. All 40 emotions resolve to exactly two harvested phrasings each — one expressed, one suppressed — drawn from only 10 distinct family-level texts. Some pairings read oddly: intoxication, pain and shame all inherit the fear family, so an intoxication direction can say “the fear plain in the voice”. This is not a bug in the generator — it is exactly what the round-1 corpus contains (425 of the intoxication rows use that phrasing verbatim), and the model was already trained on those pairings. It is kept for comparability and flagged as the first thing to revisit if intensity control still lags for those emotions.

Two bugs in this generator were found and fixed on 25 August, seventeen minutes before the round-3 run launched — one that would have made every row render the same intensity word, and one that left the untimed 30 % of samples with no direction at all. Both would have made the round meaningless without changing any training-time number. I.9.4 gives them in full.

I.8.3 What varies between training samples

Nothing is fixed; each sample draws from a seeded generator, so the model sees the whole range and never requires any single element to be present. Measured end to end on 62,672 real corpus rows through the training loader:

Sampling probabilities of the training loader, realised on 62,672 rows ([TR-0829] §9.3).
SettingProbabilityWhy
Timed script0.70The other 30 % render the same script with no numbers, so the model stays usable when a caller supplies plain text. (Realised: 71.1 %.)
At least one direction present0.85The rest train direction-free, so a plain prompt stays in distribution. (Realised: 79.5–80.2 %.)
Direction dropped per segment0.15So the model never learns that a duration tag must be preceded by one.
Bursts dropped0.10Each burst's span folds into the surrounding pause so the numbers still add up.
Reference audio used0.50Of rows that have one; the rest fall back to a speaker name or to nothing.

The Instruction field itself takes one of four shapes: gen_both (a GENERAL: line and a SCRIPT: block) at 45 % for voice profiles and 75 % for real speech; caption_tpl (a natural-language instruction instead of the GENERAL: line) at 30 % for voice profiles only; gen_general at 15 %; gen_script at 10 %. These four shapes are the prompt forms the agent of Part II tests as a Tool (II.3.4). Realised band mix over those 62,672 rows: 58.8 % extreme, 23.3 % intense, 10.1 % moderate, 7.8 % faint. That skew is not an accident — it comes from training on the “extreme subset” — and I.9.11 argues it is a prime suspect for the failure that follows.

I.8.4 The corpus

Corpus v2 and the extreme subset ([TR-0829] §9.4).
ArtefactSizeComposition
Corpus v23,147,802 rows · 128 shards · 17 GB1,200,530 voice-profile + 1,947,272 real-speech rows, with word timelines, burst spans, packed 40-emotion and 57-VoiceNet float32 blobs, blend, genuineness and emotion strength. The burst sidecar matched 1,200,530 of 1,200,530 voice-profile rows.
The “extreme” subset398,282 rows (12.65 %)Core 331,902 + a blend arm of 66,380. Voice profiles 179,617 / real speech 218,665; German 176,705 / English 221,577.

The extreme subset is worth one paragraph because of how its definition had to change. The original rule was “the top 10 % of each of the 40 emotions plus the top and bottom 10 % of each VoiceNet dimension”. Measured, that rule selects essentially the whole corpus: the emotion top-10 % union alone is 89.01 %, the VoiceNet top-plus-bottom is 99.37 %, and the combination is 99.84 % of the full corpus. With 154 overlapping axes, a per-axis decile is not a filter. It was re-implemented as a fixed quota per axis; at 6,000 rows per emotion each pick is the top 0.19 % of its axis. Band composition of the subset — 59.5 % extreme, 23.8 % intense, 9.9 % moderate, 6.8 % faint — against the full corpus's 45.6 / 32.4 / 12.8 / 9.3 %.

I.8.5 The evaluation harness

Every model comparison in I.9–I.13 uses the same harness, so the rows are comparable to each other and to nothing else:

The evaluation deliberately uses a stricter yardstick than training. Training needs a gradient, so its band score ramps; evaluation asks a different question — did the clip land in the band that was requested — and answers it with a plateau that pays nothing outside the band. Keeping the two apart is what makes numbers published before and after a reward change comparable. Burst realisation is an F1 over prompted-versus-detected bursts in which each match is weighted by how well its duration matches — a matched class scores 1.0 and a mismatched one 0.35, times exp(−|Δduration| / 0.5), with a 1.5 s tolerance on position. Recall alone would reward emitting a burst everywhere and precision alone would reward emitting none; the original listening complaint was bursts that are too long, so the duration term is what carries the signal, and a matched pair with the wrong length scores near zero. Note that 320 clips per model means the standard error on a between-model reward difference of the size seen in I.12.3 (≈0.01) is of the same order as the difference; I.12.3 says so where it bites.

I.9 Supervised rounds: timing solved, emotion lost, format regained

Three supervised rounds have run. Round 1 established the line: all 4.13 B parameters, 3 epochs over the 3,147,802-row corpus, global batch 4,096, 2,232 steps, 64 nodes × 4 GH200 in 1 h 53 min, held-out loss improving monotonically 4.7076 → 4.6314 across all twelve evaluations. Rounds 2 and 3 are much shorter runs on top, whose purpose is the prompt format rather than more data: each is 2 epochs over the 398,282-row extreme subset, 712 steps on 32 nodes. The sequence is the most instructive thing in Part I, because round 2 was a success and a failure in the same run, the failure was caused by a line that was deleted by accident, and round 3 exists only to put it back. Source: [TR-0829] §10; the round-1 configuration and the loss series are in [P§3–§5].

The three supervised rounds ([TR-0829] §10).
RunBaseStepsWallNodesLRValidation lossTensors changed vs base
SFT round 1voice-acting v22,2321 h 53 min644.7076 → 4.6314 (12 evaluations, monotone)
SFT round 2SFT-1 + DPO-1 (full)71229 min325e-64.6278 → 4.6268375 / 425, max relative 2.098e−02
SFT round 3SFT-2 export71231 min327e-64.6296 → 4.6264378 / 425, max relative 2.740e−02

A comparability note that the protocol records rather than hides: round 2 ran at learning rate 5e-6 and round 3 at 7e-6, so their validation losses are not comparable to each other; nor are they comparable across the round-2 → round-3 prompt change, which altered the objective. I.9.7 argues that in this regime the validation loss is barely comparable to itself. The round-1 learning rate is not stated in [TR-0829]; it is not reproduced here.

I.9.1 Round 2 — timing is solved, and emotion control disappears

320 generated clips per model, band intense requested:

Round-2 models, 320 clips each, pre-correction percentile scale ([TR-0829] §10.1).
ModelRewardEmotion pctQualityBurstWERDuration err. (median)Within 0.5 s
SFT-20.48800.49050.91520.37510.11230.100 s99.4 %
+ DPO-2, step 2160.47660.49870.90130.36140.12050.100 s98.8 %
+ DPO-2, step 8640.49440.50800.92820.36020.11850.080 s100.0 %

(These three rows are on the pre-correction percentile scale of I.6.7 and are therefore comparable to each other but not to the tables from I.10 onward.) Timing is solved. Median duration error 0.08 s — one audio frame — and every one of 320 clips within half a second of the length its script's own numbers add up to. That is the capability the timed script was built to deliver, it was delivered in a single 29-minute run, and nothing since has degraded it. Emotion is not. Asked for percentile 0.90–0.98, every round-2 model lands at 0.49–0.51 — the corpus median, which is what a model produces when it is not being told anything at all. Three models with three different amounts of preference tuning all sitting on the median is not a training-strength problem; it is the signature of a missing input.

I.9.2 The diagnosis: what round 2 deleted

The cause was found by reading the diff between the round-1 and round-2 prompt renderers. Round 2 had dropped the parenthesised delivery direction entirely. Round 1 had put a direction in front of the spoken text on 55 % of voice-profile samples and 85 % of real-speech samples — (intensely amused: letting it out / not hiding it, warm and open, unguarded) and the like. Round 2 replaced that slot with the timed script, and the direction did not survive the replacement. Nobody removed it deliberately; the new renderer simply did not carry it forward. What was left was description, not instruction. The emotion still appeared, in the GENERAL: line, as a clause saying what the source recording sounded like — “reads as contentment”. That tells the model what kind of clip this is. It does not tell it what to do. The model was trained for two epochs on a format in which no one ever asked it to perform anything, and it learned exactly that: produce the corpus median.

The control that proves it, and that invalidated two RL runs. If the direction is the missing input, then feeding a round-2 model a prompt that does carry one — a format it has never seen — should not merely fail to help. It should break the model, because the direction sits inside the script, in the same bracket space as the timing tags the model has learned to obey. It does.
WER with inline directions, 320 clips per row ([TR-0829] §10.2).
EvaluationModelWER with inline directions
eval_sft2_dirplain SFT-2, no adapter at all0.5122
eval_dpo2_dirSFT-2 + DPO-2, step 8640.4828
eval_grpo1_dirSFT-2 + GRPO run 10.4798
eval_grpo3_dirSFT-2 + run-1 adapter + GRPO run 30.4836
eval_sft2_dirv3SFT-2, corrected-scale rerun0.4473
eval_sft3_dirSFT-30.0987

Word error rate 0.48–0.51 for every round-2 model including the bare supervised checkpoint, against 0.11 for the same checkpoint on direction-free prompts. Because the bare model fails as badly as the tuned ones, this is not a property of any adapter. It is the format. Verbatim from the evaluation log — clip k550_age5_bg1__E__Relief__C__en.c036, on a script asking for “November was the only month the headline appeared: The last sporting taboo”: "wer": 4.545454545454546, "hyp": " Ah, ah, ah, ah, ah, ah, ah, ah, ah, ah, …". This matters well beyond the supervised round. Two GRPO runs had been spending hours optimising against exactly that collapse (I.11.1), which is how a reward computed on out-of-distribution prompts ends up measuring the prompt instead of the policy. It is also the earliest appearance of the “ah, ah, ah” babble that the agent's burst-adapter bug reproduced in September (Part II.7).

I.9.3 What round 3 changed: the direction goes back, inside the script

Round 3's entire purpose is to put the delivery direction back and train on it, in the position the timed script leaves for it: in round brackets, without a number, standing immediately before the duration tag of the segment it applies to. Nothing else about the format changed. Two consequences bound what round 3 could possibly have achieved: the direction carries the intensity word, and nothing else does — if the direction is absent, or if the band is computed wrongly, the model has no intensity signal in the prompt at all; and the direction is emitted from the row's own measured percentile, using the same cutoffs as the reward and the evaluation harness, so what the prompt asks for is exactly what gets measured.

I.9.4 Two prompt bugs, found and fixed seventeen minutes before the run started

Both were found by rendering real corpus rows through the training loader and reading the output, which is a cheap check the project now runs before every prompt-format change. Both would have made the round meaningless in a way that no training-time metric would have revealed. The backup files the fixes left behind are timestamped 09:54:21 and 09:55:19 on 25 August; the SFT-3 job started at 10:12:48 the same morning.

Bug 1 — the intensity band was read from a column the training corpus does not have. The band came from an in_extreme flag. That column exists in the preference corpus, where it records membership of the extreme subset. It does not exist in corpus_x, the supervised corpus round 3 trains on. The pre-fix line, preserved verbatim in prompt_lib2.py.bak-20260825-band: band = "intense" if row.get("in_extreme") else "moderate". On every supervised row row.get("in_extreme") returns None, so band is "moderate", and the adverb table maps moderate to “clearly”. Every single row of the round would have rendered “clearly” — a model asked to be clearly angry, two epochs running, and never once asked to be intensely or overwhelmingly anything. The fix reads the row's own emo_strength percentile with the reward's cutoffs; 83 % of rows then carry intense or extreme wording, matching what the audio actually sounds like.
Bug 2 — the untimed 30 % of samples carried no direction at all. Directions attach to duration tags. 30 % of samples are rendered without numbers on purpose, so that the model stays usable when a caller supplies plain text — and an untimed script has no duration tags to attach to. The substitution therefore found nothing and that entire arm trained direction-free. The failure mode is at inference, not in training: a caller who supplies a delivery direction but no durations — the most natural way to use the model, and the round-1 shape — would have been straight out of distribution, which I.9.2 has just shown costs a factor of five in word error rate. The fix places one direction at the head of the line for untimed renderings, which is exactly the round-1 shape. Direction coverage rose from 57.7 % to 79.5 % of samples. The 79.5 % is independently confirmed by the loader measurement of I.8.3 over 62,672 rows; the 57.7 % is not reproduced by any artefact still on disk.

The prompt-format hash moved from 64d956b8bb97f0d9 to f90a7b16b4cea093 — the hash is written into every checkpoint's state file, so a checkpoint can always be matched to the exact prompt renderer that produced its training data. That is the mechanism by which the round-2 accident was findable at all.

I.9.5 The round-3 training configuration

SFT-3 configuration ([TR-0829] §10.5).
SettingValueNote
Basethe SFT-2 exportThe full-parameter round-2 checkpoint, not a merge
Trainedall 4.13 B parametersNot a LoRA. Rounds 2 and 3 are full fine-tunes.
Corpuscorpus_x — the 398,282-row extreme subsetMaterialised as 128 parquet shards, 401,345 rows, of which 3,063 are the validation holdout carried unchanged from round 1
Epochs2
Steps712356 optimizer updates per epoch
Global batch1,024= 4 micro × 2 accumulation × 128 ranks
World128 ranks — 32 nodes × 4 GH200
Learning rate7e-6Round 2 used 5e-6
Warmup178 stepsExactly half of epoch 1
Wall clock30 min 46 sSLURM start 10:12:48, end 10:43:34. The training window inside the log is 25 min 45 s; the rest is model load and export.
Prompt hashf90a7b16b4cea093Recorded in the run header and in every checkpoint

One correction to a figure that has circulated internally: round 3 was not trained on 247,049 rows. That number is the emotion-LoRA corpus of I.12.2, built the same evening by a different selection script, and the coincidence of dates is the likely source of the conflation. Round 3 trained on corpus_x. This checkpoint — SFT3 — is the Instrument of Part II.

I.9.6 The validation series, in full

SFT-3 held-out loss by step ([TR-0829] §10.6).
Step89178267356445534623712
Held-out loss4.62964.63004.63014.62944.62894.62764.62674.6264

I.9.7 Why that series is nearly useless, and what it is good for anyway

The whole run moves the held-out loss by 0.0032 nats, and it moves the wrong way for the first three evaluations. Against a model whose negative log-likelihood is about 54 nats per position — each supervised position carries 13 channels (I.3.1) — that is a relative change of six parts in a hundred thousand. The reason is structural and it applies to every short fine-tune in this project. The loss is dominated by predicting audio codes, and a change to the prompt format changes almost nothing about how hard the audio is to predict. The whole intervention — putting a parenthesised phrase back in front of each segment — is a few dozen text tokens against several hundred audio frames per example. A metric that averages over both cannot see it. So the honest reading of the validation series is: it confirms the run did not break, and it says nothing whatever about whether the run achieved its purpose. It is a smoke test, not a result. Every claim in I.9.9 comes instead from generating 320 clips and scoring them.

I.9.8 What actually moved in the weights

The export step compares every tensor against the base checkpoint, which is a cheap check that a run did what it claimed and touched what it should have. Round 3 moved more of the model than round 2 did, both in count and in magnitude — 378 tensors against 375, and a peak relative change 31 % larger (2.740e−02 vs 2.098e−02) — which is consistent with the higher learning rate and with the fact that round 3 introduced a new token pattern into every sample rather than merely reshuffling an existing one. Roughly 47 tensors were untouched by both runs, which is the expected shape for a fine-tune that leaves embeddings and certain normalisation parameters alone. This is a sanity check, not a result; it is recorded because a later run did report a suspicious tensor count, and it was a bug (III.2.9).

I.9.9 Round 3 — training the directions in fixes the collapse

Both models on the same 320 prompts, all carrying inline directions:

SFT-2 vs SFT-3 with inline directions, n = 320 clips each ([TR-0829] §10.9).
MetricSFT-2SFT-3
Reward0.39390.4900
WER0.44730.0987
WER, English / German0.574 / 0.3200.070 / 0.127
Duration error, median0.100 s0.080 s
Within 0.5 s92.8 %100 %
Burst hit rate0.5160.666
Burst realisation0.2240.356
Blend / genuineness (raw)4.737 / 3.3674.646 / 3.304
Emotion percentile (uncorrected scale)0.54960.5095

The collapse is gone. Word error rate falls from 0.447 to 0.099 — a factor of 4.5 — and it falls on both languages, from 0.574 to 0.070 on English and from 0.320 to 0.127 on German. Timing improves at the same time (median duration error 0.100 → 0.080 s, and 100 % of clips inside half a second against 92.8 %), and so does burst realisation (0.224 → 0.356) and burst hit rate (0.516 → 0.666). One 31-minute run on 32 nodes bought all of that. It is worth being precise about what was and was not demonstrated. Round 3 did not teach the model a new capability. It taught the model a format — and the size of the effect is a measure of how completely a transformer will fail on a prompt shape it has not seen, even when every individual element of that shape is familiar. The directions round 3 trains on are the same directions round 1 trained on. Two epochs of not seeing them was enough to make them poison.

Emotion control did not improve. SFT-2's higher emotion number (0.5496 against 0.5095) is not usable as a comparison: it was producing broken audio at word error rate 0.45, and broken audio scores on emotion heads — a screaming-noise clip reads as high arousal. Rescored on the zero-point-corrected percentile scale of I.6.7, with no regeneration, the picture is worse and more honest:

Rescored on the corrected scale ([TR-0829] §10.9).
ModelEmotion pct, raw…correctedClips in the requested band, raw…corrected
SFT-30.50950.3499.4 %4.4 %
SFT-20.54960.40215.9 %10.6 %

4.4 % of clips land in the band that was requested. That number, not the composite reward, is the one that states the problem: asked 320 times for a specific intensity, the model delivers it fourteen times.

I.9.10 A check that had to be done before considering a retrain

If the zero-point defect had also corrupted the training labels — the emo_strength percentile from which each direction's intensity adverb is derived — the whole round would have needed redoing. It did not. If the labels were a floor artefact, heads with a high zero point would win almost always. They do not: the winning head has a real signal (raw > 0.01) on 99.9 % of rows, including 99.9 % of the rows labelled at percentile ≥ 0.90, and the 40 heads are evenly represented at 3–4 % each. Repeating the supervised round on corrected labels would produce the same labels.

I.9.11 The standing hypothesis for why intensity does not move

The training subset is 59.5 % extreme and 6.8 % faint. The model sees “overwhelmingly” almost always and “faintly” almost never — so there is little contrast from which to learn that the word controls anything. This is the same failure as bug 1 of I.9.4, one order of magnitude weaker: bug 1 would have given the model one intensity word on every row; the corpus as shipped gives it four words in a 59/24/10/7 mix. The full corpus holds roughly 290,000 faint and 400,000 moderate rows, so a band-balanced selection of the same size is feasible. This remains untested as of 9 September and is the most obvious unexplored lever in the project (IV.2).

I.10 Preference tuning: seven runs, four corpora, and one metric that lies

Preference tuning is the objective this project has spent the most compute on, and it is the one that produced the DPO adapter in the agent's shipped stack (weight 1.0, Part II.3). It has also produced a completely consistent negative result on the capability it was aimed at. This section covers all of it: how the objective is configured, every run that was launched, how each of the four preference corpora was constructed and what each was meant to fix, which checkpoint is best and why, and what the training telemetry does and does not say. Source: [TR-0829] §11; corpus builds in [P§6–§9].

I.10.1 How DPO is set up here

DPO trains on pairs: a chosen audio sequence and a rejected one for the same prompt, pushing up the log-probability of the chosen relative to the rejected. Two configuration choices carry most of the weight, and both were arrived at by getting them wrong first.

DPO configuration ([TR-0829] §11.1).
SettingValueWhy
Sequence scorelength-normalised log-probability85.6 % of the raw preference pairs differ from the chosen sequence mainly in length. An unnormalised summed log-probability reached 1.000 preference accuracy at step 102 of 1,194 by counting tokens — a perfect score for a model that has learned nothing.
β30The margin temperature.
Chosen-NLL anchor0.013An absolute term pinning the chosen sequence's likelihood. The model's negative log-likelihood is about 54 nats per position, because each supervised position carries 13 channels (I.3.1). A first attempt at weight 0.25 put the anchor term at ≈13.6 against a DPO term of 0.69 — the anchor simply won and the preference run became a supervised pass.
AdapterLoRA rank 64, α 128, dropout 0.05Targets the 23 modules of I.4 including all twelve audio_lm_heads, which is why it ships unmerged (III.2.2).
Learning rate1e-6
Scale8 nodes / 32 ranks, 256 preference pairs per optimizer stepConstant across every run on SFT-3, which is what makes step numbers comparable between runs: equal steps mean equal numbers of pairs seen.

I.10.2 Every preference run

Seven runs exist on disk. Two of them are not results — one trained on the wrong base and is kept deliberately (III.2.9), one was superseded before it was evaluated — and they are listed anyway, because a table of five successes is a different document from a table of seven runs.

All seven preference runs ([TR-0829] §11.2).
RunBaseCorpusStepsCheckpoints keptFinal rew_chosenOutcome
dpo2SFT-2v1864 / 864419–864−0.7797The round-2 preference model. Best round-2 checkpoint at step 864.
dpo_on_sft2_WRONGBASESFT-2v1184 / 18446–184−0.6348Ran on the wrong base — 32 nodes, global batch 1,024. Kept as a result (III.2.9).
dpo3SFT-3v1864 / 864419–864−1.3769Discarded; superseded by dpo_sft3.
dpo_sft3SFT-3v1864 / 864420–864−0.8162The v1 result. Reward 0.4668.
dpo_sft3_v2SFT-3v24,772 / 7,1353566, 3770, 3976, 4180, 4385, 4590−2.3037Wall clock. Best checkpoint was step 3216 — and its weights were deleted by checkpoint rotation (III.2.7).
dpo_sft3_cfgSFT-3CFG5,006 / 8,9533963, 4176, 4389, 4476, 4694, 4912−1.1211Six-hour timeout. Step 4912 was the best model in the project until 28 August.
dpo_sft3_p2SFT-3contrastive families5,126 / 10,3313906, 4132, 4355, 4579, 4800, 5022−1.0015Step 5022 is the best model this line produced; its adapter is the DPO entry of the agent's stack.

None of the three long runs finished an epoch. All three hit a six-hour wall clock somewhere between a third and two thirds of the way through, and in all three cases the best checkpoint was the last one taken — which means nobody knows whether the remaining half helps, and it is an open task rather than a finding (IV.2).

I.10.3 Corpus v1, and 88.9 % of it going missing

The preference corpus has five pair families. The rejected side is constructed, not sampled — the design decision that makes the whole approach cheap, since no second generation pass is needed:

Pair families available ([TR-0829] §11.3).
FamilyAvailable pairsHow the rejected side is made
vp_emotion1,064,594A stored lower-intensity take of the same voice-profile condition
vp_truncation1,186,406The chosen codes cut short
vp_continuation1,200,531The chosen codes with donor codes appended — running long
rs_truncation2,011,920Stored upstream
rs_continuation1,947,272Stored upstream

Version 1 of that corpus came out at only 236,166 pairs across 128 shards (1.7 GB), and the reason is a good example of a cost that gets stated rather than buried: the builder looked word timings up in the supervised corpus, which contains only the rows selected for supervised training. Every voice-profile pair whose chosen clip was not in that selection was dropped — 946,266 pairs, 88.9 % of the emotion arm — and because the length arm is capped to the emotion arm's size, the whole corpus shrank with it.

Corpus v1 by family ([TR-0829] §11.3).
FamilyRows in v1
vp_emotion118,328
rs_continuation32,252
rs_truncation32,074
vp_continuation27,029
vp_truncation26,483
total236,166

A correction to an earlier figure that propagated through the protocol and through the 26 August report: v1 is 236,166 pairs, not 236,656. The parquet shards and the build log both give 236,166, and the family counts force it arithmetically — 118,328 on the emotion arm against 117,838 on the length arm. It appears to be a digit transposition that was then copied forward.

I.10.4 Corpus v2 — the timings were never missing

They were in the annotation product's tar files all along: <uid>.json members carrying MMS-FA word alignments in exactly the {w, s, e} form the renderer expects. A sidecar keyed by clip id was built in three stages on a single node — plan (3,451,531 references → 2,146,797 distinct ids over 2,000 shards, 85 s), bursts (1 s per task), words (84 s per task). Coverage after the sidecar: word timings 2,137,175 of 2,146,797 (99.55 %); burst spans 2,146,797 of 2,146,797 (100 %).

Corpus v1 against v2 ([TR-0829] §11.4).
v1v2
Emotion pairs118,3281,061,613
Total pairs236,1661,853,486
Shards / size128 · 1.7 GB128 · 13 GB
Dropped for missing annotation946,266 (88.9 %)3,784 (0.36 %)
Arm ratio (emotion : length)50 : 5057 : 43

The v2 drop count splits as 2,981 rows with no annotation at all and 803 with an annotation but no usable duration; the protocol labels the 803 as the no-annotation count, which is the wrong one of the two. The 11.9 % pre-sidecar coverage figure quoted elsewhere could not be reproduced — the v1 auxiliary table was overwritten by the rebuild — so it is stated here as a protocol claim rather than a measurement. The 57 : 43 ratio is a deliberate departure from the original brief's strict equality, kept because the length arm has already done its work — timing is solved — and the emotion arm is what remains unsolved.

What eight times the data bought. Reward 0.4668 → 0.4687, emotion percentile 0.3518 → 0.3401, word error rate 0.1117 → 0.1094, burst realisation 0.3973 → 0.3929 (320 clips each). That is: a rounding-level gain in composite reward, a fall in the emotional intensity the corpus was enlarged to improve, and no material change anywhere else. Eight times the preference pairs on the arm that matters moved the target metric in the wrong direction. It is the cleanest evidence in this section that the emotion ceiling is not a data-volume problem, and it is the measurement that motivated changing the construction rather than the size.

I.10.5 The classifier-free-guidance corpus

Every contrast built so far pits an intense clip against a mild one — which a model can win by being generically expressive, without ever reading the instruction. The classifier-free-guidance (CFG) construction tries to make that impossible. (This CFG is a *data* construction; the inference-time classifier-free guidance the agent uses as a Tool is a different thing, Part II.3.2. The name collision is unfortunate and is kept because both artefacts carry it.)

A rendered CFG prompt, verbatim: <user_inst> - Reference(s): <|audio|> - Instruction: GENERAL: A voice; timbre is very dark - Tokens: 101 - Quality: None ... - Text: (very dark, and hold it there; otherwise exactly as this voice normally speaks) [5.5 seconds duration] ... [0.5 seconds pause] [2.0 seconds duration] ... </user_inst>

Composition of the CFG corpus, 2,327,904 pairs over 128 shards, 17 GB ([TR-0829] §11.5; the cfg_low/cfg_high counts are re-counted at render time from $SC/cfg_rows, 474,418 CFG pairs).
FamilyRowsShare
vp_emotion1,061,61345.60 %
cfg_low237,20910.19 %
cfg_high237,20910.19 %
rs_continuation218,6659.39 %
rs_truncation217,5369.34 %
vp_continuation179,6177.72 %
vp_truncation176,0557.56 %
CFG share474,41820.38 %

The two CFG halves are not merely balanced to within a tenth of a percentage point, as the 26 August write-up of this build reported — they are exactly equal, 237,209 each, which is what a correct flipped-role duplication must produce and is therefore a stronger check than the approximate one. The render-time recount confirms it (Appendix A.6: German 247,350 / English 227,068 pairs). The prompt-format hash moved to 073aeb09dc923376.

The model learned the task. Preference accuracy on the CFG families rose from 0.5625 over the first 40 logged batches — chance — to 0.9750 over the last 40. With the words removed and the roles flipped, the instruction is the only signal distinguishing chosen from rejected. A model ignoring it would sit at 0.5 forever. This one started there and learned its way out. Per-family means over the whole run: cfg_high 0.9093 (193 batches), cfg_low 0.9136 (189), vp_emotion 0.9518 (342), rs_truncation 0.9205 (153).

I.10.6 The contrastive-families corpus

Two further families were added on top of the CFG corpus, aimed at two measured defects rather than at a general idea. I.13 gives their construction, their selection rules and the per-family learning curves. In summary: p2_emox (309,128 rows) pits intense against intense — same speaker, matched length, emotion A against emotion B, with the instruction naming A — so that turning intensity up is worth nothing and only reading the instruction pays; p2_len (59,389 rows) keeps speaker, text and emotion fixed and varies only the duration. Mixed onto the CFG corpus: 2,327,904 + 368,517 = 2,696,421 pairs, the new families 13.7 %.

I.10.7 Which checkpoint is best, and why

All rows on the same 80 prompts, 4 completions each, inline delivery directions, zero-point-corrected percentile scale, plateau yardstick — so they are comparable to each other and to nothing else. (The full 18-run ranking table with every arm of I.11–I.12 is Appendix A.1.)

Model ranking, 320 clips per row ([TR-0829] §11.7).
ModelRewardWEREmotion pctQualityBurstHit rate
SFT-3 + DPO contrastive families, step 50220.47570.09770.35410.92080.41800.762
SFT-3 + DPO contrastive families, step 39060.47440.09160.34780.92310.41060.762
SFT-3 + DPO CFG corpus, step 49120.47080.09500.33730.92350.42710.772
SFT-3 + DPO CFG corpus, step 39630.46980.09530.34200.92350.39400.700
SFT-3 + DPO corpus v2, step 32160.46870.10940.34010.92110.39290.709
SFT-3 + DPO corpus v10.46680.11170.35180.91080.39730.694
SFT-3 + emotion LoRA r16 (joint)0.46630.12390.35720.91890.40020.684
SFT-3 + DPO corpus v2, step 45900.46330.10240.34270.90620.39300.719
SFT-3, no adapter0.45840.09870.34940.91270.35640.666
SFT-3 + GRPO v40.45510.12860.34680.89500.38370.694
SFT-3 + GRPO v50.45120.12530.32750.92820.35860.647

Step 5022 of the contrastive-families run is the best model this line produced, and the case for it rests on three things rather than on the composite reward alone, since the top four rows span 0.0059 in reward and this evaluation cannot resolve ±0.01:

Its neighbour at step 3906 is a statistical tie on reward, is better on word error rate (0.0916, the lowest of any model in the project), and is worse on emotion (0.3478) and on burst realisation. Step 5022 was chosen because emotion is the unsolved capability and 3906's WER advantage of 0.006 is inside noise. Both are published. What none of them did. The target band is 0.90–0.98. The best model reaches 0.354. Four preference corpora, seven runs and something over a hundred node-hours have moved that number from 0.3494 to 0.3541 — a gain of 0.005 on a scale where the goal is 0.55 away. The corpora are not what is wrong, and I.10.9 argues that the objective may be.

I.10.8 rew_chosen: a metric that is uninformative across runs and anti-informative within one

The DPO trainer logs two reward terms per evaluation: rew_chosen, the implicit reward on the preferred sequence, and rew_rejected, the same on the rejected one. A healthy preference run should raise the first while lowering the second. Every run, without exception, ends with a negative rew_chosen:

Final evaluation row of every preference run ([TR-0829] §11.8).
RunLast logged evalVal lossPref. accrew_chosenrew_rejected
dpo28640.74350.9864−0.7797−15.47
dpo38640.75760.9802−1.3769−18.69
dpo_on_sft2_WRONGBASE1841.02740.8184−0.6348−8.34
dpo_sft38640.75260.9805−0.8162−15.58
dpo_sft3_v235660.64991.0000−2.3037−57.85
dpo_sft3_cfg44760.71060.9844−1.1211−46.57
dpo_sft3_p225820.72770.9805−1.0015−38.85

The trainer emits an explicit <-- DAMAGE: policy suppresses the CHOSEN audio line and writes "healthy": false into eval.jsonl whenever this happens. It has happened on every evaluation row of every preference run in this project. Within a run the metric moves the wrong way as the model improves. From dpo_sft3:

dpo_sft3 evaluation series ([TR-0829] §11.8).
StepVal lossPreference accuracyrew_chosenrew_rejected
2160.94100.8652−0.2592−8.03
4320.79310.9707−0.6095−13.02
6480.75940.9766−0.8474−15.24
8640.75260.9805−0.8162−15.58

Preference accuracy climbs from 0.865 to 0.981 while rew_chosen gets three times worse — and step 864 is the checkpoint that generates best. The policy is suppressing the chosen audio while learning the ranking: it makes the bad audio catastrophically less likely and the good audio slightly less likely. Both go down; one goes down less. That satisfies the objective exactly and is not what anyone wanted.

A tempting cross-run version of this story, which is false. It is natural to compress the above into “the run whose rew_chosen was worst generated the best” — and the 26 August report said exactly that ([TR-0826] §13.1: “ends with reward(chosen) at −1.12, the worst of any run in the project, while generating the best audio”). The data refute it. The worst final rew_chosen belongs to dpo_sft3_v2 at −2.3037, whose best checkpoint scores 0.4687 and ranks fifth. The best model comes from dpo_sft3_p2, whose −1.0015 is the second least negative of the seven. Ordering the runs by final rew_chosen and ordering them by evaluated reward gives essentially unrelated sequences. The supported statement is the weaker and more useful one: rew_chosen is anti-correlated with generation quality inside a run and carries no information between runs. This retraction is listed again in III.3.

I.10.9 Why none of it moved emotion — one explanation refuted, three standing

Refuted: “the high side was only locally intense.” The obvious explanation is that pairing within a voice makes “high” a relative target — a voice profile that never gets angry has a top-1 % anger clip that is barely angry. Measured over 224,750 voice-profile rows in 500 voices, each voice's own top 1 % per emotion head scores a median 0.995 on the global zero-point-corrected scale, and 99.7 % reach the 0.90 band. The chosen side was genuinely extreme, and the training signal pointed at real extremes.

A fourth possibility has since arrived from a different direction and is recorded with the layer forensics (Part II.4): measured inside the model, the direction that means “more of this emotion” and the direction that means “higher measured quality” have a cosine of −0.746 at h20 and −0.954 at the last state before audio codes are chosen. If those two are close to opposite instructions, then any objective that rewards emotion and quality together is asking for something the representation makes difficult, and every reward in this project — including the agent's Best-of-N reward — does exactly that.

I.11 GRPO: five runs, and one structural negative result

GRPO generates G completions for the same prompt with sampling, scores each with the measurement models, normalises the rewards within the group (A = (r − mean) / std), and pushes the policy towards the completions with positive advantage. With one inner update per rollout the PPO ratio is exactly 1, the clipping term is inert, and the loss reduces to −(A · logp).mean(). The whole learning signal is therefore the spread of scores inside a group. Source: [TR-0829] §12; run logs in [P§10–§14].

reward = (0.40·r_emo + 0.40·r_qual + 0.20·r_burst) × wer_factor(WER)r_emo the band score of the percentile reached on the requested emotion head against the requested band; r_qual = 0.5·percentile(vocal-burst blend) + 0.5·percentile(genuineness); r_burst the F1 over prompted vs detected bursts, weighted by duration match; WER from Whisper-large-v3-turbo against the script's own plain text.

All five GRPO runs ([TR-0829] §12).
RunBasePrompt formatGKL βHoursStepsRollout rewardOutcome
1SFT-2round 2, no directions16off1.4750.657 → 0.555reward fell
2SFT-2round 216offDDP deadlock, 35 GB of core dumps
3SFT-2 + run-1 adapterwith directions32on2.70.409 → 0.347reward fell
4SFT-3with directions160.022.00880.503 → 0.527stable; bursts and WER improve
5SFT-3with directions160.022.021050.611 → 0.608emotion worse on the strict yardstick

I.11.1 Runs 1 and 3 were not a GRPO result

Both fell, and the cause was outside the optimiser. Run 3 was given prompts with inline delivery directions — the right idea — but the model had never been trained on that format, because round 2 had dropped directions by accident (I.9.2). It collapsed, and GRPO spent 2.7 hours optimising against a reward dominated by that collapse. The conclusion is not that GRPO needs tuning; it is that a reward computed on out-of-distribution prompts measures the prompt, not the policy.

I.11.2 Run 4 — the first run that neither deadlocks nor collapses

88 steps, 176 rollout groups, 2.00 h on one 4-GPU node, from SFT-3, KL β 0.02, corrected percentile scale. skipped = 0 for the entire run: the rank-divergence fix (III.2.1) held.

GRPO run 4 by third of the run, 176 groups ([TR-0829] §12.2).
Third of the runRewardEmotionQualityBurstWER
first0.5030.2590.8400.5610.128
middle0.5500.2800.8730.6660.103
last0.5270.2420.8340.6690.098

Burst realisation rises 0.561 → 0.669 and WER falls 0.128 → 0.098. Emotion is flat within noise. (An earlier reading of single steps had suggested an emotion collapse; the aggregate over 176 groups does not support it, and per-step numbers of two groups each are far too noisy to read as a trend — a small methodological correction that is recorded rather than quietly dropped.)

I.11.3 The decisive measurement

Splitting run 4 by the band the prompt asked for, where emotion is the band score of the percentile achieved against the percentile band requested:

GRPO run 4 split by requested band ([TR-0829] §12.3).
Band asked forEmotion, first thirdEmotion, last thirdBurst, first → last
faint (0.40–0.70)0.3690.3170.756 → 0.541
moderate (0.70–0.90)0.2790.3840.506 → 0.671
intense (0.90–0.98)0.2340.1940.543 → 0.741
extreme (0.98–1.00)0.0710.1390.592 → 0.582
The model scores three to five times better when asked for a mild emotion than for an extreme one — and that is why GRPO cannot learn emotion here. At the extreme band the band score sits on its floor for essentially every completion, so inside a group of 16 completions of the same prompt the emotion term is nearly constant — and a constant term contributes exactly nothing to a group-normalised advantage. Half the prompt mix (intense + extreme, by design) therefore carries almost no emotion gradient, while burst realisation and WER vary freely across completions and dominate what the update can learn from. Supporting numbers, run 4: mean within-group reward standard deviation 0.1063 — that is the entire learning signal per step. Across groups, corr(reward, emo) = +0.753, corr(reward, burst) = +0.310, corr(reward, wer) = −0.432, with std(emo) = 0.292, std(burst) = 0.339, std(wer) = 0.122. The reward is not ignoring emotion — the policy cannot move it where it is being asked to.

I.11.4 Run 5 — four fixes for that, and why they made it worse

Four levers were built and all four default off, so run 4 stays reproducible:

Run 4 against run 5 by third ([TR-0829] §12.4).
Third of runRewardEmotionQualityBurstWER
v4 first0.5030.2590.8400.5610.128
v4 last0.5270.2420.8340.6690.098
v5 first0.6110.4410.8530.5690.122
v5 last0.6080.4240.8340.5990.119

On the strict evaluation yardstick, run 5 is the worst of the four models on emotion (0.3275 against SFT-3's 0.3494 — see I.12.3). The training-time gain was the mechanical artefact it had been predicted to be: lever 1 inflates the number by paying for partial progress, lever 3 inflates it by asking easier questions. The realised band mix confirms the second: v4 ran 0.28 moderate / 0.28 intense / 0.22 extreme / 0.21 faint, v5 ran 0.36 / 0.30 / 0.15 / 0.20 — the curriculum moved the run onto the easier half and the run ended before it came back. Hypothesis, stated as a hypothesis: a ramp that pays for partial progress supplies a gradient but removes the incentive to arrive. A completion that moves slightly toward the target now collects reward without reaching the band, and the strict yardstick only counts arrival. Levers 1 and 3 push in the same direction and it is the wrong one. A ramp that is steeper near the band, or a rank-within-group reward that keeps the target absolute, would test this.

I.11.5 What GRPO did produce: data

Lever 4 is the one that paid. Run 5 saved the most emotional usable completion of every group: 204 clips at median emotion percentile 0.641 and median WER 0.049, median 148 frames — 73 moderate, 61 intense, 41 faint, 29 extreme. High emotion and clean speech in the same clip is the combination that has been missing from every evaluation in this project. The model can do it; it does not do it on command. Two hours of rollout on one 4-GPU node therefore yields supervised training material at roughly 100 usable clips per GPU-hour — and best-of-N supervised training on that harvest was, on 29 August, judged a more promising route than further reward shaping. The agent of Part II is the inference-time version of the same conclusion: select, do not push.

Hyperparameters, for reproduction: group size 16 (32 in run 3), 2 groups per optimizer step, micro-batch 2 completions, LoRA rank 32 / α 64 (68.7 M trainable of 4.27 B), learning rate 2e-6 constant after 20 warmup steps, KL β 0.02 with the k3 estimator against the same model with the adapter disabled (so no second copy in memory — measured KL stayed at 2–3e−4 and the constraint was never binding), sampling temperature 1.0 / top-p 0.95 / top-k 50, max 340 new frames (27.2 s), prompt pool 4,000 rows, checkpoint every 25 steps, 2-hour stopping rule. Prompt mix: 50 % carry vocal bursts, 35 % carry two emotions, 50 % carry a reference clip, 85 % carry a delivery direction. Throughput: one 4-GPU node produces 45–52 optimizer steps per hour at G = 16, i.e. 1,440–1,680 generated and fully scored clips per hour. Loading the three measurement models costs about 40 s per rank.

I.12 Specialist adapters: the only thing that moved intensity before the agent

I.12.1 Why a quality gate is not optional

Measured on a 360-clip listening sample decoded from the corpus's own audio codes — one clip per emotion per tier ([TR-0829] §13.1):

Quality against intensity tier, n = 360 clips.
TierEmotion percentile rangeBurst blend, medianGenuineness, median
top 10 %0.550 – 1.0000.4510.391
top 5 %0.782 – 1.0000.4430.360
top 1 %0.959 – 1.0000.4630.360

Genuineness falls as intensity rises, and both quality axes sit below the corpus median of 0.50 throughout. The most intense recordings are on average the least genuine. An adapter trained on intensity alone would be taught to overact.

I.12.2 The general emotion LoRA: two arms and one gate

Selection for the general emotion adapter ([TR-0829] §13.2).
StepRows
Top-1 % union, emotion heads only, no gate731,948 (23.3 % of the corpus)
Emotion arm, gated174,368
VoiceNet arm, gated111,326, of which 69,618 are new
Total243,986 (7.8 %) — materialised at 247,049 rows including the unchanged validation holdout

Two things fall out of this. First, the VoiceNet arm added about 40 % more material than the emotion heads alone would have found, and the axes that contributed are the affective ones — high valence (16,726 rows qualifying on that axis alone), explosiveness (14,437), arousal shift (9,890), valence shift (8,164), velocity flux (4,680), low arousal (3,323), low valence (2,383), high arousal (1,972). The stylistic axes — ranting, ASMR, whispering, cartoonish — contributed nothing in the top eight: their top 1 % was either already in the emotion arm or failed the quality gate. Second, the source mix flipped. The gated selection is 65.4 % real speech, 34.6 % voice profiles — the reverse of the ungated listening sample (273 synthetic clips to 87). The gate removes synthetic extremes disproportionately, which supports the reading that the synthetic “top 1 %” is often overacted. [TR-0826] §12.6 adds the bucket caveat: the largest VoiceNet buckets are S_ASMR bottom (23,817) and S_CART bottom (18,934) — “not ASMR” and “not cartoonish”, the unremarkable default, whose high genuineness (0.87) reflects that. Only the four two-sided affective axes have a meaningful lower tail.

I.12.3 Result: rank 16 wins, and it is the first thing to move the number

Three adapters on the same 247,049-row corpus, 5 epochs, cosine 1e-4 → 5e-6, α fixed at 2 × rank so the effective scaling is identical and rank is the only variable, global batch 256 (4 × 2 × 32 on 8 nodes), 4,395 steps, 2 h 35 to 2 h 40 each. Base SFT-3, evaluated alone on it.

The three ranks ([TR-0829] §13.3).
RankαTrainableShare of model
163234.4 M0.825 %
326468.7 M1.637 %
64128137.4 M3.220 %

The complete ranking of every evaluation on the corrected percentile scale — eighteen runs, all on the same 80 prompts, 4 completions per prompt, the strict plateau yardstick — is Appendix A.1. It replaces every partial comparison published earlier, including the eight-row table of [TR-0826] §12.3. Three readings of it:

Six further evaluations exist and are deliberately excluded from that ranking, because they predate the percentile correction (I.6.7): eval_sft2_dir, eval_sft2_dirv3, eval_dpo2_dir, eval_grpo1_dir, eval_grpo3_dir and eval_sft3_dir_UNCORRECTED_ECDF. Their word error rates (0.45–0.51 for the SFT-2-based ones) are comparable and genuinely bad; their reward and emotion columns are not.

I.12.4 Merge weight: two sweeps, two answers, and why both are right

An obvious next question is whether merging an adapter more strongly pushes emotion further. Two sweeps have been run, and they came out differently. Reported together, because the difference between them is the finding — and because the agent's shipped stack (Part II.3) sets adapter weights above 1.0 on the strength of sweep B.

Sweep A — one general adapter, band intense, 80 prompts: no trend. The single pooled emotion LoRA (rank 16, trained on all 247,049 gated rows) evaluated at seven merge weights on the standard harness, 320 clips each:

Sweep A ([TR-0829] §13.4).
Merge weightRewardWEREmotion pctQualityBurstBurst hit rate
0 (SFT-3)0.45840.09870.34940.91270.35640.666
0.250.45790.11740.35330.91360.38270.666
0.500.45540.13450.33650.91500.37120.666
0.750.46220.10550.33900.90520.38950.697
1.000.46630.12390.35720.91890.40020.684
1.250.45840.11300.33150.90500.39290.684
1.500.45090.13250.33580.92600.36110.631

Emotion percentile spans 0.3315–0.3572 with no monotone structure and a peak at the trained value. Reward spans 0.015 — inside what these evaluations show between arms that ought to be equivalent. On this axis, with this adapter, at this band, merge weight is not a dial for intensity. ([TR-0826] §12.4 reported this sweep first, with six weights and the note that the numbers were “measured today and reported here first”.)

Sweep B — 31 per-emotion adapters, band extreme, with bursts: a clear monotone rise. The second sweep takes the per-emotion bucket adapters of I.12.5, each on prompts naming its own emotion at the extreme band with vocal bursts requested, and scales each adapter's contribution over six weights. Design details that make it a strong measurement: the same eight prompts are used at every weight for a given adapter, the torch seed is reset to a fixed value before each weight, and weight 0 is the same loaded model with the LoRA scaling zeroed rather than a separately loaded baseline — so a difference cannot come from anything but the adapter. 31 adapters × 6 weights × 8 clips = 1,488 scored generations.

Sweep B, 248 clips per weight ([TR-0829] §13.4).
Merge weightEmotion pctGenuineness pctBlend pctBurst realisationWER (mean)WER (median)Mean |duration error|
00.40790.81690.92510.26050.16690.0000.127 s
0.250.40680.84440.95490.27340.17900.0000.118 s
0.50.42970.83310.92290.30350.14640.0000.115 s
1.00.44070.83640.95360.23070.12980.0000.133 s
1.50.47140.84600.96130.25470.09560.0000.120 s
2.00.49230.87960.96880.22270.18420.0300.129 s

Emotion percentile rises monotonically from weight 0.5 onward and is +0.084 higher at weight 2.0 than at weight 0 (paired over 248 clips, bootstrap 95 % CI [+0.048, +0.123]); at weight 1.5 the paired gain is +0.064, CI [+0.030, +0.098]. 27 of 31 adapters are higher at the top weight than at zero — the three that are not are Helplessness (−0.185), Bitterness (−0.104) and Contempt (−0.024) — and the fourth, Jealousy and Envy, is the dead metric of I.6.4 and reads 0.000 everywhere. At weight 1.5 the split is 25 up, 5 down, 1 tied. Per-adapter argmax weight: 2.0 for 13 adapters, 1.5 for 7, 1.0 for 6, 0.5 for 2, 0.25 for 2 and 0.0 for 1. Timing is untouched at every weight — mean absolute duration error stays inside 0.115–0.133 s across the whole sweep, so overdriving an adapter does not cost the capability that was actually solved.

The recommended operating point is 1.5, and here is what breaks at 2.0. Word error rate is best at weight 1.5 (mean 0.0956, the minimum of the sweep) and worst at 2.0 (0.1842). But that is not a broad collapse — it is a tail. Exactly five of 248 clips at weight 2.0 have WER ≥ 1.0, and two of them are the same failure: an Astonishment_Surprise clip and a Fear clip that both derail into “Ha ha ha ha ha ha ha …” repeat loops and score 11.11. Dropping just those two brings the mean at 2.0 to 0.0954 — indistinguishable from 1.5. Two other clips transcribe as silence and one as a single stray word. Underneath the tail: at weight 2.0 the median WER moves off zero for the first time (0.030), and the number of clips with exactly zero word errors falls to 111 from 128–135 at every other weight; meanwhile the count of moderately bad clips (WER ≥ 0.5) is at its minimum at weight 2.0 (9, against 12 at 1.5 and 25 at weight 0). So: 2.0 makes the typical clip slightly worse and adds a small population of catastrophically derailed ones; 1.5 buys most of the emotion gain with the best intelligibility in the sweep. 1.5 is the published recommendation — and it is the BURST_LAM_MAX = 1.25 / SFT3_VN_MAX = 1 ceiling of the agent's shipped stack in embryo; Part II.7 records what happened when two burst adapters at 1.5 were summed.

Why the two sweeps disagree, and which to believe for what. They are not the same experiment. Sweep A scales a single general adapter trained on all 154 axes at once, on the intense band, against a prompt set that mostly does not request bursts; sweep B scales per-emotion specialists on the extreme band, on prompts that name the adapter's own emotion and request a burst. The reconciling reading is that a general adapter has no single direction to overdrive — scaling it amplifies an average of 154 directions, which is close to noise — while a specialist adapter has exactly one, and scaling that one works. Both results stand. Coverage note: sweep B was launched when 31 of the 40 per-emotion adapters existed. Amusement, Embarrassment, Intoxication, Relief, Sadness, Sexual Lust, Shame, Thankfulness and Triumph are absent from it. Re-running it would cover all 40; as of 9 September that has not been done.

I.12.5 One adapter per emotion — and the selectivity problem

The natural next step is 40 adapters, one per emotion, rather than one general one. All 40 were trained (rank 16, α 32, 5 epochs, cosine 1e-4 → 5e-6, 34.4 M trainable parameters, each on the gated top 1 % of its own emotion, all finishing with zero non-finite batches). Training set sizes run from Sourness at 2,663 rows to Intoxication at 15,725; step counts from 3,330 to 19,660; wall time from 42 to 224 minutes per adapter. These are the emotion adapters of the agent's shipped stack (Part II.3). Source: [TR-0829] §13.5.

Seventeen of them were measured against the best general checkpoint of the time — SFT-3 + CFG-DPO — on identical prompts, identical sampling and identical seeds. The seventeen are Affection, Anger, Astonishment/Surprise, Awe, Bitterness, Concentration, Contempt, Contentment, Disgust, Elation, Emotional Numbness, Jealousy and Envy, Longing, Malevolence/Malice, Pleasure/Ecstasy, Pride and Sourness. That set is not a chosen subset: the generator shards adapters across ranks with names[RANK::WORLD], and the adapter directory was still growing while the workers launched, so ranks duplicated each other and skipped others. Seventeen is where the job landed.

The design of that evaluation is the interesting part. A specialist adapter should do two things: make its own emotion stronger when it is asked for, and change nothing when it is not. So each adapter gets ten prompts naming its emotion at the intense band and ten prompts naming no emotion at all — the corpus's own neutral direction. A rise in the first is the adapter working. A rise in the second is the adapter leaking into plain speech.

Selectivity result, 17 adapters × 10 + 10 prompts ([TR-0829] §13.5).
MeasurementValueAdapters improving
Emotion when asked for+0.0466 (se 0.0184, t = 2.53)11 of 17 (5 down, 1 exactly zero)
Emotion when not asked for+0.0327 (se 0.0259, t = 1.26)12 of 17 (4 down, 1 zero)
Selectivity ratio1.43 : 1
Paired difference (Δ asked − Δ neutral)+0.0140 (se 0.0335, t = 0.42, n = 17)
Word error rate, with adapter vs baseline0.1896 vs 0.1215
Burst realisation, with adapter vs baseline0.2547 vs 0.3252worse for 16 of 17

So the adapters do push their emotion — and they push nearly as hard when nobody asked. They also cost about half again as much transcription error and about a fifth of the burst realisation. As they stand they are not a drop-in improvement. They are tinting, not control.

Three statistical caveats on that 1.4 : 1, stated because they matter. The selectivity gap itself is not statistically significant: the paired difference is +0.0140 with a standard error of 0.0335 — t = 0.42, n = 17 adapters. The on-target lift alone is marginally resolved (t = 2.53); the off-target lift is not (t = 1.26). “1.4 : 1” is the right point estimate and it is well inside noise. The medians point the other way: median Δ asked is +0.0368 and median Δ neutral is +0.0423 — the typical adapter leaks slightly more than it steers; the 1.43 ratio is carried entirely by the means. And one adapter dominates the off-target average: Elation's neutral-prompt delta is −0.2916, roughly three times any other magnitude in either direction. Remove it and the off-target mean rises sharply — so the “leakage is nearly as strong” conclusion rests on one outlier pulling the mean down. [TR-0826] §12.5 reported the ratio without any of these caveats; they were added on 28 August.
Per-emotion selectivity, WER on matched prompts only ([TR-0829] §13.5).
EmotionΔ when askedΔ when neutralWER with LoRAWER baselineTraining rows
Elation+0.159−0.2920.2030.2956,046
Emotional Numbness+0.158+0.1290.0590.1364,633
Concentration+0.150+0.1610.2940.1675,400
Bitterness+0.131+0.0100.3340.0963,364
Pride+0.092−0.0580.0380.0866,784
Contempt+0.087+0.0530.2930.3103,105
Contentment+0.068+0.1050.4430.1703,445
Anger+0.053+0.1320.2660.0836,142
Astonishment / Surprise+0.037+0.0550.2190.0244,899
Malevolence / Malice+0.036−0.0370.1180.2126,628
Longing+0.034+0.0380.1940.0927,981
Jealousy and Envy0.0000.0000.1500.2694,503
Sourness−0.003+0.1620.0640.0102,663
Pleasure / Ecstasy−0.023+0.0420.0390.0422,976
Awe−0.056+0.0610.2090.0573,247
Disgust−0.061+0.0080.1010.0863,271
Affection−0.070−0.0130.2140.0615,235

Word error rates here are for the emotion-naming (matched) prompts only. Different adapters draw different slices of the prompt pool, so cross-adapter comparison of the WER columns is not prompt-matched. The prompts are drawn from three source corpora — the run log's “three voices” label is a fallback string, not three verified speakers — so the design controls for corpus, not for speaker identity, which is weaker than it looks. Jealousy and Envy reads 0.000 everywhere because of the key mismatch of I.6.4, not — as [TR-0826] §12.5 said — because of the percentile correction. The agent-era combination study (Part II.3.5) re-measured the emotion adapters' contribution on 399 clips and found a larger, resolved main effect (+0.077 target-z, t = 2.81) — with a different prompt set, different ears and the full shipped stack around it.

I.12.6 Is there enough data for a LoRA per bucket?

Bucket sizes over 3,144,739 rows, 154 buckets ([TR-0829] §13.6).
GroupBucketsUngatedGatedMedianSmallest
40 emotions401,258,097283,2756,7062,663
VoiceNet, expressive axes321,006,528272,0537,854426
VoiceNet, the rest822,579,188552,0665,0840

Feasible — a median of 6,700 rows is comfortable for a rank-16 adapter — but badly skewed, and skewed by exactly the effect that motivated the gate. The four smallest emotion buckets are Sourness (2,663), Pleasure/Ecstasy (2,976), Contempt (3,105) and Awe (3,247), and their genuineness median in the top 1 % is 0.23–0.30: the most intense recordings of those four are almost entirely inauthentic, and the gate removes three quarters of them. Intoxication (15,725 rows, genuineness median 0.804) and Amusement (13,272, 0.774) have ample genuine material. A 40-adapter programme therefore needs either unequal training sets or a relaxation to the top 2 % for the four problem emotions.

I.12.7 The VoiceNet axis adapters

The same bucket machinery was pointed at the VoiceNet axes on 26 August at 16:47. The policy is the one I.12.2 established: both tails for AROU, VALN, ARSH and VALS; the top only for the rest — the survey found that the largest “bottom” buckets of the style axes are simply ordinary unremarkable speech, so there is no such thing as an extreme absence of ranting. Twenty candidate buckets went in; three were dropped by a row-count floor of 1,500 applied after the top-1 % tail selection and the blend-and-genuineness quality gate, and 17 came out: 146,175 rows in total, median 8,216, minimum 1,922, maximum 22,849.

VoiceNet axis buckets, gated rows ([TR-0829] §13.7).
BucketGated rowsBucketGated rows
VALN_high22,849VULN_high8,216
EXPL_high22,378AROU_low7,492
ARSH_high13,219AROU_high4,104
VALS_high13,037EMPH_high3,306
ARSH_low10,365S_DRAM_high2,883
VALS_low10,161S_RANT_high2,608
VFLX_high9,849TENS_high2,532
VALN_low9,286VOLT_high1,968
S_ASMR_high1,922droppedRANG_high 1,147 · S_WHIS_high 939 · S_CART_high 680

The three dropped axes are exactly the ones where the quality gate bites hardest: high pitch range keeps 1,147 of 31,453 candidates, whisper-talk 939 of 31,452, and cartoonish 680 of 31,450 — a 2.2 % survival rate for cartoonish, whose genuineness median in its own top 1 % is 0.109. That is the I.12.1 finding again, in its sharpest form. Training of those 17 adapters started at 16:51 on 26 August (rank 16, α 32, 5 epochs, AdamW lr 1e-4, cosine with 10 % warmup, batch 4, non-DDP, one bucket per GPU stream). They were not published as of 29 August; the delivery-axis adapters the agent ships (SFT3_VN_LEVELS, Part II.3) are this family's descendants, and their measured WER cost (+0.0546, t = 6.91, n = 160) is in Part II.3.5.

I.12.8 Every adapter is trained on SFT-3 alone, and here is why

All 40 emotion adapters, all 500 voice adapters and the 17 VoiceNet adapters are trained on SFT-3 — never on a merged SFT-3 + DPO checkpoint. The reason is structural, not a preference: the DPO adapter targets audio_lm_heads.*, and those twelve tensors are weight-tied to audio_embeddings.*. They are literally the same tensor. Merging the head delta therefore rewrites the embedding that produced it, and the model is corrupted (III.2.2 gives the measurement). The adapters are stacked at inference instead, each with its own scaling factor. The published recommendation on 29 August was DPO at 1.0, the voice adapter at 1.0 and the emotion adapter at 1.5 — the emotion weight from sweep B; the voice weight has not been swept and 1.0 is simply its trained value. The agent's shipped stack (Part II.3) keeps DPO 1.0 and voice 1.0 but sets emotion at 1.0, with the three quality adapters and quality_dpo at 1.5 around it — a different recipe, chosen by the combination study rather than by sweep B.

I.13 Two new contrastive pair families — and the model they produced

These are aimed directly at two measured defects — the leakage of I.12.5 and an emotion-blind length signal — and they are the only intervention before the agent that raised the emotion percentile above the supervised baseline without costing word error rate. The adapter they produced is the DPO of the agent's stack. Source: [TR-0829] §15; the earlier build in [TR-0826] §16.2.

I.13.1 What they are

Family A — emotion-contrastive, for selectivity. Every contrast built before this pits an intense clip against a mild one, so a model can win by being generically expressive — which is exactly the 1.4 : 1 tinting the per-emotion adapters showed. Family A pits intense against intense: same speaker, same length, emotion A against emotion B, with the instruction naming A. Turning the intensity up is then worth nothing; only reading the instruction is. The builder's own docstring states the motivation in one line: “A ratio of 1.4 : 1 is not control, it is tinting.” The selection rule is strict on both sides. The chosen clip must sit at percentile ≥ 0.90 on the head being asked for and have that head as its own top emotion; the partner must be ≤ 0.50 on that head and ≥ 0.90 on its own top head. So both sides really are intense, on different emotions. Lengths are matched to within 10 % of frames. Each base pair is emitted twice, the mirror naming the partner's own top emotion. Because chosen and rejected are different recordings with different words and a DPO pair shares one prompt, the spoken words are removed, as in the CFG construction.

Family B — length-contrastive, conditioned on the emotion. The existing truncation and continuation families know nothing about emotion, so they teach “the right length is better” in the abstract. Family B keeps the speaker, the text and the emotion fixed and varies only the length: chosen is the clip at its true duration, rejected is the same clip cut to 50–75 % or run to 125–150 %. Here the words are in the prompt, with the full timed script, because the timing tags are exactly what the rejected side violates. Both families take the emotion from the measured value, never from the prompt that generated the clip. Constants: HI_PCT 0.90, LO_PCT 0.50, FRAME_TOL 0.10, 8 emotion-contrastive and 6 length-contrastive pairs per cell, cut range 0.50–0.75, extension range 1.25–1.50, RNG seed 4711.

I.13.2 What was built

The contrastive-families build ([TR-0829] §15.2).
QuantityValue
Emotion-contrastive cells (voice × emotion)19,884
Emotion-contrastive base pairs → rows (each emitted twice)154,564 → 309,128
Length-contrastive: cut / extended29,564 / 29,825 = 59,389
New rows written368,517 over 342,284 distinct clips
Mixed onto the CFG corpus2,327,904 + 368,517 = 2,696,421
New share13.7 %

The base is the CFG corpus rather than the older one, and the mixer says why: because the CFG corpus produced the best checkpoint measured so far. Schema equality between the new rows and the base was checked field by field before the merge — 30 fields each, none missing, none extra — and all 342,284 referenced clips were found. An earlier build with a smaller per-cell budget produced 138,088 rows (118,304 emotion-contrastive over 19,884 cells, plus 9,953 truncation and 9,831 extension pairs over 500 voices — [TR-0826] §16.2) and was superseded.

I.13.3 What the run did

The run is dpo_sft3_p2: 32 ranks over 8 nodes, LoRA rank 64 (137.4 M trainable of 4.267 B, 3.220 %), β 30, chosen-NLL anchor 0.013, learning rate 1e-6, 10,331 steps per epoch with 1,033 warmup, 256 preference pairs per step, arm-balanced sampling. It ran at about 62 pairs per second and hit its six-hour wall clock at step 5,126 of 10,331 — half an epoch. Checkpoints were kept at 3906, 4132, 4355, 4579, 4800 and 5022.

Preference accuracy per family, means over logged batches ([TR-0829] §15.3).
Windowemotion-contrastivelength-contrastivecfg_highcfg_low
steps 1–1771 (first half)0.8640.7860.8250.766
steps 1772–3543 (second half)0.9381.0000.8880.938
post-warmup (1033–3543)0.9510.9500.9130.946
cumulative to step 25820.8890.8420.8350.813
cumulative to step 35430.9020.8850.8540.851

The model learned the harder contrast. Emotion-contrastive accuracy rises from 0.864 to 0.938 across the two halves and sits at 0.951 post-warmup, above the older CFG families' 0.82–0.84 at the corresponding cumulative point. The length-contrastive family saturated at 1.000 over the second half — the task is solved and it is no longer informative, consistent with timing having been solved since round 2; it could be dropped from a future mix. And the run failed the health metric, exactly as every previous DPO run has: the mid-run evaluation at step 2582 reported val loss 0.7277, preference accuracy 0.9805, rew_chosen −1.0015 and rew_rejected −38.851, and the trainer wrote "healthy": false.

I.13.4 The result: a new best model

Step 5022 against its predecessors, 320 clips each ([TR-0829] §15.4).
ModelRewardWEREmotion pctQualityBurstHit rate
SFT-3, no adapter0.45840.09870.34940.91270.35640.666
+ DPO, CFG corpus, step 4912 (previous best)0.47080.09500.33730.92350.42710.772
+ DPO, contrastive families, step 39060.47440.09160.34780.92310.41060.762
+ DPO, contrastive families, step 50220.47570.09770.35410.92080.41800.762

Because both this run and the CFG run use a global batch of 256 preference pairs, equal step numbers mean equal numbers of pairs seen — so step 5022 here and step 4912 there are a nearly matched comparison, and the difference is the corpus rather than the budget. The emotion percentile is 0.3541 — the highest of any preference-tuned model in this line, and the first one above the supervised baseline of 0.3494. The gain is +0.005, against a requested band of 0.90–0.98. Both statements are true and the second one is the important one. I.10.9's first standing hypothesis predicts precisely this: a preference objective can teach a model to recognise which of two recordings matches an instruction without teaching it to produce one, and high emotion-contrastive accuracy is evidence of the former only. The adapter is published as laion/moss-va-sft3-dpo-lora-p2. The remaining half epoch was never run; whether it helps is unknown.

I.14 Three inference-time levers, measured end to end (28–29 August)

Sections I.8–I.13 are about making the model learn to follow an intensity instruction. That programme reached a wall: five GRPO runs and seven preference runs did not move emotional intensity. This section is about the other approach — leaving the trained weights alone and pushing the model at the moment it speaks. Three such levers exist, all three were measured across every adapter and dimension the project has, and the picture they produce is consistent enough to be a property of the model rather than of any one method. This is the study that the agent of Part II was built on; Part II.3 re-measures the same levers inside the shipped stack. Source: [TR-0829] §24 (added 28 August); the underlying study pages are cited in [P§40–§53].

The three levers ([TR-0829] §24.1).
LeverWhat it doesDialCost structure
LoRA adapterrank-16 weights trained on the top 1 % of one attribute, added at run timemerge weight w; 0 = base, 1.0 = as trainedfree at inference; must not be merged (I.14.6)
Steering vectora difference-of-means direction added to the hidden state, h ← h + α·(v/‖v‖)·‖h‖α, dimensionlessnegligible compute; breaks the output above α = 0.3
Classifier-free guidancetwo forward passes per frame, one with the emotion condition neutralised; logits = logits_uncond + g·(logits_cond − logits_uncond)g; 1.0 = ordinary sampling1.93× wall clock

Scale: 7,419 CFG cells / 29,676 clips, 5,740 adapter dose cells over 79 adapters, 10,459 combination cells / 31,377 clips over 60 attributes. Every figure is paired — the same prompts with and without the intervention, a prompt's samples averaged before averaging across prompts, so n counts prompts. Word error is Parakeet-TDT-0.6b-v3 (I.14.8).

I.14.2 The central result: manner is writable, feeling is not

Each lever at its best setting ([TR-0829] §24.2).
Lever, at its best setting40 emotion headsDelivery axesRatio
steering, α = 0.10+0.10+1.78 (S_DRAM, t 8.58)18×
adapter, at its saturation weight+0.035 (t 4.82)+0.690, 17/17 improving20×
guidance, at WER ≤ 0.20+0.0254 (g = 3.0, t 2.18)+0.0569 (g = 2.5, t 3.12)2.2×

Three mechanisms with nothing in common — a weight patch, a hidden-state nudge, a logit extrapolation — all move the delivery axes several times further than the emotion heads. The ordering matches what a probe can read from the same activations: R² 0.85 for the 57 VoiceNet dimensions against 0.47 for the 40 emotion heads (Part II.4). The model has a rich, writable representation of how a line is delivered and a thin one of what it feels, and no amount of inference-time pressure changes that ratio. Decodability nevertheless fails to predict steerability per attribute: Anger (R² 0.523) gains +0.088 from a steering push while Elation (0.503) gains +0.311, and Jealousy_and_Envy — the least decodable head at 0.209 — gains +0.070. Reading a layer and writing to it are different operations.

I.14.3 The dose-response of the adapters

79 adapters swept at six merge weights. 33 have a usable weight; none is harmful at any weight. Every failure is a failure to move the target, not a broken guardrail.

Dose-response by family ([TR-0829] §24.3).
FamilynUsableMedian safe wShape
emotion40111.0climbs to 1.0 (t 4.82) then flat: 1.5 − 1.0 = −0.0026 (t −0.40)
VoiceNet delivery17161.512 monotone, 4 saturating — the best behaved family
quality311.5right direction at every weight, misses significance
vocal burst1941.5mixed

Two cautions the table cannot carry alone. “Below resolution” is not “no effect”: the 24 quiet emotion adapters still move pooled (+0.024 at w = 1.0, t 2.70); ten prompts each is not enough to prove it one at a time. And 20 of the 71 burst adapters cannot be evaluated at all — trained almost entirely on isolated snippets under two seconds, with no speech row in the corpus carrying the class. The most actionable number: raising the delivery clamp from 0.75 to 1.5 buys +0.375 on the target axis (t 5.18, better on 15 of 17) for a word-error change of +0.003 (t 0.55, i.e. none). The clamp was chosen when several delivery adapters ran at once; it is not justified for one at a time. Note the contradiction with the agent-era measurement: with the full shipped stack around it, one delivery axis costs +0.0546 WER (t 6.91, n = 160, Part II.3.5). Both stand; the difference is the stack.

I.14.4 Guidance, and its true cost

Best guidance at a pre-registered threshold of Parakeet WER ≤ 0.20: g = 3.0 for emotion (+0.0254, t 2.18), g = 2.5 for delivery (+0.0569, t 3.12). Below g = 1 guidance actively hurts (−0.0370 at g = 0.5, t −2.56), the directional control the arm needed. Side effects are mild and partly favourable: genuineness rises with guidance on the emotion family (3.85 → 3.94), blend is flat, duration error unmoved; burst realisation falls slightly (0.414 → 0.392) and the delivery family pays intelligibility above g = 2 (share at WER ≤ 0.2 falls 0.828 → 0.744 by g = 3). g = 3.0 is the BON_GUIDANCE the agent ships (Part II.3.2).

CFG costs 1.93×, not the 1.2–1.5× the architecture suggests. Measured 1.889× at batch 1, 1.947× at batch 4. The premise was that only the semantic transformer doubles while the local (talker) transformer and the codec decode run once. The talker doubles too — twelve calls per frame per branch, because the unconditional branch needs its own local KV state to predict the next channel. Semantic 72–77 %, talker 18–19 %, heads 4–7 %, codec decode 1.6 % — and the decode is the only genuinely shared component, so sharing it saves nothing. The model is at RTF 0.66 solo; CFG consumes that headroom. Part II.3.2 collects five independent cost measurements, 1.89–2.04×.

The LoRA difference vector as a guidance signal. With SFT3+DPO-v2 as the unconditional branch and SFT3+DPO-v2+LoRA as the conditioned one, the same arithmetic applies at the logit level. At equal nominal factor, weight-scaling reaches further — but at 3.0 it hits WER 0.310 and collapses burst realisation to 0.287, where logit guidance stays at 0.111 and 0.400. At matched intelligibility the ranking reverses for the delivery axes (0.5827 against 0.5634) and ties for emotion. Two independent estimates of the adapter's own effect, through different code paths, agree to four decimals (+0.0878 vs +0.0879).

I.14.5 Combining the three

The levers are cumulative on the target and more than cumulative on the cost. Excess over the sum of the parts, over 60 attributes:

Super-additivity ([TR-0829] §24.5).
QuantityExcessRatio to sum-of-parts
target attribute+0.156 SD (t 3.21)1.24
word error+0.142 (t 15.8)1.91
genuineness loss−0.77 (t −17.4)1.80
duration error2.69

The super-additivity is specifically an emotion phenomenon (1.52); delivery axes are almost exactly additive (0.97) and the quality attributes saturate (0.42). The dominant interaction is steering × CFG (+0.208, t 6.30), and it is a wiring choice rather than a fact: steering both CFG branches instead of only the conditioned one costs 18 % of the effect and returns 0.209 of word error and 0.75 of genuineness.

The most useful single trick: subtract flatness rather than add feeling. Emotion and perceived authenticity are anti-correlated inside the representation at a cosine of −0.62 to −0.95, and causally a +0.10 emotion push costs −0.86 of genuineness out of 6. Pushing away from Emotional_Numbness at −0.10 returns +0.60 of genuineness (t 9.64, 67 of 80 prompts) at no cost in emotion when an adapter carries the emotion; on the steering-only version it removed essentially the whole cost (−0.86 → +0.04). It stacks with CFG at −0.05 (+0.68, t 10.94) but not at −0.10. Adding the genuineness vector back instead is the weaker but still favourable trade: 16 % of the emotion gain buys 56 % of the loss.

Two further results close open questions. Re-extracting the steering vector with the adapter loaded does not help: the adapter shifts the bucket mean by 3.8 % of its norm but shifts the shared reference nearly as much, so the direction survives at cosine 0.9958 — below the split-half noise floor at 60 of 60 attributes. It contributes a small correctly-signed extra only when the adapter is also loaded at generation (+0.095 z, t 2.16, against −0.041 without). Conversely steering is not made redundant by the adapter: it adds +0.864 SD (t 12.12) on top of the adapter-only cell.

Where the recipe lands. A balanced operating point exists for 53 of 60 attributes, a high-effect one for 56. Commonest balanced recipe: adapter 1.0 + CFG 2.0 (19 attributes); commonest high-effect: adapter 1.0 + steering +0.10 at one layer (12). Median balanced cell: WER 0.086, genuineness 3.92/6, blend 4.95/10, burst realisation 0.436, |duration error| 0.07 s. Two composition rules are measurements rather than preferences: ESTH and S_RANT cancel (ranting alone +0.464, t 7.01 on 12 of 12 prompts; with aesthetics added, −0.012), and the matched random direction is null alone (−0.033) but not null at the combined operating point (+0.106, t 2.78), so every recommendation must beat a random push of the same magnitude. The shipped agent chose adapter + CFG, not steering; the reason is the listening test of Part II.3.3.

I.14.6 A defect in merging adapters, found in the demo server

The demo server at LAION-AI/Humaneness-Voice-Demo-Server — the production home of the agent of Part II — loaded adapters by parsing the tensors itself and adding the product straight into each module's stored weight (lora_bank.py:274), undone by subtracting the recomputed product at end of turn (:305). PeftModel was never constructed. In this architecture that collides with tie_weights() (modeling_moss_tts.py:114): text_lm_head.weight = transformer.embed_tokens.weight, and for every codebook head.weight = embedding.weight — the same tensor object. Every adapter targets audio_lm_heads.0…11, and those twelve tensors are audio_embeddings.0…11. Three consequences, none externally visible:

Fix: load through PEFT and leave unmerged, scaling by multiplying the stored scaling — identical in effect to a merge at that weight, with the base weights never written. Keeping a custom loader requires at minimum excluding audio_lm_heads.* from the merge and applying those modules as forward hooks, and restoring from a snapshot rather than subtracting. The state of the server's loader as of 9 September is recorded in Part II.2.

I.14.7 The stop-talking bug, and what it says about duration

The same server chased a different symptom: the model said its line and then kept going, a few often-invented words on every second or third reply. Their conclusion constrains the rest of this report.

The constraint this places on everything else: the model treats the duration budget as a contract, not a ceiling. The project's own duration errors are small because its budgets are tight, not because the model stops well — and any future work that loosens a budget should expect filler rather than silence. The server's extra_w metric, counting words after the last one that still matches the reference, is a more direct measure of over-generation than anything in the harness and should be adopted.

I.14.8 The word-error instrument

Every WER in I.8–I.13 comes from openai/whisper-large-v3-turbo. From 2026-08-29 the primary instrument is nvidia/parakeet-tdt-0.6b-v3, a transducer with no language decoder, chosen on the theory that Whisper would hallucinate fluent text on over-driven audio and therefore flatter exactly the failures these levers produce. The theory was wrong, and measuring it is the useful part. Pooled offset −0.0177 over 29,676 clips (independently −0.0237 and −0.0231 in the other two studies), and the offset does not grow with guidance: slope −0.00081 per unit g, t −0.19, n = 1,863 cells, with the flattered share flat between 3.4 % and 4.5 % at every guidance value. Whisper is the harsher instrument here, worst where the audio is worst — on the most broken clips it returns mean WER above 1.0, hallucinated insertions rather than the plausible substitutions that were feared. The Whisper thresholds of I.8–I.13 were not too permissive and those tables stand as written. Parakeet remains primary for a different reason: on identical audio it is 2.4× less noisy, a real gain in resolution for any harness running few samples per cell. It is the WER ear of the agent.

I.14.9 Two things nobody had measured on 29 August

Blending two emotions. Nothing in any of the studies tested what happens when you ask for two feelings at once — bitter but amused, frightened but trying to sound calm — or for a point somewhere between them. What was known constrains how it should be attempted, and rules out the obvious version: adding two scaled directions is not interpolation (anger at 0.10 plus genuineness at 0.10 realises 0.1604 at the layer they share, because the two directions are not parallel; the correct construction is a spherical interpolation between the two normalised vectors, then a single α); some pairs are destructive (ESTH × S_RANT: +0.464 alone, −0.012 together); the response is not uniform across emotions (Sexual_Lust +0.487 and Affection +0.476 at the free operating point, but Fatigue_Exhaustion −0.113 and Doubt −0.070); and the two tails of one attribute are orthogonal, not opposite (median cosine −0.0004). The cheapest version is not a vector at all: the corpus captions already contain blends — a real GENERAL: line reads “affect is mildly positive … reads as amusement, teasing, interest” — so blending at the prompt level is in-distribution and blending at the vector level is not. As of 9 September this is still unmeasured (IV.2).

Spectral content. Every score in this project comes from a model, and not one of them would notice a gentle low-pass. The MOSS tokenizer is a residual quantiser: codebook 0 carries the coarse structure and each of the eleven that follow carries the residual of the ones before it — sibilance, breath, the band above roughly 6 kHz. Those deep channels are the most fragile: predicted last, conditioned on every earlier channel of the same frame, weakest signal, least conspicuous errors. Every adapter targets all twelve audio heads and stacked adapters sum their deltas linearly, so the load on those channels grows with the number of adapters — the same quantity the demo server measured driving word error from 0.013 to 0.258. A band-energy ratio — 6–16 kHz against 0–6 kHz on a fixed line — would be a few lines of code and would close the hole. Still unmeasured (IV.2).

I.14.10 What the lever study does not settle

I.15 The trajectory corpus

Every measurement in Part I is of one utterance in isolation. That is the right unit for “can this model perform contempt at the 95th percentile” and the wrong unit for almost everything a director actually asks for, because direction is about change: start polite and let it curdle, come down off it slowly, angrier with every sentence. None of those is a single target percentile, none can be scored by running an instrument on one clip, and nothing in this project has ever trained on or evaluated one. A trajectory is the smallest object that makes that measurable: an ordered list of clips from one speaker in which one measured dimension moves monotonically. Because every clip in the corpus already carries the full 40-emotion and 57-VoiceNet annotation of I.6, such sequences can be mined rather than recorded. Source: [TR-0829] §17 and its companion page.

The trajectory corpus ([TR-0829] §17).
Full corpusPublic subset
Trajectories10,653,7137,638,961 (71.7 %)
Distinct clips referenced24,922,34317,722,101
Speakers1,507,0641,214,241
Referenced audio114,920 h74,858 h
Clips per trajectory, mean3.523.51
Contiguous in the source59.3 %82.7 %
File size581 MB419 MB

A row is a specification and contains no audio. Twenty-four columns: the ordered clip identifiers, the speaker they share, the dimension being walked, the distance covered and the per-step size, the monotonicity and consistency thresholds and the step count, quality bounds, total duration, whether the clips are adjacent in the source recording, and selection bookkeeping. Roughly 55 bytes per trajectory, which is why 7.6 million of them fit in 419 MB.

Trajectory families ([TR-0829] §17).
FamilyRuleWhat it walksFullPublic
emotionB1one emotion head, monotonically5,069,8203,545,392
voicenetVN1one VoiceNet dimension3,072,0002,160,000
proxy_tailliftPXRa proxy dimension into its tail1,260,345977,981
emotion_twosidedAB2two emotions at once, in opposite directions1,251,548955,588

The two-sided family is the smallest and the most interesting. A single-dimension walk asks the model to do more of something; a two-sided walk asks it to trade — and every acting note about a turn, a realisation or a change of heart is a trade rather than an intensification. Trajectories are short: two to five clips, and the step histogram sums exactly to the row total in both versions, so there are none longer. The public release is a clean removal of three source datasets — podcast, snippets and evasnippets — whose transcripts are not part of the open real-speech release. A trajectory row names exactly which segments were used, which is a weaker disclosure than publishing the transcripts but is not nothing. The exclusion of snippets in particular was a cautious default rather than a settled decision.

Nothing has been trained or evaluated on it. The corpus was built and published on 28 August 2026 as laion/moss-va-trajectory-corpus (public subset, CC-BY-4.0; the unfiltered corpus is retained privately); its value is entirely unmeasured as of 9 September. It is the most direct available attack on the standing hypothesis of I.9.11 — that the model cannot learn what an intensity word controls because it has seen too little contrast. A trajectory is contrast by construction: the same speaker, at several points on the same axis, in order. Whether rendering one as a single training example teaches anything is the obvious next experiment and has not been run (IV.2).

I.16 Prehistory kept for the record: the DramaBox/LaionBox audio DiT, the search agent, and the first evaluation stack

The 16 August reference report ([REF-V1]) covered two model families under one name, and later reports dropped the second. It is kept here because the brief for this report forbids losing anything, because two of its findings (the bf16 ULP floor and the rank-buys-identity rule) are the same lessons the MOSS line learned independently, and because “DramaBox” names both the data source of voice-acting v1 (I.3.2) and this unrelated model — a name collision worth defusing once. Nothing in this section is a MOSS number, and the MOS/quality tables must not be read as such.

I.16.1 The DramaBox / LaionBox audio DiT (LTX-2.3 22B)

A different model: the LTX-2.3 22B audio-only diffusion transformer (base ltx-2.3-22b-dev-audio-only-v13-merged, 48 blocks, inner dim 2,048, 3.286 B trainable parameters) — a flow-matching speech enhancement / re-synthesis model, physically located in Voice-Acting-Pipeline/. HF lineage: laionbox-v0.1v0.3v0.5v0.6-wip. Source: [REF-V1] Part II, itself compiled from that project's protocol.md §7, §10.2, §11, RUN12_ADALN_STATUS.md, run9/run10_training_log.txt and eval_output_combined/eval_full_report.html.

DiT campaign, runs 5–17 ([REF-V1] §17).
RunMethodDiTlrEpochsHF lineage / outcome
5AdaLN-Zero speaker conditioning, DiT frozen bf16 → became the full-FT “quality champion”frozenadaln 5e-55laionbox-v0.3-wip
6Standard full FT, fp32 master, no AdaLNtrainable2e-61did not converge
7LoRA rank 64 (113 M, 3.2 % of DiT)LoRA4e-52hit the bf16 ULP floor
8Frozen LoRA-merged DiT + fresh AdaLN-Zerofrozenadaln 7e-56
9LoRA rank 128 (226 M, fp32 master) + pretrained AdaLN-ZeroLoRAdit 4e-5 / adaln 1e-53
10LoRA-only baseline r128/α128, pure flow, native ref-condLoRA1e-43v0.5-wip (step 1479)
12Frozen (run11-merged) DiT + AdaLN-Zero, 4 aux lossesfrozenadaln 5e-510best = step 210
13Frozen DiT + AdaLN-Zero (gentler MOS-push)frozenadaln 2e-52best = step 60
15LoRA-only r256/α256LoRA1e-45v0.6-wip
16v0.7 mixed 70 % German-Emolia / 30 % diverse, r256/α256LoRA1e-4, ga1620 M samplesin progress on 16 August; result never recorded in any later report
17rank-32 LoRA on a 5,596-sample high-emotion subset (from run16-merged)LoRA1e-410

Authoritative final evaluation — 72 samples (5 speakers × 6 prompts × 2 seeds, 3 EN + 3 DE), the same sidon_normalized references for every model. Metrics: MOS / UTMOS / NISQA / ScoReQ (quality), CE / PQ (AudioBox aesthetics), spk (WavLM-base-plus-sv cosine). Higher is better. n = 72 per row; no confidence intervals were recorded.

DiT final evaluation, raw output, n = 72 ([REF-V1] §17).
Model (raw output)MOSUTMOSNISQAScoReQCEPQspk
Run 5 — full-FT DiT+AdaLN (v0.3)4.6013.6474.3694.6366.1887.8960.879
Run 10 — LoRA r128 baseline (v0.5)4.4813.4484.2014.5005.9827.7930.876
Run 12 — frozen DiT+AdaLN+4aux (best)4.3783.3954.0574.4286.0237.8580.894
Run 13 — frozen DiT+AdaLN gentle (best)4.4343.4984.1384.4715.9977.7890.880
Run 15 — LoRA r256 5 ep (s2465)4.4913.4464.1994.5126.0007.8410.916
Run 15 (s2440)4.4743.3854.1844.4876.0007.8510.919

SIDON post-processing lifts every row by about +0.15–0.2 MOS (Run 5 → 4.747). ChatterboxVC → SIDON pushes speaker similarity highest: Run 15 r256 reaches spk 0.947 (against 0.920 for the full FT). Campaign verdict: (1) the full fine-tune (Run 5 / v0.3) is the quality champion at every stage (raw MOS 4.60 vs ≈4.48 for LoRA); (2) higher LoRA rank (r256) buys speaker similarity, not quality (raw spk 0.92 vs 0.88; VC→SIDON 0.947 vs 0.92); (3) AdaLN-Zero + 4 aux losses did not beat plain LoRA, and aux over-training degrades quality (Run 12 at step 1700: MOS 3.62); (4) SIDON is a reliable +0.15–0.2 MOS post-processor.

I.16.2 The first evaluation stack, the 99-vector scorer and the search agent

[REF-V1] §16 records the evaluation stack of the pre-JUPITER era, which is the ancestor of the agent's Ears: WER from Parakeet-tdt-0.6b-v3 (transformers ParakeetForTDT, fp32) — so Parakeet was the first ASR instrument, was replaced by Whisper-large-v3-turbo for the timed-script round, and came back as primary on 29 August (I.14.8); va_rescore.NinetyNineScorer — 40 EmoNet emotion experts + 57 VoiceNet regression dimensions + genuineness (slot 97) + blend (slot 98); torchmetrics DNSMOS for quality; ECAPA-TDNN and Orange/Speaker-wavLM-tbr for speaker similarity; laion/vocalburst-locator for bursts; and a Gemini LLM judge (gemini-3.6-flash via a gateway), audio-in, 0–10 verdicts.

The reward-design default of that era, user-specified: fitness = (Σ wᵢ·norm(targetᵢ) + w_g·norm(GENU) + w_b·norm(BLEND)) × (1 − WER). The (1 − WER) factor is mandatory for speech missions (it curbs overacting); GENU + BLEND together carry about the same total weight as the targets. All fitness is measured over a cohort (n = 8 default), never a single sample. Compare the agent's R_old of Part II.6 — the same three ingredients, plus CLAP, with the same (1 − WER) multiplier.

The “same audio twice” bug (found and fixed). The MOSS codec decode returns stereo (two identical channels); flattening with reshape(-1) concatenated both channels in time, so every saved sample contained its content twice back-to-back, saturating WER (≈ 1.0, the ASR heard the text twice). Correct downmix: w.mean(0). Always run a 50 ms-envelope self-repetition check before publishing audio players. This is the ancestor of the half-speed stereo bug that later cost a whole corpus (III.2.6).

Is the VoiceNet representation good? — the EXT-MLP benchmark. On the ≥ 2-rater VoiceNet-EXT cut (13,917 items / 395 prompts) the ordinal-regression MLP predictors top balanced accuracy per prompt, ahead of both VoiceCLAP contrastive readouts: VoiceCLAP-Small (110 M) contrastive cosine 0.6367 / per-prompt ρ 0.1051; VoiceCLAP-Large (7 B) 0.6510 / 0.1475; VoiceNet-Dim predictors (MLP) 0.6693 / 0.1317. Source: [REF-V1] §6, from the paper-rebuttal numbers in WIP docs/08.

The autonomous LoRA + prompt search agent (LAION-AI/voice-acting-search-agent) searched LoRA merges and prompts against these rewards on a single A100-80GB. Overnight programme (1–2 August 2026): 163 agent generations, ≈ 10 k scored takes. The Gemma-4-31B QAT brain won decisively (10/10/9 on single-dimension missions, only 22 tool calls; the 12B arm “wins” raw scores only by paying WER). Supervisor: Gemini was an effective active director (tracks the objective at r = 0.72, gives sonic-diagnostic feedback); a local MOSS-Audio-8B judge was only usable as a lenient score-only gate (active steering was harmful). The Part II agent's choice of a Gemma-4-12B local Mind is a cost decision made with this result in view; the 31B/12B gap was never re-measured on the new instrument (IV.2).

I.16.3 The gotchas list of 16 August, kept whole

The canonical documentation package of that era was the private LAION-AI/Voice-Acting-Pipeline-WIP (docs/01–23 + docs/LEARNINGS.md + code); the DiT project lived in Voice-Acting-Pipeline/. Both links were dead (404) by 26 August ([TR-0826] §18.2).

Part II — The actor: Mind, Instrument, Tools, Ears, Selection

II.0 The metaphor and its technical counterpart

Part II walks through the actor one part at a time. The project lead gave this undertaking a framing narrative, and it is not decorative: it is the architecture. An actor has a mind that reads a scene and decides how to play it; an instrument that responds without wanting anything of its own; tools with which the actor works on the instrument and produces several takes; and ears with which the actor hears those takes and selects one. The table below maps each part of the metaphor to its technical counterpart and to the section of this part that measures it.

The metaphor and its technical counterpart
MetaphorTechnical counterpartIn this report
Actor's Mind — the brainThe language model. It is given a task or a scene and plans: it writes the prompt from it. In the demo server this is the “director” model.II.1
Actor's Instrument — intuition, System 1The text-to-speech system itself, SFT3. It has no intent; it responds to the prompt.II.2
Actor's Tools — craftThe LoRA adapters, CFG/guidance and the various prompt forms. With these the actor works on the instrument and produces several takes.II.3 and II.4
Actor's Ears — perceptionVoiceCLAP specialists, the dimension predictors, genuineness, vocal-burst blend, EmoNet with 40 emotions, speaker similarity and the word error rate with Parakeet v3. From these comes the selection: best-of-N.II.5 and II.6

II.1 Actor's Mind — the directing model

The actor's mind is given the situation and writes the prompt from it. In the running server this is a language model that delivers four things in a single constrained pass: the reply (reply), the delivery prose (deliveryGENERAL), the script with the bracketed cues (scriptSCRIPT), and two real tool calls — the choice of reference voice and the choice of generation mode. This is enforced through a JSON-schema grammar, not through a second call: a separate tool turn would have put a full prefill and decode onto the time to first sound.

II.1.1 The rule that holds the whole system together

The directing model chooses what should happen — never how strongly. All dosages are measured values and are set server-side, so that an invented number cannot get through. The steering strengths, too, arrive only as words (“soft”, “medium”, “strong”) and are resolved against a fixed table. The same holds for time: the timing script computes every duration, every pause and every burst length itself.

From this follows the most dangerous unwritten rule of the system: a parenthesis with a number is a sound, one without a number is an instruction. (chuckle) is a noise, (clearly amused) a stage direction — and (clearly amused, with a small chuckle) produces no chuckle at all, because the whole parenthesis is read as an instruction. The duration is mandated by the format and forbidden in the prompt. Whoever confuses the two silently turns every stage direction into a sound.

A second precondition that long went unspoken: a cue is written in English, even when the spoken line is German. The corpus is labelled in English; its German lines read literally Das zerreisst einen einfach, weisst du? (relief sigh). A German cue lies outside the distribution and drags the words down with it.

II.1.2 The chain a turn passes through

From input to sound there are eleven stages. At several points the order is itself a result and not a convention. The count of “eleven stages” comes from the German evaluation page ~/voice_acting_agent.html; in the source itself six section markers are set and the README describes four steps — so it is a reading of the code, not the code's self-description.

The eleven stages of _turn() (app.py)
#StageWhy it sits there
1Directing model writes reply, delivery prose and scriptone pass, JSON schema enforced
2Language decisionthe declared language field is overruled when the text looks German — a turn written entirely in German had once declared itself “English” and pulled an English reference clip
3Cue anglicisationbefore retrieval, not after: with German cues the retriever falls back to the emotion *named* by the director, and that is how a horror scene came to be conditioned on Jealousy_and_Envy
4Select reference voice via codeslanguage, tempo, voice profile
5Retrieval overrides step 4searches for clip and emotion; the emotion is read from the brackets alone, because a long descriptive text otherwise drowns it out
6Build the reference stackanchor → tails of the last two turns → clip for this moment. Concatenating whole clips destroys the voice (speaker similarity falls from 0.777 to 0.280 by the fourth clip), hence only 4-second tails
7Adapter planwhich adapters, at which weight (II.3)
8Script repaira burst that was only mentioned *inside* a stage direction gets its own parenthesis — otherwise no sound is produced at all
9Burst weights and budgetcap per adapter and cap on the sum; over budget is scaled, not dropped (II.7)
10Lever selectionmust run after 7, because it can remove adapters again
11Generate, select, play outbatch path for best-of-N, guidance or SIDON; otherwise streaming. Speaker similarity comes last of all, after streaming, so that it delays nothing

II.1.3 What is actually measured about the mind

And here stands the most uncomfortable single result of the whole programme. The written stage direction moves nothing. In a fully balanced grid of 1,440 clips — 6 emotions × 5 utterances × 2 languages × 2 punctuation variants × 2 direction variants × 6 voice conditions, plus a 240-clip control — the direction changes the emotion score by -0,002 [-0,023, +0,020] (paired, 720 pairs). Not the genuineness, not the blend, not the burst realisation, not the word error rate. This is the factor around which the entire prompt surface of the third round was built.

The obvious objection was checked and does not hold: had the emotion adapter already saturated the score, the direction could add nothing more. In the control with the adapter at weight 0 the effect of the direction remains non-significant (+0.022 [−0.031, +0.077]), and switching the adapter off lowers the emotion percentile only from 0.864 to 0.821. Neither the adapter nor the direction produces the high scores.

What produces them is the text: 35,1 % of the variance of the emotion score is explained by *which sentence* it is. The utterances were written to match their emotions, and the classifier hears content, not only delivery. From this follows a consequence that reaches beyond this study: as long as the evaluation uses emotionally coloured text, no training method can show an effect — not because it does not work, but because the instrument is saturated before the model speaks.

Two things the mind very much does move come from the work on the demo server and so far appear in our protocol only as a reference: an example moves a language model, a rule does not. The rule “put pauses in the middle of the sentence” took effect roughly zero times per reply; rule plus one example roughly once; rule plus one example per feeling plus an explanation of the tempo value one to three times. And 19 prompt additions from an evolutionary search over 2,949 scored clips changed nothing — the first run flattered itself; corrected, the control wins. The take-to-take spread of the same cell there is SD 1.57 out of 15, so that a best-of-3 fitness buys +2.78 points from luck alone.

What was measured. Four-factor grid: 1,440 clips, 0 discarded, plus 240 control clips; paired main effects over 720 pairs each with 95 % bootstrap interval. Stage direction -0,002 [-0,023, +0,020] on the emotion percentile (protocol §18). Share of the variance explained by the choice of text: 35,1 %. Language is the largest effect in the grid: English +0.129 [+0.105, +0.153] over a German translation of the same sentence, at a lower word error rate. Whether the model is less expressive in German or the classifier less sensitive in German, this design does not separate.

II.2 Actor's Instrument — SFT3

The instrument has no intent. It is given a prompt and answers. The basis is laion/moss-tts-local-transformer-4.55b-voice-acting-v2 — a semantic transformer with around 4 billion parameters over 36 layers, a “talker” transformer with around 550 million and twelve audio heads, together 4.13 billion trainable parameters. The audio tokenizer is MOSS-Audio-Tokenizer-v2, 12 codebooks of 1024 at 12.5 frames per second: one frame is exactly 80 milliseconds.

II.2.1 The prompt surface, in three rounds

Round 1 carried the text in plain form, preceded by a delivery instruction in parentheses. Round 2 introduced the timing script — every speech duration, every pause from 0.20 s and every burst length are computed and written into the prompt — and dropped the delivery instruction without replacement in the process. That went unnoticed at the time and is the central negative result of that round: giving an instruction afterwards to a model that has never seen one destroys it. With a forced instruction the word error rate of all round-2 models lies at 0.48–0.51 against 0.11 without, including bare SFT-2 without any adapter.

Round 3 brought the instruction back, this time *inside* the timing script. That fixes the collapse: 0,447 -> 0,099 word error rate, with better timing fidelity and better burst realisation at the same time. Emotion control does not improve. That is the dividing line along which this programme has run ever since.

II.2.2 What the instrument can do

Two things that may count as solved
QuantityFindingEvidence
Total durationfully controlled via the Tokens budget: duration error 0,00 s in every cell of both jobs. No adapter is needed for this.protocol §39.3
Timing fidelity along the linemedian |requested − produced| 0.000 s in 13 of 14 arms of the transition study, regression slope 0.979. The only exception is the spliced arm at 0.690 s — and precisely that only this one deviates rules out a measurement artefact: if the programme were cutting to the target, it too would sit at 0.000.protocol [P§37]

II.2.3 What it cannot do

Tempo. A tighter budget forces faster speech (slope ≈ 0.66), but there is a floor at about 18 cps: below it the model does not slow down, it fills with silence. In the arm in which the budget was held fixed and only the tempo label was varied, all four slopes are zero — telling the model 14 or 26 characters per second changes nothing, it speaks 18. That holds for the adapter that was trained on exactly this, too.

Why this could not work is, in hindsight, arithmetic: the character rate of the corpus was computed from the printed duration and the printed text. In every training line the rate agrees with the duration, and the duration is in the prompt — the model never had a reason to read the rate. The label is an invertible paraphrase, not a second control channel. 121,940 lines and 1 h 47 min of training bought no new capability. The adapter is harmless and useless; stage 2 sits safely on top of it, but it must not be described as tempo control.

Emotional intensity on command. Three objectives were set on it — GRPO with group-relative reward, DPO with contrastive pairs, DPO with symmetric, instruction-conditioned pairs — and none moved it. The only method that has moved it is the emotion LoRA on the filtered top one percent (0.3572 against 0.3494 baseline): more extreme data, not a cleverer loss function.

II.2.4 The training path, in one table

The runs that shaped the instrument (protocol [P§4]–[P§17])
RunWhat it was meant to doWhat came out
SFT round 2learn the timing scripttiming solved; delivery instruction lost; emotion at 0.49–0.51, the corpus median, although percentile 0.90–0.98 was requested
GRPO runs 1 and 3optimise rewardreward fell in both. Two independent defects found: a prompt–reward conflict (the caption names the genuineness of the source recording — exactly what the reward optimises) and a DDP rank divergence that brought one run to a standstill with 35 GB of core files
SFT round 3instruction inside the timing scriptcollapse fixed (0,447 -> 0,099), emotion unchanged
GRPO runs 4 and 5four levers against the emotion weaknessv5 is the worst of the four models on emotion. The training gain was the mechanical artefact that had been predicted: a ramp that pays for partial progress supplies a gradient and removes the incentive to arrive
DPO round 2 and v2preferencesrew_chosen never turns positive; the policy pushes the *chosen* side down while it learns the ranking. Preference accuracy is not the quantity to steer by
CFG-DPOintensity via symmetric pairsthe best general model on reward, word error rate (0.0950 — better than bare SFT-3), quality and burst hit rate (0.772 against 0.709) — and it did not improve precisely the thing it was built for. The symmetry that prevents a drift also prevents the push
A side finding that affects the readability of all DPO numbers: in LoRA mode the reference distribution comes from peft.disable_adapter(), which switches off all adapters. rew_chosen therefore measures against bare SFT3 and not against the run's starting point. rc > 0 means “better than bare SFT3” and is thus a weak guard rail, not a strong one.
What was measured. Timing fidelity: 1,596 clips of the transition study, median duration error 0.000 s in 13 of 14 arms, slope 0.979, mean signed error −0.045 s. Tempo: two jobs, 240 + 160 cells, 8 texts × 5 requested rates × 3 samples, paired on a shared seed; interaction (rate − duration) = −0.09 cps; in the held-fixed arm all four slopes zero within their standard error. Emotion: 80 shared prompts, strict criterion, four models within 0.015 of one another.

II.3 Actor's Tools — adapters, guidance, prompt forms, steering

The actor works on the instrument with craft. For us that means three kinds of tool: LoRA adapters (small additional weights), guidance (the difference between “with instruction” and “without” is amplified) and the prompt forms (how a request is phrased). A fourth, the steering vectors, was built, measured — and switched off again after a listening test; it therefore gets a section of its own (II.3.2).

II.3.1 The shipped stack

What is actually inside the model when a reply is generated is not a question of documentation but of configuration. The following table is read at render time from $SC/hv_mirror/config.py.

The stack, read from config.py of the running server
AdapterName in codeWeight
Preference adapter (DPO)sft3_dpo:p21.00
Voice adapter (voice profile)sft3_voice1.00
Quality: genuinenesssft3_quality:genuineness_high0.25
Quality: burst blendsft3_quality:blend_high0.50
Quality: aestheticssft3_quality:esthetics_high0.50
Second preference adapter (quality DPO)sft3_qdpo:quality_dpo1.50
Emotion adaptersft3_emotion1.00
Sum of the permanent adapters5.75
Delivery axes (VoiceNet), at most 1 of themsft3_voicenet:…≤ 1.50
Sum before the first burst5.75 to 7.25
Burst adapters, up to 3 of themburst…:<class>each ≤ 1.25, sum ≤ 2.00

That is seven permanent adapters, i.e. 5.75 to 7.25 of merged weight, before the first burst adapter is added. This number is not a side issue; it is the actual experimental setup, and II.7 shows why.

Two precisions for verification. First: the effective voice weight is PROFILE_LORA_LAM (1.00); the constant SFT3_VOICE_LAM (1.00) sits in the same file and is read by nothing — dead configuration with the same numerical value. Second: config.py still carries, above BURST_LAM_BUDGET, the refuted comment block (“two adapters cost nothing over one … it is the weight of a single adapter that destroys a line, not the sum” (translated)) and below it the correct version. Whoever reads from the top reads the refuted justification first.

II.3.2 Guidance (CFG) — the tool that needs no training

For every frame the model is run twice: once on the full prompt, once on the same prompt with the condition removed. The audio logits are extrapolated away from the unconditional branch: logits = logits_uncond + g · (logits_cond − logits_uncond). At g = 1 the unconditional term cancels exactly — that is a genuine control which runs through the same code path, including both KV caches.

The entanglement is at channel level, not at frame level. A frame consists of twelve autoregressively sampled audio channels; both branches must continue with the same sampled token for channel *c* before *c+1* is predicted. Otherwise the difference measures how far the two branches have drifted apart, not the condition. The end-of-sequence decision is taken only from the conditional branch: it is structural, and amplifying it would change the clip length.

II.3.2.a On the emotion condition

The first lever in the whole programme that moved emotional intensity at all. 96 clips, 24 utterances × 4 strengths, on the three emotions furthest from saturation: emotion 0.809 (control g = 1.0) → 0.887 (g = 1.5) → 0.917 (g = 2.0) → 0.912 (g = 3.0). Paired against the same prompt-and-seed combination: +0.078 · +0.108 · +0.103, with 95 % intervals that all contain zero.

24 paired comparisons per strength are a feasibility proof, not a measurement. The direction is the same across all three strengths, and that is the reason to pursue it — the honest wording is “promising and not yet significant”. Timing fidelity survives intact: mean |duration error| 0.010 s at every strength, unchanged from the control. That was the main risk of an intervention at channel level, and it did not materialise.

The large follow-up study over 7,419 cells / 29,676 clips puts the magnitude firmly into perspective and confirms the direction: for emotion the best value under the word-error bound is g = 3 with +0.0254 (t +2.18, 63 of 120 improved), for the VoiceNet axes g = 2.5 with +0.0569 (t +3.12, 57 of 87). Below g = 1 guidance actively hurts: −0.0370 at g = 0.5 (t −2.56). These two measurements of the same object differ by an order of magnitude because they were measured from different baselines (0.809 versus 0.263) and one takes three hand-picked emotions, the other all forty. To this day they have not been merged into a single number, and that stands in the paper outline as an open item.

II.3.2.b The price, measured five times independently

The expectation was well below double, since only the semantic transformer runs twice. The expectation was wrong. The local “talker” transformer runs twice as well — twelve times per frame per branch, because the unconditional branch needs its own local state for the next channel's logits — and so do the twelve audio head projections. Only the codec decoding is shared, and that accounts for 1.6 % of the work. Sharing it saves nothing.

Five independent measurements of the same multiplier
MeasurementMaterialMultiplier
CFG study (origin)instrumented, batch 1 and 41.889× (batch 1) / 1.947× (batch 4)
CFG study, reported rounded7,419 cells1.93×
Crossfade v2, paired over arm × text127 cells, batch 4, one GH2001.92× median (51.5 → 98.6 ms/frame)
Crossfade v2, per condition36 arm × g cells1.90–1.94×, identical at g = 2, 3 and 4
Burst CFG, smokewith adapter switching per branch1.97× (arm E) / 1.98–2.00× (arm D)
Burst CFG, main run624 cells, 1,871 clips1.91–2.04×, mean 1.94×

The price is the second branch, not the strength. So everything that argues for guidance at all argues for running it at the best point of the band rather than the smallest. In the server it therefore does not run streaming: the realtime budget there is 1.0 and the merged baseline 0.764; 1.93× lands at ~1.47 and the player starves. A guided take is rendered in full and then played. The automatic lever selection therefore never emits guidance.

II.3.2.c On the burst — the clearest guidance finding of the programme

Here it is not the emotion condition that is neutralised but two axes at once: the unconditional branch receives a neutral prompt and has the burst adapter switched off. The amplified difference vector thus carries the prompt cue and the adapter delta together. For this to be implementable at all, the adapter state had to be switchable inside the twelve-channel loop — the twelve audio heads are LoRA targets themselves. The index is therefore split: 252 modules in the semantic transformer (once per branch and frame) and 16 in the channel loop, so that 16 rather than 268 numbers are rewritten there. Cost of the switch: below the measurement noise.

Pooled over 3 classes × 2 positions, n = 48 paired prompts (the three takes of a cell are averaged first)
Armstrictfamily-relaxedWER (Parakeet)
A — no adapter, no cue0.0000.0000.165
B — cue only0.0070.0070.148
C10 — current best case0.0420.0830.181
C15 — adapter at 1.50.0760.1880.197
D2 — both guided, g = 20.1250.2360.255
D4 — both guided, g = 40.1390.2920.404
D5 — g = 50.1040.2010.488
E2 — prompt only guided, g = 20.1040.2150.167
E4 — prompt only guided, g = 40.0900.1810.204

Paired against C10: D4 +0.0972 strict (t +2.72) and +0.2083 family-relaxed (t +4.41) — but at +0.2235 word error rate (t +5.63). E2 +0.0625 / +0.1319 (t +3.25) at ΔWER −0.0141 (t −0.77) — i.e. for free. And the decisive decomposition: D beats E on the strict metric at no g (t 0.72 / 0.94 / 1.41) and pays +0.09 to +0.20 word error rate for it at every g. The prompt carries the larger part, not the adapter delta.

The one exception is exactly where the adapter can do something: on scream / inline — 496 verified training lines — D4 − E4 is +0.333 strict (t +3.74), and the class goes 0,208 -> 0,583. That is the largest movement any lever of this programme has produced on a burst class. shriek (48 lines) and frustrated_groan do not exceed 0.083 strict under any arm. Guidance amplifies a direction that exists; it does not create one.

The wiring control is not optional and was measured rather than assumed: at g = 1 the guided path must reproduce the single-branch path exactly. Largest deviation |guided − conditional| over every channel of every frame: max 9.5·10⁻⁷, mean 1.1·10⁻⁷; the results are cell-identical, D1 − C10 = E1 − C10 = 0.0000 on every metric. Everything above g = 1 is therefore guidance and not wiring.

II.3.2.d The preference pairs from guidance contrasts

$SC/cfg_rows — 128 shards, counted at render time
FamilyRowsShare
cfg_high237,20950.0 %
cfg_low237,20950.0 %
language de247,35052.1 %
language en227,06847.9 %
total474,418100.0 %

The third application of CFG is a training corpus: contrast pairs that differ only in how strongly one dimension is expressed, with the spoken text removed so that the preference cannot go through the words. Every pair is emitted twice with swapped roles: the instruction “high” makes the intense clip the chosen one, the instruction “low” the mild one. Both clips appear as chosen and as rejected — so the model cannot win through a global preference for intense audio, but must condition on the instruction. The two halves are balanced to 0.1 percentage points; in the table above they are exactly the same size.

And exactly this symmetry is the reason it did not work. The trained model became the best general model of the programme — on reward, word error rate, quality and burst realisation — and did not move emotional intensity. The gradient towards intensity under “high” is mirrored by an equally large gradient away from it under “low”; the net movement of the unconditional distribution is zero by construction. Correct for a conditioning signal, useless for raising a ceiling. That the construction took hold is attested independently: preference accuracy on the CFG families rose from 0.562 (chance) over the first 40 logged batches to 0.975 over the last 40.

Emotion: 96 clips, 24 paired prompts per strength (feasibility test); follow-up study 2,473 conditions × 3 prompts × 4 completions = 7,419 cells, 29,676 clips. Burst: 624 cells / 1,871 clips, n = 48 paired prompts, wiring control max 9.5·10⁻⁷. Cost: five independent measurements, 1.89–2.04×, identical across g. Corpus: 474,418 preference pairs in 128 shards. Not measured: whether the guided path changes the best-of-N numbers of II.6 — see II.6, “the largest open risk”.

II.3.3 Steering vectors — the lever that won in the measurement and failed in the ear

A steering vector adds a direction directly into the hidden state: h ← h + α · unit(v) · ‖h‖. The direction is a difference of means — the mean of the activations over the clips that show an attribute strongly, minus a shared reference set. Where these vectors come from and in which layer they sit is described in II.4.

II.3.3.a What was measured

The full 2×2×2 of the three levers, target in standard deviations, paired over prompts — computed from combination_study/stats/analysis.json
Familyadapter onlysteering onlyguidance onlyadapter × steeringadapter × guidancesteering × guidancen
Emotion (40)+0.077 (t +2.81)+0.384 (t +9.41)+0.050 (t +1.84)+0.038 (t +1.36)−0.031 (t −1.22)+0.277 (t +7.51)399
Delivery axes (17)+0.377 (t +8.66)+0.614 (t +6.15)+0.026 (t +1.13)−0.164 (t −3.75)−0.125 (t −3.50)+0.144 (t +2.02)170
Quality (3)+0.399 (t +3.26)+0.006 (t +0.04)+0.032 (t +0.34)+0.259 (t +2.05)−0.060 (t −0.63)−0.334 (t −2.24)30

Three rules follow from this, and the server implements exactly them. Emotion: adapter and steering are cleanly additive — the interaction is indistinguishable from zero — and steering there is roughly five times the adapter. Delivery axes: significantly sub-additive — the two levers do the same thing, so exactly one runs. Quality: steering moves nothing, the adapter is the whole lever.

The quality null finding pools 30 cells — three attributes at ten prompts. It agrees with two independent observations (the quality directions collapse from k ≥ 2, and the geometry of the representation says the same), so the *direction* of this statement is safe. The *precision* of this null is not.

The ladder itself is the most interesting part, because it was once read completely wrong. At α = 0.10 on layer h20 the emotion percentile moves from 0.4354 to 0.5840, paired +0,1486 (t 3,17) over 24 prompts, while a matched random direction of the same length moves only +0,0135 (t 0,48). The word error rate falls in the process from 0.166 to 0.115, and the share of clips under WER 0.2 rises from 0.84 to 0.91. The price is paid monotonically in genuineness: 3.77 → 3.64 → 3.15 → 2.92 → 2.41 → 1.90 up the α ladder, while the random direction holds 3.64–4.18.

II.3.3.b And then a human listened

The shipped mode is GEN_MODE = "adapter". The automatic lever selection was once the default and was withdrawn: it chose steering in almost every turn, and a listener described the result as *“more emotional, with strange artefacts, and the timbre off”* (translated) — even though the scoring models were satisfied.

This finding is documented in five places ([P§40]; ~/reports/ensemble_protokoll.md; ~/reports/crossfade_v2.md; ~/wikiskills/REVISION_2026-08-31.md; ~/wikiskills/REVISION_2026-09-02.md). It is probably the single most important lesson about the distance between the agent's “ears” and real ears. And it carries not a single number.

Provenance of this finding, stated explicitly. There is no quantitative record of this listening test: no clip count, no listener count, no comparison protocol, no date of the session. All sources write “one listener” in the singular, without n. The only primary source named is docs/FIELD_NOTES.md of the demo server, which is not in this home directory. It is the project leader's statement from listening, and it is carried here as what it is — a judgement, not a measurement. It remains remarkable nonetheless, because it is the only lever a human has ever checked.

Two qualifications that sharpen the finding rather than explain it away. First: the server ran with α = 0.10 under a ceiling of 0.15 — i.e. inside the range the measurements call good. The report is therefore not an artefact of the far tails and does not disappear when α is lowered. Second: the listener's session concerned the emotion axes, and there the measured effect is +0.10 percentile, i.e. almost nothing — almost everything that is audible there is therefore artefact. The open question is narrow: does steering on the delivery axes sound right at α ≈ 0.1? Nobody has listened. Until someone does, GEN_MODE = adapter stays.

The consequence reached far beyond the default. When the study on emotional transitions was repeated, five of its nine arms were dropped without replacement — they *were* the steering schedule. What remained in the second version crossfaded adapter merge weights instead of steering strengths, and the result was unambiguous: not a single crossfade beats a random reordering of its own weights. The best value of a crossfade against its own shuffle control is +0.044 with 14 of 28 texts improved, p = 1.000 — as exact a null as the design can produce. The two-part cut that is already shipped, by contrast, beats them at all four guidance levels (p = 0.013–0.036).

II.3.3.c Two hard refusals that hold regardless of the setting

A third, smaller result of the same family is in the configuration: two adapters can cancel each other out. esthetics_high and S_RANT_high work individually (+0.20 to +0.32 and +0.464, t 7.01, 12 of 12) and together not at all (−0.012). The aesthetics adapter is therefore set to 0 server-side when ranting and halved on the neighbouring axis.

Factorial: 10,459 records, 31,377 clips, 60 attributes, all eight corners on the same prompts and seeds; n = 399 / 170 / 30 paired prompt-attribute cells per family. α ladder: 11 configurations, 12 emotions × 2 languages = 24 prompts, 4 samples, 264 cells = 1,056 generations, with a matched random direction at every strength. Controls of the factorial: adapter scaled to zero against the bare vector cos 1.00000; the same base condition from three independent SLURM jobs, 360 cross-comparisons, max. difference 0.000000. The listening finding carries no n — see the box above.

II.3.4 Prompt forms — the tool that costs nothing

The third kind of tool is the wording. It is cheap, it needs no training and no second forward pass — and across the whole burst programme it is the most reliable lever we have found. Fifteen classes, in the server's stack at its own doses, 5,360 cells / 32,160 clips over two jobs, paired over ten prompts per class; the pooled rows run over n = 15 classes.

The prompt forms, paired against the same strength and the same seed
FormΔ hit (family)Δ hit (strict)Δ genuinenessΔ WERΔ blend
cause sentence in the GENERAL line+0.026 (t +1.9, 9/15)+0.002 (t +0.3)−0.035 (t −1.4)−0.023 (t −1.3)−0.02
longer announced duration+0.022 (t +1.1, 6/15)+0.023 (t +1.7)−0.065 (t −2.0)+0.095 (t +5.9, 15/15)−0.57 (t −5.0)
both together+0.044 (t +1.9, 10/15)+0.030 (t +3.3, 9/15)−0.110 (t −3.8)+0.075 (t +5.4)−0.51 (t −6.5)
burst as an *action* instead of a sound−0.077 … −0.106 (t −2.1 … −3.3)rises +0.183
label mid-sentence instead of at the sentence boundary−0.073 … −0.124
asking for a neighbouring class as substitute−0.000 … −0.022−0.021 (t −2.9, 1/15)

The first two forms add up cleanly: +0.026 and +0.022 give +0.044; +0.002 and +0.023 give +0.030. Together they are the only prompt result of this study that is significant on the strict metric. The costs add up with them.

Two forms hurt, and both in an instructive way. The action form ((sie stöhnt frustriert, 0.5 Sekunden) instead of (frustrated groan, 0.5 seconds)) drives the failure rate up by +0.120 (t +3.4, 11 of 15): the model degrades to silence, not to a wrong burst. The model was trained on labels from the detector's vocabulary; an actor's direction in that field is not a label it knows. And at the same time blend (+0.183) and DNSMOS (+0.032) rise while the burst stops happening — the clearest case in the programme of a metric moving in the right direction for the wrong reason.

Mid-sentence placement drives the failure rate up by +0.306 to +0.366 (t +8.4 to +10.4), and all fifteen classes get worse. Roughly a third of all mid-sentence bursts simply do not happen. Read the other way round, that is a *positive* result about the shipped prompt: the sentence boundary is worth about 0.3 in failure rate. One class reverses itclears_throat goes from 0.250 to 0.467 mid-sentence, its best cell of all. Placement is a decision per class.

A prompt form measured at one dose is a statement about that dose. The combined form (+0.044) was measured in a block in which every class ran at its own best strength — typically just above the previous recommendation, i.e. exactly where the form pays off. Measured over 55 further classes at the same strength it is a null (−0.004, t −0.6; better for 13 of 55), and on the strict metric a null at every position (|t| ≤ 0.9). It is a property of the pair of form and dose, not of the form.

The longer-duration effect also has an uncomfortable origin, which was cleanly separated. Over 15,409 cues where anything at all was detected, the probability of a family hit rises with the detected duration — 0.282 (under 0.3 s) → 0.388 → 0.476 → 0.510 (0.8–1.2 s), then it saturates. Against an announcement 1.0 s longer, however, the delivered burst grows by only +0.054 s (t +2.3). The detector's slope times 0.054 s predicts +0.020; observed is +0.012. No residual remains that one could attribute to better realisation — the gain belongs to the measuring instrument, not to the model.

The twelve best of the 31 recipes with a recommendation, parsed from ~/wikiskills/VOCAL_BURSTS.md. N is the candidate count for a 90 % success probability; N cons. uses the hit rate minus one seed spread (0.068).
ClassArmWeightPrompt formHit (fam.)conservativestrictNN cons.
relief_sighvb0.80GENERAL-Ursache0.750.680.0023
chucklevb2.00GENERAL-Ursache + längere Dauer0.730.660.3823
contented_sighvb1.50GENERAL-Ursache + längere Dauer0.680.610.6833
soft_humvb2.00GENERAL-Ursache + längere Dauer0.680.610.3033
wistful_sighvb1.00Grundform0.550.480.0734
cacklevb1.80Grundform0.540.470.0234
clears_throatvb1.25Etikett mitten im Satz0.480.410.0045
low_mumblevb2.00GENERAL-Ursache + längere Dauer0.480.410.4545
childlike_gigglevb0.80Grundform0.470.400.0545
exasperated_sighvb1.25Grundform0.450.380.0545
breathy_gigglevb2.00Grundform0.400.330.2356
sharp_inhalevb2.30Grundform0.380.320.3857

In total 31 of 117 burst pages carry a recipe; 19 classes realise not a single hit at any weight and under any prompt form, and a further 17 stay below the threshold. The pattern is families, not classes: every mouth class and every whistle class of the bank is on one of the two lists. They are not weak adapters that would need a larger weight — the sound is missing from what the model can produce, and the way out is data, not a dial.

A contradiction that stays and is not resolved. 18 of the 31 parsed recipes name a weight above the shipped ceiling of 1.25 (up to 2.30). The two numbers are not comparable: the recipe weights come from a sweep over prompt forms in which the absolute word error rate of every cell was checked individually and passed; the ceiling comes from a paired mean over a different carrier set. Discarding measurably good recipes on the weaker evidence would be wrong — so the contradiction stands in the knowledge base, and an environment variable enforces the stricter rule.
Prompt forms: 5,360 cells / 32,160 clips, 15 classes, two jobs, paired at prompt level, pooled over n = 15 classes. The base prompt P0 was checked byte for byte against the prompt set of the predecessor study before GPU time was spent — only for that reason are absolute numbers comparable between the two studies. Seed floor: the best cell of every class was repeated under a second seed; mean difference −0.010 (t −0.6), but sd 0.068, largest single deviation 0.167. Every paired comparison is immune to this (both arms share the seed); every absolute recipe number is an argmax over roughly 35 such draws and thus about one standard deviation optimistic. That is why the table shows, next to every point value, the value at p − 0.068.

II.3.5 What an adapter actually moves — the combination study

From ~/wikiskills/coefficients.json (as of 2026-09-02), computed at render time
FamilyAttributeswith balanced operating pointwith high-effect pointwith usable adapter weight
Emotion40343711
Quality3331
Delivery axes (VoiceNet)17161616

The message of this table is the refusal: where a column says “none”, the measurement says that there is no usable setting. A consumer can no longer tell an interpolation from a measurement once it has been written down — so it is not written down.

The large dose–response study over 79 adapters × 7 merge weights (5,740 cells, 17,220 clips) found 33 of 79 adapters with a usable weight. Its thresholds are derived, not chosen: first the noise floor was measured (every w = 0 cell holds several independent clips per prompt without adapter, and any difference within it is pure sampling noise), two independent estimates agreed to within 4 %, and the threshold is then max(3σ, a substantive lower bound). The price of this is stated openly: a threshold below the noise lets honest adapters fail by chance — at 2σ that would be about 40 false failures over ~1,800 tests, at 3σ about two. Conversely, a real degradation *below* 3σ goes undetected, which is why the table says “not detectable at this resolution” and never “zero”.

Pooled per family, far better resolved than a single adapter
Familyw = 0.250.50.751.01.251.5
Emotion, Δ emotion percentile−0.0010 (t −0.16)+0.0146 (t 2.16)+0.0178 (t 2.46)+0.0352 (t 4.82)+0.0311 (t 3.85)+0.0326 (t 4.38)
VoiceNet, Δ own axis+0.0828 (t 2.59)+0.2157 (t 5.56)+0.3144 (t 6.16)+0.4481 (t 6.71)+0.5554 (t 6.02)+0.6957 (t 6.18)
Quality, Δ own target+0.0862 (t 2.98)+0.1818 (t 1.46)+0.4029 (t 1.67)+0.3696 (t 3.74)+0.4395 (t 5.01)+0.5720 (t 3.50)
Burst, Δ realisation of the class+0.0146 (t 1.11)+0.0212 (t 1.25)+0.0521 (t 2.73)+0.0843 (t 3.58)+0.0916 (t 2.18)+0.1165 (t 2.03)

Two readings sit in it. The emotion adapters are cheap, weak and stop acting at w = 1.0 — 1.5 minus 1.0 is −0.0026 (t −0.40, with only 20 of 40 adapters higher), while 1.0 minus 0.5 brings +0.0206 (t 3.65). The delivery axes are what works, monotone up to 1.5, 17 of 17 adapters improve from w = 0.5 on. And the only reliably monotone cost quantity is not naturalness but speaker similarity: 0.548 → 0.484 (emotion, t −8.3, 36 of 40 worse) and 0.559 → 0.475 (VoiceNet, t −8.5, 17 of 17), while genuineness actually *rises*. Turning an adapter up does not make the voice less genuine, it makes it less like the speaker that was asked for.

What was measured. Dose–response study: 574 conditions × 10 prompts = 5,740 cells, 17,220 clips, 12 nodes / 48 GPUs, seed depends on the prompt index alone (every level draws the same sampling noise), t over n = 10 prompts, never over clips. Controls: w = 0 through the same code path (largest deviation on any metric 0.00e+00), wrong-adapter control (+0.002 against +0.035 for the matching adapter at the same weight), half-speed canary on every line. What this does not show: the adapters were swept individually; nothing in this table predicts anything about a stack — see II.7.

II.4 Layer forensics — where in the network each attribute lives

The steering vectors of the previous section do not come from nowhere. They are the by-product of a piece of work that answers a question of its own and stands on its own: at which point in the network is which property of a recording readable at all?

II.4.1 How it was measured

The regime is the decision everything hangs on: teacher forcing, no free generation. One forward pass per corpus line, with that line's own MOSS codes on the assistant side, packed byte for byte as in training; the hidden states are captured by hook. Activation and label therefore describe the same waveform, and every target — the 40 emotion heads, the 57 VoiceNet axes, genuineness, blend — has already been measured on exactly this waveform.

Capture happens at 38 taps: h00 is the sum of the embeddings before layer 0, h01 to h35 are the outputs of decoder layers 0 to 34, h36 is the normalised output of the last layer — that is the 80 ms vector that goes into the acoustic decoder, because the global → local mapping in this checkpoint is the identity (checked at runtime) — and loc is the state of the single-layer talker transformer at the frame slot, i.e. the vector the twelve audio heads actually read. Without loc, h36 would have been a duplicate.

The data holdings ($SC/out/actforensics/release.json)
TagrequestedextractedfailuresGBpurpose
main58,80058,377423 too long (0.72 %)11.36stratified probing set
p3219,684217,5832,101 (0.96 %)42.33union of all phase-3 buckets
pairs53,32553,325010.3720,000 same-sentence groups
neutral58,80058,71486 (0.15 %)11.42main with the emotion condition removed
Total75.49
The control that carries this entire piece of work: the neutral extraction runs the same work list with the emotion condition removed — the reads as … clause, the emotion sentence of the GENERAL line and every stage direction in round brackets are gone; words, durations, pauses and burst tags remain. That costs at most 0.015 R², under 3.4 % relative on every task (emotion −0.0140, VoiceNet −0.0021, genuineness −0.0129, blend +0.0039). The probes read the audio, not the words.

II.4.2 What is readable where

One joint probe with 99 outputs per tap (not 99 individual probes), p3, speaker-disjoint, n_test 32,104
Familynbest tapR² (MLP)Ridgemetadata baselinetaps within 0.01
VoiceNet, speaking styles S_*15h120.8590.8070.3569
VoiceNet overall57h120.8460.7880.2807
Genuineness1h120.7190.5390.1906
Emotion (40 heads)40h200.4730.4370.1084
Vocal-burst blend1h250.4350.3560.1004

All 99 attributes are readable to some degree — the weakest, Jealousy_and_Envy, still reaches 0.209, against a metadata baseline of −0.003. Best readable are GEND 0.939, REGS 0.924, S_MONO 0.920; worst Jealousy_and_Envy 0.209, Awe 0.239, Thankfulness_Gratitude 0.281.

Two findings alongside, both of which count for the outlook in II.4.6. Speaker memorisation is practically absent: the difference between the random and the speaker-disjoint split is −0.0035 (emotion) and −0.0026 (VoiceNet), i.e. within the seed noise and slightly negative. The exception is genuineness at +0.042 — genuineness is partly a speaker property, and a genuineness predictor on activations is in part a speaker lookup table. And the MLP beats the ridge on all 99 attributes, median +0.043, the largest gap of all at genuineness with +0.180. That is the measurement which pulls the ground from under the premise of a *linear* steering direction: where a probe finds 0.18 R² beyond a linear map, a single direction from that layer will fall short of the probe.

The tie rule. The same probe, the same data, two seeds, over 891 pairs of individual R²: mean 0.00308, p95 0.01084, largest value 0.0432. Two taps less than about 0.01 apart are tied, not ordered. For most attributes the “best layer” is a formality on a plateau many layers wide.

II.4.3 Reading several layers, and the controls for it

Over all 38 taps, ensembles of size 2 to 6 were searched — greedy, beam and random, six attribute groups, two extractions: 4,888 fitted combinations in 13.1 GPU hours. The hidden layer stays fixed at 512, so that only the input projection grows with k.

k = 1 versus k = 6, with the group's seed band alongside
Groupbest single tapk = 1k = 6gain (4 seeds)own seed band ±above the better control
Emotion 40h200.46630.5067+0.03850.0037+0.0414
VoiceNet 57h120.85290.8695+0.01660.0003+0.0168
Speaking styles 16h120.85180.8701+0.01820.0005+0.0152
time-independent 19h120.84030.8581+0.01880.0018+0.0156
Blendh330.38410.4551+0.06160.0554+0.1240
Genuinenessh230.70860.7388+0.01810.0244+0.0677

The seed bands are the real statement here. Blend and genuineness have the largest raw gains and the largest error bars, and their gains lie at or inside those bands; emotion and VoiceNet have small gains far outside theirs. A table cannot be read at two confidence levels without saying so. Both capacity controls — the best tap duplicated k times, and the best tap plus k−1 blocks of N(0,1) noise, each at identical input width — are beaten by all 60 winners.

Selection ran on validation, never on test, even at k = 1. What that costs is directly measurable on one group: on blend the test-selected combination beats the validation-selected one by +0.0138 — pure selection overfitting. On the five other groups the same gap is 0.0000 to 0.0005.

Two side findings worth having. Redundancy does not predict complementarity: the best emotion pair h18 + h21 has a representational similarity of 0.979 and still gains +0.020. And the value of the greedy search lies in reaching across all 38 taps: h06 alone ranks 13th of 38 and would not have made any pre-selection — it enters at k = 4 and stays through k = 6.

II.4.4 What follows for vectors — and the standing puzzle

From the same activations the vector library was built: archetypes.npy, 156 × 38 × 2560 float32, from 309,631 clip means. Deliberately float32, because a steering direction is a difference of means — one to two orders of magnitude below the means themselves. The verification is strict: all 97 pre-existing dimensions × 38 taps × 3 buckets were recomputed from the raw data and agree bit-exactly (max |d| = 0).

Three measurements on this library decide how it may be used. First: the two ends of an attribute are orthogonal, not oppositecos(high − mid, −(low − mid)) has a median of −0.0004 over 2,166 pairs. “Halfway between anger and calm” is not “anger at half strength”. Every attribute steerable in both directions needs two stored vectors, and pushing down must select the low row rather than negate the high row.

Second: cos(dimension, quality axis), signed
Tap40 emotion heads57 VoiceNet axesmean cos between the dimensions
h12−0.619+0.4600.050
h20−0.746+0.4610.037
h32−0.842+0.3830.046
h36−0.862+0.3830.046
loc−0.954+0.7060.038

Emotion and perceived quality are nearly antipodal directions in the model, and they sharpen towards the output. Pushing an emotion up pushes perceived quality down in the representation itself — and the α ladder of the previous section turns that into a measured causal price: genuineness 3.77 → 1.90 up the ladder. The rightmost column is the control that makes this readable at all: it shows that these are 97 different directions and not one direction with 97 labels.

Third, and it limits the library: 34 of 156 rows carry a warning flag. 32 of the 40 emotion heads have a low end that is a tie at zero — more than one percent of the corpus sits at raw value exactly 0, 32 % in the median, 82.2 % for Awe — and for 25 of them the low end and the middle select the same clips. “Negative anger” is not defined by this library.

II.4.4.a The puzzle that stays open

Readability and leverage point in opposite directions. h36 — the final norm — and loc — one layer before the decoder — are the places where the information must be if the model is to act on it. They are at the same time the weakest taps for emotion: 0.342 and 0.339, against 0.473 at h20. The emotional information is clearest sixteen layers before the point where it would be needed.

Two readings fit equally well, and telling them apart is the most valuable open experiment of the programme. Either the model computes the emotion in the middle and then discards most of it — that would be a defect one could fix. Or the late layers compress a rich description down to the few dimensions the codec needs — that would be normal and would mean: steering must happen early or not at all.

II.4.5 Vocal bursts by layer — and the limit that determines everything

In p3, 79,402 of 217,583 clips (36.5 %) carry at least one burst; 118,660 bursts over 33 labels, median 0.28 s. The presence of a burst is best readable at h19 — ridge AUC 0.9777, balanced accuracy 0.9232 against a chance rate of 0.500 and a metadata baseline of 0.587; the identity at h18, mean AUC 0.871 over the eight classes with enough material. Both plateaus lie in taps 13 to 23 — in the emotion band, not in the VoiceNet band, while the blend scalar sits six taps deeper at h25.

Window coverage is the dominating constraint, and it is itself a finding. The window averages the last 20 % of the assistant span (median 2.41 s). A burst counts as inside the window if at least half its duration lies in it — that holds for only 18.4 % of the 118,660 bursts, and 24.7 % of the burst-carrying clips survive. The share depends strongly on the class: Yawn 49.0 %, Scream 37.0 %, Chuckle 24.1 %, Contented Sigh 8.7 %, Sharp Inhale 8.0 %. Sighs and inhales open a line, laughter and screams close it. Four of five bursts have never been seen by a probe on these taps.

The controls are unusually complete and deserve mention, because they show how to rule out a clip-level confound. The falsification — clips whose bursts all lie *outside* the window — reaches 0.778, i.e. 0.200 below the headline: that is the floor produced by the property “this clip contains a burst somewhere” alone, and exactly why the window restriction is not optional. The sharpest test — burst-carrying clips on both sides, matched on burst count — still separates at 0.968. And strictly same-speaker negatives give 0.969 against 0.978: speaker identity is not the signal.

II.4.6 Outlook (PROPOSAL): a perception module inside the instrument instead of beside it

This section is a PROPOSAL, not a result. It is written down separately so that nobody later cites it as a measurement. What is attested in it is in II.4.2 to II.4.5; what is missing is listed below in full.

The observation that carries the proposal: the probes of II.4.2 are already exactly the object in question. They are predictors on the mean of a hidden layer — not on audio — and their target is the output of the audio-side scorers themselves. The measured R² is therefore directly the share of an audio scorer's judgement that can be reproduced from activations: speaking styles 0.859, VoiceNet overall 0.846, genuineness 0.719, emotions 0.473, burst blend 0.435, and burst presence as a classifier AUC 0.978. With ensembles over several layers these values rise to 0.870 / 0.739 / 0.507 / 0.455.

A perception module that sits inside the instrument instead of beside it would build on exactly that: predictors for genuineness, vocal-burst blending, the 40 emotions, the speaking styles and speaker embeddings, read from the hidden layers while generation is running. The appeal is not the accuracy — the audio-side scorers remain the reference — but that such a module need not decode any audio: it could judge within generation rather than after it, and thereby make abort, resampling or a selection among candidates cheaper than the current chain allows.

What is already there, and could be cited in a proposal as preliminary work: the speaker-disjoint split is verified (0 of 13,757 test voices leak); the caption confound is ruled out (≤ 0.015 R²); the resolution limit is known (seed noise p95 0.011); the non-linearity is quantified (MLP over ridge on all 99, median +0.043); the capacity controls for multi-layer inputs are in place; and the costs are measured — the four extractions together needed under one GPU node hour for 388,000 lines at 75.5 GB of output, the probing 26 and 13.1 GPU hours. A single 99-output probe is thus in the range of minutes.

What is missing, in full and without embellishment. (1) Free generation. Every R² in this work is teacher-forced. A module in the generation path would see activations from free generation; that probe accuracy transfers there is not measured. That is the largest gap. (2) The usable taps of all things are the weakest. h36 and loc would be available at runtime at no extra cost and sit at 0.342 / 0.339 for emotion. (3) No temporal resolution. Everything is a mean over the last 20 %; an attribute that changes within this window is averaged away before the probe sees it — and 81.6 % of the bursts lie outside. The window has to move before bursts can be studied this way. (4) Speaker embeddings are not measured at all. There is no number in this work for how well speaker identity is predictable from activations; the nearest approximation is GEND at 0.939, and that is an instrument property, not an identity. (5) No inference cost of a probe head measured. (6) No use as a reward signal, and hence no test of whether an activation predictor can replace an audio scorer in the training loop. (7) No transfer across checkpoints — everything hangs on SFT-3 plus one particular DPO adapter. (8) For blend the floor is not settled: one evaluation names h25, the other h33, and the difference of 0.082 R² lies far outside any noise band. A blend predictor would stand on undecided ground.

A ninth point belongs here, because otherwise it makes the idea look too good: such a predictor would be a third model judge in a chain that already consists of models scoring one another. It also inherits both things measured in II.4.4 — the coupling to the quality axis and the finding that a random direction of equal length raises the emotion score more than the real vector does, once one turns the dial too far.

Extraction 388,000 lines / 75.49 GB / 38 taps, 0 duplicate uids, finite = 1.000000, no zero rows. W1: n_test 32,104, speaker-disjoint, metadata baselines and seed noise (p95 0.011) placed alongside. W2: 4,888 fitted combinations, 13.1 GPU hours, two capacity controls at identical input width, selection on validation. W3: 156 × 38 × 2560, recomputed bit-exactly against the raw data, 34 of 156 rows flagged as unreliable. W4: grouped 5-fold cross-validation on voice_key, five contrasts of 16,000 each, falsification control at 0.778. Everything teacher-forced — a high R² means the attribute is readable *while the model follows a real recording*; it is no evidence that the model uses this information when generating.

II.5 Actor's Ears — the perception module

What does the actor listen to his own takes with? Not with one model but with close to a dozen, and they have different jobs. The project leader explicitly asked for this module to be described: which models it consists of, what each of them measures, how they are combined — and what the selection actually achieves with them. The last point gets a section of its own (II.6).

A division of labour that is easy to confuse. EmoNet and the VoiceNet heads listen to the user: what comes out of the microphone is summarised and enters the directing prompt as [heard in their voice: …]. Genuineness, Blend, CLAP and the word error rate listen to the model's own output and rank the candidates. Speaker similarity runs last of all, after streaming, so that it delays nothing — by construction it can therefore take part in no selection.

II.5.1 The instruments, one by one

Each row: what the model measures, on which scale, how well it was measured, and the known flaw
Instrumentmeasuresscale / outputmeasured qualityknown flaw
VoiceCLAP (laion/voiceclap-commercial)Text–audio similarity: does this recording match this stage direction? Also the tower used for retrieval.cosine, mean-centred on both sidescentring lifts the hit rate from 0.35 to 0.61 (text) and from 0.22 to 0.44 (audio); the encoder comparison is in II.5.3no class judgement — it compares against the *whole* direction text. An enthusiastic line with the brackets “sharp inhale, delighted laugh” came back as *jealousy*, because the long descriptive text drowned them out; since then the emotion is read from the brackets alone.
Genuineness head (laion/voiceclap-commercial-genuineness)How much does this sound like a real, lived moment rather than read aloud or synthetic? Not audio quality.continuous 0–6frozen 768-d tower → MLP 768→50→1, ≈38k parameters; ~10,000 clips from various TTS systems plus Emolia, each labelled 0–6 by Gemini-3.1-Pro; held-out 140 clips (20 per level, stratified): MAE 1.00 · Pearson r 0.77 · RMSE 1.33the backbone is never fine-tuned; and the value drops on a successful burst — exactly what becomes the problem in II.6
Vocal-burst blend head (laion/voiceclap-commercial-vocalburst-blend)How naturally does a non-speech sound fit into the surrounding speech? 0 also means: no real burst present at all.continuous 0–10identical architecture to the Genuineness head, the same single embedding yields both values; validation MAE 2.057 · correlation 0.625 (against small 2.360 / 0.418 and large-v2 1.895 / 0.645)it measures whether *a* sound fits, not whether it is the right one. The model card names no training-set size, no label source and no split description — this must not be filled in
EmoNet / Empathic Insight Voice Small40 emotions in the user's voice40 head outputs, used as percentiles in the rewardreproducible from hidden layers at 0.473 R² as the target of the activation probes of II.4the percentile scale was not comparable across heads — see the journey, III.4
VoiceNet dimension predictors57 delivery and instrument properties (tempo, tension, register, style axes …)one regression each, order of magnitude 0–10reproducible from hidden layers at 0.846 R²; individual axes up to GEND 0.939EXPL is content explicitness and only 3-level, BKGN 5-level — two axes that are not delivery axes and had no business in a bucket set
Burst detector (laion/vocal-burst-detector-x2)which vocal burst lies in this span?17 classes = 16 bursts + no_burst, chance 0.05885 members averaged in probability space; see II.5.2a classifier over spans, not a localiser — the span still comes from the old localiser, only the *name* changes
Speaker similarity (ECAPA)does the voice stay the same?cosinethe only reliably monotone cost of adapter dose: 0.548 → 0.484 (emotion) and 0.559 → 0.475 (VoiceNet) from w 0.25 to 1.5runs after streaming and therefore enters no selection; the thresholds in the handbook (“regenerate below 0.58”, translated) are documentation, not behaviour
Parakeet-TDT-0.6b-v3word error rate — is the line still intelligible?WERon identical audio 2.33x less noisy than Whisper large-v3-turbo (sd 0.1340 against 0.3115; paired noise floor 0.0346 against 0.0833)see II.5.4 — the change of instrument and what it really bought

II.5.2 The burst detector and its recall floor

Only the encoder is swapped: same segments, same split, same head, same seeds. Held-out, balanced at 25 per class, chance 0.0588 — from $SC/out/vb_merge/encoder_compare.json
EncoderLicenceParametersDim.17-way, exact (real)23 groups (real)17-way (DramaBox)
voiceclap-large-v2CC BY 4.07 B + LoRA35840.4660.6050.574
voiceclap-commercialCC BY 4.0110 M7680.3930.5150.536
voiceclap-small-v2CC BY-NC 4.0110 M7680.4020.5130.551

Two things follow, and the second is the more useful one. The 7-billion encoder buys about 0.07 exact accuracy for roughly 64 times the parameters — worth it for a scorer that runs offline, probably not for something that runs per request. And the non-commercially licensed small encoder buys nothing over the commercial one — the differences lie within the noise of a 425-sample cell. That is a licensing result, not an accuracy result: there was no accuracy price for the commercially usable choice, so the licence question never had to become a trade-off.

The head that counts today is a 17-way classification — 16 burst classes that occur at least a hundred times in both halves of the corpus, plus a no_burst rejection class. Chance is 0.0588. What ships is an ensemble of five initialisations within *one* grouped split — never an ensemble across split seeds, because its members would have trained on each other's test data, and the reported held-out numbers would be partly memorisation.

The number that must be read before any hit rate. A measured hit rate measures two things at once: that the generator produces the sound, and that this detector can name it. They multiply. A class the detector recognises 20 % of the time cannot show a hit rate appreciably above 20 %, however good the adapter is. Without the recall column a mute generator and a deaf instrument are indistinguishable.
Recall per class, from production/per_class_recall.json — the six weakest and the three strongest on real speech
ClassFamilynstrict (real)95 % intervalfamily-relaxedstrict (DramaBox)reliable
Relief Sighsigh320.062[0.017, 0.202]0.3120.469yes
Heavy Breathingbreath250.120[0.042, 0.300]0.6400.280no
Wistful Sighsigh280.143[0.057, 0.315]0.3930.400no
Deep Breathbreath590.186[0.107, 0.304]0.7120.686yes
Exasperated Sighsigh530.226[0.135, 0.355]0.3400.381yes
Soft Humhum280.393[0.236, 0.576]0.5710.280no
Affirmative Gruntgroan300.700[0.521, 0.833]0.7670.680yes
Sharp Inhalebreath390.897[0.764, 0.959]0.8970.531yes
no_burstno_burst8020.980[0.968, 0.988]0.9800.964yes

5 classes lie below 25 % recall on real speech — and what they have in common is the sigh-and-breath family, i.e. exactly the classes whose acoustic difference is duration and effort rather than timbre. For them a measured hit rate below about 20 % carries no information about the generator. Two of them arithmetically cannot break a threshold of 0.15.

That is why the family-relaxed column always stands beside it: the failure is granularity, not deafness. Deep Breath goes from 0.186 strict to 0.712 at family level, Breathy Giggle from 0.469 to 0.969; the confusions lie almost entirely within the family — Breathy Giggle → Chuckle, Exhausted Groan → Frustrated Groan, Deep Breath → Sharp Inhale, Humming → Soft Hum. A recipe that reliably produces *a* groan and gets the wrong name is a working recipe.

A claim that was repeated for days and was wrong. It was said that the shipped detector uses 8 of 83 labels and never emits Shriek. That came from an audit over 60 clips with 73 events in total — at that sample size and that skew, eight labels are exactly what one expects. Over the full corpora the same detector uses 41 distinct labels on the DramaBox half and 36 on the real one, and it does emit Shriek — on 50 DramaBox clips and 49 real ones. Its distribution is extremely concentrated (Contented Sigh alone on 16,694 of 72,500 DramaBox clips), which explains the 60-clip observation. The correct statement is: its *effective* vocabulary is small, not that its emitted vocabulary has eight labels. The correction is entered in four READMEs — and in the same report it still stands uncorrected a few sections above the correction.

II.5.3 A vocabulary problem that is brand new

The new head cannot simply re-score the old inventory. Measured on a sample of 61,579 lines with 12,894 burst events from three parquet shards: the old burst columns carry 21 distinct labels, of which only 10 lie in the new 16-label vocabulary, and 83.8 % of the existing events (10,811 of 12,894) carry a label that the new head cannot emit at all. The four most frequent old labels — Low Mumble 26.3 %, Ahem 26.2 %, Contented Sigh 20.5 %, Surprised Gasp 7.2 %, together 80.2 % — do not exist in the new vocabulary. In the other direction, six new classes occur not once in the old inventory of the sample.

The 83.8 % come from three of 16,600 parquet files. The three subsets cover about four fifths of the corpus, so the weighting is not implausible — but it is a sample, and the number is carried here with that caveat. The qualitative finding holds regardless: the re-annotation is a second opinion, not a replacement. That is why all new columns get their own prefix and stand beside the old ones; no existing burst column is touched.

II.5.4 The change of instrument, and what it really bought

Since 29 August the word error rate has been measured with nvidia/parakeet-tdt-0.6b-v3, no longer with Whisper large-v3-turbo. The stated reason was hallucination: Whisper has a language decoder and writes fluent text over audio that has stopped being speech — exactly what an over-driven adapter produces.

But the strongest argument was not hallucination, it was variance. On identical audio Whisper's per-clip word error rate has a standard deviation of 0.3115 against Parakeet's 0.1340 — 2.33x as noisy — and the paired noise floor is 0.0833 against 0.0346, i.e. 2.41 times. Measuring with Whisper costs about two and a half times the error bar for the same number of takes.

And the original worry pointed the wrong way. Three independent measurements of the offset agree in sign and size: −0.0237 (5,740 cells), −0.0231 (10,459) and −0.0177 (29,676). Parakeet therefore reports fewer word errors than Whisper, not more — on this model Whisper is the stricter instrument, and strictest where things are worst (on the most broken clips it returns a mean above 1.0, i.e. hallucinated insertions). The old Whisper thresholds were therefore not too generous. The offset moreover does not depend on guidance strength: slope −0.00081 per unit g (se 0.00436, t −0.19, n = 1,863 cells). A single pooled offset would still be the wrong object, because the two instruments diverge exactly where the audio falls apart — on 24 clean corpus clips they lie within +0.0080.

II.5.5 What the ears do not hear

Two measurements determine how far this module can be trusted, and both are uncomfortable.

A third case belongs here because it shows the same mechanism from the other side: DNSMOS must not decide between burst recipes. Its overall score moves by less than 0.04 for every lever, but it moves by +0.38 (t +16.7, 9 of 9) between two script kinds that differ only in how much speech the clip contains. On this material it measures the composition of the clip.

And the outlook from II.4.6 belongs here too: a perception module that sits inside the instrument instead of beside it — predictors for Genuineness, Blend, emotions, speaking styles and speaker embeddings, read from the hidden layers during generation. The groundwork exists (0.846 for VoiceNet, 0.719 for Genuineness, 0.473 for emotion, AUC 0.978 for burst presence, all speaker-disjoint and with controls); none of it is measured in the generation path, and the proposal is to be read as a proposal.

What was measured. Genuineness: 140 held-out clips, MAE 1.00, r 0.77. Blend: validation MAE 2.057, correlation 0.625; no training-set size and no label source in the card. Detector: 17 classes, chance 0.0588, five members, recall per class with Wilson intervals at n = 25–59; 9 of 17 recall rows carry reliable: false (n < 30 in one source) and have intervals ±18 points wide. Parakeet against Whisper: three independent populations, same sign and same size. Neither of the two predictor cards names a held-out agreement with humans — exactly the gap the pairwise comparisons of IV.1 close.

II.6 Selection — Best-of-N

From the ears a single decision emerges: the server generates a reply 8 times in one batched pass and keeps one of them. The selection is made by:

R_old = ( norm(Genuineness) + norm(Blend) + 2.0 · norm(CLAP) ) · (1 − WER)

The normalisation is a min-max within the candidate set, not against an absolute scale — the three scorers have incomparable ranges (0–6, 0–10, a cosine) and are not calibrated against human judgement. If all are equal, the term collapses to a constant 0.5. The CLAP term counts double, because it alone measures *whether this is the requested performance*; the other two measure whether the take is good at all. The intelligibility factor is the raw inverse word error: the earlier flattening above 0.85 covered six of eight candidates of one run, so the factor stopped separating exactly where the candidates lay closest together.

II.6.1 No term asks for the burst — and two pull against it

None of the four terms asks whether the requested vocal burst occurs at all. Two even pull in the opposite direction: Genuineness drops when a burst succeeds, and Blend measures how well *a* non-speech sound fits, not whether it is the right one. The final report on the server put the loss at up to +0.74 delivered hit rate on scream — and said explicitly that this number is computed, not measured. This section measures it.

Result 1 — it does select bursts, but weakly. 1,940 candidate sets, 15,517 candidates, production detector on voiceclap-large-v2, localised windows, strict class match
QuantityValuetn95 % interval
Hit @ rank 0 — the delivered take0.1531+18.721,940[+0.1371, +0.1691]
Hit @ mean — what a random pick delivers0.1183+25.671,940[+0.1093, +0.1273]
Lift = rank 0 − mean+0.0348+5.471,940[+0.0223, +0.0473]
AUC − 0.5, on mixed sets only+0.0976+9.51725[+0.0775, +0.1177]
Correlation within the set+0.075815,517[+0.0591, +0.0925]

So positive, not negative and not zero: best-of-8 under the shipped reward buys about 29 % more delivered sounds than a random pick. The premise holds — it is just smaller than assumed. Caution with the AUC: its n is 725, not 1,940, because it is only defined on the sets that contain at least one hit and one miss.

And in exactly the configuration the server really runs in, the effect disappears. On the six carrier voices with their own sft3_voice adapter — the case the server is always in, because otherwise it falls back to the speaker adapter — the lift is +0.0104 (t 1.51, n 1,164), i.e. indistinguishable from zero. On the four without a voice adapter it is +0.0714 (t 5.97, n 776). The pooled positive number is carried by the half of the material that lies outside the shipping configuration.
Which term sees the burst at all — partial correlation within the candidate set, n = 15,517
Termr95 % interval
Genuineness−0.0085[−0.0253, +0.0083]
Blend−0.0839[−0.1006, −0.0672]
CLAP+0.1482[+0.1317, +0.1646]
1 − WER+0.0388[+0.0220, +0.0556]
the proposed fifth term+0.3053[+0.2900, +0.3205]

Two corrections to our own assumption. First: Genuineness is not the culprit — its interval encloses zero. The measurement that had raised the suspicion (Genuineness 3.10 solo at w = 1 against 3.23 without adapter) is an effect between conditions, and Best-of-N never chooses between conditions — it chooses within a set in which all candidates share the same configuration. Second: it is Blend — of all things the term whose name promises the blending-in of vocal bursts actively works against it.

II.6.2 The cross-validation that closes the circularity

A detector that makes the selection, and the same detector that measures the result — that would be circular reasoning. Both encoders had already scored all 15,517 candidates anyway, so the counter-check costs zero GPU minutes: rank with one, measure with the other.

Δ hit rate at λ = 3 soft, measured crosswise. Absolute hit rates are not comparable between rows — each column has a different measuring device with different recall.
ranked withmeasured withΔ hittnretention
large-v2large-v2+0.0552+7.741,940diagonal
large-v2commercial+0.0572+8.701,9401.04
commerciallarge-v2+0.0387+5.751,9400.55
commercialcommercial+0.0701+9.801,940diagonal

The gain survives the crossing. Ranked with large-v2 there is practically no shrinkage; ranked with commercial about half of the apparent advantage is device-specific. Both crossings stay significant, so the effect is about the burst and not about the detector — but the diagonal number of the commercial head must not be cited, only its crossed one.

A second built-in trap was implemented as a refusal rather than a feature: calibrating the burst term to the recall of the class raises an error, because within a candidate set the recall is a constant and cannot change the ordering. A no-op that would have looked like a result.
Result 3 — the ceiling
Selection rulestrict hit rateshare of headroom
Random pick from the 80.11830 %
R_old (shipped)0.153113.4 %
R_new, λ = 3 soft0.208234.5 %
Oracle — the best of the 80.3789100 %

The oracle is the measured version of the wiki column *P(≥ 1 hit in N)*. The entire headroom is 0.2606 — and so the gap of +0.74 on scream estimated in the final report is reachable at no λ by re-ranking. That is the value of this measurement: it replaces a calculation with an upper bound.

II.6.3 The adapter arms, and a sign that flips

Each arm paired against ship, on (class, form, prompt index)
Arm − shipΔ deliveredtnΔ set meant
bestmem — borrow the strongest adapter of the group−0.0531−2.50320−0.0492−4.70
v2 — the newer adapter set−0.0312−1.42320−0.0453−4.45
shipset — a fixed set instead of per-class resolution−0.2200−4.36100−0.1212−4.62
cap100 — ceiling lowered to 1.0−0.0400−1.42200−0.0394−3.22
cap150 — ceiling raised to 1.5+0.0050+0.15200+0.0275+2.31
ship_vn — plus one delivery axis at 1.5−0.0250−0.68160−0.0219−1.66
w0 — no burst adapter at all−0.0938−3.96320−0.0867−7.78
bestmem is a clean negative under the production stack, and that is the sharpest single result of this run. A predecessor study measures for the same strategy +0.0762 (t 3.76) there, -0.0531 (t -2.50) here. Same adapters, same prompts, same metric, same seed rule — opposite sign. The difference is the stack: four duration adapters there against seven here. The bestmem recommendation in the knowledge base is therefore to be withdrawn or qualified with its stack.

The ceiling of 1.25 is roughly right. Lowering to 1.0 costs in the set mean; raising to 1.5 gains in the set mean and is practically zero on the delivered recording — the extra burst is generated and then not selected, so it reaches no listener. That is the thesis of this section in one example.

A delivery axis at 1.5 costs, paired over 160 sets, +0.0546 (t 6.91) Parakeet WER and brings +0.121 Genuineness (t 4.61) — more word errors than the burst adapter itself, and more than half of the headroom the +0.104 gate leaves.

Intelligibility is nowhere the problem. Every arm passes both bounds, and every one has a *negative* paired word error against its own w = 0 cell; absolute on inline 0.063 to 0.087 against a bound of 0.25.

II.6.4 The largest open risk: guidance

The guided control arm: 80 candidate sets, 640 guided recordings, 4 classes, inline only, guidance 3.0. Paired on the same keys, guided minus unguided
QuantityΔtn95 % interval
R_old, delivered hit+0.1250+2.2980[+0.0182, +0.2318]
Set mean (random pick)+0.0156+0.6780[−0.0297, +0.0610]
Lift+0.1094+1.9880[+0.0013, +0.2174]
Parakeet WER−0.0009−0.0880

Guidance does not produce more bursts — it makes the shipped reward better at finding them. The set mean barely moves, the delivered rate clearly does. A selection effect, not a generation effect, and it fits the term table: guidance amplifies the delivery condition, the burst bracket is part of it, so CLAP separates the candidates better.

And with that, on this subset, it eats the advantage of the new term: +0.1125 (t 2.39) unguided against +0.0125 (t 0.30) guided. The evidence is thin — n = 80, 4 classes, one form, the decisive difference at t 1.98 — and it flips no sign. But the unguided number must not be cited as the guided one. The entire grid was generated unguided, while BON_GUIDANCE stands at 3.0. Measuring the guided path cleanly is the next thing, and it needs a batched two-branch decoder.

II.6.5 The recommendation

The recommendation, measured crosswise — ranked with large-v2, measured with commercial, λ = 2 soft, window 1.0 s at hop 0.25 s, localised
QuantityValuetn
Hit rate of the delivered recording under R_old0.1242+16.581,940
Δ hit at λ = 2+0.0454+7.601,940
— relative to the previous rate+37 %
Price: Parakeet WER (bound +0.104)+0.0040+2.491,940
Price: CLAP agreement−0.0055−6.821,940
Price: Genuineness−0.0673−7.241,940

A fifth term in the reward. The gain is measured crosswise, so the circularity is closed; the relative number refers to the rate that the same independent head measures for today's reward.

Two precisions, so that the number is cited correctly. The +37 % refer to 0.1242 — the R_old rate measured by the *commercial* head — and not to the 0.1531 of the diagonal; the two columns have different measuring devices with different recall. And the number comes from the row “ranked with large-v2, measured with commercial. The recommendation to rank with commercial in production (the tower is loaded for retrieval anyway) has its own crossed λ2 number, and that is +0.0314 (t 5.10).

This is a trade, not a free gain, and the traded quantity is CLAP: −0.0055 is 3.6 % of the within-set span of CLAP (0.1507), i.e. of the share that a re-ranking can move at all. λ = 2 rather than λ = 3, because the marginal rate breaks there: λ1→λ2 buys 3.82 hits per CLAP unit, λ2→λ3 only 2.09 and λ3→λ5 only 1.16.

Not recommended: the stepped variant (costs ΔWER +0.046, t 3.55, against +0.006 for the soft one, and brings 0.002 more hits), the agreement add-on term (moves 0.002 over the whole sweep — a clean negative) and bestmem. No recipe changes: the change concerns the selection, not the weights.

What was measured. 1,940 candidate sets, 15,517 candidates, N = 8 per prompt, 16 classes × 2 script kinds × 10 prompts, carrier set byte-identical to three predecessor studies. The seed depends on the prompt index alone, all arms draw identical sampling noise; paired on an explicit key, not on list position. Localisation fell back to the global window on 2 of 15,517 candidates. What was not measured: the total budget of the burst adapters (all 1,400 carrier scripts carry exactly one burst, so the bound never binds); the guided path at scale; and nobody has listened.

II.7 The babble bug — the fault that named the shipped configuration

A benchmark sentence — a parent shouting a child away from the window — came back as babble, word error rate 1.82. The obvious suspects were all wrong: not the brackets in the script, not the duplicated scream instruction, not the wrongly chosen emotion adapter. Four lines settle it:

Isolation experiment, one seed per line (docs/BABBLE.md)
ConditionWord error rate
bursts in script + two burst adapters at 1.5 each0.520
bursts in script, no burst adapter0.000
no burst in script, two adapters at 1.5 each0.840
bursts in script, one adapter at 1.50.000

One adapter at 1.5 destroys nothing, two at 1.5 each destroy the line — regardless of whether a burst was requested at all. Replicated over five seeds: 5 von 5 Seeds, Median-WER 0,82. The dose table over six seeds per cell, in the full shipped stack, shows the cliff between 2.5 and 3.0: sum 1.0 → median word error 0.000 · 1.5 → 0.000 · 2.0 → 0.010 · 2.5 → 0.110 · 3.0 → 0.820.

It was the sum of the adapters, not the bracket tags. And the reason the previous justification in the code claimed the opposite and cited a real measurement for it is the stack: that measurement had measured burst adapters on a bare model. Here 5.75 to 7.25 of weight already sit in front of them.

What followed from it (commit 08caa4b “Fix the babble”, follow-up commit 15c9a8c)
ConstantBeforeNowReason
BURST_LAM_BUDGET3.02.00last value that is clean on every arrangement checked
BURST_LAM_MAX1.51.25cap per single adapter
behaviour over budgetdropscalea scream that is silently not loaded is worse than one at two thirds weight — and dropping made the result depend on the order of the brackets in the script

Three further faults were found in the same pass, and the second is the most instructive. The double scream: the guard against duplicate insertion recognised only brackets *without* a comma and *without* a digit. (scream) matched it, (scream, 0.7 seconds) did not — and it is exactly this second form that the prompt has demanded since 5 September. From that day on, every director who wrote the burst correctly had a second bare (scream) inserted. End trimming showed the wrong copy: the candidates were encoded before trimming, and only the streamed take was trimmed — one reply showed a player at 18.4 s next to a stream of 16.5 s; since then every candidate is trimmed before judging, so that the reward scores what is heard. And the budget scaling was undone again right afterwards, because the code looked the weight up a second time in the recipe after scaling: sum 2.42 against a budget of 2.0. Fixed by a final overall clamping step.

The second commit deals with the same sentence in German, which kept babbling for a different reason. Only the brackets translated, every German word of the line untouched: German cues median word error 0.267 (2 of 4 takes unusable), English cues on the same German words 0.000 (0 of 4). The corpus is labelled in English. Since then the server rewrites the cues server-side — before the retrieval stage, because with German cues the retriever falls back to the emotion named by the directing model, and that is exactly how a horror scene came to be conditioned on Jealousy_and_Envy.

What was measured. Isolation experiment 4 conditions × 1 seed; replication 5 seeds, 5 of 5 takes broken at sum 3.0, median word error 0.82; dose table 6 seeds per cell in the full stack. German cues: 4 seeds per condition, median 0.267 against 0.000. Wrong emotion adapter costs 0.067 against 0.011 — real, but an order of magnitude smaller than the bracket language. Source: docs/BABBLE.md of the demo server; cross-check of the constants from config.py at render time.

II.8 The tools programme after 29 August, part 1 — conditioning datasets, transitions, tempo, crossfade, the quality adapters (protocol §32–§42)

This chapter transcribes eleven protocol sections written between 29 and 31 August 2026. They fall into four strands: three conditioning datasets built from the voice-profile corpus (§32) and the rate adapter that was trained on one of them and then measured to do nothing (§37–§39); a spectral check of adapter stacking (§33) and the protocol audit that followed (§34); the shipping of steering and guidance into the demo server with its WikiSkill layer (§35); and two rounds of the emotional-transition study (§36, §40) with the quality adapters' evaluation (§41) and a baseline error in the DPO trainer (§42). The sections are presented in numerical order, which is not the order they occupy in the file. [P§32–§42]

II.8.1 2026-08-29 — Three conditioning datasets: quality, speaking rate and speed — [P§32]

Agent A4, 2026-08-29. Code ~/quality_speed_datasets/code/, report ~/quality_speed_datasets.html, mirror research-log-2026-08/quality-speed-datasets/. The brief asked for three training sets built from the voice-profile corpus in the project's own prompt format (prompt_lib2, PROMPT_FORMAT_HASH 073aeb09dc923376): a quality preference set of about 50,000 pairs at about 100 per voice profile across all 500, an SFT set conditioned on speaking rate, and a speed preference set. It described the ~500 profiles as having been repaired by Chatterbox voice conversion and by Sidon restoration, each carrying DNSMOS before and after. [P§32.1]

The corpus is real and is exactly as described in size: $NB/train2/corpus/shard-*.parquet filtered to src == 0 is 1,200,530 takes over exactly 500 voice profiles, a mean of 2,401 takes per profile, 607,399 German and 593,131 English, mean duration 10.52 s. The repair, however, is not there at that scale, and this governs everything that follows. Three repair artefacts exist and none of them is "the 500 profiles' generated takes, scored before and after". [P§32.1]

Repair artefacts found in the tree (§32.1)
armfilecoverageDNSMOS before/after
Avprof/work/voices.parquet, 6,064 rowsone *reference recording* per voice; dnsmos_orig, dnsmos_sidon, dnsmos_cbxyes, but 500 clips in total
Bvprof/dsbuild/feat/feat-0[0-3].parquet, 106,424 rowsgenerated takes in 4 renderings, but one voice (k325_age3_bg1), 832 conditions × 32 candidatesno — audiobox-aesthetics only
Cvprof/work/vc_results.parquet, 832 chosenthe published take, VC only, one voiceyes

pvcodes.py states the reason in its own docstring: the pilot and 490-voice generation runs have "exactly one rendering: the take as generated". vprof/vc500/ does cover 500 voices at 15 TB, but it is cross-voice attribute transfer — it carries target_voice and target_kind — not repair: its quality falls by 0.282 on average and rises on 1.0 % of rows. [P§32.1]

The repair mostly makes things worse, and the confound is fatal. Measured, not assumed: on arm A (the 500 selected profiles) Sidon changes DNSMOS by −0.011 on average and improves 48.4 % of clips; Chatterbox by +0.004, improving 51.4 %. The best of the three options is "do not repair" for 173 of the 500 voices (Chatterbox 171, Sidon 156). On arm C: DNSMOS −0.038 mean, 28.0 % improve, and 51 of 832 takes clear +0.10. On arm B, best of the three repairs against the raw take: +0.0014 mean, and only 3.67 % of 26,594 candidates gain more than 0.10 (no uncertainty recorded for any of these means). [P§32.2]

The confound the brief warned about is not marginal but disqualifying for voice conversion. Against the raw take, VC moves ECAPA speaker similarity by +0.259 on average and stays inside ±0.02 on 0.08 % of candidates; it is built to move the speaker toward a reference. Sidon moves it by +0.003 and stays inside ±0.02 on 36.5 %, ±0.05 on 75.8 %. On the 40-dim emotion vector, cosine to the original is 0.924 for VC (51.3 % at or above 0.95) against 0.965 for Sidon (79.6 %). "So the guards do not trim a voice-conversion dataset, they delete it, and what survives is the one repair that preserves identity and, for that same reason, barely improves quality." An independent confirmation already sat in the tree: vcbon/code/bench_pipeline.py records "SIDON on converted output measured NEGATIVE on quality" and ships sidon_out=False. [P§32.2]

The premise of dataset 1 — that a repaired 500-profile corpus with DNSMOS before and after existed — did not survive contact with the tree. The repair artefacts cover one voice (arm B) or one reference clip per voice (arm A), and on measurement the repair helps roughly half of the takes and hurts the other half. [P§32.1, P§32.2]

What was done about it. Two things rather than reporting a small dataset and stopping. First, DNSMOS was measured where it was missing: dnsmos_pass.py scored all 106,424 arm-B clips with speechmos on the 16 kHz decode, the same scorer vpvc.py used. An operational note is recorded for whoever runs it next: speechmos builds its onnxruntime sessions with default options, and onnxruntime's default intra-op thread count is the core count, ignoring OMP_NUM_THREADS; sixteen workers asked for about 768 threads on 48 cores, load average hit 224 and nothing finished. Pinning intra_op_num_threads = 1 and chunking each tar four ways is what made it run. Second, the repair was run rather than inherited: Sidon costs about 40 ms per clip on a GPU, so mksidon.py runs it over a condition-spread sample of every one of the 500 profiles and measures DNSMOS on both sides plus an ECAPA speaker cosine per pair. Only Sidon is run, for the reason above. [P§32.3]

Dataset 2 — SFT conditioned on speaking rate. This is the one dataset the corpus supports at full scale unchanged, because it needs no repair: every take already has a duration, a transcript and word-level timings. $SC/a4_datasets/ds2_rate_sft/, 128 shards, 1,200,530 rows, all 500 voices, the corpus schema plus seven descriptive columns (rate_mode, n_rate_seg, n_rate_fallback, cps_median, cps_min, cps_max, rate_hash). No existing column is altered or dropped. The rate cannot ride in the data: both trainers render the prompt at training time (sft_train2.py:153 and dpo_train2.py:211 both call prompt_lib.user_message(row, rng)) and prompt_lib2.render_prompt builds the Text: field from text, words_json, the burst spans and dur_s. So the deliverable is the data plus prompt_lib2_rate.py, a drop-in module exporting user_message, make_rng and PROMPT_FORMAT_HASH; the change at the call site is one line. This is a new prompt format: PROMPT_FORMAT_HASH = 090c4ca315519a57, against the project's 073aeb09dc923376. No existing adapter has seen a characters-per-second tag. [P§32.4]

Three surfaces are mixed in equal thirds per row: dur[4.7 seconds duration], byte-identical to today; rate[14.2 characters per second]; both[4.7 seconds duration; 14.2 characters per second]. Decisions were each settled by measurement. Characters are len(chunk) of the emitted text, spaces and punctuation included: over 50,112 corpus segments the three candidate character rules are within 1 % of each other on stability (coefficient of variation 0.3150 counting everything, 0.3132 stripping spaces, 0.3159 alphanumerics only; counting words is worse at 0.3572), and len(chunk) is the only definition recomputable from the prompt itself. The rate is per segment and is denominated by the printed duration, rounded to the one decimal the tag shows, because computing it from the underlying float differs by up to 0.6 cps and would put a number in the prompt that the prompt cannot reproduce. [P§32.4]

A segment gets no rate below 1.0 s, below 10 characters, or outside 4–50 cps, and falls back to [D seconds duration] in every mode. The rate of a short segment is boundary noise: coefficient of variation by segment length is 3.178 under 0.5 s (standard deviation 113 cps; the 99th percentile is 128 cps, which is not speech) against 0.239 at 2–3 s and 0.207 at 3–5 s. The rule costs 6.88 % of segments — 4.53 % under a second, 0.62 % with too little text, 1.64 % physically implausible — and leaves 0.79 % of rows with no rate at all. Language is not corrected for: English runs 22.28 cps and German 20.73 (medians 21.45 and 20.15); the Language: field already states which. [P§32.4]

A bug this design avoids: directions._SEG_RX is \[(\d+(?:\.\d+)?) seconds duration\] and requires the bracket immediately after "duration", so it matches neither extended tag; rendering in rate mode and then annotating would take the untimed branch and put one direction at the head of the line instead of one per segment, silently making direction placement a function of rate mode on two thirds of rows. The prompt is therefore rendered by the production path in dur mode, with prompt_lib2, anno and directions untouched, and the finished string's tags are rewritten afterwards. Controls (selftest.py, 3,000 corpus rows, all pass): dur-mode prompts byte-identical to prompt_lib2; two independent implementations of the rate transform agreeing on 6,000 renders; tags-stripped content identical before and after; every both tag's rate equal to len(chunk) over the printed duration; and 3,091 directions inserted in each of the three modes. Shipped mix: 33.34 % dur, 33.38 % rate, 33.29 % both, and within every voice the three stay between 0.301 and 0.371. [P§32.4]

Dataset 3 — speed preference. chosen is the take as generated, rejected the same take with one speech segment time-stretched by 20–30 % up or down, prompt byte-identical on both sides. The stretch is librosa.effects.time_stretch, a phase vocoder; audiostretchy is not installed and nothing was installed to get it. A deviation from the brief is stated plainly: the brief said to build this only on the good-quality material from dataset 1, which would have yielded a few hundred single-voice pairs. Quality is instead held constant directly: every corpus take has an audiobox-aesthetics score in $NB/sftdpo/out/sft/part-*.parquet, uid-keyed with 100 % coverage of the 1,200,530 rows, and this dataset takes only qual_overall >= 3.2 — 303,483 takes, the top 25 %, across all 500 voices. [P§32.5]

Three faults were found by measurement, all of which would have shipped silently. (1) The corpus's stored codes are not the codes the tokeniser produces from the corresponding audio — 74.7 % of frames match exactly, 92.3 % per codebook — so taking chosen from the corpus would have made the negative differ by an mp3 decode, a splice and a stretch; both sides now go through the identical decode / cut / crossfade / re-encode path with the stretch factor 1.0 on the chosen side. (2) Total clip length gave the answer away: stretching one segment moved the whole clip's length by 9.35 % on average (max 23.2 %) while the shared prompt's Tokens: field states the chosen frame count. The time is now taken back out of the silence, proportionally across the pauses; residual length difference is 0.003 s, 0.03 % of the clip, 100 % of pairs compensated, at the cost of a direction imbalance of 30 % slower / 70 % faster at that stage. (3) The compensation clicked, twice: np.tile room tone put the rejected clip's worst-discontinuity ratio at 7.14 on faster pairs against 1.76 for the chosen clip; phase-vocoder re-timing without boundary crossfades put both directions at 4.85; re-timing with crossfaded boundaries lands at 1.93 against 1.76. The stretch itself measured 1.53 throughout, i.e. clean. [P§32.5]

A hard per-pair gate drops any pair whose rejected clip exceeds twice the chosen clip's discontinuity ratio (or 4.0 absolute), or whose residual length exceeds 2 %. The half-speed canary holds: samples_per_frame is exactly 3840.0 = 48000/12.5 on every original clip and 3860 on stretched clips from frame padding — not the 7680 that would mean a stereo frame read as mono. Seam clicks with the 5 ms crossfade sit at 0.178× the local 99.9th-percentile jump (max 0.98); without it, mean 1.123, 95th percentile 5.587, max 35.4. [P§32.5]

Setting the speaker guard against a null, not against a number. The first Sidon run over the 500 profiles measured ECAPA cosine between each original take and its repaired version at 0.844 on average, with nothing reaching the 0.95 that had been assumed as a guard. Read as an absolute threshold that says Sidon destroys speaker identity, and it would have deleted the dataset. It says no such thing: Sidon resynthesises through a DAC decoder, so the embedding moves even where the speaker plainly does not. The baseline adopted is the ECAPA cosine between different original takes of the same voice, and the guard is a margin against it; every row carries speaker_cosine_orig_vs_repaired, speaker_cosine_same_voice_baseline and their difference. The protocol calls this "the same shape of error as the 'steering does not work' false negative in section 26: a threshold read off an absolute scale, with no control to say what the scale means." [P§32.6]

The sampling band. The first pass took the lowest-quality takes of each condition and got takes averaging qual_overall 2.36 against a corpus mean of 3.09 and a 25th percentile of 3.00 — broken audio as much as dull audio. The band was moved to the 5th to 50th percentile within each voice, and sampling is spread over conditions by taking one take per condition before a second from any. [P§32.7]

What makes each of the three a worse idea than it sounds. Dataset 1 is built on a selected subset — the takes where DNSMOS clearly improved — so its pairs are systematically the ones where the original was worst, and the emotion guard is inherited rather than measured on the 500-voice arm (null emotion_cosine_to_original). Dataset 2 is a new prompt format that nothing has been trained on; the rate surface removes the per-segment duration, and characters per second cannot tell a fast talker from a dense sentence read at normal pace. Dataset 3's negative is synthetic: the identity control shows that running the whole path with stretch factor 1.0 still moves 62 % of frames, the direction imbalance is a real skew, and a compensated pair's pause tags no longer exactly match its audio. "Of the three, dataset 2 is the one whose material genuinely supports what was asked; dataset 3 is sound but synthetic; dataset 1 is the one whose premise did not survive contact with the tree." [P§32.8]

A second rewrite bug (§32.9). A rendered user message carries the script twice — inside Instruction: as SCRIPT: ... and as the Text: field. rewrite_tags bounded a chunk at the next [ or (, so the last segment of the SCRIPT: copy swallowed \n- Tokens:\n166\n- Quality:\nNone... into its character count and the last segment of the Text: copy swallowed </user_inst>. Every unit-level control passed; what exposed it was rendering shipped rows through the real training path and comparing fallback rates: 12.93 % against 6.88 % on the same rows. With the newline as chunk terminator the training path reads 6.91 %, matching the renderer. No shipped data was affected, since per-row statistics are computed by render(), not by the rewrite; selftest.py now renders through prompt_lib2_rate.user_message. [P§32.9]

A caveat on the speaker null. The baseline takes differ in words and performance as well as recording, so it measures "same speaker, different content". Over 82 voices it sits at 0.391, while original-vs-repaired sits at 0.892 — a margin of +0.50 with 100 % of pairs clearing it. That is a permissive reference and a lower bound on identity preservation; a stricter null (the same take repaired two different ways) was not run. [P§32.10]

Final composition (§32.11)
datasetpairs / rowsvoicesper voicecontainer
1 — quality preference24,199 pairs49548.9 meana4_datasets/ds1_quality_dpo_final, 64 shards
2 — rate-conditioned SFT1,200,530 rows5002,401 meana4_datasets/ds2_rate_sft, 128 shards
3 — speed preference42,301 pairs50084.6 meana4_datasets/ds3_speed_dpo_final, 128 shards

Dataset 1 covers 130 distinct conditions, mean DNSMOS gain among the selected pairs +0.253, English 13,872 / German 10,327. Dataset 3 covers 130 conditions, 58.8 % faster / 41.2 % slower, German 24,845 / English 17,456. The speaker guard excluded exactly one pair of 24,200; it functioned as a check that passed, not as a filter (margin +0.49, nothing near the null). Every string column is written as arrow string and every shard at the corpus's own row group size (200 for the SFT container, 150 for the DPO ones), because corpus_ds2.list_units enumerates (file, row_group) pairs: a single-row-group write would have offered 128 units instead of about 15,900 and pulled 9,339 rows per read instead of about 200. [P§32.11]

The guard, quantified on the arm where both halves can be measured (this subsection sits physically inside §35 in the protocol file). dnsmos_pass.py was run to completion over the one-voice dsbuild arm — all 106,424 clips, 100 % coverage, 443 of them (0.4 %) failing to score and excluded rather than imputed — and the guards re-applied on DNSMOS. 26,594 candidates carry all four variants. [P§32.12]

Guard survival on the four-variant arm (§32.12)
filterpairsSidon aloneanything with voice conversion
DNSMOS gain > 0.105,1251,6853,440
… and speaker shift within ±0.051,306costs 3,819 pairs, 75 % of everything that passed
… and emotion cosine ≥ 0.9593791819
survival54 %0.6 %

The converted renderings enter with 3,440 pairs that genuinely improved DNSMOS and leave with 19; Sidon alone enters with 1,685 and keeps 918. What only the four-variant view shows: vc_sidon — Sidon applied on top of converted audio — is the most frequent winner on raw DNSMOS, 2,482 of the 5,125, more than Sidon alone (1,685); it also inherits voice conversion's speaker shift (+0.232 mean, inside ±0.05 on 1.66 % of takes), so those 2,482 become 18. "The best-looking repair by DNSMOS is the one that must not be used." [P§32.12]

An earlier revision of §32.12, computed while vc_sidon was still scoring, reported the comparison over three variants and therefore missed the vc_sidon finding entirely; the protocol corrects it in place. Separately, §32.11 states the guard excluded one pair "of 24,200" while §32.12 says "one pair of 24,380" — the two denominators are not reconciled in the protocol. [P§32.11, P§32.12]

II.8.2 Does stacking adapters dull the sound? No — it destroys the words instead — [P§33]

Prompted by a listening impression that the demo audio had lost some high end. 128 conditions × 2 samples = 256 clips; 16 fixed prompts across the emotion heads, English, seed 1234, generated through the same generate_steered path with an empty injector that the steering study used. Code code/spectral_stack.py, raw cells stats/spectral_stack.jsonl, report status/status_2026-08-29_spectral.html. A new instrument was needed because every scorer in the project — genuineness, blend, aesthetics, WER — is blind to a gentle low-pass. Added: long-term average power spectrum per clip reduced to band-energy ratios, spectral centroid, 85 % and 95 % rolloff, and the log-spectrum slope from 2 to 16 kHz; ratios and frequencies only, never absolute level, because adapters change loudness. [P§33]

The hypothesis was that the codec's deep codebooks carry the fine detail and are the most fragile, every adapter targets all twelve audio heads, stacked deltas sum, and a degraded fine channel would read as dullness. It does not. [P§33]

Spectral and intelligibility measures along the stack (§33)
conditionadaptersenergy > 6 kHzrolloff 85 %WER (Parakeet)genuineness
bare SFT3 + DPO-v200.00361088 Hz0.0634.16
+ emotion (Anger 1.0)10.004710710.0654.01
+ voice profile20.003012760.0513.93
+ delivery S_DRAM 1.530.010114810.0833.97
+ delivery AROU 1.540.006113100.4163.86
+ genuineness 0.2550.006012470.4314.01
+ blend 0.560.008513730.5544.35
+ aesthetics 0.570.013111840.4964.04

High-frequency energy rises with the stack, +0.0095 paired at seven adapters (t +2.2); the centroid and rolloff rise with it. What collapses is intelligibility: 0.063 → 0.554. Genuineness barely moves (4.16 → 4.04) and blend rises (6.59 → 7.17) — "the stack makes the model say the wrong words while still sounding fine to the quality scorers." This reproduces, on a different prompt set and a different harness, the demo server's 0.013 → 0.258 over its own stack. [P§33]

The stated hypothesis and mechanism (dullness from degraded deep codebooks) turned out to be wrong: the spectrum gets brighter, not duller, and the actual damage is word error rising from 0.063 to 0.554 across seven adapters. [P§33]
Control: real and synthetic material through the identical codec path (code/spectral_ref.py, §33)
materialenergy > 6 kHzrolloff 85 %rolloff 95 %
real recordings (src 1) through the codec0.00621042 Hz1872 Hz
generated, bare0.003610881945
synthetic voice profiles (src 0) through the codec0.00306321273
generated, seven adapters0.013111842427

The model's output sits between real speech and its own synthetic training material, closer to the synthetic side. If the audio sounds veiled next to a real recording, the candidate mechanism is the corpus rather than the adapters — a data property every checkpoint trained on it inherits — and the lever would be corpus weighting or a high-frequency filter on the synthetic arm, not an inference setting. Limitations: nobody listened; a spectrum that is objectively brighter can still sound worse if the added energy is noise, which at WER 0.55 some of it very likely is. One stack order, one emotion, one voice profile, 16 prompts × 2 samples per condition. DNSMOS was not used (no local implementation in this tree). The white-noise arm of the control failed on a codec API detail; the corpus arms completed. [P§33]

II.8.3 2026-08-29 — The protocol audit, and the state of the archive — [P§34]

Agent A5. Three jobs: check every number in the protocol against the file it came from, make the material findable for a paper, and get everything that exists in only one place into the repository. Deliverables: ~/protocol_audit.md, ~/PAPER_INDEX.md, ~/paper/ICLR_OUTLINE.md. Sections 22–33 were audited claim by claim (every cited path resolved or reported dead, every table recomputed from raw rows where raw rows exist); sections 1–21 were spot-checked. The asymmetry is deliberate and is itself a finding. [P§34]

What the audit found (§34.1)
categoryn
contradictions — the same quantity given two values5
contradictions — two different measurements presented as the same quantity2
outright wrong numbers, contradicted by their own artefact4
dead file references1
claims sourced to no file at all4
claims stated more confidently than their evidence supports6

Everything recomputed from raw rows reproduced exactly: all of 29.2 (seven metrics × nine columns plus all three family rows), 29.6's cross-check, 28.2's pooled ladders and shape classification, both tables of 33, 30's block arithmetic, 26.1's eight headline findings and three baselines, and 22.1's four dataset tags. "The tables are sound. The errors are all in prose that quotes another section." [P§34.1]

Blend does not have one home layer in the protocol. 23.1 (W1) says h25, 23.2 (W2) says h33, both labelled p3, both transcribed correctly. W1 fits a joint 99-output probe and selects on test; W2 fits a per-group probe and selects on validation — and neither section said so. For genuineness (h12 vs h23) this is a tie under W1's own rule; for blend it is not: W1 scores h33 at 0.352853, 0.082 below its own best. Both sections are now labelled; the substantive question is not resolved. The same defect, milder, sits between 22's Phase-2 R² (main, n_test 7,975) and 23.1's (p3, n_test 32,104). [P§34.2]
A hardcoded number propagated into three further documents. ld_report.py states the emotion adapter's saturation step as −0.007 (t −0.90, 19 of 40) in three hardcoded strings; recomputed from coefficients.json it is −0.0026 (t −0.40, 20 of 40), which is what 28.2 says. The wrong figure travelled into lora_dose_study.html, the CFG study's narrative and cfg_study.html, and into §30, where it is the stated justification for running the entire q2a guidance ladder at w = 1.0. Same generator, line 484: the pooled VoiceNet gain is hardcoded as +0.690 (t 6.0) where the data gives +0.6957 (t 6.18). [P§34.2]

The lesson the audit keeps: the project's correction discipline is good — every documented correction landed on every page, old values shown struck rather than silently replaced — but nothing re-derives a number that another document quotes. "Every quoted figure should be computed at render time or carry its source path." [P§34.2]

Fixes applied. Thirteen in the protocol, no prose rewritten: 22 (Phase 2 labelled main); 22.1 (58,800 → 58,714 rows; 58,800 is worklist_main.parquet, a real number on the wrong object); 23.1 and 23.2 (probe and selection rule labelled); 23.7 (264 generations → 264 cells = 1,056 generations); 26 (the 3,034 cells credited to jobs 1527888/89/90, which were cancelled at 0 cells; the real runs are 1527978 / 1527979 / 1527908); 26.1 (3,034 configurations → 170 configurations over 3,034 cells); 28 (cell arithmetic; Whisper offset at w=0 0.0285 → 0.0265); 28.5 (two delivery axes 0.041 → 0.060); 29.5 (recommendations.jsoncomb_recommendations.json); 30 (the saturation figure); 31.3 (the delivery clamp: +0.375 / t 5.18 / 15 of 17 / +0.003 / t 0.55 → +0.381 / t 5.42 / 16 of 17 / +0.0041 / t 0.72 — the section that makes a recommendation to another team was wrong on all five figures). Twenty-four replacements across nine report and generator files, home and repo copies alike; the generators were fixed as well as their output. [P§34.3]

Four claims with no file behind them, none changed: 22.1's "285 files" (release.json names 167 artefacts; 276 predate its mtime); 18's raw variance components 0.04719 and 0.02553 (the derived 35.1 % is sourced to status_2026-08-26_evening.html, its denominator is not); 8's "all 1,041,100 dpo_emotion rows"; 21's first burst gate "1,124,863". The larger point is the citation chain: for sections 1–21 it very often terminates at a narrative HTML (technical_report.html, a status page) rather than at data. 3,147,802, 243,986, 247,049, 1,362,516, 1,061,613, 1,853,486 and 2,327,904 are all corroborated, but by a second prose document written from the same run. If the paper draws on corpus construction or training, that chain needs rebuilding first. [P§34.4]

One claim that looked false and is true. 26 says p3_vectors_ext.npz adds q:genuineness and q:blend; a top-level key check returns "absent" (eleven keys, neither among them), and that reading is wrong: the 99 dimensions are addressed as kind:name from two parallel arrays (actf_steer.py:226), and dim_kind == "q" with dim_name genuineness and blend are entries 97 and 98. Recorded because the false reading is the natural one and would have retracted a correct claim. [P§34.5]

Archive state. Six additions to research-log-2026-08: actforensics/ (protocol.md, report.html, README — this study was in no repository and is the upstream of all four layer-forensics workstreams and every steering result); combination-study/vectors/per_attribute_manifests/ (60 JSON, 643 KB); combination-study/grids/attrs_all.json and the 13 raw/*_r000.json seed records; layer-forensics/PLAN.md (the pre-registration that 23 opens by correcting); stats/voice_profile_names.md; the three audit deliverables. Confirmed already present: layer-forensics/w3/archetypes.npy (156 × 38 × 2560, byte-identical to the home copy), both 2026-08-29 status pages, all five studies' code/ directories, lora_dose_study at 113/113 files, and 80 of 126 top-level HTML pages byte for byte. Deliberately left out: the 741 MB full vector set (the committed topk/ form is 8.6 MB), p3_vectors_ext.npz at 112 MB, and 37 audio-grid HTML pages totalling 978 MB. ~/quality_speed_datasets.html was left alone as A4's in-flight deliverable. [P§34.6]

The 145 GB activation dataset on $SC/out/actforensics has no retention guarantee, every number in 22 through 26 depends on it, re-extraction is a multi-day GPU job, and it exists in exactly one place. Second: no job ID is persisted in any output artefact in the programme — the only study-to-job link is a log filename, which is how 26 came to credit three cancelled jobs with 3,034 cells. [P§34.7]

For the paper. ~/paper/ICLR_OUTLINE.md carries the section plan, eight figures (seven drawable today from committed files) and the gap list. Three things come first: (1) there is no listening test, and this is submission-blocking — §33 showed the quality scorers cannot see a failure that takes word error from 0.063 to 0.554, §18 showed the emotion instrument is saturated by text content; a usable blind A/B harness exists (~/listening_test.html, built 2026-08-09) and comb_recommendations.json names the operating points. (2) Three re-analyses needing no new generation: the matched-total-magnitude re-analysis of the k ladder; the reconciliation of 20's CFG effect with 30.1's, which differ by an order of magnitude off different baselines; and the blend tap of 34.2. (3) Five single-seed results are written as findings — 22's tasks A/D/E on one probe seed, 15's "rank 16 wins", the GRPO v4/v5 comparison, every steering configuration, and 29's factorial — while the tie standard of 23.1 ("treat a difference below ~0.01 as a tie") is applied rigorously in the layer work and not at all elsewhere. [P§34.8]

II.8.4 2026-08-29 — Shipping the three levers into the demo server, and the first WikiSkill draft — [P§35]

LAION-AI/Humaneness-Voice-Demo-Server had exactly one way to shape a performance: load LoRA adapters and write a good prompt. The combination study had measured two more levers and how all three combine, so the job was to give the server the other two as switchable modes, let its voice-acting agent choose between them by tool call, tell it in its own prompt when to use which, and produce the first draft of the WikiSkill knowledge layer. What shipped: pull request #3 against that repository, branched from main at 44ea226. New files levers.py (the policy), steer_engine.py (the vectors and injection points), setup/build_steering_pack.py, setup/check_levers.py, docs/LEVERS.md; changed config.py, tts_engine.py, llm_agent.py, app.py, timed_script.py and three docs. And ~/wikiskills/, mirrored to research-log-2026-08/wikiskills/: 60 pattern pages, an index, an interactions page and coefficients.json, all generated from the measured JSON by code/build_wikiskills.py. [P§35]

The policy's evidence: combination-study/stats/analysis.json, t2.by_family, target in SD units (§35.1)
familyadaptertsteeringtguidancetn
emotion+0.0772.81+0.3849.41+0.0501.84399
delivery+0.3777.10+0.6149.62+0.0260.44170
quality+0.3995.99+0.0060.01+0.0620.9430

The default mode is auto and resolves per family, because the best single lever flips by family. Emotion → adapter+steer: adapter × steering is +0.038 at t 1.36, additive. Delivery → adapter: adapter × steering is −0.164 (t −3.7) and adapter × guidance −0.125 (t −3.5), both significantly sub-additive, so exactly one lever runs. Quality → adapter: steering does not move the quality axes. auto never spends guidance — 1.93× is not a default. Subtracting Emotional_Numbness at −0.10 is attached automatically to any steered emotion; when steering and guidance are both on, both branches are steered. [P§35.1]

Three design decisions. Guidance ships non-streaming: that repository's docs/ADAPTERS.md §1 records realtime factor 1.0 as the streaming budget and 0.764 as the live merged baseline, and forward hooks for the whole adapter stack were rejected at 1.75×; guidance at 1.93× lands at about 1.47 and the player starves, so a guided take is rendered whole, the payload says "streaming": false, and batching the two branches into one forward pass is left as the documented next step. The vectors are an asset, not a commit: the research library is 112 MB, the server needs 99 × 5 × 2560 float32 = 5.3 MB, which setup/build_steering_pack.py distils and MOSS_STEER_PACK points at; the distilled pack is committed to the research log at research-log-2026-08/wikiskills/vectors/. Refusal beats interpolation: seven of sixty attributes have no configuration that clears the balanced guardrails, four delivery _low tails have no measured steering route at all (the two tails are orthogonal, not opposite, median cos −0.0004), and in both cases the server runs the adapter, records why, and the wiki page says No usable setting. [P§35.2]

Two offline checks written expecting a pass failed (§35.3). setup/check_levers.py runs 37 checks with no GPU. First, the realised magnitude at a layer is neither the nominal alpha nor the quadrature sum: components sharing a layer are summed and re-normalised, and the numbness subtraction is nearly anti-parallel to the emotion it is attached to (cos(emotion direction, quality axis) runs −0.62 to −0.95), so subtracting numbness adds almost entirely along the emotion. Over the forty emotion recipes the realised magnitude reaches 0.1926 (emo:Interest at h20; Elation 0.1907, Amusement 0.1781); a ceiling at 0.15 would have refused compositions the combination study itself ran. There are now two ceilings, 0.15 per component and 0.25 realised per layer, both below the 0.30 at which steering collapses; past the ceiling a composition is refused, not trimmed. Second, tap *t* hooks layers[t-1], not layers[t] — the test asserted the wrong module first; the port was right and the check was not. [P§35.3]

What is too thin to build a default on. The quality-axis null (+0.006, t 0.01) pools 30 cells — three attributes at ten prompts; the direction is safe, the precision of that zero is not. Guidance has no significant main effect on delivery or quality (+0.026 t 0.44, +0.062 t 0.94); the shipped g = 2.5 for delivery comes from the CFG study's own arm, not from the factorial. strength: gentle halves the recipe's alpha to 0.05 and is an interpolation — only four cells in the recommendation table use alpha = 0.05 — labelled as such in config.STRENGTH_ALPHA_SCALE. DELIVERY_LEVER = "adapter" is conservatism, not arithmetic: on a delivery axis steering measures stronger than the adapter (+0.614 against +0.377), but the adapter family was swept at scale on the demo box's own hardware and the vectors have not been anywhere. Per-attribute interaction rows are ten prompts. The random-direction floor is per-attribute for only 20 of 60; the other forty fall back to the pooled −0.033. 1.93× is the study's hardware. And 29 of 40 emotion adapters have no usable merge weight in the dose sweep, yet the server loads them at 1.0 — defensible because no adapter in the 5,740-cell sweep was harmful at any weight, but the adapter mode on most emotions is doing less than its name suggests. [P§35.4]

What could not be tested. Nothing in the pull request has run on the demo server's hardware. The offline check passes in both the assets-present and assets-absent configurations. The PR asks for six smoke-test items: that the check script runs there, that adapter mode is bit-identical to main, what the steering hooks cost in realtime factor, that a guided take completes and is intelligible, that the neutralised branch looks right in end.prompt_unc, and that the director uses the new tool call without reaching for guidance on a cheerful one-liner. The steering hooks' cost is arithmetic, not a measurement: one to five hooks, each a norm and a fused multiply-add on a [1, 1, 2560] slice, against the 536 kernel launches per token that took that server's realtime factor from 0.737 to 1.29. Nobody has measured it. [P§35.5]

The follow-up (§35.6): a gate that coupled two assets, and a link that only worked for the author. The demo team merged PR #3 and smoke-tested it the same evening. They found a real gating bug: guidance was refused whenever the coefficient table was missing, even though guidance uses no steering vectors and g has a measured family default — the 0.4 MB file was holding up the lever that does not need the 5.3 MB one, so two thirds of the feature was unreachable on their box. PR #4 splits it: guidance needs neither asset, steering needs the vector pack only, the coefficient table is what auto reads; an explicitly requested mode runs on the family default and labels it operating_point: "family_default". And the raw.githubusercontent URLs sent to them pointed into a private repository (assets in the research log since e0dc027), returning 404 with a 304 KB HTML error page that curl -o would have written over the target path; caught only by re-running the instructions on a shell with credentials stripped. Both files are now release assets on the demo server's public repository, verified unauthenticated: HTTP 200, 324,511 and 5,334,470 bytes, hashes matching, check_levers.py passing against the downloaded copies. [P§35.6]

"The instruction you hand someone else is an artefact and needs a test like any other." One thing the demo team flagged about themselves: they had *not* byte-compared the adapter path against the pre-modes tree. setup/ab_codes.py now does that comparison — fixed seed, prompt and adapter set, comparing generated RVQ code tensors rather than audio, importing nothing that generation modes added so it runs unchanged on the older checkout. That test still has not been run; it is now possible to run it. [P§35.6]

II.8.5 Organic emotional transitions — making a performance turn inside one utterance — [P§36]

The question: a person who is grieving and then finds something bitterly funny does not switch; the grief is still under the amusement at the end. Everything the model could do before was a step — prompt sentence one for A, sentence two for B, generate, concatenate. This round builds the fade nine ways and measures which construction reads as one performance. Design: ten emotion pairs × three purpose-written texts × nine methods, plus two controls and two anchors: 399 cells at 4 samples each, zero failures. Texts byte-identical across arms, the voice pinned by reference codes, the seed a function of the text slot only, so arms are paired sample for sample. [P§36]

Machinery: the decode loop is actf_steer.generate_steered with one line added, inj.set_frame(sched.weights(t)) at the top of each frame; the injection arithmetic h <- h + α · unit(v) · ||h|| is unchanged. The three-branch guidance arm keeps cfg_run.generate_cfg's channel-level interleave verbatim and extends it to l_unc + g_A(t)·(l_A - l_unc) + g_B(t)·(l_B - l_unc) with g_A + g_B = 3, the CFG study's WER-constrained optimum for emotion (§30, Q1). Both emotion directions are injected at one shared layer per pair, chosen as the layer minimising the sum of the two attributes' held-out R² ranks. [P§36]

The result that governs every other one: the constant-emotion control came back flat. The same sentence, voice and seed generated with "reads as intense sadness" and with "reads as intense amusement" differs by 0.12 and 0.12 standard deviations of its own within-condition scatter; 0 of 30 texts clear a 1-SD bar; on full clips only 3 of 40 emotion heads and 1 of 57 VoiceNet dims differ at |t| > 2, and the largest single difference is on *Affection*, which is neither briefed emotion. The planned fraction-of-the-journey scale could not be built — its denominator is noise, and dividing by it produced trajectories in the hundreds. Everything is reported instead in pooled within-anchor SD units, on which the floor (end-to-start swing of a clip that does not turn) is 1.82, sd within a clip 1.06, n = 60. A follow-up regenerated the anchors four ways (reference on/off × prompt-only / prompt + own adapter at w = 1.0), scored on full clips: separation 0.21/0.47 SD with reference and prompt only, 0.39/0.22 with the reference removed, 0.35/0.31 with the adapter, 0.49/0.63 with both changes. pick_reference selects by duration alone and handed this study an Anger take, a slurping-noises take and a hiss take; removing it does not rescue the separation. [P§36]

Result. Ranked on the median turn per text, not the mean, because the paired differences are heavy-tailed (M3's mean Δr is 5.41 with a median of 2.32; M9's are 14.46 and 2.71). Judgement is the paired sign test against the random-schedule control. M6 (multi-rate) has the largest typical turn, Δr 2.53 per text (mean 3.93), against the matched control median +0.74 (mean +2.53, 21/30 up, sign p = 0.043); runner-up M4 at 2.38. Four arms beat the random-schedule control consistently (M6, M4, M3, M2). No arm significantly beats the step baseline — the best is M6: median +0.22 (mean +2.22, 18/30 up, sign p = 0.362); M4: median +0.52 (mean +1.91, 20/30 up, sign p = 0.099). The gains are of order 0.80 where the scorer's own scatter on a non-changing clip is 1.82: "a real effect on a blunt instrument, not a solved problem." [P§36]

Two of eleven methods sit below the random-schedule control: M1 and M7. For M1 that is correct behaviour — it holds both adapters on at a constant weight and has nothing time-varying in it. M7 is the informative one: its only lever on the emotion is the wording of the brief, and the anchor result is precisely that the brief is not a lever. The smoothest arm that beat the control was M2 (1.73); the arm retaining most A at the end was M6 (0.13), the only positive value in the column. The random-schedule control — same vectors, same total magnitude, same multiset of weights, A-to-B order destroyed — travelled Δr 1.40, inside the floor, i.e. null: the movement in the working arms comes from the ordered structure of the schedule. [P§36]

The magnitude artefact: the plan predicted a mid-clip bump and what a crossfade actually produces is a dip. Every arm logged its realised per-frame α: M2 (naive linear crossfade) runs at 0.1000 at the edges and 0.0900 in the middle, M3 (equal-power) is flat at 0.1000. Section 2's 0.1604 figure is for two components at simultaneous full strength, which a crossfade never produces; the two emotion directions sit about 120 degrees apart at the shared layer, so at the crossover they partly cancel. The equal-power correction is still the right thing to do — paired, M2 minus M3 is median −0.14 (mean −3.27, 14/30 up, sign p = 0.856) — but for the opposite reason to the one recorded. [P§36]

Adapter re-weighting mid-generation was tested (M10 against M1, the same two adapters held constant): Δr median −0.58 (mean +0.32, 12/30 up, sign p = 0.362) against the step, median WER 0.018 against 0.017. It did not damage the audio, but neither arm cleared the floor — "a weak acquittal rather than a licence." [P§36]

Five predictions recorded before the run; 2 of 5 were wrong. (1) PARTLY — M0 (the step) would not be as bad as it sounds, with the seam showing in smoothness more than the endpoints: M0's typical turn is Δr 1.92 against 1.59 for the random-schedule control and 1.36 for a clip held at one emotion, so it is not separable from the controls, and its smoothness 1.69 is *better* than the working fades (1.73). (2) PARTLY — M2 would show a mid-clip bump and be beaten by M3: the bump is a 20 % dip, and M3 does beat M2 pairwise (median −0.14, mean −3.27, 14/30 up, p = 0.856). (3) WRONG — M7 (prompt-arc) would do better than its cost suggests as the only in-distribution arm: M7 is the worst arm, Δr 0.67, below the control (1.59); against the control median −0.24 (mean −0.38, 14/30 up, p = 0.856). (4) RIGHT — M6 is the one to build a product on: largest typical turn (2.53, ranked 1 of 10), beats the control (p = 0.043), the only arm whose A-ness at the end is positive (0.13); but it does not separate from M3 pairwise (median −0.48, mean −1.49, 12/30 up, p = 0.362), costs the most word error of the vector arms (0.040 median), and does not significantly beat the step. (5) WRONG — the floor in M4 would matter more than the fade shape: A-ness in the final window M3 −0.14, M4 −0.16, M5 −0.02, M6 0.13; adding the floor to M3 moved retained A-ness by median +0.00 (mean −0.02, 14/30 up, p = 0.856) and the distance by median +0.01 (mean −1.80, 15/30 up, p = 1.000). [P§36]

Honest limits. Nobody has listened blind; the grid at ~/emotional_transitions.html exists so that they can. Every number is a scoring model used at a duration (3 s) it was never validated at. Ten pairs are dramatic archetypes, one language, the fade schedule pinned to script frames (safe because mean duration error was −0.08 s). Next, in order: run M9 (three-branch CFG) on the full material — the only arm that moved decisively and it was starved of cells; fix the instrument before trusting further trajectory work; choose reference clips by content, or drop reference conditioning from emotion studies entirely. [P§36]

II.8.6 2026-08-30 — Duration control is already solved; what is not controlled is tempo — [P§37]

Measured while designing the evaluation for the rate adapter, and it revises the adapter's stated purpose. Duration obedience from the 1,596 clips of the transition study (SFT3 + DPO-v2 + LoRA, no rate adapter, requested totals 9.6–12.0 s): [P§37]

Duration obedience (§37)
quantityvalue
median |requested − produced|0.000 s in 13 of 14 arms
regression slope requested → produced0.979
mean signed error−0.045 s
only exceptionM0, 0.690 s — the spliced arm, a different mechanism

That M0 alone deviates is what rules out a measurement artefact: if the harness were padding or truncating to the target, M0 would read 0.000 too. How the obedience is achieved was measured by energy-based voicing analysis of 133 of the study's 399 published clips (20 ms hops, threshold 6 % of the 95th-percentile hop RMS, the same heuristic family as et_gen.trim_edges). [P§37]

Voicing analysis, 133 clips (§37)
quantitymedian
total duration10.40 s
voiced duration6.56 s
voiced fraction0.621
chars / voiced second21.0
chars / total second13.0

The model speaks at 21.0 characters per voiced second and pads the rest with silence; it does not slow down to fill the budget, it stops talking. 38 % of every clip in the transition study is silence. Two independent pipelines agree on the speaking rate: 21.0 chars/voiced-s from decoded audio versus cps_median = 20.70 computed by anno_rate over the 121,940 rows of ds2_rate_sft_120k. This also explains et_prompts.CPS = 14.9, a chars-per-total-second constant (13.0 measured) that was never the speaking rate. [P§37]

Consequences as written at the time: (1) the dur arm of the rate corpus teaches nothing the base model cannot already do — it is a control, and the rate and both arms are the only ones that address an uncontrolled quantity; (2) the evaluation must measure chars per voiced second, with the prediction to falsify that the base model's tempo stays flat near 21 across a requested 14–26 cps sweep while the adapter moves it; (3) every clip in the transition study asked for about 38 % silence, a specific candidate for the "two performances rather than one" impression. [P§37]

Consequence (1) and the prediction in (2) were both withdrawn by §38 the same day: the flat-tempo prediction was falsified on the first partial data, and the claim that rate/both address an uncontrolled quantity was corrected once the rate tag was traced to a re-encoding of the printed duration. The measured 21.0 cps and 0.621 voiced fraction stand. [P§37, P§38]

II.8.7 2026-08-30 — The rate tag is a re-encoding of the duration, not a second control — [P§38]

Written 2026-08-30 14:25, before the paired half of the tempo evaluation (job 1540103) finished, so the prediction is on record ahead of its test. §37's prediction is already falsified for the base condition on partial data, n ≈ 20 cells per mode, base only: [P§38.1]

Partial base-only data, n ≈ 20 cells per mode (§38.1)
modeslope±setempo @14tempo @26silencedur_errWER
dur0.6050.11617.625.30.2550.000.066
rate0.6580.10018.126.40.2990.000.048
both0.6730.11717.024.50.2560.000.069

Why the prediction was unsound: §37's 21.0 cps came from the transition study, which built every prompt with the same constant (et_prompts.CPS = 14.9), so there was no variation in the requested rate at all; one operating point cannot establish a slope over a range the data never covered. What the base actually does is compress the range: it will not go slower than about 17.6 cps however much time it is given, but nearly reaches the request at the fast end. [P§38.1]

All three modes carry the same signal because the harness varies the budget. The base model has never seen a rate tag — it only knows 073aeb09dc923376 — so its response in rate mode is to the Tokens: budget, which varies with the requested rate in every mode. The discriminating quantity is therefore the interaction (adapter rate − adapter dur) − (base rate − base dur), not the slope. [P§38.2]

The prediction: that interaction is about zero, by construction of the corpus. anno_rate.seg_rate computes cps = n / dur_s from the printed chunk text and the printed duration, both already in the prompt. In both mode the rate tag is strictly redundant with the duration tag beside it; in rate mode it is an invertible re-encoding — D = n / R with n visible — at the cost of making the model count characters, so rate is harder than dur, not more expressive. No row in ds2_rate_sft_120k can have a rate that disagrees with its duration. The corpus teaches an alternative surface form, not a second control axis. What the adapter is then worth is an interface gain: a voice-acting agent can ask for "26 characters per second" without computing a duration. Falsification: if the interaction term is clearly positive, this analysis is wrong. [P§38.3]

§38.3 explicitly corrects §37's consequence (1): rate and both address the same quantity as dur in a different notation. "The error in both cases was the same: inferring capability from surface form without checking where the number comes from." [P§38.3]

II.8.8 Result: the rate adapter changes nothing measurable — [P§39]

Job 1540103 (budget follows the request, 240 cells / 120 paired) and job 1540486 (budget pinned, 160 cells / 80 paired); base = SFT3 + DPO-v2, adapter = ratelora_sft_r16/step3788 at w = 1.0; 8 texts × 5 requested rates × 3 samples, Parakeet WER, paired within cell on a shared seed. [P§39]

Budget follows the request (§39.1)
modecondnslope±setempo@14tempo@26silencedur_errWER
durbase400.6590.07217.625.80.2710.000.074
duradapter400.5890.08318.125.40.2920.000.071
ratebase400.6640.06617.626.40.2990.000.049
rateadapter400.6560.07518.026.40.3030.000.049
bothbase400.7740.07216.325.80.2670.000.067
bothadapter400.7780.06618.127.70.2860.000.067

Paired adapter − base: dur +0.45 (t +1.63), rate +0.36 (t +1.10), both +0.21 (t +0.66). Interaction (rate − dur) = −0.09 cps. The tag form makes no difference. [P§39.1]

Budget pinned at 14 cps, only the tag varies — the decisive arm (§39.2)
modecondnslope±setempo@14tempo@26silence
ratebase400.0080.06918.518.30.354
rateadapter40−0.0340.06918.418.10.357
bothbase400.1040.06617.918.40.358
bothadapter40−0.0060.05517.317.70.343

All four slopes are zero. Told 14 or told 26, the model speaks at 18. Paired adapter − base: +0.20 (t +0.64) and −0.49 (t −1.47). With the budget removed as a channel, the rate tag is inert — for the base as expected, and for the adapter that was trained on it. [P§39.2]

What is and is not controllable: (1) total duration is fully controlled by the Tokens: budget, dur_err 0.00 s in every cell of both jobs, needing no adapter; (2) tempo only indirectly — a tighter budget forces faster speech (slope about 0.66) but there is a floor at about 18 cps below which the model pads (silence fraction 0.35 in the pinned arm); (3) the rate tag is inert, confirming §38.3 in the one configuration where it could have carried information. [P§39.3]

The adapter is harmless but useless: WER unchanged (0.071/0.074, 0.049/0.049, 0.067/0.067), duration still exact. "121,940 rows and 1 h 47 min of training bought no capability. Stage 2 (quality DPO) sits on it safely, but it should not be described as a rate control." [P§39.3]

Why it could not have worked: anno_rate.seg_rate sets R = n / D from the printed chunk and the printed duration, so in every training row the rate agrees with the duration and the model was never given a reason to read R. A corpus that teaches tempo must vary R independently of the natural duration — the same text, a generous fixed budget, differently time-stretched audio, and the tag naming the stretched rate. That data already exists: the 42,301 compensated=True pairs in ds3_speed_dpo_final, each with its stretch_factor (the protocol calls them "audiostretchy pairs" here, although §32.5 records that the stretch was librosa.effects.time_stretch and audiostretchy was not installed). A second attempt should be built from those. [P§39.4]

II.8.9 Crossfade v2: the same performance turn, without the steering vectors — [P§40]

Why it was repeated: §36 ranked nine constructions and then reported that its own denominator was noise (anchors 0.12 SD apart, 0 of 30 texts clearing a unit). Two things then changed the design space. Steering vectors were dropped from the demo server (GEN_MODE back to adapter) after a listener called their output "more emotional, with strange artefacts, and the timbre off", which removes five of §36's nine arms (M2–M6) outright. And classifier-free guidance came in, with a field ceiling of about g = 4.0 against the CFG study's WER-constrained optimum of g = 3.0 (§30). Design: 13 arms × 4 guidance levels (0, 2, 3, 4; CFG3 has no g = 0) × 30 texts × 4 samples = 1,530 cells / 6,120 clips, on v1's material byte for byte ($SC/etrans/out/prompts.json), the turn falling on an explicit 0.5 s pause. v2 fades adapter merge weights — XF, XF_FLOOR, XF_ASYM, ARC_XF, BOTH, ARC, STEP — plus CFG3. Adapters are faded, never merged; the audio heads are weight-tied to the embeddings and merging destroys them. [P§40]

Repairing the instrument. Four anchors instead of two: ANC_A/ANC_B hold one emotion end to end with that emotion's own adapter, ANC_Ap/ANC_Bp with the brief alone. Separation is read on four windows (first 3 s, first 6 s, first sentence, full clip) and three head sets (own, av, loo), and the anchors are generated first across all texts. Re-analysis of v1's stored records had already shown the 0.12 was mostly the 3 s window — brief-only with reference on reaches 0.890 SD on full clips, 12/30 texts ≥ 1 SD on loo, and brief + own adapter 0.762 on own — so v1's headline understated the model by roughly 2.5×. A prediction was recorded before v2 ran: with adapters own should beat loo; brief-only should reverse. [P§40]

What guidance costs. Measured as t_gen_s / frames_tot, batch of 4 on one GH200, paired on arm × text over 127 cells: pooled medians 51.5 → 98.6 ms/frame, a 1.92× multiplier, independently replicating the CFG study's 1.93× (§30, Q4). The multiplier is identical at g = 2, 3 and 4 — the price is the second branch, not the strength. Report ms/frame as medians: ARC_XF's g = 0 mean is 70.1 against a median of 55.5 because one warm-up cell hit 106.8. [P§40]

Engineering, dearly bought. (a) sbatch --export splits its value on commas and an --arms list is commas — v1 lost 13 of its 14 arms to exactly this; v2 passes --export=ALL plus an args file. (b) import numpy from the study venv on /e/data1 was measured at 244 s and again at 448 s, jobs spending 82–105 minutes in imports; the analysis half (xf_agg.py, xf_report.py, xf_page.py) was moved to the EasyBuild stack on /e/software, where import numpy takes 0.087 s and full aggregation takes 0.24 s on the login node. (c) Copying the venv to scratch failed with Unexpected EOF in archive. (d) Two of the first three 4-node jobs produced zero cells, one stalling 94 minutes in d_alloc_parallel/cxiWaitEventWait; the mitigation is a longer wall clock, not more nodes. (e) Resume is by cell key (--tag <new> --resume-tags x1,x3). [P§40]

Result — the anchor separation first. On the pre-declared primary readout (loo on sent), the brief-only anchors stand 0.797 / 1.013 / 1.934 / 2.239 SD apart at g = 0 / 2 / 3 / 4, with 23 of 30 texts clearing a full SD at g = 4 (on the full clip, 3.428 SD and 25/30). Against v1's 0.12 SD, 0/30 that is about a 19× improvement, decomposable into three causes: the window (v1's own settings reproduce here at 0.188 — its 3 s slice of a 30 s zero-padded scorer input was 10 % signal); guidance, which raises separation monotonically on every contrast × segmentation × anchor set, making g = 4 the measured optimum; and the adapters, which turned out to hurt. [P§40]

The pre-run prediction is refuted, informatively. loo beats own in 15 of 16 adapter-anchor cells and in all 16 brief-only cells, and the brief-only anchors separate better on the adapters' own emotion heads (2.057 on full clips) than the adapter anchors do (1.227). The adapters compress the contrast they exist to create: within-anchor scatter is flat across all eight conditions (2.16–3.13) while the distance between the two emotional endpoints is 2.3× larger without the adapter at g = 4 (4.587 vs 1.982); guidance amplifies the brief 2.6× (1.753 → 4.587) and the adapter path only 2.1× and non-monotonically. [P§40]

The grid, read against a floor. An anchor does not turn, so its |dr| is pure noise: 0.71–0.88 on loo/sent. STEP — the two-take splice already shipping — is the only arm that clears it, at every g (median dr 1.421 / 1.306 / 1.945 / 2.623; rho6 rising 0.36 → 0.73). The best non-STEP cell is CFG3@g4 at 0.818, below the 0.838 a non-turning anchor scores at the same g. Every XF* arm, BOTH, ARC and ARC_XF land in −0.41..+0.70, the range the anchors occupy — and so does the shuffled control, at 0.599 and 0.668, above nine of the twelve genuine XF* cells. Paired against STEP, every arm loses at every g, and the deficit grows with guidance (g = 0 mostly n.s.; g = 3–4 almost all p < 0.01). [P§40]

Against the shuffled control the question closes. C_SHUF runs XF's own weights with the A→B order destroyed. Across 48 arm × g comparisons the only arm that beats it is STEP, at all four guidance levels (+0.536 to +1.736, 20–21 of 28 texts up, p = 0.013–0.036); the fading arms' best showing is XF@g0 at +0.165 (16/28, p = 0.57), and XF@g4 sits at +0.044, 14 of 28 texts up, p = 1.000. This strengthened at every refresh: at 82 % of the grid STEP's advantage was p ≈ 0.06–0.08 on 16–19 texts, at 92 % p ≈ 0.036 on 26–28, at 98 % p = 0.013 at g = 4, with no fading arm moving off zero. "The fade shape does not matter, because the fading carrier does not move the emotion." [P§40]

Costs and safety. Guidance multiplier confirmed per arm at 1.90–1.94×, identical at g = 2, 3, 4, 30/30 texts up, p < 0.001 in every arm. Guidance costs zero WER (median Δ +0.000 in 33 of 36 arm × g cells). Adapters cost more than guidance does unguided: 30 ms/frame with none, 39 with one, 48–49 with a live fading pair; STEP is simultaneously the cheapest arm (30 unguided / 57 guided) and the best-scoring; CFG3's third branch is nearly free (83 ms). v1's M10 acquittal holds at 4× the scale: changing LoRA scaling mid-generation against a KV cache built at other weights does no measurable damage (XF* WER 0.015–0.037). STEP's one real cost is duration, −0.33 to −0.54 s, the eaten pause at the splice. Half-speed canary: 0 cells. [P§40]

An aggregator bug. The projection standardises each of 97 score dimensions by the anchor pair's within-anchor scatter, and 1.7 % of dimensions (2.4 % on full clips) have exactly zero scatter; flooring them at 1e-9 acted as a 10⁹ amplifier — single cells reached |dr| ≈ 10⁶ and arm means printed as −843803.6 beside a sane median of −0.693. Degenerate dimensions are now dropped. The anchor separation was never affected, but the entire arm table was; the median column was correct throughout and exposed the bug within a minute. A second fault, the same as the audio-export-collision note: cells produced under two tags took metrics from the last tag and audio from the first, so the listening page could pair one run's clip with another's numbers. [P§40]

Recommendation: ship STEP at g = 3 (median dr 1.945) — the only method that measurably turns, the only one that beats its own shuffled control, the cheapest in the grid, and what already ships. g = 4 scores higher at identical cost but sits at the listener-reported edge. Do not ship an adapter crossfade in any shape; do not ship ARC/ARC_XF. What the metric cannot decide: STEP wins by construction — two independent takes are maximally different — and whether the splice sounds like one performance or two clips is a listening question. The listening page (508 clips, one binary per clip) asks exactly that. Status at writing: 1,476 of 1,500 cells (98.4 %), all 480 anchor cells complete, 24 of the 51 arm × g cells at the full n = 30; every number was computed three times — at 82 %, 92 % and 98 % of the grid — and no conclusion moved. [P§40]

Whether STEP's splice reads as one performance or two is untested by any listener; if it fails, the remaining candidate is the prompt lever alone at g = 4 on an instrument that can now resolve 2.2 SD. [P§40]

II.8.10 2026-08-31 — The quality adapters: a null measurement, then a listener who disagreed — [P§41]

(The protocol numbers this section's subsections 40.1–40.3; they are cited here as §41.1–§41.3.) Two evaluations of the same two adapters, five hours apart, disagree, and the difference is what they were stacked on. Evaluation A ("§39 harness", 16 prompts, adapters on bare SFT3): no adapter beat the base model on DNSMOS, best +0.017 (t +0.45), worst −0.130 (t −2.05); all ten evaluated states raised word error slightly, +0.001 to +0.042, none significant, ten of ten in the same direction. Evaluation B: 60 clips, adapters on the full demo-server stack (SFT3 + DPO-v2 + speaker + emotion + genuineness/blend/esthetics, doses from the server's own config.py, prompts built by its own timed_script). [P§41.1]

Evaluation B, adapters on the full server stack, n = 20 per condition (§41.1)
conditionWERgenuinenessburst blenddur err
A — server today0.0581.482.080.00 s
B — + mix12250.0301.432.530.00 s
C — + qual15040.0741.902.470.00 s

Paired within (emotion, take) on a shared seed, n = 20 per condition: B lowers word error by 0.027 at unchanged genuineness; C raises genuineness by 0.42 (up in 13 of 20 pairs) and pays +0.016 word error (no t or CI recorded for the paired differences). The two adapters pull in opposite directions — B toward cleanliness, C toward expression. Three explanations that cannot be separated: different text material; n = 20 is small; or an adapter behaves differently inside a five-adapter stack than alone on the base model, and evaluation A never tested the deployed configuration. "Measure the stack you ship." [P§41.1]

A listener overturned the null result. Reported 2026-08-31 after listening to the 60 clips at huggingface.co/spaces/laion/moss-quality-adapter-listening: "qual1504 actually sounds best, with somewhat more WER" (translated). This is the first listening evidence on these adapters, and it agrees with the only metric that moved (genuineness +0.42) while accepting the cost that metric predicted (+0.016 WER). Both written recommendations were wrong in the same direction: the earlier recommendation of qual376 and mix245 — the early checkpoints, chosen because the preference task saturates by step 376 — ranked qual1504 last on DNSMOS (−0.119, t −2.14); and the author had told the user none of the three should be merged, offering mix1225 if one had to be heard. The one signal that pointed at qual1504 was there and under-weighted: the highest genuineness of any condition (+0.197, t +1.90), the only near-significant positive in the whole test, recorded as "a curiosity" instead of as the finding. "DNSMOS was the wrong primary metric for this question." [P§41.2]
The merge weight was never swept: every number has qual1504 at w = 1.0, whose own scaling is alpha/r = 32/16 = 2.0. Job 1557968 sweeps w ∈ {0, 0.25, 0.5, 0.75, 1.0, 1.25, 1.5} on the same ten emotions and seeds, with w = 0 reproducing the bare server stack through the identical code path. Also open: combining qual1504 with the burst+stop adapter then training (§42) — they target different failures. [P§41.3]

II.8.11 2026-08-31 — rew_chosen is measured against bare SFT3, not against the run's own starting point — [P§42]

Found by the burst+stop agent while deciding how to select a checkpoint; it changes how every DPO run in the log should be read, including §41's. In LoRA mode dpo_train2.py never materialises a separate reference model: _REF stays None and the reference log-probabilities come from peft.disable_adapter(). That disables all adapters, so the reference policy is bare SFT3 — not the adapter the run was initialised from. [P§42]

Consequences: (1) rew_chosen at step 0 is not zero — the burst+stop run starts from ratelora_dpo_mix_r16/step245, so its rc = +1.74 at step 1 is that checkpoint's accumulated gain over SFT3. (2) rc > 0 is a weak guardrail: a checkpoint can be degrading relative to the checkpoint it started from while rc stays positive; the reward-collapse case it was introduced to catch (0.206 → −0.830 in an early run) is still caught, but only in its extreme form. (3) The §41 figures are unaffected in value but need the baseline named: quality-only reached rc = +29.9 and mix 75/25 plateaued at +12.3; both are drift from bare SFT3, and part of each was already present before either run began, since both started from the stage-1 rate adapter. [P§42]

What to use instead: per-family preference accuracy on a held-out slice, scored offline through each kept checkpoint with the same pairs. The in-training eval cannot do this — it is capped at 288 pairs and reports pooled numbers only, and fam_acc in loss.jsonl is one rank's 8 pairs per heartbeat. famscore.py (+ bsfam.sbatch) does it forward-only on one GPU and writes one record per pair per checkpoint, so a later listening test or a different metric can join against it without regenerating anything. [P§42]

A second plateau called too early. The same agent read pref_acc 0.8403 at two consecutive evals as a plateau and as over-optimisation. Step 224 came in at 0.8785. Two evaluation points are not a plateau; the protocol records it as the same failure mode as §41.2 — reaching for the over-optimisation story before the data supports it — and notes it happened twice in one day, in two different runs, to two different readers. [P§42.1]
Corrections and retractions recorded in §32–§42: §32.1–§32.2 — the brief's premise (a repaired 500-profile corpus with DNSMOS before/after) did not exist; the repair helps about half of takes and hurts the other half. §32.5 — three faults caught before shipping: corpus codes differ from re-tokenised audio (74.7 % frame match), clip length gave the pair away (9.35 %), and the silence compensation clicked twice (7.14, then 4.85, finally 1.93). §32.6 — the 0.95 absolute speaker-cosine guard (measured 0.844) would have deleted the dataset; replaced by a same-voice null. §32.7 — the first sampling pass took the bottom of the quality distribution (2.36); moved to the 5th–50th percentile. §32.9 — rewrite_tags swallowed field labels into character counts (12.93 % vs 6.88 % fallback); no shipped data affected. §32.11/§32.12 — the guard's denominator is given as 24,200 in one place and 24,380 in another; an earlier revision of §32.12 computed over three variants missed that vc_sidon is the most frequent DNSMOS winner, corrected in place. §33 — the dullness hypothesis was wrong: stacking brightens the spectrum and destroys intelligibility (WER 0.063 → 0.554). §34.2–§34.3 — thirteen numeric corrections to earlier protocol sections (22.1 58,714; 23.7 1,056 generations; 26's cancelled jobs; 28 Whisper offset 0.0265; 28.5 0.060; 29.5 filename; 30 saturation −0.0026 not −0.007; 31.3 all five clamp figures; VoiceNet gain +0.6957 not +0.690) plus the unresolved blend home-layer contradiction (h25 vs h33). §34.5 — a false 'absent' reading of p3_vectors_ext.npz was itself retracted before it retracted a correct claim. §35.3 — the single alpha ceiling was wrong (realised magnitude 0.1926); tap t hooks layers[t-1], the test was wrong. §35.6 — PR #3's guidance gate wrongly depended on the coefficient table (fixed in PR #4); asset links pointed into a private repository. §36 — the constant-emotion control was flat (0.12 SD), the planned scale could not be built; the predicted mid-clip magnitude bump is a 20 % dip; predictions 3 and 5 wrong, 1 and 2 partly wrong. §37/§38 — §37's flat-tempo prediction was falsified and its consequence (1) corrected: the rate tag re-encodes the duration. §39 — the rate adapter changes nothing measurable and must not be described as a rate control. §40 — v1's 0.12 SD headline understated the instrument about 2.5× (window artefact); the 'own beats loo' prediction was refuted; the 1e-9 scatter floor amplified degenerate dimensions by 10⁹ (fixed by dropping them); a tag collision paired one run's audio with another's metrics; steering vectors were dropped from the demo server after a listener report. §41.2 — the DNSMOS-based recommendation (qual376, mix245; qual1504 ranked last) and the advice to merge none were both overturned by a listener; DNSMOS was the wrong primary metric. §42 — rew_chosen is measured against bare SFT3, so rc > 0 is a weak guardrail and §41's rc figures need their baseline named. §42.1 — a plateau at pref_acc 0.8403 was called on two points; step 224 read 0.8785.

II.9 The tools programme after 29 August, part 2 — the vocal-burst adapters: dose, labels, manufactured data, guidance on the burst (protocol §43–§56)

This chapter follows the vocal-burst adapters from a field report ('four classes do not come out') through fourteen studies. The arc is: the merge weight is swept and is not the lever (§43); a preference run teaches stopping and burst class on different schedules (§44); the quality adapter's weight sweep cannot separate doses (§45); a listening page ships dead (§46); the training labels are audited against the detector and sixty adapters turn out to have been trained on a label the detector contradicts (§47); re-filtering leaves 9.8 % of the rows (§48); manufactured bursts are built and score 0.000 on the strict metric while working at family level (§49, §50); the family-relaxed hit rate is adopted as the metric (§50a–§50c); the other levers are stacked (§51, §52); classifier-free guidance triples the hit rate (§53); an external generator and an external annotator are tried against the detector's blind spots (§54, §55); and adapter rank is shown not to be the bottleneck (§56). Sections are presented in numeric order; the protocol's file order differs and is noted in the module docstring. [P§43–§56]

II.9.1 The vocal-burst merge weight: a dose sweep under the real stack, and a metric that could not see the answer [P§43]

Full write-up ~/reports/burst_dose_study.md; data $SC/out/burst_dose/bd_r000..015.jsonl (2,549 rows), aggregate bd_agg.json, code $SC/code/burst_dose/, with SC=/e/scratch/reformo/schuhmann1_moss. A field report said four burst classes — scream, shriek, frustrated_groan, exhausted_groan — do not come out, and the suspected lever was the burst adapter's merge weight. Reading lora_bank.py and config.py before touching a GPU turned up three discrepancies: (1) the documentation and the code disagree by a factor of two — the docstring and the plan_blend() comment quote the manual's 0.5 inline / 0.75 solo, while config.py ships BURST_LAM = 0.25 / BURST_LAM_INTENSE = 0.5, and neither halving is recorded in docs/; (2) two different rules decide 'solo' — lora_bank.plan() uses < 14 plain words, plan_blend() uses < 10, and docs/ADAPTERS.md documents only the 14-word rule; (3) 'the burst stands as its own beat' is implemented purely as 'the script is short', so a burst at the end of a long line is always dosed as inline. §28 had swept this family against the bare SFT3 export, and its own §28.5 says 'nothing here predicts a stack.' [P§43.1]

Documentation/code mismatch: lora_bank.py documents burst weights 0.5 inline / 0.75 solo; config.py ships 0.25 / 0.5. Two undocumented 'solo' rules (< 14 vs < 10 words) coexist. Neither is recorded in docs/. [P§43.1]

Design. 264 conditions × 8–10 prompts = 2,420 cells, four arms crossed with six classes (the four named failures plus chuckle and contented_sigh as working controls) × two script kinds (inline = a 23–57-word corpus line; solo = a 5-word lead plus the burst) × 8–10 held-in-common prompts. Every d is a within-prompt paired difference against the same condition at w = 0 on the same seed; n = 55 pooled prompt-level pairs. The decision rule — largest pooled gain in r_burst_cls among weights that are significant (t ≥ 2.00) and hold all six per-class guardrails — was fixed in bd_agg.py before the grid ran, with §28's 3σ guardrails and an absolute WER_pk ≤ 0.25 cap on inline. [P§43.2]

§43 arms
armmodelweights
stackSFT3 + DPO 1.0 + per-class emotion 1.0 + genuineness_high 0.25 + blend_high 0.5 + burst @ w0, 0.2, 0.25, 0.4, 0.5, 0.6, 0.8, 1.0, 1.2, 1.5
alonebare SFT3 + burst @ w (§28's configuration)same 10
mismatchstack, but the next class's burst adapter @ 1.01.0
stack_nolorastack with the burst adapter absent from the model

Controls, all passing. The stack_nolora arm gives the zero-sharing check: a five-adapter stack with the burst adapter loaded and scaled to zero is bit-identical (0.000e+00 on every metric, all 12 prompt sets) to the same stack with no burst adapter in the model, which licenses w = 0 as a baseline. The half-speed canary (samples_per_frame = 3840, §21) was asserted per clip: 0 violations in 7,260 takes. The mismatch control at w = 1.0 gives d r_burst_cls −0.044 (t −2.19) inline — a wrong-class rank-16 perturbation does not raise detections. [P§43.2]

§43.3 primary result: best pooled d r_burst_cls per panel
arm / scriptbest pooled dmax |t|pre-registered rule
stack / inline (ships at 0.25)+0.0170.79none qualifies
stack / solo (ships at 0.5)+0.0241.52none qualifies
alone / inline+0.0471.50none qualifies
alone / solo+0.0512.39w = 0.8

The single qualifying weight in 36 tests is in the arm the demo server never runs, and the same weight in the corresponding stack panel is −0.0022 (t −0.14): §28.5's warning, confirmed with a control. The study is powered: the SD of the paired differences gives a minimum detectable pooled d of +0.031 to +0.050 depending on panel, and §28's bare-stack burst curve rose about +0.12 across the same range — 'That effect size is comfortably inside this study's resolution and it did not reproduce in the stack.' [P§43.3]

The methodological result. On r_burst_cls, stack/solo shows nothing (|d| ≤ 0.024, |t| ≤ 1.52). On hit-rate — the fraction of prompted target-class cues that came back as a detection of the right class — the same 7,260 takes give +0.061 at w = 0.5 (t 2.32), +0.061 at 1.0 (t 2.46), +0.091 at 1.5 (t 2.98), rising with dose. Reading reward.burst_realisation explains it: it is an F1, so every extra detection enters the precision denominator; a matched detection of the wrong class still earns 0.35 credit, so converting a wrong burst into a right one moves the score by 0.65 of a duration term; and that term exp(-|dur - prompted| / 0.5) drives a right-class burst of the wrong length toward zero. The dose does exactly the two things the metric punishes. [P§43.4]

§43.4 what the dose does, stack arm, per clip
scriptquantityw = 0w = 0.5w = 1.0w = 1.5
solodetected bursts / clip0.8730.9581.0061.139
solomean detected duration (s)0.3730.4270.4400.461
inlinedetected bursts / clip1.2731.3701.4551.533
The pre-registered primary metric r_burst_cls was built to be blind to the manipulation under test: the knob works (about 30 % more bursts, about 24 % longer) and the metric is near-neutral to that combination by construction. stack/inline/frustrated_groan at w = 0 reads r_burst_cls 0.113 with a hit-rate of exactly 0.000 — every point of the 0.113 is wrong-class partial credit. The hit-rate reading is a post-hoc re-analysis and is labelled as such. [P§43.4]
§43.5 absolute hit-rate over every arm, kind and weight
classprompted target cuescorrect detectionshit rate
chuckle13205880.446
contented_sigh13204490.340
exhausted_groan11881960.165
scream1186900.076
shriek1045110.011
frustrated_groan118800.000

frustrated_groan produced zero correct bursts in 1,188 prompted cues, every arm, both kinds, all ten weights; its d hit is exactly +0.000 at all nine weights because the numerator is zero at both ends. 'The dose parameter cannot lift a floor of zero.' The failures are not silence but substitution down the arousal axis: inline frustrated_groan → Contented Sigh ×127, Wistful Sigh ×14; inline exhausted_groan → Low Mumble ×135, Ahem ×25; inline shriek → Surprised Gasp ×86, Contented Sigh ×29; inline scream → Surprised Gasp ×75, Contented Sigh ×29; solo exhausted_groan → Low Mumble ×148; solo shriek → Contented Sigh ×47. no_burst appears only in single digits. Genuineness at w = 0 tracks the split: chuckle 4.68/6 and exhausted_groan 4.88/6 against frustrated_groan 1.49/6 and shriek 1.79/6. The protocol calls this an adapter-quality or label-quality problem, not a merge-weight problem, noting that frustrated_groan's labels came from a sidecar rescan that never entered the corpus burst_labels column. [P§43.5]

Recommendation: keep BURST_LAM = 0.25 and BURST_LAM_INTENSE = 0.5. Paired directly against what ships: in stack/solo no weight is significantly better than 0.5 on hit-rate while turning the adapter off is significantly worse (−0.061, t −2.32); in stack/inline nothing separates from 0.25, including w = 0. Raising the dose has one-directional costs (stack/solo d blend falls monotonically to −0.427 and d |dur err| rises to +0.017 s by w = 1.5). It does not differ by class: the per-class best w column scatters over the whole range with 8 stars in 108 cells against about 5 expected by chance, two of them negative. inline is flat on everything (max |t| 0.79, high confidence); solo is the only place anything moves (moderate confidence). The solo hit-rate maximum sits at w = 1.5, the top of the swept range, and 'is not an optimum' — the design cannot distinguish a peak from the last point before one, the same failure mode as §45. Priority follow-ups: a forced-choice listening test on stack/solo at 0.5 / 1.0 / 1.5; re-scoring with hit-rate primary, labelled post hoc; auditing frustrated_groan's and shriek's adapters against their training data; adding esthetics_high. [P§43.6]

Limits (four of eleven). No human heard any of the 7,260 takes. The stack measured is five adapters while production runs six (esthetics_high @ 0.5 absent; §33 places the intelligibility collapse at the addition of a further adapter). 'Standalone' cannot be separated from 'short script'. And 'the adapter is weak' cannot be distinguished from 'the detector cannot hear it' — every outcome passes through one burst detector never validated against human labels. The one positive claim: in a five-adapter production-shaped stack, the burst merge weight is not the lever that fixes the field report's complaint. [P§43.7]

False alarm, kept on purpose: mid-run 205 cells looked missing, concentrated in stack/inline for shriek (78) and frustrated_groan (53), and were flagged as non-random. It was a scheduling artefact — bd_run.plan() assigns whole chunks to ranks least-loaded-first, so stack/inline/shriek lived entirely on ranks 000 and 001 (the last two of sixteen to finish) and stack/inline/frustrated_groan on ranks 008/009. Both conditions are now 80/80 and 90/90; the grid is complete at 2,420 of 2,420 cells with 129 duplicate rows collapsed last-write-wins. [P§43.8]

Operational. Both jobs ran the same Python 3.11 env_transcribe venv, so the grid is single-version; a fallback launcher bdrun4.sbatch (Python 3.13.5 / torch 2.9.1 / transformers 5.14.1) was prepared and never submitted. import transformers off /e/data1 cost job 1546767 3 h 53 m and timed it out; the same import cost 38 min the next morning and about 13 s under source /e/scratch/reformo/schuhmann1_moss/code/fastgen.sh; full aggregation of 2,549 rows takes 0.7 s that way against 244–448 s for import numpy from the venv. MOSS_IMPORT_STAGGER does nothing against this: job 1550920's four ranks staggered 0/20/40/60 s all finished import numpy at exactly 604.4 s. The HF caches must stay split (HF_HOME on scratch, HF_HUB_CACHE on /e/data1). Cancel-and-resubmit churn (bdose3/4/5) produced nothing; 2,291 cells took 2 h 57 m wall at about 74 cells/min on 16 ranks. [P§43.9]

II.9.2 31 Aug 2026 — Burst realisation and stop/no-improvisation: one DPO corpus, and a per-pair scoring harness that answers 'which checkpoint' [P§44]

Stage 3 of the rank-16 chain: ratelora_sft_r16/step3788ratelora_dpo_mix_r16/step245 → this run. Full report ~/reports/burst_stop_dpo.md; all numbers measured on JUPITER, 2026-08-31. Corpus $SC/a4_datasets/ds_mix_dpo_bs: 14,932 pairs in 16 shard-*.parquet — quality 4,478 (30.00 %, reused ds1_quality_dpo_final), speed 2,985 (19.99 %, reused ds3_speed_dpo_final), burst 3,734 (25.01 %) and stop 3,732 (24.99 %), both new. Inside the burst arm: burst_swap 1,476 / burst_silence 1,310 / burst_excise 948 = 39.5 / 35.1 / 25.4 % (planned .40/.35/.25); stop kinds equalised at 1,244 each. Two design invariants: the swap negative is a burst of a different class by the same speaker, re-timed and RMS-matched (silence-only negatives would train a presence detector rather than class control; donors are 100.0 % cross-semantic-group over 170 target→donor class pairs and all 500 voice-profile speakers); and stop_twosent states the full original duration while the script holds only sentence one, verified on 410/410 timed renders (one speech segment each, segment sum equals full length, residual max 0.22 s). The burst arm is voice-profile-only by necessity — on the real-speech half voice_key is per-utterance (277,280 distinct 'voices' for 819,313 low-mumble detections), so no same-speaker donor exists. Gates applied before training: burst_no_free_slot 1,345, twosent_bad_split 673, truncate_bad_span 431, caption_leak 188, not_two_segments 144, first_seg_not_sentence 132; plus 18 of 44 detected burst labels dropped for under 40 candidate clips. caption_leak matters: on 188 stop rows caption_general contained a word from the cut part of the script. [P§44.1]

Parquet row-group defect: finalize.py wrote row_group_size=128 into about 933-row shards, giving exactly 8 row groups per shard; corpus_ds2 hands rank R the units [R::4], so all sixteen remainder groups landed on rank 3 (pools 4096 / 4095 / 4095 / 2643). Because dpo_train2.py takes the global MIN of epoch_batches, ranks 0–2 would have left about 35 % of their pool untouched every epoch. Rewriting the shards with 8 near-equal row groups (117/117/…/114) moved the pools to 3743 / 3744 / 3743 / 3699 and the run from 660 to 902 steps. Rule: the number of row groups per shard must not be a multiple of the rank count W. finalize.py still has ROWGROUP=128. [P§44.2]

The run. Job 1557904, 4 GPU, [done] steps=902, EXIT=0, 11:43 → 12:26 (42 min of a 6 h wall). From ratelora_dpo_mix_r16/step245, base $SC/out/sft3/export, rank 16 / alpha 32, β = 30.0, length-normalised, lr 1e-6, warmup 90, 2 epochs, 32 global pairs/step, --val-frac 0.02, prompt_lib2_rate (hash 090c4ca315519a57); 17 checkpoints kept, step56 → step902. Run on fastgen ($SC/code/fastgen.sh: py 3.13.5, torch 2.9.1+cu128, transformers 5.14.1, peft 0.20.0): import gate 24.4 s and in srun 26 s after job start, against 604 s for import numpy alone on the venv. rew_chosen is flat for the entire run (3.29 to 3.49) while rew_rejected falls monotonically from +0.295 to −3.87: every bit of margin comes from pushing the rejected side down. Read with §42 (rc is measured against bare SFT3 via peft.disable_adapter()), the flatness says no checkpoint moved the chosen-side likelihood away from where stage 2 left it. [P§44.3]

famscore.py. Job 1562661, 1 GPU, 14 min, 40 s per checkpoint: the same 317 held-out pairs through the init adapter and all 17 checkpoints — 18 conditions, 5,706 records in $SC/work_bs/famscore/per_pair.jsonl, plus summary.json and paired.json. Identical pairs across conditions make the test paired, which is what makes n = 22–32 per family usable. Three implementation points: reuse the trainer's own code (load_dpo_units, build_pair_examples, DpoStep imported from dpo_train2); assert the adapter name-mapping hit every tensor (PEFT names …lora_A.default.weight, a saved adapter …lora_A.weight; a silent miss scores the base model 18 times and reads as 'training changed nothing'); pack one row at a time, because build_pair_examples drops undecodable rows and its meta is not index-aligned with its input. Cross-check: famscore puts step56 at 0.7476 on 317 pairs; the in-training eval put it at 0.7743 on its 288-pair subset. [P§44.4]

Result. Held-out n: quality_repair_sidon 95, speed_stretch 55, burst_swap 32, stop_truncate 31, stop_twosent 31, burst_silence 27, burst_excise 24, stop_append 22. (1) stop_twosent was inverted by the stage-2 adapter: the init adapter scores 0.258 (8 right, 23 wrong) — 'the model does not fail to stop, it actively prefers the take that keeps talking' — fixed by step 112 (0.871) and ≥ 0.968 thereafter; the pooled pref_acc could never have shown this (31 of 288 pairs moving 0.258 → 1.000 is worth 0.08 on the pooled number). (2) The stop and burst arms learn on different schedules: all three stop families are finished by step 280 (stop_truncate 1.000 from step 168, stop_append 1.000 from 280); burst_silence is at 0.370 at step 224, below its init of 0.407, first significantly positive at step 336 (d_logit +0.479, t +2.90) and peaks at step 784 (0.889, +1.338, t +7.62); burst_swap sits at init through step 168 and settles at 0.844 only from step 560. 'A checkpoint chosen on the stop arm alone would be chosen 500 steps before the burst arm was done.' (3) Nothing was damaged: quality_repair_sidon dips insignificantly at step56 (−0.164, t −1.51) and is significantly positive from step280 (+1.398, t +9.43 at step896); speed_stretch reaches 1.000 by step504. burst_excise is unsolved: 0.458 → 0.625, best 0.708, paired d_logit significant late (+0.865, t +3.84 at step896); the suspected cause is that it is the only burst family whose two sides differ in length, overlapping the speed_stretch objective. [P§44.5]

Recommendation: the range step 336 → step 896, and nobody has listened. Lower end step 336: first checkpoint where all eight families are at or above init on paired logit and the burst arm has moved (burst_silence +0.479, t +2.90; quality recovered +0.564, t +4.04), stop arm at 1.000 / 1.000 / 0.968. Upper end step 896: largest paired gain on burst_excise (+0.865, t +3.84) and quality_repair_sidon (+1.398, t +9.43), tied for best pooled accuracy (0.9432), no family below init; step902 is the six-step tail and the only late checkpoint with a reversal (burst_swap 0.844 → 0.688, 5 pairs). Explicitly not selected: best-by-val_loss (step 840 — for DPO a lower val_loss is a larger margin and picks the most over-optimised checkpoint) and best-by-pref_acc (step 616 — swings 7 pairs between evals and ranks seventh on the 317-pair ladder). [P§44.6]

'No clip from this run has been heard by anyone.' Every §44 number is a forward-only preference logit on base + one adapter: no demo-server stack, no CFG, no steering vector, no merge weight, no vocoder. The proposed listening test spans four conditions — ratelora_dpo_mix_r16/step245 (control, 0.258 on stop_twosent), bs_r16/step336, bs_r16/step616, bs_r16/step896 — with two separate judgements per clip. per_pair.jsonl joins by uid. (§51.1 later hears it: a null on burst realisation.) [P§44.6]

II.9.3 The qual1504 weight sweep cannot separate the doses [P§45]

Job 1557968, prompted by the listening result in §41.2. Seven weights on the same ten emotions, same voice, same seeds, in the full demo-server stack; w = 0 is the control through the identical code path. n = 20 per weight (10 emotions × 2 takes), paired within (emotion, take). The adapter's own scaling is alpha/r = 32/16 = 2.0, so w multiplies that. Duration error is 0.00 s at every weight. [P§45]

§45 qual1504 weight sweep, n = 20 per weight
wWERΔ vs 0tgenuinenessΔ vs 0tgenu up
0.000.0581.48
0.250.139+0.081+1.631.78+0.30+0.8113/20
0.500.102+0.045+0.861.55+0.07+0.2212/20
0.750.123+0.065+1.351.89+0.41+0.9511/20
1.000.074+0.016+0.721.90+0.42+1.2513/20
1.250.045−0.013−0.431.49+0.01+0.029/20
1.500.064+0.006+0.271.96+0.48+1.3713/20

No weight is distinguishable from the control. The largest |t| anywhere is 1.63. More informative than under-powering is that the response is not monotone: genuineness runs +0.30, +0.07, +0.41, +0.42, +0.01, +0.48, with the collapse at w = 1.25 sitting between two of the three largest effects — 'A real dose-response cannot do that.' Either the effect is smaller than the noise at n = 20 or it is genuinely not smooth. What survives: the three largest genuineness gains cluster at w ∈ {0.75, 1.00, 1.50} (+0.41 / +0.42 / +0.48), and among those w = 1.50 dominates on both axes — largest genuineness gain (+0.48) and smallest word-error cost (+0.006, against +0.016 at w = 1.00 and +0.065 at w = 0.75). [P§45]

Correction: the first write-up said w = 1.00 was 'the cheapest of the three in word error (+0.016 against +0.065 and +0.006)', contradicting its own parenthesis since 0.006 < 0.016. The user caught it from the table. The error produced a recommendation to keep the weight already in use when the swept data pointed the other way. [P§45]

Extended to w = 3.0 (job 1563064). w = 1.50 was re-run as the join and reproduces exactly (WER 0.064 / 0.064, genuineness 1.96 / 1.96 across two independent jobs), so the two tables compose. [P§45.1]

§45.1 extension, n = 20 per weight
w1.501.752.002.503.00
Δ genuineness+0.48+0.58+0.31+0.47+0.52
t1.371.770.921.631.83
Δ WER+0.006+0.145+0.054+0.148+0.039
t0.271.771.392.080.69
1.50 is not a peak — genuineness keeps creeping up — but the only significant result in the entire eleven-weight sweep is a cost, not a benefit: WER at w = 2.50, t 2.08. Genuineness never reaches significance (largest |t| 1.83). The response stays ragged (+0.42, +0.01, +0.48, +0.58, +0.31, +0.47, +0.52): 'the grid was widened where it should have been deepened'; n = 20 per weight is the binding constraint. [P§45.1]

Practical reading: w = 1.5. It takes almost all of the available genuineness gain (+0.48 against a best of +0.58) for essentially no intelligibility cost (+0.006); w = 1.0 remains the conservative alternative and the difference is not demonstrable from these data. 'This is a null on a metric, not on the phenomenon' — §41.2's precedent is that a listener overturned a null DNSMOS result on these same adapters the same day. All clips are on the listening page (spaces/laion/moss-quality-adapter-listening, Part 2), seven weights side by side per emotion. (The protocol gives both '220 clips' and '140 clips' for the page; 140 is the first seven-weight sweep, 220 the eleven-weight total.) [P§45.1]

II.9.4 31 Aug 2026 — The listening page shipped with its audio players dead, and nobody would have known [P§46]

Found 2026-08-31 only because the user said (translated) 'I cannot find the audio here — there are only tables.' The page returned HTTP 200, every one of the 140 OGG files returned 200, scale.js parsed as valid JSON, and the tables rendered. 'Everything I had verified was true, and the thing the page exists for did not work.' [P§46]

Cause. Successive edits left a stale copy of the Part 2 renderer in front of the current one inside the same <script> block. The stale copy read a data shape that no longer existed, threw Cannot read properties of undefined, and because an uncaught exception aborts the entire script block, the correct copy below it never ran. Part 1's 60 players were unaffected; Part 2's 140 silently rendered nothing. Verification had checked status codes, byte counts, JSON validity and the presence of id="app2" — none of which touch whether the JavaScript runs. [P§46]

Fix and tool. nodejs is in the EasyBuild stack. A 40-line simulator — a stub document/localStorage/CSS, then eval of the real data.js, scale.js and each inline block — reports per block whether it threw and how many <audio> elements each container ended up with: before, n1.js THREW and app2 <audio> count: 0; after, n1.js ran OK and app2 <audio> count: 140. It also caught a second latent fault: Part 2 read the saved-ratings object from Part 1's top-level const; Part 2 now reads localStorage itself inside a try/catch that prints the failure into the page. Rule: 'A published page is verified when its scripts have been executed against its real data, not when its files return 200.' Kept at /tmp/sim.js. Every listening page in the project so far — including the 399-clip transition study — had been published on HTTP checks alone. [P§46]

II.9.5 The burst adapters do not fail because of a weight: sixty of them were trained on a label the detector contradicts [P§47]

§43 measured six of the seventy-one burst adapters and left one question open (three classes work and three do not, but the working three have both more rows and a strict embedded_quality_gate, so six cases cannot separate the factors) and one unasked: is the burst even in the training audio? Full write-up ~/reports/burst_data_audit.md; dose results ~/reports/burst_dose_extended.md. [P§47]

The harness on trial, 13/13. frustrated_groan at 0.000 at every one of ten weights is as consistent with a harness that never applies the adapter as with an adapter that does nothing. $SC/code/vb_audit/va_verify.py (job 1574096, one GPU) ran on the identical code path bd_run.py uses: peft is the real 0.20.0, not the two-file stub that shadowed it earlier the same day; the adapter attaches to 268 LoRA modules at base scaling 2.0; set_lora_scale is exact and idempotent (0.5 applied twice gives 1.0, not 0.25 — a cumulative implementation would have made the ladder decay monotonically); per-class fingerprints differ; a deterministic forward moves with the weight (max |Δlogit| = 56.75 from w = 0 to 1.5); w = 0 with the adapter loaded is bitwise identical to the adapter not loaded (Δ = 0.0 on logits, same sha on the code stream), which is why round 2 dropped stack_nolora; replaying bd_run's per-cell arm sequence leaves the intended adapters active; w = 0, w = 1.5 and mismatch give three different code streams. On CPU: the prompt surface form is not the problem — the exact timed cue (<class>, D seconds) appears on 60–65 % of renders for every class (0.637 chuckle, 0.631 frustrated_groan) and prompt_format_hash 073aeb09dc923376 is identical in the manifest and both prompt-set files. 'So section 43's numbers are real.' [P§47.1]

embedded_quality_gate was never a burst-presence gate. $SC/code/vb_buckets.py walks for gate in (0.80, 0.50, 0.00) over has & (blend_pct >= gate) & (gen_pct >= gate) and stops at the first gate yielding 800 rows; the gate varies blend and genuineness percentile, not presence. Over the 3,144,739-row corpus scan only 44 of the detector's 83 labels ever fire — low mumble 846,299 times and frustrated groan 19 times. The eight classes at gate = 0.80 are exactly the eight most frequent detector labels; the gate is a proxy for corpus frequency. Twenty-seven classes were built from scripted cues: vb_buckets.py --cues admits a row when the script names the burst and the detector found some burst at the aligned index — explicitly ~has, meaning the detector called that span something else, and the scripted name is then written over the detector's label. Splitting into det / cue / std: chuckle 1600 = 800 det + 0 cue + 800 std; scream 891 = 800 + 0 + 91; shriek 386 = 174 + 0 + 212; frustrated_groan 568 = 19 + 339 + 210; fast_breathing 1600 = 0 + 1600 + 0. Twelve classes have no embedded rows at all. [P§47.2, P§47.3]

The measurement. $SC/code/vb_audit/va_audit.py, job 1574095, 4 nodes / 16 GPUs, 9 minutes, 43,139 rows — every embedded row of all 70 buckets plus up to 400 standalone per class. Each row's own MOSS codes are decoded with train_grpo.decode_waves and moss_pack.codes_from_blob (imported; the source comments record three GRPO runs lost to a hand-rolled decode), resampled to 16 kHz, and handed to the same reward.RewardModel.bursts instance that scored §43's generations. p_same = fraction of rows where the detector found a burst of the labelled class within 1.5 s of the labelled position. [P§47.4]

§47.4 result 1: detector-sourced rows, p_same by gate
gate 0.80p_samegate 0.50/0.00, same 800-row poolp_same
low_mumble0.94wistful_sigh (0.50)0.72
contented_sigh0.86resonant_hum (0.50)0.73
ahem0.78yawn (0.00)0.73
chuckle0.73deep_breath (0.00)0.74
exhausted_groan0.69scream (0.00)0.69
surprised_gasp0.56soft_hum (0.00)0.39

Result 2 — cue rows realise the label at 0.00, and the substitution is the finding: fast_breathing 1,600 cue rows → Contented Sigh 1,330; normal_breathing 1,163 → Contented Sigh 962; heavy_breathing 1,162 → Contented Sigh 874; sobs 661 → Contented Sigh 427; frustrated_groan 339 → Contented Sigh 159, Wistful Sigh 103. 'The instrument that scores this entire pipeline hears a sigh where the script wrote a groan, a pant or a sob.' Result 3 — the standalone column is an instrument reading: contented_sigh (0.86 on embedded rows) scores 0.05 on its own curated isolated sighs; the locator is a Whisper-small segmenter and the classifier's embedder a 30-second encoder, so a 1.1 s clip padded to 30 s is out of distribution for both. Standalone p_same is flat across clip length (0.046 at 0.7 s, 0.093 at 9 s); 33.4 % of standalone rows produce no located event against 2.7 % of embedded rows, and when the locator fires the classifier names the wrong class nine times out of ten (3,423 other to 397 same). The standalone rows were admitted on a different classifier's verdict (clf_is_burst in vblora_perclass), leaving about 40,000 rows — 59 % of the burst training set — unadjudicated. [P§47.4]

§47.5 verified rows against §43 outcome
class§43 outcomeverified rows in the whole bucket
chuckleworks (0.57 inline @0.25)832
contented_sighworks (0.80 inline @1.2)727
exhausted_groanworks solo only (0.30 @1.5)690
screamweak (0.22 inline @1.5)556
shrieknearly dead (0.083)65
frustrated_groandead, 0.000 everywhere8

'An adapter trained 2,000 steps on 568 rows of which eight contain the burst it is named after has been trained to reproduce the other 560, and the other 560 are sighs.' The exception is scream: 556 verified rows and still weak, so bucket size (891 rows / 1,114 steps against chuckle's 1,600 / 2,000) is the remaining candidate. Re-filter: keeping a row only if every labelled span of the target class has a same-class detection within 1.5 s, 15 classes have ≥ 300 verified rows (breathy_giggle 1026, chuckle 832, low_mumble 798, wistful_sigh 757, ahem 727, contented_sigh 727, exhausted_groan 690, yawn 667, resonant_hum 636, surprised_gasp 608, deep_breath 593, childlike_giggle 576, scream 556, sharp_inhale 550, soft_hum 376); 54 classes < 100; 32 classes = 0. A probability floor does not help (at θ = 0.3 every class halves, at 0.5 everything collapses including the working classes); the ranking is θ-invariant. The defensible statement is 'the training label and the evaluation instrument disagree on twenty-seven classes, and the model learned the instrument's answer.' [P§47.5, P§47.6]

The instrument on trial (§47.7). Re-reading the 44,344 scored rows: the detector fires on isolated bursts and is more confident there (median 0.489 on isolated chuckles vs 0.258 embedded; only 3.3 % no detection) but renames them one step inside the family — isolated contented_sigh → Wistful Sigh 55 / Contented Sigh 5; isolated chuckle → Breathy Giggle 48 / Chuckle 30. reward._same_class is strip-non-letters then equality or substring, so 'contented sigh' vs 'wistful sigh' scores 0, the same 0 as silence. At family level over eight families: det 0.666 → 0.766, std 0.068 → 0.221, cue 0.000 → 0.028; ten classes reach ≥ 0.60 (guffaw 1.000, cackle 0.955, chuckle 0.901, humming 0.842 from an exact rate of 0.000) — 4,644 rows certifiable with no new work. Crop experiment (job 1575339, vb_deaf/vd_probe.py): ten contented_sigh rows embedded in speech that the full pipeline named correctly, cut to the burst alone (0.24–0.55 s). [P§47.7]

§47.7 crop experiment, n = 10
full clip 5.8–12.2 scut to the burst alone
full pipeline (locator → classifier)10 / 104 / 10 — 5 locate nothing, 1 → Ahem
classifier alone, locator bypassed10 / 10 Contented Sigh

Isolation costs the pipeline six of ten correct answers and every loss is the locator's; the cheap fix is to bypass the locator for any clip already of burst length. The fix does not rescue the standalone rows: on the curated isolated sighs both stages agree (Wistful Sigh 8/10 with the locator, 8/10 without), so the residue is the source, not the length. For frustrated_groan, on twenty cue rows with the locator bypassed the classifier says Contented Sigh 9, Ahem 3, Wistful Sigh 3, Low Mumble 3, Surprised Gasp 2 — Frustrated Groan zero, at confidences 0.11–0.67; 46 clips are staged at $SC/pages/deaf/. The standalone caveat narrows from a blind instrument to a near-sighted one whose error is directional and localised to a named stage. Rule: 'find the smallest edit to the input that turns one reading into the other', and 'when a metric reports 0, check whether its comparison function can express nearly right.' [P§47.7]

Round 2, the full-coverage dose sweep (§47.8). All 70 adapters on one common carrier — the ten neutral held-out prompts from §28, re-timed per class with that class's cue in the same slot, requested duration set to the class's median labelled burst duration — configuration stack@carrier-v3 (§43's is stack@class-rows-v2; absolute levels are not quotable between them). Arms stack / alone / mismatch, kinds inline / solo, weights 0 / 0.25 / 0.5 / 0.8 / 1.0 / 1.25 / 1.5, 10 prompts × group 3: 15,400 cells / 46,200 takes, 8 nodes × 4 GPUs, job 1574185. The job stopped on TIMEOUT at 4 h: 14,972 / 15,400 cells (97.2 %), 1,501 conditions, 0 canary violations, 0 duplicates; the loss lands in the diagnostic arms (alone inline 1,752/2,100, mismatch inline 636/700). The stack arm is 9,784 / 9,800 — 69 of 70 classes complete; whispered_mumble's inline w = 0.8 cell is not measured and w = 0.5 is 4/10, marked as such rather than omitted. [P§47.8]

Pooled: no global weight is recommendable, and that is not a null. The inline effect is significant at four of six weights (peak Δ r_burst_cls +0.0115, t = 3.59 at w = 0.5), but the count of classes inside the §28 guardrails decays monotonically — 69/70 at w = 0.25, 66 at 0.5, 55 at 0.8, 47 at 1.0, 40 at 1.25, 36 at 1.5 (solo: 65 → 25). BURST_LAM 0.25 is the only inline weight where 69 of 70 classes stay inside the 3σ bands; round 2 gives no reason to move either default. The two target metrics disagree about the dose in opposite directions: r_burst_cls peaks at w = 0.5 inline and is gone by 1.5; hit_rate climbs monotonically to w = 1.25 inline (t = 4.68) and 1.5 solo (t = 4.02), exactly where the guardrails have given out (Δ genuineness −0.19, Δ blend −0.83 at 1.5). The mismatch control moves r_burst_cls by −0.0033 (t = −0.99) inline, +0.0007 (t = +0.20) solo. Per class: 17 of 70 have a weight that separates from zero — 3 in both kinds (childlike_giggle, effort_grunt, deep_breath), 9 inline only (deep_breathing, coughing, sharp_inhale, heavy_breathing, frustrated_groan, wolf_whistle, pleasure_moan, yawn, cough), 5 solo only (scream, soft_hum, hiccup, snorting_giggle, drinking_noises). [P§47.8]

Only five of the seventeen are audible: at the recommended weight hit_rate is non-zero for childlike_giggle (0.17 / 0.13), sharp_inhale (0.17), scream (0.07), soft_hum (0.07) and deep_breath (0.03), and 0.00 for the other twelve against a 0.00 baseline — r_burst_cls can rise because a burst-shaped event of about the right length appeared while still being named something else. Four of the seventeen (wolf_whistle, hiccup, snorting_giggle, drinking_noises) have zero verified training rows in §47.5. The audit's best-supported classes (chuckle, breathy_giggle, low_mumble, contented_sigh, wistful_sigh) mostly show no dose effect because their w = 0 baseline is already high (contented_sigh 0.37, low_mumble 0.33, chuckle 0.23): 'A flat dose curve is evidence about headroom, never about capability.' Rule: report coverage before results, and give a class with no cells the row 'not measured'. [P§47.8]

II.9.6 Re-filtering the burst buckets: 9.8 % of the training rows survive, and 34 of 70 classes have nothing to re-train on [P§48]

Companion to §47, run in parallel by a second agent: it builds the filter, the decision table and the re-training, and adds one failure mode §47 did not look for. The defect no detector was needed to find. prompt_lib2.render_prompt writes every burst into the script as (<label>, <D> seconds) with D = burst_end − burst_start, while dur_s also enters the GENERAL caption and the Tokens: budget; the stated burst length is a literal instruction. Comparing every row's dur_s against its own audio (frames / 12.5) over all 67,540 rows of the 70 buckets: [P§48.1]

§48.1 stated minus true duration, by row kind
row kindnp50p90max> 0.25 s> 1 s
embedded27,562+0.000+0.052+0.0790.0 %0.0 %
standalone39,978+0.128+3.903+9.89722.9 %12.4 %
9,170 rows — 13.6 % of the whole vocal-burst training set — state a burst longer than the entire clip they are paired with, every one a standalone row. The real prompt for vb::shriek::s1_sample015499_13180 reads GENERAL: … 5.3s, EN. / SCRIPT: (shriek, 5.3 seconds) / Tokens: 4 against 0.32 s of target audio: 'it teaches that a five-second shriek means emitting almost nothing.' [P§48.1]

The filter. A row survives iff Gate A |dur_s − frames/12.5| ≤ 0.15 s, no span ending past the audio, no empty spans (0.15 s because anno._fmt rounds to one decimal and the worst embedded row is off by 0.079 s — the tolerance admits every embedded row by construction); Gate B every target-class span has a same-class detection within 1.5 s (reward._same_class, burst_realisation's own tolerance) with the weakest probability ≥ θ = 0.174; and not is_val. θ is calibrated, not chosen: the 10th percentile of the weakest same-class detection probability over the 1,891 Gate-A-clean rows of the three adapters that demonstrably work (p10 by class 0.199 / 0.162 / 0.171). $SC/code/vb_filter/vf_score.py is a pinned fork of §47's va_audit.py scoring all 67,540 rows; two jobs, about 20 minutes on 12 GPUs. [P§48.2]

§48.3 verdicts over 70 classes
verdictclasses
retrainable (≥ 200 surviving rows)12
retrainable but thin (40–199)2
too few rows (1–39)22
no usable data (0 rows)34

6,615 of 67,540 rows survive (9.8 %). The twelve: low_mumble 624, chuckle 590, contented_sigh 582, deep_breath 540, exhausted_groan 530, yawn 504, scream 496, resonant_hum 477, breathy_giggle 456, childlike_giggle 412, wistful_sigh 405, sharp_inhale 345; survivors are 97 % embedded rows. The three field-report classes split three ways: scream is not a data problem (800 embedded rows realise at 68.6 %, indistinguishable from exhausted_groan's 68.2 %; 496 survive at 55.7 %); shriek yields 48 rows and is instrument-limited on top (only 36.8 % of corpus-labelled Shriek spans still read as a shriek after the MOSS codec round trip — they come back as Scream, Surprised Gasp, Childlike Giggle); frustrated_groan yields 4 rows — all 339 scripted-cue rows fail, 155 contain a Contented Sigh at the labelled position, its 191 standalone rows realise at 1.0 % — recommendation: withdraw the adapter, not re-train it. [P§48.3]

The standalone corpus is the single largest source of bad supervision. Realisation on isolated-burst rows: chuckle 30 %, exhausted_groan 13 %, scream 6 %, contented_sigh 5 %, shriek 0 % of 158, frustrated_groan 1 %; 60–87 % of them fail Gate A. A dead end recorded: 'the isolated recordings are known-good positives, so the detector's hit rate on them measures the detector' does not work — it reads 30 % / 5 % / 13 % on three working classes. The check that works is redet (re-detection of prov = det rows through the codec): 65–89 % on every class with a real pool, 36.8 % on shriek, undefined on frustrated_groan. §47's independent audit lands in the same place (chuckle 832 / scream 556 / shriek 65 / frustrated_groan 8 against 590 / 496 / 48 / 4; 15 rescued vs 12 solid + 2 thin). Repairing Gate-A failures by clamping dur_s is not worth it: only 335 rows across all 70 classes fail A and would pass B, none from shriek, scream or frustrated_groan. [P§48.4]

Re-training. train_bucket_loras.py --buckets $SC/buckets_strict --kind all --out $SC/out/bucket_loras_strict with the hf_burst/manifest.json recipe verbatim: rank 16, α 32, 5 epochs, lr 1e-4, batch 4, base sft3 export, prompt_hash 073aeb09dc923376. Holding the recipe fixed shrinks the step count — chuckle 2,000 → 738, scream 1,114 → 620, shriek 485 → 60. Operational: bd_run.py gained --new-loras / --weights / --classes / --arms and arms stack_new / alone_new (NEW::-prefixed adapter strings), proven behaviour-neutral (identical 264 conditions and 16-rank plan with no new flags); two agents worked the area at once, kept clean by pinned forks (vf_score.py, bd_run3.py) and separate output roots; filtered buckets must not be written under /e/data1 (inode quota). [P§48.5, P§48.6]

Correction (§48.8 to §48.5): the claim that a fixed recipe biases 'against' the re-trained adapter is not one-directional — an under-trained adapter is also a smaller perturbation of the base weights, so at a high merge weight the fixed-recipe arm flatters the new adapter rather than penalising it. [P§48.5, P§48.8]

The instrument has a prior (§48.7). Over the 36,793 scored training rows the detector emitted 43,238 detections using only 44 of its 83 labels, and three labels are 58 % of them — Contented Sigh 28.3 %, Low Mumble 16.1 %, Ahem 13.5 %. Those three plus Surprised Gasp, Breathy Giggle, Wistful Sigh, Childlike Giggle, Exhausted Groan and Chuckle are the detector's working vocabulary and exactly the eight classes at embedded_quality_gate = 0.80, i.e. the eight whose adapters work. For each of the 34 zero-survivor classes one substitute label dominates at 19–88 %; for 16 it is Contented Sigh (fast_breathing 88 %, normal_breathing 86 %, relief_sigh 86 %, panting 86 %, slow_breathing 81 %, gurgling 66 %, sobs 64 %, wolf_whistle 62 %); gulps → Low Mumble 81 %, swallows → Low Mumble 68 %. Four are near-synonym confusions inside the detector's own label set (relief_sigh/Contented Sigh, snorting_giggle/Breathy Giggle, fearful_gasp/Surprised Gasp, whispered_mumble/Low Mumble). [P§48.7]

'The works / dead split is aligned with which classes the detector over-predicts, and nothing in this project has yet separated those two explanations.' The retrainable column is safe (positives only); the no-usable-data column is safe only as a statement about verifiability. Claims that survive intact: Gate A (arithmetic, 9,170 rows), scream (a positive), frustrated_groan (19 detector-labelled rows in a 3.1 M-row corpus and 0.000 at every weight under a harness verified 13/13). [P§48.7]

Before/after (§48.8). Two complete grids: 5 classes × {inline, solo} × {stack, stack_new, alone, alone_new} × w ∈ {0, 0.5, 1.0, 1.5} = 160 conditions, 46 prompts per (kind, arm, weight) — prompt sets are 9/8/10/10/9, so 1472 records is the full grid, not a shortfall of 1600. br swaps in adapters re-trained with the recipe held fixed; bs with the optimiser step count held fixed. Paired at the prompt level, each prompt's 3 samples averaged first, t over n = prompts. [P§48.8]

§48.8 pooled Δ hit rate, new − old
panelw = 0.5w = 1.0w = 1.5
br stack inline−0.029+0.051+0.087 (t +1.73, ns)
br stack solo−0.014+0.029+0.014
bs stack inline−0.022−0.022−0.094, ΔWER +0.328
bs stack solo+0.072* (t +2.66)+0.022−0.174* (t −4.70), ΔWER +0.329
Under a fixed recipe the filter buys nothing measurable: two of eighty class cells clear |t| > t₀.₀₅(n−1), one favourable and one adverse — chance. The only new number is shriek 0.000 → 0.167 at w = 1.0 inline from 48 training rows, n = 8. The step-matched arm holds both real results: w = 0.5 solo +0.072 (t +2.66), carried by chuckle 0.533 → 0.733 (t +2.71, wrong rate 0.467 → 0.200); and at w = 1.5 the same adapters break the model — chuckle solo 0.633 → 0.267 with WER 0.413 → 0.920, exhausted_groan solo 0.259 → 0.037, chuckle inline WER 0.097 → 0.632, shriek inline WER 0.278 → 0.783; guardrail breaches go from 9 in br to 30 in bs, four caused by the swap. Step-matching a bucket that shrank 5–10× multiplies epochs (shriek 40 epochs over 48 rows, chuckle 14 over 738), so the two arms bracket the truth and 'neither an under-trained nor an over-trained adapter off buckets_strict is usable' at w = 1.5. [P§48.8]
§48.7 eats most of the win: the chuckle cell carrying +0.072 is +0.200 strict and +0.000 family-relaxed (0.800 → 0.800); pooled +0.072* strict against +0.029 (ns) relaxed. The re-trained adapter made the model produce laughs the detector names correctly (Breathy Giggle → Chuckle), not more laughs. The adverse w = 1.5 cells have no such escape: −0.367 strict and −0.400 relaxed on chuckle solo. Rule: 'A same-recipe ablation on a filtered dataset is not one factor, it is two. Run both or claim neither.' [P§48.8]
Internal contradiction noted: §48.8 states '§43 set the burst merge weight at 1.5 by dose response on the original adapters', whereas §43.6's recorded recommendation is to keep BURST_LAM = 0.25 / BURST_LAM_INTENSE = 0.5 and explicitly not to ship 1.5. §48.8's conclusion that anything trained on filtered buckets must be re-dosed, with w = 0.5 on isolated bursts as the only operating point with positive evidence, stands independently of that sentence. [P§48.8, P§43.6]

Two aggregation defects that print as the same character (§48.9). bd_agg.tstat returned nan for n < 2 (one paired observation; now dropped, counted and listed by name — both final runs report 0 of 80 dropped, and on partial data one n = 1 cell had produced a 't' that was pure noise) and for sd = 0 with mean exactly 0 — the run's built-in correctness check at w = 0, where stack and stack_new load the same model; it now prints 0 (ctrl), whereas before 'the run's own proof of validity was rendered as its own failure mode.' sd = 0 with mean ≠ 0 printed +inf and is now an exact sign test, p = 2^−(n−1). tcrit is now exact (scipy.stats.t.ppf(0.975, n−1), scipy 1.18.1) — the previous lookup rounded up to the next tabulated n, giving n = 11 the value 2.201 instead of df = 10's 2.228; the fallback now rounds down. Displayed old→new levels are means over the same prompts the difference is taken on. Twenty-three synthetic checks cover the guard. Rule: 'A sentinel value that your correctness check and your failure mode both render as is not a sentinel.' [P§48.9]

II.9.7 Manufacturing the missing bursts: voice conversion fails the gate, and the corpus that works scores 0.000 [P§49]

Rather than discard rows, build them: splice a real, human-recorded vocal burst from laion/vocal-bursts-clean into the middle of a real SFT-corpus utterance, annotate the seam exactly, and train on the result. frustrated_groan is the test class (§43 measured it at 0.000; neither detector stage hears a groan on any of 20 real cue clips). Three gates: A — does Chatterbox voice conversion preserve the burst's identity while moving it into the surrounding speaker's voice; B — does the splice leave an audible seam; C — does the annotation parse back to the burst that is there. B and C passed at full scale. [P§49]

The gate reported a pass having compared nothing. The first pipeline run exited 0 with use_vc: false, proceed: true, and the build, training and comparison all ran on that verdict; n_bursts was 0. Proximate cause: the conversion worker wrote atomically to "<out>.wav.tmp<pid>" and soundfile infers the container from the extension, so all 288 calls raised TypeError: No format specified …, were caught, printed FAIL, and the worker exited 0 with ok=0 fail=288. Distal cause: every rate was a ratio over n = max(tot["n"], 1), so with no rows raw_heard was 1.0 and proceed = 1.0 >= 0.5 was True, while use_vc came out false — the same output shape as a gate that had run and found voice conversion unsafe. The gate now computes enough_data = n >= --min-n (default 100 of 264) and emits verdict ∈ {NO_DATA, PASS, FAIL_VC, FAIL_SOURCE}; re-run on the identical empty inputs, proceed flips true → false. Originals kept as gateA/*.void.json. Rule: 'A denominator floored to avoid a crash is a silent default answer.' [P§49.1]
§49.2 Gate A, 264 burst pairs and 24 speech controls
rawconvertedthreshold
names the source dataset's label exactly0.0340.008
names something in the right family0.3410.254
identity preserved0.360≥ 0.45
family preserved0.470≥ 0.60
heard as a burst at all1.0000.996
located by the full pipeline0.7200.610
speech-control WER, n = 240.1220.267
GATE A: FAIL. With the format fixed, 288/288 clips converted, and both identity thresholds are missed. The converted clip is still heard as a vocal burst 99.6 % of the time — it is a different one; a gate asking only 'did a burst survive' would have passed it. The dominant transformation is toward mumbled speech (Ahem → Low Mumble 15, Exhausted Groan → Low Mumble 10, Childlike Giggle → Low Mumble 9); survival tracks how speech-like the burst is (soft_whistle 0.71, snicker 0.67 against heavy_breathing 0.12, spitting 0.17, tongue_click 0.21; frustrated_groan 0.29). The corpus is therefore built from the clean burst unconverted, in a voice that is not the surrounding speaker's — the model can in principle learn 'the burst is where the voice changes'. Mitigation: independent levelling with a random −5…+2 dB jitter, and 1,003 of 2,000 rows carry a same-speaker reference clip. [P§49.2]

The raw arm bounds the approach. Scoring the classifier alone on the 264 raw clips (0.6–2 s) against vocal-bursts-clean's labels: strict 0.034, family 0.341, heard-as-a-burst 1.000, located by the full pipeline 0.720. §47's locator finding reproduces at n = 264 (classifier hears a burst on 264/264; full pipeline finds one on 190). The disagreement is systematic — Shriek → Scream 20/24, Soft Whistle → Low Mumble 17/24, Frustrated Groan → Exhausted Groan 14/24, Clicks Tongue → Ahem 13/24 — and five classes score 0.000 strict and 0.000 family. 'The ceiling on strict r_burst_cls for a perfect splice of a real burst is 0.034.' [P§49.3]

The adapter produces the burst and the metric scores 0.000. Gate B on the delivered shards: 8,000 junctions, seam-click ratio median 0.0036, fraction > 1 = 0.00088. Gate C: 1,956 / 2,000 = 97.8 % parse back exactly; 1,897 rows survive into training; 2,375 steps, 5 epochs. Against the shipped adapter on the burst_dose set, 9 paired prompts per cell, stack vs stack_new, w = 0 → 1.5; the w = 0 control is exact in all four panels. Strict: +0.000 at every weight in both kinds; old 0.000, new 0.000. Family-relaxed, solo: 0.000 → 0.259 at w = 1.0 (n = 9, t +2.80*, 5 of 9 up) and 0.000 → 0.296 at w = 1.5 (t +4.44*, 7 of 9 up) — while the family-relaxed wrong rate at the same cell rises 0.074 → 0.556 (+0.481, t 4.27). Inline is weaker and not significant (+0.111 at w = 1.0, t 1.41). At w = 0.5 the hit rate does not move on either metric while the wrong rate already rises +0.370. t = 2.80 at n = 9 is p ≈ 0.023. [P§49.4]

The alternative 'more bursts of any kind, some landing in the family by chance' was tested: P(family | wrong burst) is 0/17 for the old adapter and 10/40 for the new at w ≤ 1.0, Fisher p = 0.025 (all weights 2/116 vs 35/142, p = 2.0e-08). The label it emits is Exhausted Groan (7 of its 22 wrong hits at solo w = 1.0) — the same label the detector put on 14 of the 24 real clean Frustrated Groans the corpus was built from. ΔWER stays inside §28's 3σ floor (+0.104) for every w ≤ 1.0 and breaches it at 1.25 and 1.5. 'The synthetic data did what it was built to do, and the metric was asked the wrong question.' On the unbiased set burst_dose2 (a neutral carrier sharing no prompts, 10 per cell), inline w = 1.0: family-relaxed 0.000 → 0.533, +0.533, t 7.24, 10 of 10 prompts up, while the family-relaxed wrong rate falls −0.200; strict +0.000 at every weight. Both sets agree that w = 0.5 does nothing (+0.033 and +0.000) and that the effect lives at w = 1.0, the recommended merge weight. This is the mirror image of §48.7 (chuckle +0.200 strict, +0.000 relaxed): 'Either metric read alone calls one of these two studies a success and the other a failure, and in both cases backwards.' [P§49.4]

Rule and fix. 'When you manufacture training data for a metric, label it in the metric's vocabulary, not the source's.' The remaining ten classes should be built with each spliced burst labelled by the detector's own top-1; five of the eleven should not be built until then because the detector has no family-level word for them. What that does not buy: 'Aligning the corpus to the detector's vocabulary buys agreement, not ground truth' — the instrument loses 74 of 264 burst-length clips to its own locator and disagrees with the source dataset 96.6 % of the time. The deliverable is a listening page, old and new side by side at solo w = 1.0. [P§49.4]

II.9.8 Five more classes from manufactured data: the family-level effect replicates, the strict metric still cannot see it, and five classes are unmeasurable in principle [P§50]

Five new classes built with §49's fix (source bursts filtered so the detector's top-1 falls in the class's family), trained on §49's recipe verbatim, compared against the shipped adapters on burst_dose2 reduced to the six built classes, then aggregated: 960 rows, 48 (class, arm, weight) cells, 6 classes × 2 arms × 2 cues × 4 weights × 10 paired prompts, complete. Coverage was checked before any average. The class list is chosen by the instrument. vs_clsburst.py ran the detector over every clean source burst with the locator bypassed. Family consistency of the source pool — built: snicker 159/160 (0.994), cackle 376/403 (0.933), shriek 215/254 (0.846), frustrated_groan 138/227 (0.608), cough 83/348 (0.239), sniff 33/186 (0.177, built, borderline); attempted then skipped: heavy_breathing 9/200 (0.045), soft_whistle 1/446 (0.002), tongue_click 0/330, clicks_tongue 0/221, spitting 0/139. Reuse was capped at 10× (n_rows = min(2000, aligned × 10)), so cough has 830 rows and sniff 330. frustrated_groan was deliberately not rebuilt, so its numbers are the conservative, non-filtered ones. [P§50, P§50.1]

Five classes are unmeasurable in principle: vf_compare.family_of returns None for clicks_tongue, heavy_breathing, soft_whistle, spitting, tongue_click, so the family-relaxed metric has no family to score them in and the strict metric a name the detector never emits. 'A perfectly built spitting corpus would read +0.000 on both. Reporting a null there reports the instrument.' The detector hears soft_whistle → Low Mumble 327/446, tongue_click → Childlike Giggle 81 / Breathy Giggle 76 / Low Mumble 68, clicks_tongue → Ahem 81, heavy_breathing → Contented Sigh 67 / Exhausted Groan 45, spitting → no_burst 24 / Breathy Giggle 24. [P§50.1]
§50.2 stack_new vs stack, inline cue, w = 1.0, n = 10 paired prompts per cell
classstrict old→newΔ strictfamily old→newΔ familytwrong burst old→newΔWER
frustrated_groan0.000→0.000+0.0000.000→0.533+0.533*+7.240.633→0.967 (+0.333)+0.050
cackle0.000→0.033+0.0330.300→0.733+0.433*+2.900.700→0.867 (+0.167)+0.037
snicker0.000→0.000+0.0000.267→0.667+0.400*+2.570.767→0.800 (+0.033)+0.071
sniff0.000→0.033+0.0330.067→0.300+0.233*+2.690.733→0.900 (+0.167)+0.043
cough0.000→0.000+0.0000.100→0.200+0.100+1.150.800→0.867 (+0.067)+0.117
shriek0.000→0.000+0.0000.000→0.033+0.033+1.000.467→0.833 (+0.367)+0.062
POOLED+0.011+0.289*+6.04+0.189+0.063

On the solo cue the same six classes give pooled family +0.128 (t +3.10, n = 60), with only cough +0.133* and shriek +0.133* individually starred and frustrated_groan +0.200 (t 1.96, ns). Strict is +0.011 pooled, t +1.43, on both cues — driven entirely by one cackle prompt and one sniff prompt. 'Four of six classes reproduce §49.4's effect at family level, and two do not', and the pooled number is dominated by the two strong classes. [P§50.2]

Every family-level gain is bought with more wrong bursts. In four inline cells the gain exceeds the cost (frustrated_groan +0.533 vs +0.333, cackle +0.433 vs +0.167, snicker +0.400 vs +0.033, sniff +0.233 vs +0.167); elsewhere it does not — shriek gains +0.033 and pays +0.367 in wrong bursts, eleven times the gain; cough solo +0.133 vs +0.367; frustrated_groan solo +0.200 vs +0.333; §49.4's solo cell +0.259 vs +0.481. 'Making it emit the right burst is achieved in four of twelve class × cue cells.' The family-aware wrong rate moves the other way for the classes that worked (inline w = 1.0: cackle 0.400 → 0.167, snicker 0.500 → 0.133, frustrated_groan 0.633 → 0.433). [P§50.3]
§50.4 which label grows at w = 1.0, pooled over both cues, old → new
classlabel that grewold → newin the target family?
frustrated_groanExhausted Groan3 → 25yes — the label the detector gives the real clean Frustrated Groans
cackleChuckle / Childlike Giggle11 → 21 / 5 → 10yes (laugh); Guffaw 0 → 1
snickerChuckle / Childlike Giggle8 → 15 / 0 → 9yes (laugh)
coughAhem / Exhausted Groan3 → 10 / 3 → 8Ahem yes (throat), Exhausted Groan no
shriekSurprised Gasp2 → 21no — breath, not scream; Scream itself 0 → 4
sniffno_burst / Sharp Inhale5 → 16 / 5 → 8the biggest move is toward no burst at all
shriek is the diagnostic case: its source pool was 85 % aligned, its corpus the largest built, its training clean — and the adapter learned to emit a gasp. 'Detector alignment of the training data does not guarantee detector alignment of the model's output', a limit on §49's recommended fix that one class could not have shown. [P§50.4]
§50.5 §49.4 on two prompt sets
prompt setcuefamily old→newΔtn
burst_dose (sy1, train_seen, biased toward the OLD adapter)inline0.000→0.111+0.111+1.419
burst_dose (sy1)solo0.000→0.259+0.259*+2.809
burst_dose2 (sy2, neutral carrier)inline0.000→0.533+0.533*+7.2410
burst_dose2 (sy2)solo0.100→0.300+0.200+1.9610

All four cells are positive and the star moves between cues: 'a positive family-level shift of roughly +0.1 to +0.5, direction reproduced, magnitude not pinned down.' Quoting +0.533 (t 7.24) as the effect size would be quoting the luckier of two measurements. [P§50.5]

§50.6 dose, pooled over six classes, n = 60
cuew = 0.25w = 0.5w = 1.0
inline Δ family+0.011 (t 0.41)+0.050 (t 1.84)+0.289 (t 6.04)
inline ΔWER−0.007+0.016+0.063
solo Δ family+0.028 (t 0.93)−0.006 (t −0.23)+0.128 (t 3.10)
solo ΔWER+0.048+0.156+0.141
The effect is a step function at w = 1.0 on both cues. Inline WER stays inside §28's 3σ paired floor (+0.104) at every w ≤ 1.0; solo breaches it at both w = 0.5 and w = 1.0 — a guardrail failure the one-class study could not see. Thirteen cells breach a guardrail, twelve of them ΔWER and all but one on solo. Recommendation: w = 1.0 for bursts inside speech, no recommendation for bursts generated alone. w > 1.0 was not swept (§48: w = 1.5 unusable, pooled −0.174, t −4.70, ΔWER +0.329). [P§50.6]
Two family maps, and a comment that says they are one. vs_clsburst.py's map is headed 'same family map the gate-A report and vf_compare use' — it is not: the selection map carries mouth and nose; the scoring map in vf_compare.py carries neither and puts sniff in breath. sniff's pool was filtered under the narrow nose family (33 of 186) and scored under the wider breath family; the filter was stricter than the metric, so +0.233 is not inflated, but 33 source bursts are the whole of that class's evidence and the maps must be unified. 'A comment asserting an invariant that the code does not hold is worse than no comment.' [P§50.7]

What ships. $SC/publish/hf_burst_synth/ — six adapters as adapters/<class>/{adapter_config.json,adapter_model.safetensors}, plus manifest.json, README.md, RESULTS.md; 761 MB; not uploaded, staged for a human to push. The README carries the before/after table with n on every row; a 'never merge' warning with the three-line data_ptr() check showing audio_lm_heads.N.weight and audio_embeddings.N.weight are one tensor, and the set_weight helper through PEFT's scaling; that Chatterbox VC was tested and rejected at gate A (identity 0.360 vs 0.45, family 0.470 vs 0.60, speech-control WER 0.122 → 0.267) so bursts carry a different voice than the surrounding speech; the circularity caveat; and the recommended weight. Listening page $SC/pages/synth6/: 148 OGG, tables rendered statically into the HTML (§46's lesson: verify_page.js fails any container holding zero <audio>, and a stats-only container is exactly that); verify_page.js exits 0, 148 audio srcs, 0 missing. Rules added: aligning training data to the detector does not align the model's output to it; a class the metric has no family for is no result, not a negative one; check coverage per cell before averaging. [P§50.8, P§50.9]

II.9.9 2 Sep 2026 — The family-relaxed hit rate becomes the burst metric, and the table that carries it [P§50a, protocol heading '25.']

Decision, 2026-09-02, project owner, taken while round 3 was still generating — a pre-registration, not a metric chosen after seeing an outcome. (Translated:) 'the relaxed class boundaries are entirely fine.' A near-miss inside the same burst family counts as a hit; strict same-class hit stays as the secondary column. That promoted vf_compare.FAMILY from a footnote to the quantity everything is judged on, and it was not fit for the job: over the 64 classes still to be scored, 33 mapped to no family at all, so their 'relaxed' rate was silently identical to the strict rate. [P§50a]

Two defects in family_of, both fixed: tokenisation, not taxonomy — it required an exact token, so coughing missed cough, humming missed hum, panting missed pant, breathing missed breath, gulps missed gulp (light stemming with a doubled-consonant rule repairs it); and three missing families — sob (sob, cry, weep), whistle, and mouth (smack, lip, lick, kiss, chew, slurp, suck, drink, swallow, gulp, spit, tongue, click, tsk, hiccup); growl and snarl joined groan. Coverage 31 → 62 of 64; gurgling and hiss stay unmapped on purpose because a family of one would dress the strict rate up as a relaxed one. [P§50a]

Structural change. The table moved to code/burst_dose/burst_family.py (defining it in vf_compare or bd_agg made an import cycle). bd_agg.cell_metrics now emits hit_rate_fam, wrong_rate_fam and fam_known beside the strict pair, carried in METRICS so every paired t-test and pooled roll-up gets them. Re-aggregating round 2 unchanged except for the new columns: cells with any non-zero hit at any weight go from 37 to 90 of 140; the old aggregate is kept at bd2_agg.json.bak_pre_fam. Round 2's n is 10 prompts per cell, so these are indicative only. Round 3 (job 1619887, 16 nodes / 64 ranks, launched 2026-09-02): the 64 class adapters that have never had a weight recommendation (71 published minus 6 reported minus blend_genuine_emotional, which has no single target burst); the same ten carrier prompts as round 2, byte-identical; weights 0 / 0.25 / 0.5 / 1.0 / 1.25 / 1.5; both placements; nine samples per prompt instead of three — 90 generations per cell, standard error about 0.05 against round 2's 0.13. 'At n = 10 a hit rate of 0.17 and one of 0.27 are the same number.' [P§50a.1]

II.9.10 Two framings corrected by the owner [P§50b, protocol heading '26.']

Correction 1 — the emotion adapters and transitions. §35 reports that the emotion adapters compress endpoint contrast 4.587 → 1.982 and concludes they are the wrong tool for a turn. The measurement stands; the conclusion does not follow: for a transition the goal is a smooth path frame by frame across the 80 ms vectors, and endpoint contrast does not measure that — a trajectory that snaps between two states scores perfectly on contrast and is exactly what a transition must not do. The same re-reading applies to the two-take splice, which won on that yardstick and is by construction the maximally discontinuous answer. No continuity metric exists in the project, so both results are open, not negative. Candidates needing only re-scoring of clips on disk: per-frame trajectory monotonicity, largest single-frame jump, whether the path stays inside the convex hull of the two endpoints. [P§50b.1]
Correction 2 — the retrained burst adapters. The split verdict (strict +0.011, family-relaxed +0.289, t 6.04, frustrated_groan 0.000 → 0.533) is settled by the metric decision of §50a: on the accepted metric the retrain worked, and the strict column is the detector's naming resolution, not the adapter's failure. [P§50b.2]

II.9.11 2 Sep 2026 — The quality DPO adapter ships at 1.5 [P§50c, protocol heading '27.']

laion/moss-va-sft3-quality-dpo-lora (checkpoint 1504), merge weight 1.5, chosen by listening on 2026-09-02. The sweep could not choose: over 220 takes at eleven weights no genuineness delta reached significance (+0.42 at 1.0, +0.48 at 1.5, +0.58 at 1.75) and the only significant effect anywhere was a harm, word error +0.148 at 2.5. 1.5 is also the last step before the word-error cost appears (+0.006 at 1.5, +0.145 at 1.75). Recorded in wikiskills/coefficients.json -> global.quality_dpo with chosen_by: listening, so the layer distinguishes a scorer's verdict from an ear's. 'Second time on this adapter that an ear decided what a table could not: a DNSMOS ranking put 1504 last of ten states before a listener picked it as best.' [P§50c]

II.9.12 2 Sep 2026 — The other levers, stacked: the DPO adapter is a null, the dose ladder had not ended, and two prompt forms add up [P§51]

Agent vb_lever, 2026-09-02. Full report ~/reports/burst_levers.md, German status page ~/status_vocal_burst_levers.html, data $SC/out/vb_lever/vl_agg.json, code $SC/code/vb_lever/. (The section notes that the running series ended at 50 and a second block numbered 25–27 from another agent's series follows it, so 51 was free.) Fifteen classes spanning §43's tiers (five tier A, five tier B/C, five tier D including frustrated_groan, shriek, scream, one mouth class and one whistle class), in the demo server's own stack at its own doses: 5,360 cells / 32,160 clips over two jobs (1622728, 16 nodes, blocks A+B; 1623266, 8 nodes, block C at a second seed). Nothing merged; every adapter attached with PEFT and weighted through module.scaling[name]. Zero half-speed-canary violations. P0 was asserted byte-identical to §43's prompt strings on all 24 P0 sets before any GPU time. [P§51]

Lever 1 — the burst+stop DPO adapter (§44) is a null on burst realisation, heard for the first time. All three ends of its recommended 336 → 896 range, at weight 1.0, in the full stack: +0.007 (t +0.8), +0.017 (t +1.4), +0.006 (t +0.4) family-relaxed, n = 15 classes. Not a harm: WER falls slightly, DNSMOS rises slightly, genuineness flat. Step 616 cuts the miss rate by 0.033 (t −2.7, 13/15) — a burst happens more often, not more often the right one; step 896 buys genuineness +0.044 (t +1.9) and DNSMOS +0.026 (t +1.6). 'Audio does not narrow §44's range either. If it ships, ship 896, and ship it for genuineness.' [P§51.1]

Lever 2 — the combination is redundant. both_S − dose is +0.001 … +0.006, all |t| ≤ 0.4; both_S − dpo_S is +0.077 … +0.092 (t 3.5–4.2). All of the gain belongs to the burst adapter; confirmed in block C (P0 + DPO-BS 896 on the best cell: −0.010, t −0.7). The one measurable interaction is a small harm to duration control (|dur err| +0.020 s, t +2.9). [P§51.2]

The dose ladder had not ended. §43 stopped at 1.5. One more level per class (w*+0.5) is worth +0.024 family (t +1.3) and +0.022 strict (t +2.3) — as much strict accuracy again as the whole step from no adapter to §43's recommendation — at WER +0.130 (t +2.2, 14/15). Block A put eleven of fifteen classes at its own new top, so block C carried the ladder to w*+2.0, capped 3.0. Now the optimum is interior for eleven classes, typically w*+0.5 … +1.0, and beyond it the curve falls hard (relief_sigh 0.717 → 0.033, clears_throat 0.283 → 0.000, cackle 0.540 → 0.267). Three (soft_hum, surprised_gasp, frustrated_groan) are still rising at 3.0. Hard ceiling: nine of 400 ladder cells produced no decodable audio at all, every one at w ≥ 2.3. [P§51.3]

Lever 3 — prompt forms. Cause sentence in the GENERAL line (g): +0.026 (t +1.9) alone, free apart from −0.035 genuineness. Longer stated duration (l): +0.022 (t +1.1) alone; costs WER +0.095 (t +5.9) and blend −0.57 (t −5.0). Both together (C_gl): +0.044 family (t +1.9, 10/15) and +0.030 strict (t +3.3, 9/15) — the only prompt result significant on the strict metric; +0.026 and +0.022 make +0.044, +0.002 and +0.023 make +0.030, cleanly additive; costs add too (genuineness −0.110, t −3.8; WER +0.075; blend −0.51). The cue as an action ((she groans in frustration, 0.5 seconds)): −0.077 … −0.106 (t −2.1 … −3.3), miss rate +0.120 (t +3.4) — the model degrades to silence, while blend rises +0.183 and DNSMOS +0.032: 'the study's clearest case of a metric moving the right way for the wrong reason.' Mid-clause placement: −0.073 … −0.124, miss rate +0.306 … +0.366 (t +8.4 … +10.4, 15/15 worse) — read the other way, the clause boundary is worth about 0.3 of miss rate; one class inverts it, clears_throat 0.250 → 0.483 mid-clause. Neighbour substitution: null on family (−0.000 … −0.022), a harm on strict (−0.021, t −2.9, 1/15); for frustrated_groan, asking for an exhausted groan gives 0.133 against 0.117. 'The cheapest hoped-for fix for the dead classes does not work.' [P§51.4]

The longer-duration gain belongs to the detector, not to the model. Over 15,409 cues where something was detected, P(family) rises with the detected duration — 0.282 (< 0.3 s) → 0.388 → 0.476 → 0.510 (0.8–1.2 s), then saturates — while P(strict) rises by a third as much (0.121 → 0.177). Against a +1.0 s longer request the delivered burst grows by only +0.054 s (t +2.3); the slope × 0.054 s predicts +0.020 and the observed gain is +0.012, so there is no residual to attribute to better realisation; across classes the correlation between how much longer the burst got and how much the hit rate rose is +0.53. [P§51.5]

Seed floor. Block C re-ran block A's best cell for every class at a second seed: mean difference −0.010 (t −0.6), no drift, but sd 0.068, max |diff| 0.167. Every paired comparison is immune; every absolute recipe number is the argmax over about 35 draws and is optimistic by roughly one sd. Recipes are reported at the point estimate and at p − 0.068. What comes out instead: 14,451 wrong-label events; Contented Sigh (26.3 %), Low Mumble (20.0 %) and Ahem (12.3 %) are 58.6 % of them, the most frequent burst classes in the corpus — a prior collapse, of which §43's 'down the arousal axis' reading is a special case. The adapter moves the substitution toward the target's family without reaching it: clears_throat Low Mumble → Ahem, shriek Low Mumble → Scream, cackle Contented Sigh → Breathy Giggle + Chuckle. 'The adapter buys the family, not the member.' [P§51.6, P§51.7]

Recipes. Best cells (family-relaxed, control in brackets): relief_sigh 0.750 (0.383, cause, w 0.8), chuckle 0.725 (0.333, C_gl, w 2.0), contented_sigh 0.683 (0.450, C_gl, w 1.5), soft_hum 0.683 (0.217, C_gl, w 2.0 + DPO-BS 896), cackle 0.540, clears_throat 0.483 (mid-clause), low_mumble 0.483, childlike_giggle 0.467, frustrated_groan 0.267 (0.017), exhausted_groan 0.217, surprised_gasp 0.183, scream 0.167 (0.000, t +4.7). Twelve of fifteen cross 0.15; nine still cross it after the seed floor. scream's effect is the most reliable in the table (t +4.7 against a control of exactly 0.000) even though its level is not. [P§51.8]

No recipe: shriek 0.117 (best-of-47), lip_smack 0.017 (one family hit in 2,159 cues over 36 conditions), sharp_whistle 0.000 — not one hit in 2,160 cues across 36 conditions, weights 0 to 3.0, both script kinds, all seven prompt forms, both seeds. 'The mouth and whistle families are absent, not weak. Tier D is reachable only where a dose past §43's range reaches it; where it is not, it needs data, not a knob.' [P§51.8]
DNSMOS should not be used to choose between these recipes: its total moves by < 0.04 for every lever and its only significant component under dose is p808 (−0.081, t −2.5), but it moves +0.38 (t +16.7, 9/9) between two script kinds that differ only in how much speech is in the clip — on this material it measures clip composition. Blend is the consistent casualty (0.4–0.6 of ten for every lever that raises the hit rate); word error is the price of length, not of dose (§43's dose +0.003, one more level +0.130, a longer cue +0.095). Rules: a best value at the edge is not an optimum, and a new best at the new edge is not a fix; quote an absolute hit rate with its seed floor; a metric that improves while the difficult thing stops happening is not a win. [P§51.9, P§51.10]

II.9.13 2 Sep 2026 — The other fifty-five burst classes: nineteen recipes, nineteen classes that do not exist for this model, and a prompt result that did not generalise [P§52]

Code $SC/code/vb_ext/, data $SC/out/vb_ext/ve_agg.json, tables $SC/out/vb_ext/tables.md, report ~/reports/burst_ext.md, listening page $SC/hf_vb/ (staged, not uploaded). Applied to the fifty-five classes §51 did not cover: nineteen cross the 0.15 family-relaxed bar, sixteen still cross it after the seed floor, seventeen realise but stay below it, nineteen never realise at all; with §51 the model has a written operating point for 31 of its 70 burst classes. Grid: nine conditions per class — a control at burst weight 0, four burst weights, and two prompt forms (P0, C_gl) crossed with those weights; ten carriers, six takes — 487 conditions, 4,870 cells, 28,692 takes, one job on 16 nodes, 1 h 14 m. Not measured again: the DPO adapter, the action cue, mid-clause placement, neighbour substitution. The weight ladder is per class: for the 35 classes with a non-zero §43 curve max(0.25, w*−0.5), w*, w*+0.5, w*+1.0; for the 20 whose §43 curve is zero at every weight an absolute ladder 1.0 / 1.5 / 2.0 / 2.5. The seed floor sd = 0.068 is inherited from §51; every recipe is quoted at point − 0.068. All 55 P0 prompt sets are asserted byte-identical to §43's; a cell with no decodable audio is written with n_ok = 0 rather than dropped (§51 dropped them). [P§52, P§52.1]

§52.2 dose pooled over 55 classes, paired against w = 0
ladder positionΔ hit famtΔ hit stricttΔ blendΔ WER
one level below §43+0.020+3.6+0.004+1.5−0.31+0.008
§43's own weight+0.024+3.0+0.010+2.6−0.66+0.085
w*+0.5+0.022+2.9+0.012+2.2−1.04+0.187
w*+1.0+0.033+2.9+0.021+2.4−1.21+0.284

§43's weight is below the optimum for thirteen of the nineteen working classes; eleven have an interior optimum, six sit at the top of their ladder, two at the bottom. The costs are §51's and scale with dose rather than switching on at a threshold. [P§52.2]

C_gl does not generalise — it is an interaction with dose, not a lever, which corrects §51's recommendation. Over fifty-five classes: best C_gl cell against best P0 cell −0.004 (t −0.6), C_gl better for 13 of 55; paired at matched weight −0.011 (t −2.4) one level below §43, +0.000 at §43's weight, +0.016 (t +1.7) at w*+0.5 (+0.057, t +2.2, restricted to the nineteen that work), −0.001 at w*+1.0; on strict a null at every position (|t| ≤ 0.9) — §51's +0.030 strict does not appear. §51's block C ran every prompt form at its block-A best weight, near w*+0.5, exactly where the form pays. 'A prompt form measured at one dose is a claim about that dose.' [P§52.3]
Six winning cells are the model going silent: displeased_grunt at 2.0 C_gl has a Parakeet WER of 1.000 and a duration error of −2.14 s; cough 0.854, fearful_gasp 0.833 (−1.26 s), humming 0.827, ahem 0.692, coughing 0.657, against controls of 0.27–0.41. A speech-safety gate is added: a cell qualifies if its Parakeet WER is at most 0.10 above the class's own w = 0 control. Sixteen of nineteen keep a safe recipe at or above 0.15 (four move down the ladder: cough 0.408 → 0.233, fearful_gasp 0.429 → 0.283, humming 0.350 → 0.233, breathy_giggle 0.417 → 0.400); ahem, coughing and displeased_grunt have no safe cell and ship flagged. Across all 432 adapter-on cells the correlation between Δ hit and Δ WER is −0.033 — a concentrated failure, not a general trade. [P§52.4]

Decodability ceiling. 42 of 4,870 cells produced no decodable audio (0.9 %). Nothing fails below w = 1.25; 1.25 → 0.006, 1.5 → 0.007, 1.8 → 0.023, 2.0 → 0.011, 2.3 → 0.037, 2.5 → 0.032. displeased_grunt alone accounts for 25 of the 42. 'A recipe above w = 1.8 needs a retry, because roughly one take in thirty comes back empty.' [P§52.5]

Nineteen classes that do not exist for this model: clicks_tongue, convulsive_sob, gulps, gurgling, hiccup, hiccups, hiss, nervous_gulp, person_whistling_to_get_attention, quiet_sob, smack_one_s_lips, smacks_lips, sobs, soft_whistle, spitting, swallows, tongue_click, tsk, wolf_whistle produced not one hit in the 531–540 scored cues each received, at every weight 0.25 to 2.5, both prompt forms, through the code path that produced 0.550 for wistful_sigh. Of thirteen mouth classes eleven are at exactly zero and the best is 0.017; all three whistle and all three sob classes are at zero. Every other family has a working member: sigh 0.550, breath 0.429, throat 0.408, hum 0.367, laugh 0.333, groan 0.250. With §51's lip_smack and sharp_whistle, none of the fourteen mouth or four whistle classes exceeds 0.017 at any weight either study tried — the same conclusion §48 and §49 reached from the data side. [P§52.6]

Listening page. $SC/hf_vb/ holds 31 classes — §51's twelve at the recipe its WikiSkill pages publish and §52's nineteen at their shipped recipe — six takes at the recipe and two controls at burst weight 0, each clip carrying the detector's strict and family verdict, at a different seed from the measurement (each recipe is an argmax and carries a winner's curse). OGG Vorbis through qa_listen.to_audio_b64 (torchaudio 2.9.1 routes save through TorchCodec, which has no working FFmpeg here). ve_pagesim.js evaluates the real inline script against the real data.js. Final state: 238 of 238 players render, every file present, plus ten cards for takes with no decodable audio, shown rather than omitted. [P§52.7]

ve_pagesim.js's first run failed: data.js carried no takes array and the renderer threw before drawing anything — §46 repeating itself, caught before a single clip had been generated. Two earlier audio versions were discarded as true numbers reading as false ones: v1 put all six takes on one carrier sentence (cackle 1/6 against a measured 0.54); v2 spread them over carriers 0–5, the harder half (per-carrier family rate 0.034 … 0.160, 0–5 averaging 0.073 against 0.101 for 6–9). Restricted to the same six carriers the recipe arm was only −0.069 (t −1.9) and the control arm −0.070 (t −1.9), so the generation path was sound; v3 uses stride [0, 2, 3, 5, 7, 8], fixed before any hit rate was looked at. [P§52.7]

The page as a free replication. 58 of 186 recipe takes land in the right family (31 %), 17 strict (9 %), against 4 of 62 for the controls (6 %). Per class, observed minus measured is −0.096 (t −3.2) for §51's twelve recipes, −0.022 (t −0.6) for §52's nineteen, −0.051 (t −2.0) over all 31: the study that selected over about 35 cells per class is optimistic by about four times as much as the one that selected over 9, so the point − 0.068 convention 'is about right for a nine-cell argmax and about a factor two too generous for a thirty-five-cell one.' Rules added: a prompt form measured at one dose is a claim about that dose; gate a recipe on the speech before shipping it; record silence as a measurement; do not bracket a meaningless argmax; a demo sample must be drawn the way the measurement was. [P§52.7, P§52.8]

II.9.14 Classifier-free guidance on the burst: it works, g = 4 triples the hit rate, and the prompt is doing most of the work [P§53]

Study code/vb_cfg/, data out/vb_cfg/run1, page pages/cfgburst/, report ~/reports/burst_cfg.md. Job 1632593, 4 nodes × 4 GPU = 16 ranks, 32 min 41 s, EXIT=0, 624 cells / 1,871 clips, every clip kept as OGG. Base out/sft3/export + out/dpo_sft3_p2/ckpt-step5022/adapter, on the stack the demo server ships. The existing CFG (cfg_gen.py) neutralises the emotion condition in the unconditional branch; this tests a proposal that moves two axes at once — unconditional = neutral prompt AND burst adapter off, conditional = cued prompt AND adapter on. cfg_gen.py could not be reused (no per-branch adapter state); of three wirings the chosen one toggles scaling["tgt"] between branches, since a LoRA layer computes h + scaling[name]·B(A(x)). Because audio_lm_heads.0…11 are LoRA targets in both the burst adapters and the DPO adapter, the toggle must reach inside the 12-channel loop: the index is split into 252 semantic transformer modules (once per branch per frame) and 16 (local_transformer.* + the 12 heads, per channel). Cost of the toggle is below noise (arm E 1.98×, arm D 1.92×). [P§53, P§53.1, P§53.2]

Wiring control. At g = 1 the unconditional term cancels, so D1 and E1 must reproduce C10; lu + 1.0·(lc − lu) is not lc in float32, so it was measured: g1_max_logit_dev is max 9.5e-07, mean 1.1e-07 over 96 D1/E1 cells, and D1 − C10 = E1 − C10 = 0.0000 on every metric with identical frame counts. Design. Same 8 prompts, same seeds (1234*1000003 + j), 3 takes pooled per cell, n = 48 paired prompts. Arms: A adapter off / neutral; B adapter off / cued; C10 adapter 1.0 / cued; C15 adapter 1.5 / cued; D{g} cond (1.0, cued) vs unc (0.0, neutral), g = 1…5; E{g} cond (1.0, cued) vs unc (1.0, neutral), g = 1…4 — E holds the adapter on in both branches so only the prompt cue is amplified. Classes scream (496 verified rows), shriek (48), frustrated_groan using the synthetic re-train (bucket_loras_synth/vbr/). NEUTRAL replaces (scream, 0.6 seconds) with [0.6 seconds pause]; CUED adds one GENERAL-line sentence naming the sound. [P§53.3, P§53.4]

The shipped frustrated_groan adapter was checked first and scores 0.000 strict and 0.000 family at all ten weights of §43 in both positions — 'amplifying an identically-zero difference vector is a test of nothing' — so the synthetic re-train was used for that class. [P§53.4]
§53.5 pooled strict / family-relaxed hit rate and Parakeet WER
armstrictfamilyWER
A0.0000.0000.165
B0.0070.0070.148
C100.0420.0830.181
C150.0760.1880.197
D20.1250.2360.255
D30.1250.2290.340
D40.1390.2920.404
D50.1040.2010.488
E20.1040.2150.167
E40.0900.1810.204

Paired against C10 (n = 48): D4 +0.0972 strict (t +2.72), +0.2083 family (t +4.41), WER +0.2235 (t +5.63). D2 +0.0833 / +0.1528 at WER +0.0741. E2 +0.0625 / +0.1319 (t +3.25) at WER −0.0141 (t −0.77) — free. C15 +0.0347 / +0.1042. g = 5 degenerates: D5 is worse than D4 on the burst metric and +0.084 worse on WER, confirming the field ceiling of 4.0. [P§53.5]

§53.6 D against E
D2 − E2D3 − E3D4 − E4
d strict+0.021 (t 0.72)+0.028 (t 0.94)+0.049 (t 1.41)
d family+0.021 (t 0.50)+0.049 (t 1.12)+0.111 (t 2.22)
d WER+0.088 (t 3.44)+0.147 (t 4.24)+0.200 (t 5.35)

D never beats E on strict, beats it on family only at g = 4, and pays +0.09…+0.20 WER at every g. The exception is where a well-trained adapter meets an in-line cue: on scream / inline, D4 − E4 = +0.333 strict, t +3.74. scream / inline (496 rows) goes 0.208 → 0.583 strict at D4 — 'the largest movement any lever in this project has produced on a burst class'; shriek (48 rows) and frustrated_groan never cross 0.083 strict under any arm, though frustrated_groan / inline goes 0.042 → 0.458 family at D4. 'Guidance amplifies a direction that exists; it does not create one.' Guardrails: the wrong-burst rate falls (D4 −0.097 against C10) and wrong labels move toward the right family (C10's top three Surprised Gasp 30 / Contented Sigh 17 / Wistful Sigh 17; D4's Exhausted Groan 19 / Surprised Gasp 13 / Low Mumble 11); length control survives (|dur err| 0.057 → 0.067 s; no D cell reached the 340-frame cap); the WER cost is voice damage on the failures — D4 scores WER 0.163 on takes that hit and 0.443 on the rest (C10: 0.178 / 0.181), while E shows the opposite smaller pattern (0.297 / 0.195), the burst confusing the transcriber. [P§53.6, P§53.7, P§53.8]

§53.9 D{g} − C15 (guidance against weight-scaling)
D2D3D4D5
d strict+0.0486 (t 1.36)+0.0486 (t 1.31)+0.0625 (t 1.64)+0.0278 (t 0.78)
d family+0.0486 (t 0.93)+0.0417 (t 0.81)+0.1042 (t 1.91)+0.0139 (t 0.26)
d WER+0.0574 (t 2.11)+0.1425 (t 4.32)+0.2068 (t 6.16)+0.2909 (t 6.99)
Against §30 Q3 the sign flips, but only on the target metric. Q3 found logit-scaling lost to weight-scaling, worse as g rose (emo g 1.5 −0.0191 t −1.88; emo g 3.0 −0.0844 t −3.68; vn g 3.0 −0.0729 t −2.34); here every guided arm is above weight-scaling at every g. Both reasons to doubt the transfer were verified: (a) the baseline Q3 lost to is broken for these classes — C15 buys only +0.035 strict (t 1.53, 6 of 48 prompts) over C10; (b) varying the prompt too is a different and larger signal — arm E recovers +0.132 family (t 3.25) of D4's +0.208 at zero WER cost, while D − E is not significant on strict at any g. Q3's direction survives on the cost axis: WER rises monotonically (+0.074, +0.159, +0.224, +0.308). [P§53.9]

Wall clock. 0.261–0.265 s per generated frame single-branch, 0.505–0.540 two-branch — 1.91–2.04×, mean 1.94× over 624 cells, inside the project's four previous 1.92–1.96× measurements. Recommendation: ship arm E at g = 2 — cue the burst in the script and on the GENERAL line, keep the burst adapter on in both branches, guide at 2.0: +0.132 family-relaxed (t 3.25) at WER −0.014, at 1.9× wall clock. Reserve arm D at g = 3–4 for classes with a well-trained adapter and an in-line cue (scream: +0.375 strict at t 2.55, WER +0.059 t 1.07, not significant). Do not use g = 5. Not settled: the detector remains the instrument, naming the dataset's own label on 3.4 % of 264 clean source clips against 34.1 % for the right family; 701 clips are staged at pages/cfgburst/ (18 cards, verify_page.js, 0 audio srcs missing) so the 0.208 → 0.583 claim can be checked by ear. [P§53.10, P§53.11, P§53.12]

II.9.15 DramaBox can be asked for a scream in plain English; our detector cannot be asked whether it got one [P§54]

Question. Our burst adapters barely produce screams (5 of 50 clips on the weight ladder) and never shrieks (0 of 50), and the manufactured corpus that fixed frustrated_groan numerically was judged on listening not to sound like the real thing. Can ResembleAI/Dramabox — a 3.3 B IC-LoRA fine-tune of the LTX-2.3 audio branch — produce these bursts natively, well enough to become a data source? Design. 20 screenplay-format scenes (10 shriek, 10 scream; five male and five female each; sixteen distinct situations plus the owner's four worked examples verbatim), three seeds each = 60 clips. Fixed settings recorded on every clip: CFG 2.5 → cfg_scale, STG 1.5 → stg_scale, breathing 1.10 → duration_multiplier, fixed duration 0.0 → gen_duration (auto), reference window 10.0 → ref_duration (inert: voice_ref=None). Watermarking off, generate() in place of generate_to_file(). Scored with reward.RewardModel(...).bursts(). Generation is solved: 60/60 clips, 7.0–8.1 s per about 11 s clip on one GH200; four GPUs produced the grid in about six minutes including a 68 s model load. [P§54]

§54 detector result
strictfamily-relaxed
shriek, 30 clips0 (0 %)6 (20 %)
scream, 30 clips3 (10 %)3 (10 %)
all 603 (5 %)9 (15 %)
Compared with the burst adapters (scream 5/50 = 10 %, shriek 0/50 = 0 %) the result is identical, and the obvious reading — DramaBox is no better, the plan is dead — is wrong. Across all 60 clips the detector emitted 73 events carrying 8 distinct labels of its 83, and Shriek was not among them, zero times; two labels (Surprised Gasp, Contented Sigh) account for 70 % of everything said; 9 of 60 clips yielded no located span at all, including one whose loudest 300 ms sits 9.4 dB above its own conversational median; median confidence over all 73 events is 0.319; the locator proposed 1.2 spans per clip on eleven seconds of scripted screaming. shriek 0/30 measures the detector, not the model. Where the detector does fire on Scream it fires in the right place (all ten events between t = 0.02 s and 4.36 s) with the highest median confidence of the three main labels (0.440 against 0.333 and 0.300). [P§54]
§54 detector-independent acoustics, peak 300 ms window against the clip's own median
npeak-over-median dBsustain sHF ratio at peak
DramaBox scream306.14.700.029
DramaBox shriek307.93.950.036
MOSS adapters, scream inline259.83.050.011
MOSS adapters, scream solo258.41.300.007

DramaBox puts 3–5× more energy above 2 kHz at its loudest moment — the spectral signature of vocal strain (shriek vs MOSS inline p = 0.02, rank-biserial +0.37) — and honours the shriek/scream distinction the detector cannot represent: its shrieks are shorter than its screams, 3.95 s vs 4.70 s sustain, p = 0.014. 'The two systems that scored identically at 10 % are not producing the same sound, and the metric that equated them is the one that cannot tell them apart.' [P§54]

Findings that cut against the plan: DramaBox has lower dynamic contrast than our adapters (6.1 dB vs 9.8 dB, p = 4×10⁻⁵) because it performs the whole scene at high intensity; only 4 of 60 clips reach 10 dB over their own median, so training on such scenes would teach 'this whole utterance is loud' — a defect of the scene design (no calm lead-in in any of the 20). A seed effect larger than the class effect: all three strict scream hits are seed 9012, four of six shriek family hits are seed 1234 (about a 1-in-9 coincidence; suggestive, not established). Licensing: the LTX-2 Community License defines Derivatives to include 'methods based on the generation of synthetic data by LTX-2 for training the other model' and §6(b) requires any Derivative to be distributed exclusively under that licence, with a paid commercial licence for entities with revenue ≥ $10 M; a burst adapter trained on DramaBox output would ship under those terms — an owner decision. [P§54]

Verdict. 'Not yet, and not for the reason the run was built to test.' Whether the audio sounds right is a listening judgement staged at $SC/pages/dramabox/ (60 clips with prompt, seed, settings and detection per card, verified). What is settled is that the filter, not the generator, blocks the data-source plan: at 5 % strict yield one generates 20 clips per usable row, and for shriek no quantity produces rows because the label is never emitted. Lesson, met for the third time: 'Before a measurement decides something, check that the measurement's own vocabulary contains the answer' — one line, count how often the target label appears anywhere in the run, and it inverts the conclusion. [P§54]

II.9.16 A multimodal model annotates the bursts our detector cannot name: 88 % against 5 %, zero invented labels, and spans too wide to train a locator on [P§55]

Method. All 60 DramaBox clips to gemini-3.8-flash through hyprlab, one call per clip: POST /v1beta/models/gemini-3.8-flash:generateContent, audio as inline_data (base64 OGG), the full 83-class taxonomy in the prompt, a typed responseSchema asking for every burst as start/end in seconds, 1–3 labels most likely first, a confidence and a free-text description, with explicit permission to answer no_burst; temperature: 0. 60/60 parsed first try. Code $SC/code/vb_gemini/, responses cached at $SC/out/vb_gemini/cache/. gemini-3.8-flash does not appear in GET /v1beta/models (newest listed is 3.7-flash) but the POST works and modelVersion echoes 3.8 on every call. The request was blind: the prompt never contains 'scream' or 'shriek' outside the taxonomy, never says the clips were generated to contain a burst, never carries the scene; the requested class was joined afterwards. Labels were typed as free strings rather than an enum, so inventions would be visible: 0 off-taxonomy labels in 445 emissions across 191 events. [P§55]

§55 Gemini against the project detector
Gemini top labelGemini any of 1–3our detector, strict
shriek scenes (30)46.7 %76.7 %0 %
scream scenes (30)93.3 %100 %10 %
all 6070.0 %88.3 %5.0 %

88.3 % against 5.0 %, Fisher p = 1×10⁻¹³. Shriek is emitted 68 times where the detector emits it never; 31 of 83 classes are used against the detector's 8. The shriek/scream distinction carries signal: Shriek is the top label on 14/30 shriek scenes and 2/30 scream scenes, Fisher p = 0.00091, in the direction §54's acoustics predicted. Allowing 1–3 labels is worth 30 points and was the owner's call: 188 of 191 events came back multi-label, Scream + Shriek the most common pair (64); forcing a single label costs shriek recall 76.7 % → 46.7 %. Where the two systems disagree they disagree totally: 68 of the detector's 73 events overlap a Gemini span, but only 9 of 68 top labels match; where the detector says Contented Sigh, Gemini says Scream 15 times, where Surprised Gasp, Scream 12 times; where the detector manages Scream the two agree 8/10. There is no clip where the detector finds a scream-family burst and Gemini does not; Gemini alone finds one on 50 of 60. [P§55]

§55 timing, detector-independent (loudest 300 ms window inside an annotated span)
union coverageloudest moment insidelift over chance
Gemini, all spans38.2 %48.3 %1.27×
Gemini, scream-family only28.5 %46.7 %1.64×
our detector6.5 %23.3 %3.57×
The result that cuts against the plan: Gemini annotates 38 % of every clip, 3.18 bursts per clip against 1.22, and beats chance placement by ten points; the detector — wrong about what almost every time — is nearly three times more precise about where per second it commits to. The labels are a large improvement, the spans are not, and training the locator on these would teach wider onsets. [P§55]
The negative control: Gemini set no_burst 0/60 — worth nothing on material written to contain a burst — so 20 hard negatives were cut from the longest stretches Gemini itself left unannotated and re-submitted. It answered 'no burst' on 1 of 20, and 11 of 20 came back with Scream or Shriek as the top label, in audio the same model had just declared empty: with the real scream present it reserves Scream for it; with the scream cut away, the shouted dialogue becomes the scream. 'The labels are context-dependent and not stable under re-segmentation.' Binding constraint: annotate whole clips and cut rows out of the spans; never annotate rows. [P§55]

Cost. $0.1361 for 60 clips = $0.00227/clip, $2.27 per 1,000; $153 for all 67,540 training rows, $98 for the 43,139 the audit scored. Thinking tokens are 59 % of the bill (715/clip against 278 visible); 810 of 1,082 input tokens are the fixed taxonomy, so short rows cost about 80 % of an 11 s clip. Throughput 8.8 clips/min at 3 workers, no 429 seen. Staged at $SC/pages/gemini_annot/ — per clip: player, Parakeet transcript (generated for this), the DramaBox prompt, Gemini's spans, the detector's events, the loudest moment; sorted most-disagreement-first; verify_page.js exit 0, 60 players, 0 missing. Report ~/reports/gemini_annot.md. Verdict: good enough to retrain the classifier on, pending listening; not good enough to retrain the locator on without a span-tightening pass; not safe to run on short rows at all. Caveat: every number here is an agreement statistic, and Gemini and DramaBox are both large generative models trained on dramatic speech, so 'the prompt said scream, the audio sounds like screaming, the label is Scream' can be one shared prior — 'when the new instrument agrees with the thing it was pointed at, check that they are not the same instrument.' [P§55]

II.9.17 Adapter capacity is not the bottleneck: rank 32 and rank 64 on the five classes that have data and still will not fire [P§56]

Five classes have verified training data (141–540 strict rows) and still barely produce their burst at any merge weight in the §51 sweep: yawn, scream, wistful_sigh, deep_breath, surprised_gasp. Every burst adapter is r = 16 / alpha = 32; this section varies it. shriek is excluded (48 strict rows, below the 100-row threshold). Design, fixed before any number was looked at. Seven adapter roots × 5 classes: r16 (the deployed incumbent), r16r (rank 16 retrained — the training-noise control), r32, r64 on the incumbent's own rows, and r16s, r32s, r64s on the §48 strictly filtered rows. alpha/r held at 2.0 (train_bucket_loras.train_one sets lora_alpha = 2 * rank); optimiser steps equal across ranks by construction (total = ceil(rows/batch)·epochs). Dose re-swept per root over w = 0, 0.25, 0.5, 0.8, 1.0, 1.25, 1.5, 2.0 inside the deployed stack (DPO 1.0 + emotion 1.0 + genuineness_high 0.25 + blend_high 0.5 + burst @ w), with a wrong-class adapter control at every rank and an explicit donor map (bd2's --mm-offset 35 would have paired deep_breath with fast_breathing). 790 conditions, 7,900 cells, 23,700 generations, SLURM 1632827, COMPLETED in 2 h 51 m on 32 GPUs; 7,900 / 7,900 rows, 0 empty cells, 0 canary violations. [P§56]

Three checks passed first. (i) Zero check: at w = 0 all seven roots are the same model — max cross-root deviation over 60 panels 0.000e+00. (ii) Replication: stack@r16 on bd2's own prompt file at bd2's seeds reproduces bd2's stack arm on 70 of 70 shared cells exactly, on strict hit rate, family hit rate and a floating-point WER. (iii) Mismatch control: no wrong-class cell exceeds 0.020 strict and there is no rank trend, retiring the confound 'a bigger adapter is a bigger perturbation and the detector answers to size'. The treatment took: end-of-training loss falls monotonically with rank in 10 of 10 class-ladders (original rows r32 −0.103, r64 −0.275 against r16r; strict rows −0.082 and −0.217), while at step 50 the three ranks agree to 0.010. [P§56]

§56 paired contrasts, pooled over 2 prompt kinds and 7 non-zero weights, n = 700
contrastd strict hit95 % CItd family hitd WER (Parakeet)
r16rr16 (same recipe, new init draw)−0.004[−0.015, +0.007]−0.76−0.003+0.009 n.s.
r32r16+0.009[−0.003, +0.021]+1.50+0.009+0.003 n.s.
r64r16+0.010[−0.003, +0.024]+1.57−0.001+0.050, t = +2.21
r32sr16s+0.014[+0.003, +0.026]+2.44+0.014+0.007 n.s.
r64sr16s+0.007[−0.005, +0.020]+1.12+0.002+0.027, t = +3.88
r16sr16+0.001[−0.011, +0.014]+0.23+0.002+0.014 n.s.
Nothing on the original ladder is significant, nothing survives correcting for six comparisons, and 'rank 64 — four times the parameters, a training loss 0.275 lower — is indistinguishable from rank 32 and from zero while being the only arm that costs word error rate.' r16sr16 = +0.001 independently replicates §48's null on four classes it had never measured; the capacity × label-noise interaction comes back flat (+0.014 on clean labels against +0.009 on the original rows). [P§56]
The control this section would have been a false positive without: train_bucket_loras.py seeds the data order and not torch, so PEFT's Kaiming draw for LoRA A is fresh on every run. r16r measures that for the first time: pooled −0.004, but per class as large as the rank change — re-running the incumbent's own recipe moved wistful_sigh's held-out best from 0.033 to 0.200 (rank 32 gave 0.233, rank 64 0.267) and scream's from 0.033 to 0.100, where rank 64 gave 0.000. 'Standing rule from here on: no claim that adapter A beats adapter B without a re-run of B.' [P§56.1]

The dose did not move with rank. A rank-64 adapter puts roughly twice the norm into the residual stream at the same merge weight (3.27 against 1.53 on a toy layer), predicting an earlier peak; instead all seven roots rise monotonically and peak at the same weight, w = 2.0 — pooled d hit at w = 2.0 is +0.100 / +0.127 / +0.123 for r16 / r32 / r64. The shipped BURST_LAM 0.25 / BURST_LAM_INTENSE 0.5 are not wrong for a higher-rank adapter but sit at the bottom of the ladder where every root is flat; the weight is roughly an order of magnitude the larger lever (w = 0 → 2.0 moves the pooled hit rate +0.100 to +0.127, against +0.016 for the largest rank effect at matched weight). Where rank does anything it renames: yawn +0.000 at both r32 and r64 over all fourteen cells (504 rows); deep_breath −0.002 / −0.000; scream +0.017 / −0.010, wrong sign at the higher rank; surprised_gasp +0.024 at r64 only; wistful_sigh the largest (+0.033 / +0.038 strict) with family-relaxed +0.007 / −0.026 — its family rate is already 0.533 at w = 0. 'Rank is a bigger multiplier on a direction that must already exist. It does not create one' — §53.7's finding by an unrelated lever; 'treat it as established.' [P§56.2, P§56.3]

Guardrails and the one clean rule. WER: rank 32 is free (+0.003 pooled, n.s.; dWER ≤ 0 at every weight up to 1.5), rank 64 costs (+0.050, t = +2.21; +0.041 on the strict ladder, t = +2.92), reaching dWER +0.327 in individual cells against §28's +0.104 gate. 'If capacity is ever added to this slot, add 32, not 64.' Genuineness bounds the incumbent's range (pooled −0.224 at w = 2.0 for r16, −0.276 for r32, past §28's −0.341 floor in most cells) and r64 does not breach it (−0.022; r64s positive). Winner's curse, measured a third time: held-out selection gives mean selection-half 0.193 against held-out 0.095, a curse of +0.098 over 35 (root, class) pairs — against §52's +0.096 for a 35-cell argmax. The held-out table carries its own warning: rank 32 gives scream 0.267 and rank 64 0.000 on the same rows and prompts, which no monotone capacity effect can do. [P§56.4, P§56.5]

Verdict: adapter capacity is not the bottleneck for these five classes. Quadrupling the adapter moves the pooled strict hit rate by +0.009 / +0.010 — inside the band a re-seed of the incumbent already occupies — and costs WER at rank 64. Three interventions have now been tried and none is about the adapter: strict re-filtering (§48, null, collapse at w = 1.5 — pooled −0.174, dWER +0.329), higher rank (null), and only merge weight and prompt move the number, with weight running out of room at the genuineness guardrail. 'The remaining suspects are the data and the detector.' [P§56]

Compatibility, settled positively. Ranks stack: r16 and r64 loaded together, deltas add exactly (max abs error 5.96e-08); module.scaling[name] is per-adapter and rank-independent; smoke tests 1632431 / 1632475 confirmed rank-32 and rank-64 adapters loading into the same PEFT tgt slot. Staged: $SC/publish/burst_rank_r32/ — the five rank-32 adapters on the original rows, README first screen the negative result; not a recommended upgrade; rank-64 deliberately not staged; all 25 under $SC/out/bucket_loras_rank/, keep r16r as the project's only measurement of initialisation noise; nothing uploaded. Not settled: shriek; LoRA placement (all 25 target the same 23 modules); ranks above 64 (the quantity monotone in rank is the price: benefit +0.009, +0.010, cost 0, +0.050); and the detector, which §54/§55 established is wrong about what a burst is on most clean clips. Artifacts: ~/reports/burst_rank.md; $SC/code/vb_rank/ (br_run.py, a pinned fork of $SC/code/vb_dose2/bd_run3.py; br_agg.py with --ref; br_report.py, br_replicate.py, br_curves.py); data $SC/out/vb_rank/{br_r*.jsonl, br_grid.json, br_rank.json, br_rank_refs.json, br_replication.json, curves.json}. [P§56]

Corrections and retractions recorded in §43–§56: §43.1 — the shipped burst weights (0.25 / 0.5) are half what lora_bank.py's documentation states (0.5 / 0.75), and two undocumented 'solo' rules coexist. §43.4 — the pre-registered primary metric r_burst_cls is structurally blind to the dose; the hit-rate reading is a labelled post-hoc re-analysis. §43.8 — a mid-run alarm about 205 'missing' cells concentrated in two conditions was a per-rank scheduling artefact, not a failure. §44.2 — a parquet row-group defect would have sized the run by rank 3 and wasted about 35 % of three ranks' pools; fixed by rewriting the shards (660 → 902 steps). §44.6 — no clip of the burst+stop DPO run had been heard; §51.1 later measures it as a null on burst realisation. §45 — the first write-up said w = 1.00 was the cheapest in word error; the table shows w = 1.50 is (0.006 < 0.016), reversing the recommendation. §45.1 — the extension to w = 3.0 shows 1.5 is not a peak and that the only significant result in eleven weights is a WER cost at w = 2.5. §46 — the quality-adapter listening page shipped with Part 2's 140 players dead despite every HTTP check passing. §47.1 — a two-file peft stub had shadowed the real package earlier the same day. §47.2 — embedded_quality_gate was never a burst-presence gate; §47.3 — twenty-seven classes were built from scripted cues over which the detector's contrary label was overwritten. §47.7 — the standalone-row caveat narrows from a blind instrument to a locator bug plus a real source disagreement. §48.5/§48.8 — the claim that a fixed recipe biases only against the re-trained adapter is withdrawn; the bias runs both ways and neither retrain is usable at w = 1.5. §48.8 — the one strict-metric win (chuckle +0.200) is +0.000 family-relaxed: a renaming, not a new burst; and §48.8's sentence that §43 set the merge weight at 1.5 contradicts §43.6's recorded recommendation to keep 0.25 / 0.5. §48.9 — nan printed for both the correctness control and degenerate cells; tcrit rounded the wrong way; displayed levels contradicted the printed Δ; all fixed. §49.1 — Gate A reported proceed: true on zero converted clips because of a floored denominator; a NO_DATA verdict was added. §49.2 — Chatterbox voice conversion fails Gate A (identity 0.360 < 0.45, family 0.470 < 0.60) and is rejected. §49.4 — the strict metric scores the synthetic adapter 0.000 at every weight while the family-relaxed metric shows the effect; the metric, not the data, was wrong. §50.4 — shriek's aligned corpus produced an out-of-family gasp, limiting §49's fix. §50.6 — solo breaches the WER guardrail at w = 0.5 and 1.0; no solo recommendation. §50.7 — a comment claiming vs_clsburst.py and vf_compare.py share one family map is false; sniff was filtered and scored under different families. §50a — family_of mapped 33 of 64 classes to no family (tokenisation and three missing families); fixed, coverage 31 → 62. §50b.1 — §35's conclusion that the emotion adapters are the wrong tool for a transition is downgraded to open; endpoint contrast does not measure continuity. §50b.2 — the retrained burst adapters' split verdict is resolved as 'worked' on the accepted family-relaxed metric. §50c — the quality DPO adapter ships at 1.5 by listening, superseding the table-derived 'keep 1.0'. §51.1 — the burst+stop DPO adapter is a null on realisation at all three recommended checkpoints. §51.3 — §43's dose ladder had not ended; two extensions were needed and three classes are still rising at 3.0. §51.5 — the longer-duration prompt gain belongs to the detector's duration sensitivity, not to better realisation. §51.7 — §43's 'substitution down the arousal axis' is reframed as a corpus-prior collapse. §51.9 — DNSMOS measures clip composition on this material and must not choose recipes. §52.1 — §51 dropped no-audio cells; §52 records them as n_ok = 0. §52.3 — §51's C_gl prompt recommendation does not generalise; it is an interaction with dose. §52.4 — six of nineteen winning cells were the model going silent; a speech-safety gate demotes four recipes and flags three. §52.7 — the page simulator's first run failed (missing takes array); v1 and v2 of the demo audio were discarded for biased carrier sampling; the point − 0.068 convention is about twofold too generous for a 35-cell argmax. §53.4 — the shipped frustrated_groan adapter is 0.000 on both metrics at all weights, so the synthetic re-train was substituted. §53.9 — §30 Q3's finding that guidance loses to weight-scaling reverses on burst classes, with both reasons verified. §54 — the headline 'DramaBox 10 %, adapters 10 %' is inverted by the fact that the detector never emits Shriek; DramaBox's lower dynamic contrast and the LTX-2 licence cut against the data-source plan. §55 — Gemini's no_burst control fails on re-segmented negatives (11 of 20 relabelled as scream/shriek), its spans are too wide for locator training, and gemini-3.8-flash is absent from the model listing. §56 — higher rank is a null; the initialisation re-run r16r shows per-class swings as large as any rank effect; rank 64 costs WER; the shipped weights sit where every root is flat.

II.10 The tools and ears programme after 29 August, part 3 — real burst data, the 17-class detector, the scream gain, consolidation, and two predictors put to a human vote (protocol §57–§70)

II.10.1 The calm lead-in doubles DramaBox's dynamic contrast: 11.5 dB against §54's 6.1 dB, and a 24,500-prompt corpus balanced 50/50 by gender across three age bands (no date stated) [P§57]

§54 had established that DramaBox TTS produces a scream when asked in plain English, but that it shouts the whole scene: measured as the loudest 300 ms window over the clip's own median window, its scenes sat at 6.1 dB (scream, n = 30) and 7.9 dB (shriek, n = 30) against 9.8 dB for the project's own MOSS adapters, with only 4 of 60 clips reaching 10 dB. §54 diagnosed a design defect in its own prompts — none of its 20 scenes contained a calm lead-in — and §57 tests that diagnosis by building a corpus that fixes it. [P§57]

The corpus. 500 scene prompts per vocal-burst class that the adapters cannot fire, every prompt opening on a level spoken line and placing the burst as a departure from it. The class list is 49, not the 51 the brief nominally assumed: wolf_whistle is a lip whistle and tsk a tongue click, and both fall inside the source list's own exclusion set for non-phonated sounds. 49 × 500 = 24,500 prompts, × 3 seeds = 73,500 clips. Speaker demographics are assigned, not sampled, walking a six-cell schedule; the realised counts were taken off the written files rather than from the sampler's intent. [P§57.1]

Realised demographic counts of the 24,500-prompt corpus (§57.1)
malefemaletotal
young_adult4,1164,1168,232
middle_aged4,0674,0678,134
older_adult4,0674,0678,134
total12,25012,25024,500

Gender is exactly 50/50; the bands differ by 98 records in 24,500 (0.4 %), the remainder of 500 not dividing by six. All 49 classes carry an identical signature — 250/250 by gender, 168/166/166 by band — so any per-class slice is as balanced as the whole. The balance had to reach the prose, not just a metadata column, because DramaBox reads the voice out of the description. Verified: 0 prompts whose assigned age phrase is absent from the text, 0 with no gender word in the first 400 characters, and 0 records where the top-level gender/age_band disagree with the copy inside the spec. [P§57.1]

The result: the diagnosis was right. 600 pilot clips, five classes spanning the intensity range, 40 prompts each, 3 seeds; the statistic is copied verbatim from the §54 detector, so the two runs are the same instrument. [P§57.2]

Peak-over-median contrast, this corpus against §54 and the MOSS adapters (§57.2)
groupnmedian dB>= 10 dB
this corpus, all60011.5419 (70 %)
this corpus, loud classes24012.0171 (71 %)
this corpus, mid classes24011.2166 (69 %)
this corpus, quiet classes12011.582 (68 %)
§54 DramaBox scream, no lead-in306.10 (0 %)
§54 DramaBox shriek, no lead-in307.94 (13 %)
MOSS adapters, scream inline259.812 (48 %)

scream goes from 6.1 dB to 12.0 dB — Mann-Whitney p = 2.8e-14, rank-biserial +0.90. Across all loud classes, 12.0 dB against 7.0 dB over all 60 §54 clips (p ~ 0, rank-biserial +0.84). The corpus also clears the MOSS adapters' 9.8 dB (p = 0.0025, rank-biserial +0.37), the first time a DramaBox prompt form has beaten the trained adapters on this statistic. 70 % of clips reach 10 dB, against 4 of 60 (7 %) in §54. The protocol's summary: "A single sentence of calm before the burst is worth about six decibels of contrast. §54's defect was real, and it was in the prompt, not the model." [P§57.2]

What the number does not say (§57.3). The gain is nearly uniform across intensity bands — quiet 11.5 dB, mid 11.2 dB, loud 12.0 dB. Peak-over-median measures departure from the lead-in, which is exactly the mechanism installed, but it follows that the statistic no longer separates a soft_hum from a scream: a soft hum after a level spoken line is also a large relative excursion. The number validates the lead-in; it is not a loudness ranking, and selecting loud clips by it will not work. [P§57.3]

Two operational findings. (1) "A proof whose failure is not fatal is not a proof." The audio job carried a line meant to prove the MP3 encoder before 128 ranks discovered it missing two hours in — but it was written ffmpeg160 -version | head -1, whose exit status the pipe discards. The first pilot exited with status 0 after 57 seconds having written zero clips. It is now a real encode-and-probe round trip under set -euo pipefail, called as || exit 1. (2) "Prove a binary's dependencies on the node type that will run it." The underlying failure was libXv.so.1, a fourth undeclared FFmpeg dependency after libbz2/libx264/libx265, invisible on the login node because it resolves from the OS /usr/lib64; the compute-node image carries no X11 libraries at all. [P§57.4]

II.10.2 Real burst training data with a blind multimodal second opinion: 3,598 utterances over 49 classes, family agreement 64.8 % against exact-class 19.0 %, and MP3 rejected for a 23 ms delay (no date stated) [P§58]

Two WebDataset releases of real audio carrying a blind multimodal burst annotation, staged and not uploaded: burst_gemini_utterances (3,598 whole utterances, 49 classes, 4 shards) and burst_gemini_segments (5,161 segments cut out of them, 6 shards). Every row is real recorded or real-speaker-derived audio; nothing is spliced, concatenated or voice-converted. [P§58]

The ladder. Owner amendment: ≥ 200 verified rows available → take 200; 100–199 → take 100; below 100 → take everything, counting availability after the publication and manufactured-row exclusions. 36 of 49 classes are exhausted rather than capped; the capped? column in every table keeps a class at 100 because the target stopped it distinguishable from one at 100 because that is all there is, and the two are never summed. selected versus shipped is the second distinction: the 69-row gap is the §7 round trip refusing rows it could not parse back. [P§58.1]

Selection tiers of the real burst release (§58.1)
tierclassesselectedshipped
full-200 capped at 200112,2002,163
reduced-100 capped at 100110093
below-100 exhausted, strict only5169165
below-100+relaxed strict exhausted, topped up18769754
relaxed-only no strict row exists at all14429423
An arbitrary line in the selection, found and moved (§58.2). gt_select.py skipped any class whose buckets_strict/<kind>/<cls>.parquet did not exist, although the family-relaxed gate is computed from the same stored scores. Effect: displeased_grunt (0 strict, 42 relaxed) was admitted because its bucket happened to leave a strict file behind, while nervous_giggle (0 strict, 88 relaxed) was not — "same evidence, opposite outcome, decided by a filename." Corrected and verified purely additive against the previous 3,280-row selection: 0 rows dropped, 0 rows changed, 387 added over 13 classes; selection went 3,280/36 classes → 3,667/49, and displeased_grunt moved to relaxed-only. [P§58.2]

Verification. strict 2,727 shipped rows (same-class detection within 1.5 s, weakest clears theta = 0.174); family_relaxed 871 (identical gate except a same-family detection counts — same theta, same tolerance, same geometry gate, no new compute). Relaxed rows were admitted only where strict rows could not reach the tier, are tagged on every row, and the strict filter was never loosened to hit a target. [P§58.3]

Publication. 1,797 podcast rows excluded. The ~937 mediathek_* rows do ship: they are src == 0 voice profiles — DramaBox-written text spoken by a synthetic voice built from a mediathek speaker — and laion/laion-voice-profiles-sft (public, CC-BY-4.0) already ships exactly these, its part-00377 being 100 % mediathek_# voices. The corpus contains no real mediathek recordings or transcripts. Attribution is an allowlist over uid + src + voice_key: 0 unattributable rows. Shipped sources: emolia 1,038, vp/emolia 835, vp/mediathek 627, vp/kseries 523, vocal_bursts_clean 352, vp/refvoice 118, vp/anime 77, kartoffelphon 28. MLS is absent, not excluded. [P§58.4]

What the annotation says: granularity, not vocabulary. Blind, whole utterances only, closed 83-class taxonomy in systemInstruction, temperature 0, 1–3 labels most-likely-first, explicit licence to answer "nothing here". 3,598 calls, 0 failures, ~87 clips/min, 5,161 events (1.43/row). 70 distinct labels used, 59 as a top label, against the detector's 8, and 0 labels outside the taxonomy. [P§58.5]

Agreement between Gemini's labels and the corpus class (§58.5)
levelper rowrestricted to events overlapping the corpus span
top-1 = corpus class19.0 %17.0 %
corpus class in the 1–3 labels37.3 %33.4 %
same burst family64.8 %59.0 %
no burst returned8.2 %
any event overlaps the corpus span80.9 %

Inside a family the model reaches for the prototype: every laugh collapses onto Chuckle (Breathy Giggle 106, Cackle 59, Childlike Giggle 49, Nervous Giggle 31), every hum onto Humming/Soft Hum (Resonant Hum 86 + 44), Shriek onto Scream (47). Commonest label pairs: Deep Breath + Sharp Inhale 596, Breathy Giggle + Chuckle 557, Humming + Soft Hum 264, Scream + Shriek 150. "So it agrees about what family of sound is present roughly two thirds of the time and does not corroborate our fine-grained class boundaries." Whether the project's distinctions are too fine or the model's too coarse, this data cannot say; every number is an agreement between two models both downstream of training on expressive speech. [P§58.5]

§55's span-width result does not replicate (§58.6). §55 measured 38.2 % clip coverage for Gemini against the detector's 6.5 %; on this corpus coverage is a median 10 % and a mean 17.6 % — less than half — most likely because these are whole utterances with a median length of 11.8 s. Median event 0.76 s against the detector's 0.36 s. The warning stands anyway: coverage is not boundary accuracy, onset/offset error against a human-marked burst has never been measured. Dataset B cuts the asserted span plus 50 ms, ships an energy-tightened window as metadata (nucleus_start_s/nucleus_end_s) and a speech_overlap_frac on every segment (median 0.0, p90 0.60). Use B for the label; do not use it to train a locator. [P§58.6]

The round trip, validated against itself. Gemini's spans replace the detector's and the row is re-rendered into general + script; general is not rewritten. The clock is frames / 12.5, not dur_s. Every script is parsed back with the GRPO reward parser and dropped unless labels return in order, nothing runs past the audio, text survives verbatim and the budget sums within 0.50 s (a tolerance calibrated on 792 corpus rows). The corpus's own detector spans convert 3,638/3,667 = 99.2 %; Gemini's spans 3,598/3,667 = 98.1 % (39 residual: 19 text, 9 past-end, 2 unparsed) — a 1.1-point format penalty. [P§58.7]

MP3 works here and is still the wrong choice. lameenc 1.8.4 encodes faithfully (corr 0.9998 on round trip), but on all six clips tested the decoded MP3 leads the input by exactly 1,105 samples = 23.0 ms at 48 kHz and runs 32–40 ms long at the tail — LAME's encoder delay, which lameenc cannot record in a Xing/LAME frame. Read with soundfile or librosa, every span would have sat 23 ms off its own audio. OGG/Vorbis via libsndfile is sample-exact (0 samples added, 0 lag) and ~40 % smaller, so the release ships 48 kHz mono OGG; MP3_DELAY_SAMPLES_48K = 1105 records the compensation an MP3 release would need. Half-speed canary (samples_per_frame == 3840) asserted on every row in both decode passes and packing: 0 failures. [P§58.8]

Not settled (§58.10). No human listening pass — every number is model-versus-model; the page at $SC/pages/gemini_traindata/ (verify_page.js exit 0, 143 players, 0 missing sources) is the instrument for changing that. 34 of the 83 taxonomy classes appear in no release row; the bucket corpus covers 70, and for 21 of those no row survived even the family-relaxed gate. 1,797 podcast rows are the largest verified pool that cannot be used. Segment boundaries remain unmeasured. [P§58.10]

Artifacts. Report ~/reports/gemini_traindata.md; status page ~/dataset_status_burst_gemini.html (self-contained); releases $SC/publish/burst_gemini_{utterances,segments}/; code $SC/code/vb_gtrain/ (gt_common, gt_audio, gt_decode, gt_select, gt_annotate, gt_convert, gt_score, gt_pack, gt_page, gt_readme, gt_status, gt_report, gt_protocol, run_rest.sh); data $SC/out/vb_gtrain/{selection,decoded,converted,selftest,scores,pack_summary}.json, cache/ (3,667 responses), ogg/. Nothing uploaded. [P§58.11]

II.10.3 Two instruments instead of one: keeping the eval audio so a multimodal model can score it, and tuning the classifier head on the 11 classes Gemini confirms (design section, no date stated) [P§59]

A design section written at submission time: what will be measured is fixed before any number lands, so the design cannot be reconstructed to fit whatever comes out. The problem: every burst number in the programme comes from vocal-burst-detector-v2, which knows 83 classes, emits 8, and has never once emitted Shriek; §58 showed a second model using 70 distinct labels on the same corpus. A null on the detector's strict hit rate is therefore ambiguous, not negative, and two independent readouts are the minimum needed to tell "the model did not produce the burst" from "the detector cannot see it". [P§59.1]

The eval was throwing the evidence away. gl_run.py scored each generated wave in memory and dropped it. The patch writes every clip as 48 kHz mono OGG Vorbis, peak-normalised to 0.9 — byte-for-byte the shape gt_decode.py writes the real training rows — and records ogg_key on the row. Cost ~37 KB per 20.8 s clip, ~0.8 GB for the gem arm, no change to any existing number; written through .tmp and os.replace. The structural constraint: compute nodes have no internet, so an API annotator cannot run inside the job. [P§59.2]

Priority order. ~21,600 clips at ~2,000 tokens each. gl_gemeval.priority() fetches the decisive contrast first — alone / alone_new at w = 0.0 and w = 2.0 (the §56 rank-sweep optimum) — then the stacked arms, then w = 1.0, then the rest of the ladder, so a stopped run is a usable experiment rather than a random prefix. Blind, identical to §58's pass: the request never names the prompted class, carries no transcript, and the system instruction is byte-identical to gt_annotate.SYSTEM. [P§59.3] [P§59.4]

Four rates. strict (top-1 label is the class), top3 (class anywhere in the 1–3 labels), family, and no_burst; wrong — an event heard but nothing in the right family — is reported separately. Averaged within a prompt first, n = prompts (§41). [P§59.5]

The classifier head, tuned on what Gemini confirms. Same frozen 768-d embedder, shipped architecture, trunk warm-started from vocal_burst_mlp_v2.pt, new output layer. The filter speech_overlap_frac == 0.0 costs five classes: 16 classes reach 100 Gemini-confirmed segments, 11 after the filter (Wistful Sigh misses by one row, at 99); the surviving 11 include Scream at 155. The split is by speaker. Balance lives in the sampler (K indices per class with replacement per epoch; truncating to the smallest class, 104, would discard 58 % of the data); no class weights on top. Test set balanced, so accuracy compares directly to 1/11 = 0.091. Two baselines reported: the shipped head's argmax over all 83 classes and restricted to these 11; a --no-warmstart arm asks whether the shipped trunk contributes anything. [P§59.6]

What this cannot answer (§59.7). The tuned head classifies a span a locator has already called a burst; it has no reject class and cannot decide whether a burst is present. The locator remains untouched and unmeasured. Gemini and the detector may share a prior; agreement between them is not ground truth, and no listening test has settled which resolution is right. [P§59.7]

Artifacts. Code $SC/code/vb_chain/{gl_gemeval,gl_gemscore,gemeval_watch.sh}, $SC/code/vb_clf/{bc_feat,bc_train,s_feat.sbatch,s_train.sbatch}, patch in gl_run.py (write_ogg, --ogg-dir; pre-patch copy gl_run.py.pre_ogg) and s3_eval.sbatch. Data $SC/out/vb_chain/eval_*/{ogg,gemcache,gem_readout.json}, $SC/out/vb_clf/{emb.npy,meta.json,train_report_*.json,vocal_burst_mlp_{gem,cold}_s*.pt}. [P§59.8]

II.10.4 Re-labelling the same audio with a better annotator: burst LoRAs trained on Gemini's spans against the ones that ship (no date stated) [P§60]

Every shipped burst adapter was trained on rows the project's own detector labelled, and §55 measured that detector to be the bottleneck (over 60 DramaBox clips: Shriek zero times, 8 of 83 labels used, 70 % of detections on two labels, requested burst in top-3 on 3/60 against Gemini 3.8 Flash's 53/60). This section re-trains the shipped recipe on the same real audio with Gemini's labels and spans substituted. Held constant: rank 16, α 32, 5 epochs, lr 1e-4, batch 4, base sft3/export, PROMPT_FORMAT_HASH 073aeb09dc923376, the 70-class carrier prompt set, group 3, and the prompt-index seed rule. Changed: burst_starts / burst_ends / burst_labels, and dur_s replaced by frames / 12.5 (§48.1: 13.6 % of corpus rows state a burst longer than their own clip). [P§60]

Re-bucketing is part of the change. A row joins the bucket of every class its own rendered script names; keeping it in the detector's bucket would train an adapter named affirmative_grunt on the cue (exhausted groan, 3.9 seconds). Round trip, twice: gt_convert.py dropped 69 of 3,667 rows; this chain repeats the check on the values a trainer actually reads (float32 parquet columns, prompt_lib2.render_prompt over all five epochs' RNG draws, anno.parse_script) and rejects 1 more ({'residual': 1}), leaving 3,538 clips. Trained: 15 classes at a floor of 100 rows (affirmative_grunt, breathy_giggle, chuckle, deep_breath, exasperated_sigh, exhausted_groan, frustrated_groan, heavy_breathing, humming, panting, relief_sigh, scream, sharp_inhale, wistful_sigh, yawn); 40 classes excluded below it. A detector-labelled control — same clips, converter, trainer and row budget, own labels — was trained for 9 of them, because Gemini − shipped carries a row-count confound (shipped buckets 111–1,600 rows, these 110–675) and Gemini − detector-control does not. [P§60]

Paired evaluation, Gemini-labelled adapter minus shipped adapter, pooled over classes on per-prompt differences (§60)
kindwd hit (strict)tprompts upd hit (family)d wrongd WER_pkd genuineness
inline0.0+0.0000--0/150+0.0000+0.0000+0.0000+0.000
inline0.25-0.0156-2.142/150+0.0133+0.0644+0.0051+0.020
inline0.5-0.0089-1.074/150-0.0356-0.0644+0.0314+0.098
inline1.0-0.0222-1.913/150-0.0533-0.0711+0.0300+0.114
inline1.5-0.0244-2.722/150-0.0089-0.0200+0.0241+0.156
inline2.0-0.0533-3.087/150-0.0533+0.0489-0.0259+0.187
solo0.0+0.0000--0/150+0.0000+0.0000+0.0000+0.000
solo0.25-0.0044-0.504/150-0.0222-0.0244-0.0467+0.044
solo0.5+0.0067+0.585/150-0.0089-0.0178-0.0324+0.095
solo1.0-0.0178-2.023/150-0.0622-0.1000+0.0342+0.157
solo1.5-0.0333-2.588/150-0.0511-0.0333-0.0242+0.220
solo2.0-0.0405-2.776/148-0.0631-0.0586-0.0910+0.207
Single-variable arm: Gemini-labelled adapter minus detector-labelled control (§60)
kindwd hit (strict)tprompts upd hit (family)d wrongd WER_pkd genuineness
inline0.0+0.0000--0/90+0.0000+0.0000+0.0000+0.000
inline0.25-0.0259-1.624/90+0.0074+0.0444-0.0087+0.003
inline0.5-0.0111-0.734/90-0.0222+0.0148+0.0092+0.072
inline1.0-0.0111-0.695/90-0.0481-0.0556+0.0079-0.026
inline1.5-0.0667-3.622/90-0.0926+0.0296+0.0052-0.036
inline2.0-0.0519-1.937/90-0.0556-0.0185-0.0052+0.021
solo0.0+0.0000--0/90+0.0000+0.0000+0.0000+0.000
solo0.25-0.0074-0.582/90-0.0037-0.0519+0.0022-0.054
solo0.5+0.0111+0.556/90+0.0000-0.0185-0.0415+0.075
solo1.0-0.0296-1.583/90-0.0296+0.0000+0.0570+0.027
solo1.5-0.0481-2.027/90-0.0667-0.0222-0.1000-0.038
solo2.0-0.0444-1.839/90-0.0370-0.0370-0.0156+0.014
Negative result, with the caveat that decides how to read it (§60). Re-labelling with Gemini does not raise the strict hit rate at any weight against either the shipped adapter or the detector-labelled control; the strict deltas are null or negative (worst inline w = 2.0 −0.0533, t −3.08, 7/150 prompts up; inline w = 1.5 −0.0667, t −3.62 against the control). The evaluation detector is the same instrument whose labels are being replaced and scores a hit only when it can name the burst, so a null on strict hit rate is ambiguous rather than negative; the family-relaxed rate is reported beside it but is the same instrument. A positive result would carry no such ambiguity; none was found. The listening page /e/scratch/reformo/schuhmann1_moss/pages/gemini_loras carries the weight: five scenes per class, shipped and Gemini adapter at w = 0.25 / 1.0 / 2.0 on the same scene at the same seed. [P§60]

Reproduce. $SC/code/vb_chain/: gl_trigger.py (artefact-based wait on the sibling's converted.json), gl_buckets.py, train_bucket_loras.py unchanged, gl_run.py (pinned fork of bd_run3.py, md5 ea434661931509ac0db3a856663c1d30), gl_agg.py, gl_listen.py (fork of vb_ladder.py), gl_page.py; submitted as a --dependency=afterok: graph. [P§60]

II.10.5 Does DramaBox actually produce the burst it was asked for? 73,500 clips, annotated whole (no date stated) [P§61]

This section closes the §54/§55 loop at corpus scale: 500 prompts × 3 seeds × 49 classes = 73,500 clips, every one annotated by Gemini 3.8 Flash through the Batch API, then asked the only question that decides whether the corpus is worth training on — given a prompt that asked for class X, is an X in the clip? Annotation: gemini-3.8-flash, thinkingLevel: "low", temperature 0, JSON schema enforced, 37 batch jobs of ≤ 2,000 clips. 73,500 responses, 0 errors, 0 clips without a response, 0 unparsable. 123,004 events (1.674 per clip), 97 distinct labels, 394 clips where the model heard nothing, median event 1.27 s. Cost 100.8 M input, 12.6 M output, and 2,362,228 thinking tokens — thinkingLevel: "low" is not zero, and it is billed; the check is now in the merge step because the first run assumed it rather than measured it. [P§61]

Each clip is shown whole, never pre-cut. On excerpts it had itself left unannotated, the model returned Scream or Shriek eleven times out of twenty; labels are context-dependent, so cutting first and labelling second manufactures bursts. Spans past the end of the audio are clamped and marked out_of_range, never dropped: 0 clamped here. Measurement: bx_hitrate.py counts clips, not events; any = the requested class appears among the 1–3 labels of any event; top-1 = it is an event's first choice; family = some event's label shares the final token. [P§61]

Per-class rate at which Gemini finds the requested burst in DramaBox output, 1,500 clips per class (§61)
classanytop-1familyno burst
humming99.7 %98.1 %99.7 %0
guffaw97.7 %51.1 %97.7 %0
exasperated_sigh95.9 %88.3 %96.9 %0
soft_hum95.5 %14.9 %97.9 %0
relief_sigh93.5 %77.9 %97.7 %2
cackle89.2 %66.7 %89.2 %1
scream88.6 %85.7 %88.6 %3
shriek88.3 %31.0 %88.3 %13
growl86.1 %74.2 %86.1 %7
hiss83.1 %77.7 %83.1 %19
coughing79.3 %65.2 %79.3 %9
cough75.5 %46.6 %75.5 %0
affirmative_grunt74.7 %61.2 %75.0 %5
deep_breath73.6 %26.2 %73.6 %1
wistful_sigh72.3 %40.4 %95.0 %0
childlike_giggle69.7 %50.4 %95.1 %2
snicker68.8 %10.0 %68.8 %0
hiccups65.3 %39.1 %65.3 %9
frustrated_groan65.0 %35.7 %75.5 %2
fearful_gasp60.3 %8.3 %72.8 %8
surprised_gasp52.8 %44.8 %53.1 %35
displeased_grunt49.3 %32.9 %50.8 %3
effort_grunt48.1 %40.5 %61.7 %1
hiccup46.7 %46.1 %46.7 %13
yawn44.5 %40.5 %44.5 %1
sniff42.1 %34.6 %42.1 %15
clears_throat32.1 %8.2 %32.1 %1
heavy_breathing32.0 %15.1 %43.1 %0
sharp_inhale30.3 %12.1 %30.3 %1
swallows29.4 %3.4 %29.4 %24
pleasure_moan29.4 %14.8 %34.5 %0
snort27.6 %16.3 %27.6 %32
pain_moan26.5 %16.5 %30.7 %0
panting25.5 %13.5 %25.5 %1
sobs24.1 %12.3 %24.1 %1
gulps23.9 %18.9 %23.9 %62
nervous_giggle22.6 %5.3 %87.1 %0
mournful_wail22.5 %12.0 %22.5 %4
purr21.9 %15.8 %21.9 %25
fast_breathing20.6 %5.1 %39.9 %8
deep_breathing19.5 %8.0 %41.6 %3
normal_breathing16.7 %1.9 %30.3 %17
slow_breathing10.7 %0.9 %30.7 %3
convulsive_sob7.4 %0.3 %9.7 %1
nervous_gulp7.1 %1.3 %7.1 %52
snorting_giggle4.6 %1.3 %81.0 %0
trembling_whimper1.5 %0.8 %1.5 %5
whispered_mumble1.3 %0.0 %11.9 %5
quiet_sob0.3 %0.1 %0.3 %0

What this changes. The four classes the programme has failed at hardest come out of DramaBox at rates the project's own stack has never approached: shriek present in 88.3 % of clips that asked for it (the detector emitted Shriek zero times in §55; the shriek adapter's best hit rate over 70 classes was 0.00), scream 88.6 % against an adapter peak of 0.07, frustrated_groan 65.0 % against 0 correct bursts in 1,188 prompted cues (§43). "The generator was never the problem for these classes; the data was." (No uncertainty recorded on the per-class rates.) [P§61]

It does not rescue everything (§61). quiet_sob 0.3 %, whispered_mumble 1.3 %, trembling_whimper 1.5 %, convulsive_sob 7.4 % — asking a TTS system in English for a quiet sob does not produce one, at any of three seeds, in 1,500 attempts. These classes have no source of data yet, synthetic or real. [P§61]

Two rates that must not be confused. top-1 is far below any for classes whose neighbours dominate the prior: soft_hum 14.9 % top-1 against 95.5 % any (first choice Humming), shriek 31.0 % against 88.3 % (first choice Scream), deep_breath 26.2 % against 73.6 %. Any pipeline that takes labels[0] alone throws most of the signal away. This is one model's opinion at scale, not ground truth. No human has listened to any of it; the prompt and the annotation are not independent evidence. What the table licenses is a priority order for human checking, not a training set. [P§61]

Artifacts. laion/dramabox-burst-audio (74 WebDataset shards, 1,000 clips each, 160 kbit/s mono 48 kHz MP3 as written by the generator and never re-encoded, metadata.parquet, annotation_stats.json), laion/dramabox-burst-prompts. Code $SC/code/vb_batch/bx_make.py, bx_upload.py, bx_submit.py, bx_drive.py (submit/poll/collect with 429 backoff — the Files API rate-limits hard enough that submission is the slow step), bx_merge.py, bx_hitrate.py. Rates in $SC/out/vb_batch/hitrate.json. Licence: the audio is DramaBox output; the LTX-2 Community Licence is unassessed and this corpus must not be trained on until it is. [P§61]

II.10.6 Two sources, one cut policy: a burst dataset with verified negatives, a detector graded across generators, and an annotator that confabulates on silence (no date stated) [P§62]

§58 cut 5,161 burst segments out of real speech and §61 annotated 73,500 DramaBox clips whole. This section joins them into one dataset with one cut policy and asks what neither could ask alone: does a burst head trained on one generator work on the other? Real speech and DramaBox TTS are acoustically unrelated; a head that transfers between them has learned the burst rather than the generator. bs_common.py copies PAD, cut_span, nucleus, rms_env and speech_overlap out of gt_pack.py md5 2ec167f75251ce92a8c5be5eb6eff72f; run over the same 3,598 real utterances the fork reproduces 5,161 segments — exactly the published count with clamped_span 0. [P§62]

The timeline. The DramaBox corpus ships as MP3 and Gemini timestamped it with its own decoder; the shipped bytes open with ID3v2 and contain no Xing, Info or LAME marker, so no decoder can strip the delay, annotator and cutter see the same stream, and no offset is applied. Decoded length runs ~38 ms past the stated dur_s; every clamp is against the decoded length. [P§62]

The negatives. A negative window must clear two instruments and an acoustic condition: ≥ 0.5 s from every Gemini span, ≥ 0.5 s from every span laion/vocalburst-locator v2 + vocal-burst-detector-v2 found (§55: 3.57× lift for containing the loudest moment against Gemini's 1.27×), and one of two tagged sub-types: speech (≥ 60 % covered by capped Parakeet word extents, ≥ 1 whole alphabetic word, RMS ≥ 0.15× the clip's) or silence (no capped word extent within 0.2 s, RMS ≤ 0.10× the clip's). They are never pooled. The mix is 75/25 speech:silence. [P§62]

Two silent instrument bugs (§62). (1) A Parakeet-TDT token's duration is the encoder advance to the next token, so word spans come out contiguous and every pause is attributed to the word before it: raw spans made speech_overlap non-zero on 91.3 % of real segments against §59's forced-alignment ~53 %, and the first real shard returned 278 speech negatives and 0 silence, which would have let a detector learn silence == dramabox. Both readings now ship (speech_overlap_frac_asr raw, speech_overlap_frac_word capped at 0.30 s). (2) The real half's keys contain dots, so splitting a tar member on the first . read a 1,000-sample shard as 96 clips; the DramaBox half is immune, which is why a DramaBox smoke test could not see it. [P§62]
Verification pass, 1,050 Gemini calls, 0 unparsable (§62)
armnburst returned
speech negatives, shipped3004.3 % (dramabox 4.0, real 4.7)
silence negatives, shipped30099.0 %
naive-gap control, not shipped30016.7 %
burst positive control15091.3 %
The annotator confabulates on silence (§62). 99 % on silence has two readings the verification pass cannot separate, so bs_nullctl.py sent synthetic audio through the identical call: 40 clips of exact digital zeros, and white noise at −60, −45 and −30 dBFS. All four arms came back at 100.0 % "burst returned", mean confidence 0.88, with confident prose descriptions of coughs, sniffs and gasps. The excerpt-level instrument is unusable below speech level; the 99 % measures the instrument, and this also re-frames §61's own 11-of-20 control: part of it is likely the same confabulation. What can still be said is physical: median peak −52.3 dBFS (dramabox) / −42.4 (real) for silence against −8.0 / −6.5 for speech, 58 % of DramaBox silence windows below −50 dBFS. [P§62]
The filter is worth less than the raw arms suggest (§62). Bucketed by the excerpt's own peak level, naive-gap windows at speech level come back at 5.0 % (n = 240) against 4.1 % (n = 293) for the filtered ones, while below −30 dBFS the naive arm is at 95–100 %. At matched level the ASR filter buys almost nothing; what it guarantees is that the negatives sit in the regime the detector operates in — a smaller claim than "16.7 % → 4.3 %". [P§62]
Three incidental corrections (§62). The DramaBox annotation pass returned labels outside its own closed 83-class taxonomy on 0.28 % of events as top-1 and 1.05 % in any slot (giggle, sigh, groan, gasp), where §58 reported 0 for the real pass. clamped_span fires on 0.22 % of DramaBox segments, every one a ≤ 4.0 ms overrun from end_s rounded to two decimals, so §61's "0 clamped" reproduces under any tolerance ≥ 5 ms. And the shipped detector's vocabulary is not 8 labels: over these two corpora it uses 41 distinct labels on the DramaBox half and 36 on the real half, and it does emit Shriek (50 and 49 clips); its distribution is extremely concentrated — Contented Sigh alone on 16,694 of 72,500 DramaBox clips — which is presumably why §55's 60-clip audit saw 8. [P§62]

The split is the experiment. Six cells with the test set for a source fixed per seed and reused by every arm. Grouping is by speaker in the real half and by prompt in the DramaBox half (a held-out prompt is the analogue of a held-out speaker; a held-out seed is not). The no_burst test quota is split per sub-type. §59's speech_overlap_frac == 0 filter is not carried over: it needs forced alignment the DramaBox half lacks, and this head is trained against speech negatives, so a burst over speech is exactly the discrimination it must learn; a <= 0.2 arm is a sensitivity check. [P§62]

Cross-source 2×2, five seeds, 17 classes including no_burst, chance 5.88 % (§62)
train -> testbalanced accburst vs no-burstneg speechneg silenceshipped restr.shipped 83-way
real -> real43.4 % +- 0.497.6 %96.1 %70.4 %26.6 %14.8 %
real -> dramabox34.2 % +- 1.592.8 %95.3 %83.9 %25.5 %14.8 %
dramabox -> real34.3 % +- 1.094.7 %95.8 %26.4 %26.6 %14.8 %
dramabox -> dramabox50.4 % +- 0.397.4 %97.9 %82.0 %25.5 %14.8 %
both -> real38.2 % +- 1.796.3 %96.7 %28.0 %26.6 %14.8 %
both -> dramabox51.2 % +- 1.697.2 %98.2 %79.7 %25.5 %14.8 %

The head transfers: both cross-source cells sit at ~5.8× chance and above the shipped detector restricted to the same classes on the same test sets (26.6 / 25.5 %). The failure is granularity, not deafness: at burst-family level the same predictions gain 16–21 points in every cell (dramabox->real 34.3 % → 55.7 %), with confusion almost entirely within family (Breathy Giggle → Chuckle 43 %, Exhausted Groan → Frustrated Groan 62 %, Deep Breath → Sharp Inhale 48 %, Humming → Soft Hum 29 %). The warm start from vocal_burst_mlp_v2.pt is a clean null (43.3 % against 43.4 % on real->real). [P§62]

Two negatives (§62). Adding DramaBox makes the real-speech head worse — both->real 38.2 % against real->real 43.4 %; DramaBox outnumbers the real half 127,682 : 10,922 usable rows, and the harm disappears once speech-overlapping segments are removed. And the silence negatives do not transfer: dramabox->real 26.4 % against dramabox->dramabox 82.0 %, while speech negatives transfer at 95–98 % everywhere — silence is digital zero in TTS output and a room noise floor in a recording. [P§62]

A second encoder as a single-variable arm, on owner request: laion/voiceclap-large-v2 (3584-d, last-token pooled, L2-normalised; loaded without sentence-transformers) instead of the 768-d extractor. It wins every cell by +1.4 to +6.8 points — dramabox->dramabox 56.0 %, both->dramabox 57.6 %, real->dramabox 40.6 % — and removes most of the mixing penalty: both->real 44.9 % against real->real 45.6 %, so that harm was largely an encoder limitation. One regression: a VoiceCLAP head trained on DramaBox alone recognises real-speech negatives less well (84.1 % against 95.8 %). Cost 15.65 segments/s on one GPU, 3.1 GPU-hours for all 175k items. Its adapter needed a key remap — the shipped LoRA carries one extra model. level, and PEFT loads it silently without matching a single tensor; 0 of 1,232 matched before the remap, 1,232 of 1,232 after. [P§62]

What this cannot answer (§62). No human has listened to any of it. Cross-source transfer rules out "learned the generator"; it does not rule out "learned the annotator" — the same model annotated both halves, and the real half's voice-profile rows are themselves DramaBox TTS over real speakers. The spans remain unsuitable for training a locator, and the silence sub-type is dominated by near-digital-zero DramaBox windows. [P§62]

The dataset. laion/vocal-bursts-segments: 128,165 burst segments (real 5,161 + dramabox 123,004, exactly §61's event count) and 46,494 verified no-burst segments (36,400 speech / 10,094 silence), 58.2 h, subtrees real/ and dramabox/, configs real / dramabox / all / negatives, 1,000 per shard, metadata/*.parquet including the balanced manifest (16 classes × 1,302 burst rows and the same number of no-burst rows). laion/vocal-burst-detector-x2 ships the head, five seeds, VoiceCLAP heads under voiceclap/. Code $SC/code/vb_seg2/, report ~/reports/burst_detector2.md, listening page $SC/pages/vb_seg2/ (255 players). Licence: the dramabox/ half is DramaBox TTS output and the LTX-2 Community Licence remains unassessed. [P§62]

II.10.7 Shipping the burst detector, and the recall floor that makes a hit-rate table readable (no date stated) [P§63]

§62 showed that a burst head transfers between generators but shipped no instrument. This section ships one production detector and, more importantly, the number to read before reading anything the detector says. The configuration was chosen by the 2×2: laion/voiceclap-large-v2 beats the shipped 768-d FastScorer extractor in every cell and shrinks the penalty for mixing the two corpora from 5.2 points to 0.7. [P§63]

Encoder choice on the cross-source 2×2 (§63)
train → testFastScorer 768-dVoiceCLAP 3584-dchance 5.88 %
real → real43.4 %45.7 % ± 1.9
real → dramabox34.2 %40.6 % ± 1.96.9× chance
dramabox → real34.3 %35.7 % ± 1.56.1× chance
dramabox → dramabox50.4 %56.0 % ± 2.0
both → real38.2 %44.9 % ± 2.3
both → dramabox51.3 %57.6 % ± 1.2

both wins the DramaBox column outright and is within noise of the best on the real column; what the detector re-scores is generated audio, so the DramaBox column governs. The shipped head is VoiceCLAP, trained on both halves, 3584 → 256 → 17 (16 burst classes present ≥ 100× in both halves, plus no_burst), negatives from both speech and silence at 75/25, kept tagged. The PEFT failure mode is named: keys carrying an extra model. level load cleanly and leave the base model unchanged, and the visible symptom is a result ("VoiceCLAP is no better"), not an error; bs_vclap.load asserts every adapter tensor found a home (1,232 of 1,232 after the remap, 0 before). [P§63]

What ships is an ensemble of five initialisations inside one grouped split — never an ensemble over split seeds, because seed k's member has trained on items in seed j's test set. The five-split spread is carried alongside as the stability estimate. [P§63]

Production detector on held-out sets (§63; no uncertainty recorded)
held-out setnexactfamily-relaxed
real, all held-out1,38875.4 %84.8 %
dramabox, all held-out60063.3 %79.3 %
real, balanced 25/class42546.6 %68.7 %
dramabox, balanced 25/class42557.4 %75.8 %

The recall floor. A recipe table saying "class X is produced at hit rate H" measures two things at once — the generator producing the sound and this detector naming it — and they multiply. A class this detector recalls at 20 % cannot show a hit rate meaningfully above 20 %, however good the adapter is. Per-class recall is therefore published per class and per source, strict and family-relaxed, with Wilson intervals and an explicit reliable flag (production/per_class_recall.json). [P§63]

Recall floors, strict (§63)
realdramabox
bestSharp Inhale 89.7 %Scream 96.3 %
worstRelief Sigh 6.2 %Soft Hum / Heavy Breathing 28.0 %
no_burst98.0 %96.4 %
Five classes on the real half sit below 25 % strict recall (§63): Relief Sigh 6.2, Heavy Breathing 12.0, Wistful Sigh 14.3, Deep Breath 18.6, Exasperated Sigh 22.6 — the sigh/breath family. For those, a measured hit rate under about 20 % carries no information about the generator at all. Family-relaxed recall ships in the same table because the failure is granularity: Deep Breath 18.6 % → 71.2 % at family level, Breathy Giggle 46.9 % → 96.9 %. [P§63]

The detector is a classifier over spans, not a locator: the locator stays laion/vocalburst-locator v2 with unchanged constants, so only the name changes. Every label it was trained on is gemini-3.8-flash's, so any number measured against that annotator carries home advantage. The shipped contract is tested as shipped: bs_scoretest.py imports the copy inside the model repo using the README's two lines, asserts the adapter matched, and compares the inference-time encoder's vectors against the training features — cosine 0.99907 minimum, 0.99989 mean over 240 segments. [P§63]

The first version of that test was wrong (§63). It asserted exact argmax agreement and came back 237/240; the three disagreements were bf16 batching-by-length effects (vectors agree to ~1e-3) on near-ties, not a fault. The claim is now two conditions jointly stronger than the original: the probability of the reference class must agree to < 0.05 (measured max 0.0440, mean 0.0090), and every argmax flip must be a near-tie in the reference distribution (measured gaps 0.0044, 0.0424, 0.0005). "When a tolerance test fails marginally, the repair is to state the invariant you actually mean, not to widen the tolerance until it passes." [P§63]

Model https://huggingface.co/laion/vocal-burst-detector-x2 (production/ holds the five members, the recall table as JSON and CSV, the confusion matrices and the three source files needed to call it). Dataset https://huggingface.co/datasets/laion/vocal-bursts-segments. Code $SC/code/vb_seg2/ (bs_prod.py, bs_score.py, bs_scoretest.py). Report ~/reports/burst_detector2.md. [P§63]

II.10.8 The scream gain at 45-class scale: a family-level effect, a dose curve that turns on at 25 %, and a recipe table merged rather than replaced (no date stated) [P§64]

The preceding ablation left one question: scream gained a great deal from DramaBox-trained data and exasperated_sigh gained nothing — is the effect a class or a property? This section answers it on 45 classes, adds a dose ladder, and rewrites the WikiSkills operating points under a changed selection rule. The selection rule changed first: the owner ruled that for a burst class a genuineness fall is the expected price, so d genuineness >= -0.341 no longer rejects a cell or picks a weight; it is printed and decides nothing. WER remains the gate: paired d WER_parakeet <= +0.104 against the arm's own w = 0 cell, plus absolute WER_pk <= 0.25 on inline. The old BURST_LAM = 0.25 was capped by the genuineness guardrail; the per-class optima now sit at 1.0–2.0, a consequence of the rule change rather than a new measurement. [P§64]

The scale test. 45 classes × 2 script kinds × 6 weights × 10 prompts × 3 samples, paired at the prompt level (n is the number of prompts). 49,680 cell rows, 332 with no audio, 0 half-speed-canary violations. Pooled over all 45 classes, bulk (retrained on real + DramaBox rank ≤ 1) minus shipped: [P§64]

Pooled gain of the DramaBox-retrained adapters over the shipped ones, 45 classes (§64)
metricinline w=1.0inline w=1.5solo w=2.0
hit rate, strict-0.0015 (t -0.29, 12/450)+0.0081 (t +1.43, 22/450)-0.0186 (t -2.72, 17/438)
hit rate, family-relaxed+0.0356 (t +2.67, 93/450)+0.0274 (t +2.20, 89/450)-0.0506 (t -3.52, 60/438)
the same, adapter alone+0.0385 (t +3.34, 72/450)+0.0381 (t +2.94, 85/446)-0.0088 (t -0.66, 57/379)

The strict rate is a null and the family-relaxed rate is not. "DramaBox audio teaches the model to produce a sound of the right family more reliably; it does not teach it to honour the sub-label." The gain lives on inline at w = 1.0–1.5 and is gone by w = 2.0 and on solo. Per class, the mean best gated strict hit is 0.0552 for the shipped adapter and 0.0415 for the retrained one. Of 450 cells, four clear the two-sided 5 % point of t at 9 d.f.; ~23 are expected by chance. The two that are large as well as significant: surprised_gasp inline w = 1.5, +0.367, t +3.50, 8/10 prompts up (0.433 against 0.133) and childlike_giggle solo w = 1.5, +0.300, t +2.86. Three classes lose significantly — deep_breath, sharp_inhale, breathy_giggle — all breath/laugh classes where DramaBox's own delivery overwrites the target. Rank 0 against rank ≤ 1 is a null (largest |t| over all cells 1.92 on the twelve classes with the largest rank gap). (Section II.10.9 below records that the pooled family-level gain did not survive a second instrument.) [P§64]

Dose ladder, scream, inline, best gated hit rate (§64; no uncertainty recorded)
DramaBox fraction0 %10 %25 %50 %89 %100 %(shipped)
hit rate0.0330.0330.3330.2670.2000.4000.100

The curve turns on between 10 % and 25 % and is then flat to 50 %. Twenty-five per cent buys essentially the whole effect at WER_pk 0.099 against a 0.120 baseline. The WER gate breaks only at the top of the ladder — DramaBox-only arms reach WER_pk 0.775 (inline) and 2.240 (solo) at w = 2.0. Operating point: 25 % DramaBox at w = 1.5–2.0; genuineness there is 1.38 against a 2.62 baseline, reported and not disqualifying. exasperated_sigh stays at 0.000–0.033 strict in all ten arms at every weight, which is not evidence about DramaBox: family-relaxed it reaches 0.833, and the recall floor explains the strict number. [P§64]

The recall floor and the two thresholds. The production detector publishes per-class recall for 16 of these 45 classes. Five rows are floored (Relief Sigh 6.2 %, Heavy Breathing 12.0 %, Wistful Sigh 14.3 %, Deep Breath 18.6 %, Exasperated Sigh 22.6 %); two are mathematically incapable of clearing a 0.15 strict threshold and are carried at family level with the reason printed. Nine of seventeen recall rows are flagged reliable: false (n < 30 in one source). 9 classes clear 0.15 strict; 30 clear 0.15 family-relaxed. Strict list: breathy_giggle, childlike_giggle, chuckle, deep_breath, exhausted_groan, scream, sharp_inhale, surprised_gasp, wistful_sigh. [P§64]

The instrument trap, closed rather than flagged. The existing wikiskills/VOCAL_BURSTS.md recipes varied prompt form; this study varies adapter and weight. Both score through reward.RewardModellaion/vocal-burst-detector-v2, and vb_lever/vl_prompts.py builds its P0 to reproduce burst_dose2/prompt_sets.json byte for byte, so an old row and a new row compare two recipes as a whole; the two sets are never sorted against each other. Merge rule: an old recipe is replaced only where the new one beats its family-relaxed hit rate by more than one seed-noise standard deviation (0.068); without the margin one class would have switched on +0.003. Result: 14 old recipes stand, 12 are replaced, 5 classes were not re-measured, 19 classes get a first recipe. The bulk evaluation retained all 63,600 generated OGGs; the production detector was run over the 22,825 whose class it knows (locator agrees: 143 spans against 144 on the first batch). [P§64]

Cost and product. 51,840 optimiser steps over 45 + 12 + 3 × 2 adapters, 0 partial and 0 non-finite losses; the bulk-arm evaluation took 2 h 06 on 16 nodes. Dataset laion/vocal-bursts-per-class (45 configs, 33,991 clips, 14.09 GB), a strict subset of what was trained on: 607 mediathek_* clips are in the pools and not publication-cleared. The LTX-2 Community Licence has not been assessed for the DramaBox half; every arm using it is marked and a real-only fallback exists. German report with 1,905 embedded listening takes ~/reports/bericht_vokale_bursts.html; code $SC/code/vb_cls2/ (vc_agg.py, vc_recipe.py, vc_rescore.py, vc_bericht.py, vc_wiki.py). [P§64]

II.10.9 Addendum to §64: a second detector over the same audio, and the half of the claim that does not survive it (no date stated) [P§65]

The re-score promised at the end of §64 landed, "and it does not confirm the section's headline — it splits it." The bulk evaluation had kept all 63,600 OGGs, so the production detector was run over the identical audio: 22,825 clips (16 of 45 classes). Same locator, so the same spans — 22,379 against 22,453, ratio 1.0033; only the naming differs. The metric is span-free (any_hit) and computed identically for both instruments. [P§65]

The pooled family-level gain of §64 does not survive (§65). The +0.0356 (t +2.67, 93/450) at inline w = 1.0 divides along the new detector's label space: on the 16 verifiable classes it is −0.0292 (t −1.28, 24/160) at w = 1.0 and −0.0667 (t −2.78, 26/160) at w = 2.0 — negative — and the second instrument confirms independently on the same audio (−0.0083, t −0.31 at w = 1.0; −0.1313, t −4.33 at w = 2.0); on the 29 classes it does not know it is +0.0713 (t +4.44, 69/290) with no second instrument. "The family-level gain lives entirely in the rarer classes where the shipped adapter could do almost nothing, and it is not independently confirmed there; where a second opinion exists, both instruments agree the retrain is neutral to harmful." One exception is large and identical on both: scream inline w = 1.5, +0.267 (t +2.75) on v2 and +0.400 (t +4.13) on x2. [P§65]
A retracted assertion (§65). "Does a recommendation change? I asserted no before measuring it, and that was false." The argmax cell moves for 14 of 16 verifiable classes. The decision-relevant quantity is regret: keeping this study's cell costs, judged by the other instrument, mean 0.106, median 0.083 (max 0.300); taking the other instrument's cell costs 0.100 on this one. The swap is symmetric — neither instrument has the better recipe — and the median sits barely above one seed-noise standard deviation (0.068). Only 2 (v2) and 4 (x2) of ~24 gated cells lie within one such standard deviation of the top: the peak is real, its identity is not. No recommendation is withdrawn; every row now reads "a good setting" rather than "the best", and best-of-N is named as the larger lever. Lesson: a second instrument is worth more than a bigger n. [P§65]
Two provenance corrections (§65). *The partition:* the family-relaxed column comes from burst_family.py (md5 19a0607b), coarser than the programme's published 23-group scheme (vm_groups.py md5 f83e3850, vocal_burst_groups.json md5 3e774204) and crossing its boundaries (groan spans grunt, sigh_neg, moan, growl; breath spans gasp, breath_calm, breath_fast, sniff). Re-scored under the published map, 21 classes clear 0.15 rather than 30, 21 of 45 recipes are unchanged, and exasperated_sigh falls from 0.833 family to 0.200 group; both numbers now appear in every table. *The strict column:* reward._same_class is a substring test, so a Coughing detection counts as a strict hit for cough; over 97,349 target cues it promotes 12 of 3,269 counted detections (0.37 %) and 0 cross a 23-group boundary — a footnote, but "strict" now says so. [P§65]

Scheduling note: the booster queue was saturated all afternoon and the job started only after being reshaped from four nodes to one; two sibling studies made the same change the same day. Code $SC/code/vb_cls2/vc_rescore.py, vc_rescore_agg.py, vc_group.py, vc_stability.py, vc_prov.py; artefacts rescore_agg.json, group_recipes.json, stability.json, provenance.json. [P§65]

II.10.10 Combining the per-class burst adapters we already have: a clean negative at equal weight budget (no date stated) [P§66]

The question. Seventy per-class burst LoRA adapters exist and the grouping folds the 49 requested classes into 23 groups. Does loading a group's member adapters together at inference evoke the group's sound better than the best single member adapter? Nothing is trained here. The one rule that makes the comparison mean anything: a LoRA layer computes h + scaling·B(A(x)) with scaling = (alpha/r)·w = 2.0·w, so N adapters at weight w each is an N-fold larger perturbation; every arm spends the same total weight budget W = Σ wᵢ, verified to have zero budget mismatches before a single clip was generated. Exactly one arm (stackfull, all n members at w = 1.0) breaks the rule on purpose and is reported as the trap. [P§66]

The grid. 15 target classes over 9 groups inside the production detector's 16-label space. 812 of 816 conditions carry data, 8,160 generated cells, 22,860 clips, 46 distinct adapters, both prompt kinds, budgets 0.50, 1.00, 1.50, 2.00, arm alone only (burst adapters on the bare SFT3 export). 0 half-speed-canary violations. Never merged: audio_lm_heads.N.weight IS audio_embeddings.N.weight — weight-tied — so a merge writes the burst delta into the embedding table; everything is done through scaling. [P§66]

The answer is a null, and a clean one (§66). In none of the 16 (form × budget × metric) fields does the combination beat the best single member adapter at equal budget with |t| ≥ 2; the most favourable field is inline, W = 0.50, v2_strict: −0.009 at t = −0.89 — and the member it is compared against was chosen on the same ten prompts it is then measured on, so the selection handicaps the opponent and the combination still does not clear it. In 14 fields the combination is significantly worse at equal budget. The coherent group beats the mismatch control in 5 fields, so the grouping contributes something — just not enough. [P§66]
Pooled result on the production detector x2, strict hit; unit is one prompt, three samples averaged; u-best is the paired difference against the best single member at the same budget (§66)
formWbaseshipped ownbest singleensembleuniformownhalfu-bestt
inline0.500.0470.0490.0960.0580.0510.059-0.044-2.68
inline1.000.0470.0780.1070.0630.0710.062-0.036-2.29
inline1.500.0470.1130.1290.0730.0760.110-0.053-2.74
inline2.000.0470.1420.1800.0850.0830.101-0.095-4.21
solo0.500.0330.0490.1000.0630.0560.072-0.044-3.56
solo1.000.0330.1060.1280.0670.0710.064-0.057-3.34
solo1.500.0340.1190.1240.0730.0940.102-0.029-1.77
solo2.000.0340.1090.1270.0660.0660.081-0.059-2.42

The mechanism is the one the budget rule predicts: the single adapter's hit rate rises with budget while the combination stays flat; splitting W over n adapters turns each down to W/n, including the only one that learned the requested class. ownhalf sits between the two; the ensemble column (one member picked at random per sample, computed) tracks the combination, not the best member. [P§66]

Group level against both random controls; net is against the size-matched control (§66)
formWuniform grouprand (dist)rand (size)uniform netbest single grouprand (size)best single net
inline0.500.0870.0780.080+0.0070.1440.126+0.018
inline1.000.1220.1010.105+0.0170.1690.138+0.031
inline1.500.1440.1060.110+0.0350.1840.157+0.028
inline2.000.1470.1060.110+0.0370.2380.208+0.031
solo0.500.0780.0800.079-0.0020.1310.123+0.009
solo1.000.1180.0940.095+0.0230.2000.154+0.047
solo1.500.1520.1150.119+0.0330.1930.149+0.045
solo2.000.1320.0870.090+0.0420.1950.150+0.044

The combination does clear its random control, but so does the best single adapter, and by more. The trap, measured: stackfull spends budget n instead of W; 20 of 26 of its cells fail the WER gate, 7 at WER_parakeet = 1.000 — complete transcription failure — against 0.08–0.52 for the same class's own adapter at w = 2, and most of those cells score 0.000 strict. [P§66]

Substring collision, settled by counting. vk_agg.py takes n_exact for every *_strict column; the substring reproduction is carried in a separate *_lax column. For the 15 target classes the primary instrument has 0 substring collisions in its declared label space; detector-v2 has exactly 1 (deep_breathDeep Breathing), within a group. Over the 117-label union, deduplicated by normalisation, there is 1 cross-group substring pair with both members emittable, touching no target class; two cross-group pairs in the annotator's free vocabulary would bite any _same_class-scored metric over annotator labels: exasperated_sigh (sigh_neg) ← sigh (sigh_pos), heavy_breathing (breath_fast) ← breathing (breath_calm). [P§66]

A grouping must cover the instrument's label space (§66). An emittable label the map does not know makes a Wolf Whistle detection a strict hit for whistle but no group hit, so the relaxed rate falls below the strict one and a metric monotone by construction stops being monotone. Found as a divergence between two counts of the same census (1 pair when both labels must be mapped, 3 when group_of merely differs) and fixed at the root: all 82 emittable burst labels mapped, 0 unmapped, exporter asserts coverage. Post-fix both counts converge at 1 — snort (sniff) × snorting_giggle (laugh_soft), documented rather than fixed. [P§66]

Provenance. vk_metric.norm and vm_groups.snake merge and split no label differently over the union, so the strict columns here and in the grouping analysis are the same equality relation; the grouping analysis's net-over-random gain is measured on 123,004 real annotated events with annotator labels, every number here on generated audio scored by a detector — they may not share a sorted table. Shared protocol with the pooled-group sibling: same carrier set (burst_dose2/prompt_sets.json, 10 prompts × 2 kinds), same prompt-only seeding, budgets, WER gate and instruments; vk_metric.py is a byte-identical pinned copy of the sibling's vg_metric.py. The random control deals only the labels the scheme maps, leaving the unmapped set out of both arms (measured: prod 16 labels into 16 slots, rest 0; v2 33 into 33, rest 0, 1 left out of both arms). [P§66]

A portability trap (§66). Nested braces in an f-string expression are a syntax error before Python 3.12; the login node runs 3.9.25 and the compute stack 3.13.5, and every upload and page build runs only on the login node, so the defect surfaces at the very end of a run. Every file in this chain is now parsed against both interpreters. [P§66]

Artefacts. Code $SC/code/vb_comb/, data $SC/out/vb_comb/, German report ~/kombinierte_adapter.html, state ~/reports/STATE_burst_comb_lora.md. [P§66]

II.10.11 Consolidation: what the vocal-burst programme established, what it disproved, and eight ways a green exit code lied (grouping fix dated 2026-09-04) [P§67]

This section consolidates the programme, records the three studies that closed after the preceding sections were written, and gives the methodological failures their own place; "the negative results are stated as results, not gathered into an appendix of things that did not work." [P§67]

II.10.11.a The standing caveats, stated once [P§67]

No human has listened to any of this — not the training data, not the generated audio, not the clips behind any number. Every hit rate and recall is one model's judgement of another model's output. The instruments disagree: two independently trained detectors scored the same generated audio and agreed only moderately — cell-level Pearson 0.51 strict, 0.21 family over 768 cells — and one names the target class more than twice as often as the other (0.135 against 0.061 over 22,825 clips). "Absolute rates in this programme are properties of a measurement chain, not of the model." The LTX-2 Community Licence position on the DramaBox audio is unassessed, and it bears on the one robust positive: scream, inline, w = 1.5, +0.267 on the old detector and +0.400 (t +4.13) on the production detector, comes from a DramaBox-trained adapter; the real-audio-only fallback measures at roughly a quarter of its hit rate. [P§67]

II.10.11.b The arc: what was tried, and what holds [P§67]

Interventions and their standing verdict (§67)
interventionverdict
per-class LoRA adapters on real audio with detector labels (the shipped bank)works for a minority of classes; the baseline everything else is measured against
re-labelling the same audio with a better annotator (§60)null at every weight
more adapter capacity, rank 32 and 64 (§56)null — capacity is not the bottleneck
prompt form: cause in the GENERAL line, longer stated duration, placement (§51–52)works, and is additive; the largest reliable lever found
classifier-free guidance on the burst (§53)works, g = 4 triples the hit rate — but the prompt is doing most of it
burst + stop DPO adapternull on burst realisation; ship it for genuineness or not at all
neighbour-class substitutionnull on family, significant harm on strict
training on a second TTS corpus, DramaBox (§64–65)one large class-level win; the pooled gain does not survive a second instrument
combining member adapters at inference (§66)negative in 8/8 fields, 0 of 120 per-class wins
pooled group trainingnegative against every baseline
re-using a group's strongest member adapter across the grouppositive, and free
scaling weightmatters; the optimum moved to 1.0–2.0 once genuineness stopped gating
best-of-N candidatesthe largest practical lever, and the one every recipe now leads with

"Almost everything that changes the model fails, and what works are ways of asking" — prompt form, guidance, weight, and drawing more samples. The one training intervention that produced a large win produced it on a single class, in a way a second instrument could not confirm anywhere else. [P§67]

II.10.11.c Pooled group training: a clean negative, with a free positive beside it [P§67]

Design. 23 groups sorted into 10 primary (≥ 100 real rows), 10 DramaBox-only (≥ 100 DramaBox rows, fewer than 100 real) and 3 untrainable (whistle 18 real rows, mouth 15, misc_body 1; skipped and stated). 30 adapters over the 20 trainable groups in two recipes: grp_mix_full (all real + all DramaBox at rank ≤ 1, dose-matched, 20 buckets) and grp_mix25 (DramaBox subsampled to 25 %, 10 buckets). Pooled rows keep their own member label as the script cue. Deduplication on uid/dkey collapsed 95 duplicate real rows and 0 DramaBox duplicates. 50,205 optimiser steps, 628 GPU-minutes, 30/30 adapters complete, 0 partial, 0 non-finite losses. Five arms: shipped, percls, grpfull, grp25, and bestmem — the group's strongest member adapter loaded for every member. Carrier burst_dose2/prompt_sets.json, 10 prompts × 2 kinds, 3 samples per prompt, seeds depending on the prompt index alone, n is prompts. Five weights (0, 0.5, 1.0, 1.5, 2.0). 41,800 cell rows, 166 empty, 83,444 clips re-scored, 0 half-speed-canary violations. Primary instrument vocal-burst-detector-x2 (16 labels covering exactly the 10 trainable primary groups — 28 measurable classes; the other 10 trainable groups cannot be measured on it at all), detector-v2 alongside. Gate: paired Δ WER_parakeet ≤ +0.104 plus absolute WER_pk ≤ 0.25 on inline; genuineness gates nothing. [P§67]

Pooled group training loses against every baseline (§67). At group level on the primary instrument, over the 28 classes in the ten two-source groups, mean difference against the best baseline d = −0.090 (better 6, worse 18, tied 4); the broad instrument agrees at −0.098 (better 3, worse 17, tied 8). [P§67]
Matched paired tests, differenced per prompt at n = 280 (§67)
comparisonformwdt
grpfull − perclsinline1.5−0.0500−2.79
grp25 − perclsinline1.5−0.0595−3.16
grpfull − shippedsolo2.0−0.0791−4.50
grp25 − shippedsolo1.0−0.0502−3.16
grpfull − bestmeminline1.5−0.0869−4.65
grp25 − bestmeminline1.5−0.0964−4.80
bestmem beats the shelf at zero GPU cost (§67)
comparisonformwdt
bestmem − shippedinline1.5+0.0762+3.76
bestmem − perclsinline1.0+0.0357+2.32

Two things make bestmem credible: it changes only which existing adapter is loaded, so there is no new training run whose noise could be mistaken for an effect, and the member was chosen from the sibling study's finished table before any of this study's numbers existed. "The group structure is operationally worth something, but through adapter selection, not through pooled training." The DramaBox-only block is a null: mean d = −0.003 (better 1, worse 4, tied 12) over 17 classes. The gate bites at w = 2.0: pooled over the 28 primary classes on inline, all four adapter arms fail — percls +0.167, bestmem +0.168, grpfull +0.123, grp25 +0.109 against the +0.104 bound; only shipped survives at +0.038. [P§67]

II.10.11.d The class grouping both studies rest on [P§67]

23 groups over 117 label strings, seeded semantically and checked against the annotator's measured confusion. Validate by own-group lift, not own-group share: two attractor labels (deep_breath, exasperated_sigh) are top-1 for 29 % of all events regardless of what was requested; deep_breath sits at 14 % own-group share and plainly belongs where it is. The directed lift P(label | a member of its group was requested) / P(label) divides the attractor frequency out; rule lift ≥ 1.5, at least 20 emissions. 58 of 58 testable members are enriched in their own group. Over 123,004 annotated events across 49 classes: mean hit rate 0.302 strict → 0.537 grouped — but a random grouping with identical sizes already scores 0.355. Net +0.182; of the +0.235 raw gain, +0.054 is pure arithmetic. Classes at or above 0.15 go 28 → 44, against 34.1 for the random control. [P§67]

Members removed on evidence. sigh failed the lift test (0.68 over 116 emissions) and groan marginally (1.46 over 37); breathing, grunt, screaming_yawn are bare stems straddling a group boundary. Removing sigh disarms the collision that would have merged sigh_neg into sigh_pos. One cross-group pair survives, documented: snort (sniff) × snorting_giggle (laugh_soft). Four labels were unmapped until 2026-09-04; all 82 emittable burst labels are now mapped. Grouping creates material per unit, not units: at a floor of 100 segments in both halves, 16 units clear per class against 10 grouped; usable segments rise 90,924 → 104,694 of 126,512; hiccup has 2 real segments, swallow 0, hiss 3. [P§67]

A caveat the grouping records against itself (§67). The group sob is semantically correct and empirically a failure — its members score 0.001–0.008 and the group scores below its own random control. "Grouping does not substitute for the model being able to make the sound." The grouping earns its relaxation on sigh_pos, laugh_soft, hum and throat, and buys nothing at all on twelve groups. [P§67]

II.10.11.e The encoder comparison: a licence result, not an accuracy one [P§67]

Encoder comparison, held-out real speech balanced 25 per class, chance 0.0588 on 17-way (§67; no uncertainty recorded)
encoderlicenceparamsdim17-way exact23-groupDramaBox 17-way
voiceclap-large-v2CC BY 4.07 B + LoRA35840.4660.6050.574
voiceclap-commercialCC BY 4.0110 M7680.3930.5150.537
voiceclap-small-v2CC BY-NC 4.0110 M7680.4020.5130.551

The 7 B encoder buys 0.073 exact accuracy for roughly 64× the parameters — worth it for an offline scorer, probably not per request. The non-commercially licensed small encoder buys nothing over the commercial one: 0.402 against 0.393 exact and 0.513 against 0.515 grouped, inside the noise of a 425-sample cell. "There was no accuracy price for choosing the commercially usable model, so the licence question never had to become a trade-off." [P§67]

II.10.11.f Methodology: eight ways an exit code said fine and only the output disagreed [P§67]

"This is the transferable part. Each item below was found the hard way within about a day, and all but one are independent of anything specific to this programme." [P§67]

  1. Slurm records COMPLETED 0:0 on a job that lost a quarter of its work. Job 1667415, node jpbo-063-01: task 2 died on torch.AcceleratorError: CUDA error: an illegal memory access was encountered inside peft/tuners/lora/layer.py forward after writing 1 of its 100 owed cells; srun printed task 2: Exited with exit code 1; sacct recorded COMPLETED 0:0. Result 301 of 400 cells, per-rank {0:100, 1:100, 2:1, 3:100}, and every aggregate would have silently averaged over three quarters of the prompts. The only reliable gate is counting owed cells against present output lines, per rank; a refill must be resubmitted at identical width because the runner partitions work by WORLD. "If one item is taken from this list, take this one." [P§67]
  2. A second instrument was worth more than a bigger n. §65 restated: +0.0356 at t = +2.67 over 45 classes was negative on the 16 verifiable classes (−0.029 at w = 1.0, −0.067 at w = 2.0, second instrument −0.131, t −4.33) and lived entirely in the 29 classes with no second opinion (+0.0713, t +4.44). The corollary cost the author an assertion: "no recommendation changed" was written before it was measured; the argmax moved for 14 of 16 classes, regret mean 0.106 / median 0.083 against reverse regret 0.100. Report regret, not argmax agreement — and do not write down the result of a check you have not run. [P§67]
  3. Mid-run rates flatter, in more than one way. Four figures for the same quantity (cells/min/rank): planning 1.87 (a sibling's estimate); mid-run "correction" 2.76 (aggregate at 78 min); instantaneous log sample 6.94 (20 cells / 173 s on the cheapest condition); truth 1.61 (full run, 5,591 cells / 145 min / 24 ranks). Acting on the mid-run figure, a job was shortened from 5:00 to 2:30 and timed out at 5,591 of 5,600 cells — 99.8 %, nine cells short. A walltime planner solved against the wrong job model returned a −5.1 minute load budget (an infeasible plan is visible; a slightly optimistic one is not). [P§67]
  4. Pairing by list position silently corrupts. In the combination study, 77 of 816 conditions had fewer than 10 usable prompts, so position i in one list and position i in the other were different prompts; the t values still looked plausible. Pair on an explicit key, take the intersection, print n. Evaluations are split by class, not by rank, and a directory's tag must be unique. [P§67]
  5. Set-but-not-exported shell variables killed two completeness checks, silently. A child python3 -c or srun step sees nothing and "nothing owed" reads as "complete". One instance died with KeyError after every one of eight training jobs while each still printed DONE: the check that existed to catch a missing adapter was itself missing, eight times. ${VAR:?} is one character of insurance; suppress a known-benign traceback by its exact signature, never by ignoring tracebacks. [P§67]
  6. A grep that ignores binary files makes "no match" and "never examined" indistinguishable. A job log with tqdm bars is classified non-text (314 CRs and 8 other control bytes in its first 64 kB), as is any log with a NUL byte from a killed rank. On the actual CUDA-fault log the wrapped grep -Fc "illegal memory access" printed nothing and returned 1; grep -a -Fc printed 1; real GNU grep -Fc printed 1. The interactive tooling wraps grep with -I by default. Scan logs from Python (open() + re), or at minimum grep -a. Also: grep -c on no match prints 0 and exits 1, so c=$(grep -c … || echo 0) yields "0\n0"; multi-file grep -c prints file:count per file. [P§67]
  7. Any grouping raises a hit rate, so a size-matched random control is mandatory. Of the grouping's +0.235 raw gain, +0.054 was arithmetic. The control must deal only the labels the scheme maps, into the scheme's own group sizes; an earlier version dealt a fixed slot budget over the whole vocabulary and was both weaker (5 unhittable labels against 3) and stronger (the shared bucket acted as a real group for ~1.6 requested classes per draw), so the sign had to be measured: 0.0016, overstating the gain. At 300 draws the SE on the control mean is 0.00095, enough to flip the sign of a correction that size; raised to 1,000. A per-class control drawing exactly the class's own group size is the honest primary. [P§67]
  8. A script can do nothing and exit 0. A path helper took one directory level where it needed three, matched the literal set {"cur"}, reported "1 adapters, 0 links, 0 evaluable classes" and exited 0 (correct: 20 adapters, 57 links, 45 evaluable classes). A collision enumerator fed already-normalised labels into a matcher that normalises differently, so every multi-word name resolved to None and two Nones compare equal. An over-broad text replacement in a report builder silently deleted three content blocks and two command-line options; the page was 44 kB instead of 52 kB and was caught only by rendering and diffing headings. Check the artefact, not the return code. [P§67]

A ninth, from §66, of the same family: nested braces inside an f-string expression are a syntax error before Python 3.12; the login node runs 3.9 and the compute stack 3.13. The common shape: "each of these is a check that reports success by doing nothing." The habit that catches all of them: make every check state the quantity it examined (cells counted, files read, labels dealt, prompts paired, headings rendered) and treat a check that cannot say what it looked at as a check that did not run. [P§67]

II.10.11.g Artefacts [P§67]

Grouping $SC/out/vb_merge/vocal_burst_groups.json (vm_groups.py md5 f83e3850…, groups JSON md5 3e774204…), published in seven repositories; encoder comparison $SC/out/vb_merge/encoder_compare.json. Pooled-group study $SC/code/vb_grp/, $SC/out/vb_grp/ (verdict_x2_grp.json, verdict_v2_grp.json, checks.json, eval_complete.json, bestmem.json), German report ~/gruppen_adapter.html, state ~/reports/STATE_burst_group_lora.md. Adapter combination §66, $SC/code/vb_comb/, ~/kombinierte_adapter.html, ~/reports/STATE_burst_comb_lora.md. Per-class scale and dose study §64–65, $SC/code/vb_cls2/, ~/reports/vb_cls2.md, ~/reports/bericht_vokale_bursts.html (1,905 embedded takes). Detector §62–63, laion/vocal-burst-detector-x2. Datasets laion/vocal-bursts-per-class, laion/vocal-bursts-segments. [P§67]

II.10.12 Where the work ended up, and the one convention every consumer of it needs (published 2026-09-05) [P§68]

Sections 61–67 record what was measured; this one records where it was put, because "a result nobody can fetch is not a result", and three of the artefacts below did not exist in a reachable form until they were published on 2026-09-05. The convention, stated once because it is stated in six places: a cue is written in English even when the spoken line is German, verified against the training corpus (German rows read Das zerreisst einen einfach, weisst du? (relief sigh) and ... sehen sie mich vielleicht endlich. (person whistling to get attention)). A German cue is out of distribution. A round bracket containing a number is a burst, not a direction: the director writes prose and never a number — (chuckle) is a sound, (clearly amused) an instruction, and (clearly amused, with a small chuckle) produces no chuckle at all; the timed-script layer computes every number afterwards — [N.N seconds duration], [N.N seconds pause] for gaps ≥ 0.2 s, (label, N.N seconds) per burst, summed at 12.5 frames per second; burst default 0.28 s, admissible 0.14–1.2 s. "The duration is mandatory in the format and forbidden in the prompt." [P§68]

What was published (§68)
whatwhere
105 burst adapters — 45 per class, 30 per group, plus the ablation and dose arms as evidencelaion/moss-va-sft3-vocal-burst-lora-adapters-v2
the small commercially-licensable burst classifierlaion/vocal-burst-detector-commercial
the production classifier, with the 23-group scheme addedlaion/vocal-burst-detector-x2
117 recipe pages, one per burst label a caller can ask forwikiskills/
the pre-revision recipe tree, kept so an older result stays traceablewikiskills_legacy/
the prompting contract for the directing language modeldocs/DIRECTOR.md
how the parts fit togetherdocs/ENSEMBLE.md

The 23-group scheme also went into four dataset cards (vocal-bursts-segments, vocal-bursts-per-class, vocal-bursts-gemini-segments, dramabox-burst-audio), and the cue convention into six model cards (the SFT3 base, the VoiceNet, emotion, quality, voice and burst adapter sets) and two repositories (voice-acting-search-agent, moss-voiceacting-manual). [P§68]

A ninth way a green exit code lied (§68): the recipes named adapters nobody could fetch. bestmem, grpfull, bulk_mix_full are study arms, not artefacts; the trained weights lived only on a scratch filesystem deleted on a schedule, and not one of the 117 pages carried a hyperlink. A tenth: timed_script.BURST_LABELS held 22 entries against the recipes' 117; a label outside that list is parsed as a delivery direction and no sound is produced. Measured, 77 of 117 labels and 9 of the 36 classes the server actually offers were affected, including guffaw at hit rate 0.633 — the fourth-best recipe in the bank, which could not fire. The vocabulary is now sourced from the recipe directory with a test asserting the two cannot diverge. [P§68]
One disagreement left standing (§68). Twelve served recipes name a weight above 1.5, up to 2.3, while the 2026-09-05 addendum states no recipe should exceed it (all four adapter arms broke the WER gate at w = 2.0, +0.109 to +0.168 against +0.104). The two are not comparable: the table's weights come from a prompt-form sweep in which each cell's own absolute WER passed (sharp_inhale at w = 2.3 measures 0.091), while the addendum is a paired average over a different carrier set. chuckle at w = 2.0 is the best recipe in the bank at 0.73. The disagreement is written into VOCAL_BURSTS.md, and MOSS_BURST_LAM_MAX lets an operator enforce the stricter rule. [P§68]
Still unresolved, and larger (§68): the recipe pages name best-of-N as the more effective lever than any weight change and give a candidate count per class, but the server generates exactly one take per turn. "The programme's central operational advice has no implementation." [P§68]

II.10.13 Two predictors put to a human vote, and the scale that nearly got the question wrong (no date stated) [P§69]

Nothing in the programme had yet asked a person whether the genuineness score and the vocal-burst blend score mean what they claim. Both are load-bearing: two of the four GRPO reward terms, and the corpus quality gate "blend and genuineness both above median". They had been trusted on held-out MAE against Gemini labels — agreement with another model, not with a listener. What is put to the vote: 2,000 paired comparisons over 4,000 distinct clips, four sets of 500 — genuineness easy and hard, blend easy and hard — half from synthetic voice profiles and half from real recordings, published as 8 WebDataset tars (2,000 pairs packed). By whom, and how many votes: the section builds the material for the vote; it names no voters and records no votes cast (n votes not recorded). [P§69]

The thresholds were specified in units nobody had checked (§69). The brief asked for a difference of "at least 2" on easy pairs and "at least 1, at most 1.5" on hard ones; the corpus columns genuineness and blend run 0 to 1. The obvious guess — a 0–10 scale, so 0.20 and 0.10–0.15 — is wrong in two separate ways. (1) The two predictors are not on the same scale: the scorer clamps clamp(0,6) for genuineness and clamp(0,10) for blend and names its keys genuineness_0_6 and blend_0_10; both model cards and the caption template (genuineness x/6, vocal-burst blend x/10) say so. (2) The corpus columns are not those scores at all but quantile ranks — per-voice for the synthetic half, per-(dataset × language) ECDF for the real half. On one source file of 2,393 rows from a single voice, native genuineness runs 0.00–3.20 (median 1.25) while the corpus column runs 0.09–1.00 (median 0.83); because the rank is within a speaker, every voice spans the full 0–1 range. The fix: join the raw genuineness_0_6 and blend_0_10 back by uid3,147,802 rows, zero unjoined — and apply the thresholds literally; feasibility counted first, 120,765 to 461,165 disjoint pairs per cell against a requirement of 250. [P§69]
Native-score selection recomputed on the rank columns (§69)
setnative delta medianrank delta medianrank rule would keeprank ORDER agrees
A genuineness easy2.330.498471/500 (94.2 %)500/500 (100 %)
B genuineness hard1.040.20356/500 (11.2 %)500/500 (100 %)
C blend easy3.380.192241/500 (48.2 %)499/500 (99.8 %)
D blend hard1.210.06257/500 (11.4 %)468/500 (93.6 %)

Which clip is better is almost scale-independent — the two readings order the pair identically in every genuineness pair and in 93.6–99.8 % of blend pairs. Which pairs get built is not: on the hard sets the two readings overlap by about a ninth. Set C runs the wrong way from intuition: pairs whose native blend difference has a median of 3.38 points out of 10 sit at a rank difference of 0.192, below the 0.20 cut that was supposed to mean "easy"; half of the largest real score differences would have been discarded. On set D, 32 of 500 pairs are ordered oppositely by the two readings — there the answer key itself would have been inverted. [P§69]

The dataset column said emolia and meant podcast (§69). The real-speech source's dataset column has exactly the values emolia, kartoffelphon and mls, all cleared for release. Classified by uid, the rows labelled emolia are 531,459 emolia and 984,122 podcast, whose transcripts and audio are explicitly not publishable. 929,825 corpus rows were excluded on the uid rule, 914,287 for this reason. A second exclusion: 15,538 rows labelled mls have uids shaped 2037_10292_000302, outside the verified #####_#####_###### pattern; the standing instruction is not to guess, so they were dropped and the 11,900 that match were kept. [P§69]

What the corpus could not supply. There is no declared age or gender for any real recording; only VoiceNet's predicted perceptual dimensions exist, used to stratify and match and labelled in the README as predictions. voice_key is not a speaker for two of the three real sources: all 505,546 kartoffelphon rows share one placeholder key, and emolia's keys are near-unique per clip (median group size 1). Pairs matched on predicted age and gender band carry speaker_matched: false; 24 of 2,000 final pairs are of that kind. The first selection run hung for over ten minutes because a duration-rescue scan inside that 505,546-row group is quadratic. [P§69]

A ninth exit code that said fine (§69), and it propagated. The decode job was submitted as a smoke test with --dependency=afterok to the full run. The smoke job died in its first second on a SyntaxError, wrote no output, and sacct reported State=COMPLETED, ExitCode=0:0; the full job was released onto the same broken code. Cause: the batch script's last line echo "... EXIT=$?" — the echo succeeds, so the script's exit status is the echo's. Every sbatch in the project ended that way. Now $? is captured into RC and the script ends with exit $RC; decode_pack.py counts clips owed against decoded and pairs owed against written per tar. Three smaller traps: a helper named select.py on sys.path shadowed the standard library's select (surfacing as a circular import inside pyarrow.lib); grouping rows by boolean mask per group is quadratic and hung the feasibility count over 50,327 speakers; pkill -f count.py matched the session's own shell wrapper. [P§69]

Artefacts. Code $SC/code/vand_pairs/; selection $SC/out/vand_pairs/pairs.json; feasibility count $SC/out/vand_pairs/count.log; tars and README $SC/publish/vand_pairs/. Report ~/reports/vand_pairs.md, German page ~/bericht_stimmpaare.html. [P§69]

II.10.14 Re-verification and extension of voice-annotation-data-v2 (no date stated; German in the protocol, translated here) [P§70]

The lead finding (§70). Of 9,390 positives in the dataset, only 32.1 % survive re-verification with the full score of 2; 50.1 % receive a 0, "does not describe this recording at all". These candidates had already passed the question once — Gemini 2.0 Flash was shown exactly this bucket definition and said it fits. "Two thirds of that a better model does not uphold" (translated). [P§70]

This decides more than it seems. laion/emolia-voicenet-gemini-annotations covers 334 of the 395 addressable buckets with ≥ 25 clips, a number that suggests re-verification could be skipped. But those clips were never asked whether a description fits them — they were *classified*, and the bucket is where the classification landed. "85 % covered" silently assumes classified ≈ confirmed, and exactly that assumption is what the measurement above breaks. A sample of 606 of its candidates through the same prompt reaches 28.5 % at 2. Every emitted row therefore carries an evidence field (confirmed_v2 / scored_evga / predicted_vn); no column is called "confirmed" where the model was shown no description. [P§70]

The dataset and the predictors share their definitions. All 399 of 399 bucket texts in the README are byte-identical to the levels in headspub/bench/emolia/repo/dataset/emolia-dim/variables.json; tokcorpus/code/caption2.py calls exactly these texts the "multi-sentence TRAINING ANCHOR for the predicted level" and quotes AGEV bucket 4, which is literally README AGEV-4. Hence the dataset's bucket index *is* the head's ordinal index, vn57 estimates exactly that index (which is why it lies on 0–6), and polarity is correct for all 57 heads by construction. The known BKGN inversion concerns the caption prose, not the numbers. [P§70]

A single scale contradiction, hiding in a coverage number (§70). Across taxonomy, v2 data and the existing annotation, 56 shared dimensions agree exactly in level count. BKGN does not: five levels there, seven here. BKGN/5 and BKGN/6 lie outside that annotation's scale and are not "unobserved" but unreachable — the reason the honest count of empty buckets is 17 and not 15. LANG is the only dimension with no head (57 heads against 58 directories) and cannot be extended; ACNT is in the README but not in the data. [P§70]
The round column is not a repeat measurement (§70). 122 pairs (clip × dimension) were annotated twice; of the 57 with two valid values, all 57 agree exactly. At temperature 0 and thinkingBudget = 0 this *must* be so. "That is determinism, not reliability — reported as '100 % agreement' it would be the most misleading number of this study" (translated). [P§70]

Re-verification was blind and in both directions. 18,632 clips went to gemini-3.8-flash with their *own* definition and without the bucket assignment, temperature 0, schema enforced, Batch API; the question is the same for positives and negatives. 437 negatives reach the 2 — labelling errors of the source, listed individually with path. 462 clips contain no speech at all, against a dataset card that declares all samples speech-checked. thinkingLevel: "low" is not zero here: 651,232 thinking tokens on 18,632 calls (35.0 per call), unlike what §61 measured for the burst prompt. Measured token law: input = 25.06 · dur_s + 352.7; $0.192 per 1,000 requests in batch. [P§70]

Extension over the raw scale, not invented percentile bands. Bucket *b* is the interval [b−0.5; b+0.5) on vn57. The heads are strongly compressed towards the middle, so 129 of the 395 buckets receive monotone, disjoint cut points. The first draft widened each starving bucket around its own centre and produced R_CHST/5 = [4.35; 5.50) beside R_CHST/6 = [4.35; ∞) — the same rows for two levels meant to differ; two adjacent starving buckets always collide this way and it is invisible in the counts. Checked: 0 gaps, 0 overlaps, 0 empty intervals. 39,389 candidates, 393 of 395 buckets full at 100; not full: METL/6 = 61, SMTH/6 = 28, where the publishable corpus is exhausted. 20.1 % reach the 2 (39,387 checked). Spread over 8 source families, both languages in every bucket, 23–97 speakers per bucket (median 86). 969,249 rows were discarded beforehand (podcast and ZH; allowlist by uid pattern, unknown counts as not publishable). Unlike the existing annotation, which is Emolia alone, the draw covers eight sources and two languages — the two complement each other. [P§70]

Audio came from the source, not from the codec. The corpus carries only codes; the original MP3 exists on disk for every row (inline in rcsft/out/sft/*.parquet, or via (shard, audio_key) in the vprof_base tars). That saves 19–28 GPU-hours and delivers the recording rather than a reconstruction at 0.92–0.98 log-mel correlation. [P§70]

Second independent sighting of the §67 pattern — a stage reports success for work that did not happen (§70). Of 2,000 vprof_base tars the index points to, only 1,907 are on disk; that strands 54,962 of 1,200,531 src=0 uids (4.58 %). Audio fetching lost 1,647 clips and reported success, because it ended with "did it write anything, then it is done". Found not from the return value but by counting owed against actual per bucket: median 96 instead of 100. The same figure of thought as the echo "… EXIT=$?" at the end of 95 of the project's 115 sbatch scripts — here without SLURM. Fixed: non-zero return on any shortfall, a fetch_report.json with owed/actual/missing uids, and unfetchable uids excluded *before* the draw. Second run: 37,765 clips, 0 missing. "The only check that finds this is both times the same: count owed against actual" (translated). [P§70]

Data status of the section: re-verification evaluated over 18,631 of 18,632 clips (100.0 %); shares stable across the first subsets (32.5 % → 32.1 % at 2). Extension: 39,387 of 39,389 candidates scored, sample of the existing annotation 606 of 606. The Batch API delivers without commitment; the drivers keep running and vbk_protocol.py --update rewrites the section in place with the final numbers. Evidence: ~/reports/vand_buckets.md and .html, CSV per bucket under $SC/out/vand_buckets/csv/{DIM}/{VALUE}.csv, code $SC/code/vand_buckets/. [P§70]

Corrections and retractions recorded in §57–§70: §57 — the §54 dynamic-contrast defect is attributed to the prompt (no calm lead-in), not the model, and the peak-over-median statistic is declared unusable as a loudness ranking. §57.4 — an MP3-encoder proof written as ffmpeg160 -version | head -1 let a pilot exit 0 after writing zero clips; replaced by a real round-trip check. §58.2 — gt_select.py admitted or excluded classes by the presence of a strict-bucket file; corrected additively (+387 rows over 13 classes, 0 changed, 0 dropped). §58.6 — §55's Gemini span-coverage figure (38.2 %) does not replicate (median 10 %, mean 17.6 %); the locator warning stands regardless. §58.8 — MP3 via lameenc was dropped for a 1,105-sample (23.0 ms) encoder delay; OGG ships instead. §58.9 — a hard-coded protocol number was taken twice by other agents; numbering now re-reads under flock. §60 — re-labelling with Gemini is a null or negative at every weight; the null is flagged ambiguous because the scoring detector is the instrument being replaced. §61 — thinkingLevel: "low" was assumed to cost nothing and billed 2,362,228 thinking tokens; the check moved into the merge step. §62 — §58's 0 out-of-taxonomy labels is contradicted for the DramaBox pass (0.28 % top-1, 1.05 % any slot); §61's "0 clamped" holds only at tolerance ≥ 5 ms; the shipped detector emits 41/36 labels, not 8, and does emit Shriek. §62 — the annotator returns bursts on digital zeros at 100 %, so the 99 % silence-negative figure and part of §61's 11-of-20 control measure confabulation; the ASR negative filter buys ~1 point at matched level, not 16.7 % → 4.3 %. §62 — two silent cutter bugs (Parakeet-TDT contiguous spans; dot-split tar keys reading 1,000 clips as 96) were fixed before release. §62/§63 — PEFT silently loaded a VoiceCLAP LoRA matching 0 of 1,232 tensors; fixed by key remap and an assertion. §63 — the scorer contract test's exact-argmax assertion (237/240) was replaced by a probability-agreement plus near-tie invariant. §64 — the genuineness guardrail no longer gates burst recipes, moving recommended weights from 0.25 to 1.0–2.0 by rule change, not measurement. §65 — §64's pooled family-level gain (+0.0356, t +2.67) is retracted as a general claim: negative on the 16 verifiable classes, unconfirmed on the 29 others. §65 — the author's pre-measurement assertion that no recommendation changes was false (argmax moved for 14 of 16); recipes are relabelled "a good setting", not "the best". §65 — burst_family.py is coarser than the published 23-group scheme (21 classes clear 0.15, not 30; exasperated_sigh 0.833 → 0.200), and reward._same_class is a substring test (12 of 3,269 promotions). §66 — combining member adapters at equal budget is a clean negative; the grouping map had four unmapped emittable labels breaking monotonicity, fixed by full coverage (82/82). §67 — mid-run rate estimates (2.76, 6.94) against a true 1.61 cells/min/rank cost a job nine cells short at 99.8 %; positional pairing corrupted 77 of 816 conditions; an early random-grouping control was mis-built (sign measured at 0.0016). §67 — Slurm recorded COMPLETED 0:0 on a job that lost 99 of 400 cells; completeness is now counted per rank. §68 — recipe pages named adapters (bestmem, grpfull, bulk_mix_full) that were unreachable study arms; timed_script.BURST_LABELS held 22 of 117 labels, silencing 9 of 36 served classes including guffaw. §69 — the pair-selection thresholds had been read on the wrong scale twice over (mixed 0–6 / 0–10 predictors; corpus columns are within-speaker ranks); selection was redone on raw scores. §69 — the dataset column labelled 984,122 podcast rows as emolia; excluded by uid allowlist. §69 — the sbatch trailing echo "... EXIT=$?" made a failed smoke test report COMPLETED 0:0 and release the full run; fixed with exit $RC. §70 — only 32.1 % of 9,390 previously confirmed positives survive re-verification; 437 negatives are positives; 462 clips contain no speech despite the card; BKGN has 7 levels against 5 in the existing annotation (17 empty buckets, not 15); the round column's 57/57 agreement is determinism, not reliability; the first bucket-widening draft produced overlapping intervals; the audio fetch reported success while losing 1,647 clips (1,907 of 2,000 tars present), fixed by owed-against-actual counting.

Part III — The journey: failures, retractions and measurement errors, in order

Part I and Part II report what the system is and what it measures. This Part reports how it got there, and it is organised around the things that went wrong — because in this project the diagnosed failures changed how the work is evaluated at all. III.1 is the six cases in which a training metric was clear, consistent and wrong. III.2 is the engineering failures that looked like something else. III.3 is every claim that a version of this report or the protocol made and later withdrew. III.4 is the agent era — the four measurement errors of September, the babble bug and the eight ways a green exit code lied. III.5 records every contradiction between sources that this report found and could not resolve. III.6 is a chronology.

III.1 When a training metric lies — six cases before the agent

This project selects checkpoints, ranks adapters and decides whether a run worked by generating audio and scoring it, never by the training objective. That is more expensive — a 320-clip evaluation costs a GPU-hour or so — and it is not a stylistic preference. It is the conclusion of six separate cases in which a number was clear, consistent, and wrong — five of them training-time metrics, and one a measurement instrument saturating on audio that was not speech. Source: [TR-0829] §18; the first five in [TR-0826] §13.

III.1.1 A preference-run “damage” indicator at its worst while the model generated best

Preference accuracy is the metric a DPO run naturally reports, and it rises monotonically in every run here — 0.865 → 0.981 in one, up to 0.994 in another. Over the same steps reward(chosen), the implicit reward on the preferred sequence, falls monotonically. The round-1 model card states the consequence outright: “preference accuracy is not a health metric here. It was highest exactly where the model was most degraded.” And the inverse holds too, within a run. The classifier-free-guidance run of I.10.5 ends with reward(chosen) at −1.12 — three times worse than at its first evaluation — while generating the best audio in the project at the time: lowest word error rate, highest quality, best bursts. Had that indicator been used to select, the best checkpoint would have been discarded.

One caveat that the 26 August version of this section got wrong, corrected on 28 August rather than quietly: −1.12 is not the worst reward(chosen) in the project, and the compressed cross-run form of this story — “the worst indicator produced the best model” — is false. The corpus-v2 run ends at −2.30 and produced the fifth-best model. I.10.8 gives all seven runs. The lesson survives in its within-run form, which is the form in which it was actually observed. A further correction arrived on 31 August ([P§42]): rew_chosen is measured against bare SFT3, not against the run's own starting point, which changes how the absolute values of different runs may be compared at all — II.8 carries it.

III.1.2 A validation loss that could not move, because the holdout was the wrong set

The three emotion adapters of I.12.3 differ by a factor of four in trainable parameters and by 0.036 in emotion percentile on generation. Their validation losses all sat at 4.652–4.654, flat from step 1,225 onward and separated only in the fourth decimal. The reason is structural: the holdout is the general corpus holdout, not the extreme subset, so an adapter specialising in intense emotion cannot improve it. The loss was not noisy — it was measuring a different question.

III.1.3 A validation loss that was consistent, monotone, replicated ten times — and still wrong

The cleanest case is from the voice-profile rank ablation. Held-out validation loss separates the ranks monotonically — r16 < r8 < r4 — for every one of the ten voices, without exception. On loss alone you would ship rank 16 ten times out of ten, with clean, consistent, unanimous evidence.

Rank ablation, loss side ([TR-0829] §18.3).
QuantityValue
Base validation loss3.94 – 4.20 nats
Rank 16, stage 23.73 – 3.95 nats
Mean base → r16 improvement0.217 nats
Mean r4 − r16 gap+0.0084 nats
Gap as a share of the gain3.9 % (range 3.2–5.3 % across the ten voices)

The generation-based evaluation turns that 3.9 % into no measurable difference at all on 1,920 held-out clips per arm (p = 0.22 on speaker similarity, 0.94 on reward, 0.32 on WER). The two measurements agree on the ordering and disagree only on whether the remaining gap is worth paying for. Selecting on loss would have cost four times the adapter parameters to buy nothing measurable. The honest limit: the generation evaluation is itself a learned scorer, so what has really been shown is that the gap is below the resolution of every instrument that was pointed at it. No human was asked. The manual reports the same failure a third time, from a different direction: at rank 64 the validation loss rose from 4.4237 at epoch 1 to 5.4369 at epoch 8, while listeners rated epoch 1 and epoch 8 identically (1.969 both).

III.1.4 A reward that rose because the reward function had changed

GRPO run 5 reported a training-time emotion term of 0.441 against run 4's 0.259 — a 70 % improvement. It was not one. Two of the four changes made for that run (I.11.4) inflate that number directly: the ramp pays for partial progress that the strict yardstick does not count, and the curriculum asks easier questions. On the unchanged evaluation, run 5 is the worst of the four models on emotion. A training reward is only comparable across runs if the reward function did not change; when it does, the only honest comparison is on a yardstick that did not.

III.1.5 And one case where the comparison itself was on two scales

The round-3 supervised model was evaluated before the zero-point percentile correction of I.6.7 and the preference and RL models after it. For a period the emotion column — and through it the composite reward — was on two different scales, and SFT-3's 0.4900 was inflated relative to the preference model's 0.4668. Every model was re-evaluated on the identical prompt set with the corrected scale; the superseded directory is kept under an explicit _UNCORRECTED_ECDF name. The two files are byte-identical on every field except four, which is what makes the size of the distortion exactly measurable rather than estimated. The rule the project now works to: ablate on the output, not on the objective — and when the objective changes, keep one yardstick that does not.

III.1.6 A quality scorer reporting a perfect score on audio that was not speech

The sixth case is not a training metric but an evaluation instrument, and it is the most extreme of the six. In the steering grid of 26 August (II.3.3), the condition that injects the quality direction at strength 2.0 reports 1.000 on both genuineness and vocal-burst blend — the top of both reward terms — and the highest emotion percentile of the entire grid. Every other cell in that column reports 1.000 as well, at every strength and every tap. The same clips transcribe at word error rate 1.030 and run 13.76 seconds away from the length their own script asks for. The underlying raw scores are genuineness 1.31 out of 6 and blend 1.97 out of 10 — near the bottom of both scales. The percentile mapping that turns a raw score into a reward term was fitted on a corpus of speech; handed something that is not speech, it has no basis for a judgement and returns the top of its range.

The generalisable form: a learned scorer's output is only meaningful inside the distribution it was fitted on, and a percentile mapping actively hides the moment you leave that distribution — a raw score of 1.31/6 looks obviously wrong, and its percentile of 1.000 looks like a triumph. Two cheap defences caught it here, and both are structural rather than clever: a word-error control and a duration control run beside every quality number, and a constancy check (a metric reading exactly 1.000 in twelve consecutive cells is reporting a saturation, not a result). The agent-era version of the same defence is the human-proxy target of II.6, against which every reward term is correlated before it is trusted.

III.2 Engineering failures worth publishing

Several of these cost node-days, and all of them are the kind of failure that looks like something else. They are recorded because the diagnosis, not the fix, is the transferable part. Source: [TR-0829] §19; [TR-0826] §14; the agent-era additions are in III.4 and II.8–II.10.

III.2.1 Distributed rank divergence — a hang that looks like a hardware fault

In the GRPO loop each rank decided independently whether to skip a rollout group: if len(recs) < 4: skipped += 1; continue · if r.std() < 1e-6: skipped += 1; continue · if nseq == 0: continue # skips opt.step() too. So one rank ran three backward passes while another ran two, and the gradient all-reduces went out of step. The watchdog fired after ten minutes: BROADCAST, NumelIn=104 … ran for 600066 milliseconds, exit code 134, and 35 GB of core files. The fix is to make the number of backward passes a constant of the configuration rather than of the data: a degenerate group is kept with advantage 0, every group is padded to the full group size, and padding entries are excluded from the KL term so they contribute no gradient. Runs 4 and 5 report skipped = 0 across 176 and 210 groups. The same failure class had already appeared once as a heartbeat failure in a preference run, and was reproduced by hand a third time while rewriting the bucket-adapter trainer. The general form: any per-rank continue in a data-parallel training loop is a deadlock waiting for the right batch. The bucket-adapter trainer of I.12.5 therefore uses no distributed data parallelism at all.

III.2.2 Weight tying makes the preference adapter unmergeable

The published DPO adapter ships unmerged. merge_and_unload() corrupts the model, because audio_lm_heads.N.weight and audio_embeddings.N.weight are the same tensor (they are tied at load), so merging the head delta also rewrites the embedding that produced it. Measured: both tensors changed by exactly 6.103515625e−05 while the text embedding changed by 0.0, and a logit-equivalence probe reported hidden: max|diff| = 3.33, rms = 1.07e−01 against a signal rms of 2.52. That constant is 2⁻¹⁴ — the bf16 quantum at that magnitude — which is exactly the signature of a single shared tensor being written once and read twice. Three independent corroborations sit on disk: the adapter configuration targets audio_lm_heads.0…11 and carries "ensure_weight_tying": false; every training log reports audio_lm_heads.{0…11}.weight | MISSING in the load report, because those tensors are not serialised separately; and at save time PEFT prints Removed shared tensor …audio_lm_heads.3.lora_A…. The adapters are stacked at inference instead, which is also why every adapter in I.12 is trained on the supervised checkpoint rather than on the preference model — and why the demo server's own merge-into-weights loader was a defect (I.14.6).

III.2.3 A padding token that reaches the wrong embedding table — three times

Padding every channel with the text pad token puts an id around 151,643 into the audio channels, whose embedding tables hold 1,025 entries. The gather runs out of bounds and reports device-side assert triggered — vectorized gather kernel index out of bounds 19 seconds into the first rollout, taking the ranks down one by one and looking exactly like a hardware fault. Text and audio channels must be padded separately. The identical bug was then reproduced twice more, in different files. Once by omitting the label masking when the batch packing was reimplemented — an 8-node job died in three minutes on every rank. And again on the morning of 26 August, when the bucket-LoRA trainer was launched for the first time: all 32 ranks died simultaneously at 08:51 with the same assert. It is now copied verbatim with a comment saying why it is not optional, and validated on a single-node smoke test before spending an 8-node allocation.

III.2.4 A finite loss with a non-finite gradient, and an adapter that reported DONE after zero steps

After the padding fix, the bucket-LoRA jobs restarted and produced 1,305 loss=nan lines across two 8-node jobs before they were cancelled 25 minutes in. A single-GPU smoke test diagnosed it in six minutes — 21 non-finite losses in 1 steps — refusing to train on garbage — and a second smoke test caught the worse variant: a finite loss with a non-finite gradient, reported as non-finite grad norm at step 0, step skipped. That run then printed DONE emotion/Affection rows=5235 steps=0 17.9m. An adapter had been trained for eighteen minutes, taken zero optimizer steps, and reported success. The root cause: the vendored model's own gradient-checkpointing path is broken by late binding — every layer's wrapper closes over the last layer, so the recomputed activations do not match the forward pass and the gradients come back inf or NaN. The fix is to disable the vendored path entirely and wrap each decoder layer with a non-reentrant checkpoint wrapper. This is the third distinct gradient-checkpointing failure in this project (I.3.5 has the first). Two guards now stand in the trainer: a non-finite-loss counter that refuses to continue, and a non-finite-gradient check that skips the step. The second one is what turned a silent zero-step success into a visible failure — but note that it still printed “DONE”. A guard that skips is not a guard that reports.

III.2.5 The prompt contradicted the reward

Real-speech captions state the source clip's measured genuineness X/6 and vocal-burst blend Y/10 — exactly the two quantities the reward's quality term optimises. On a low-scoring source row, the model was therefore rewarded for contradicting its own instruction. Fixed by stripping those clauses in the RL prompt only; this is not out of distribution, because most voice-profile caption templates carry no such clause.

III.2.6 A shared-mode error that a consistency check could not see

The most instructive of the lot, and it predates the timed-script round. The annotated corpus shipped for weeks with the invariant moss_frames == round(dur_s × 12.5) passing on 100 % of rows. It passed because both terms were doubled together: a decoder bug had been flattening stereo frames, producing half-speed audio, and the duration, the frame count, every emotion score, all 57 VoiceNet dimensions, the WER and the word timestamps were all computed on it. An earlier verification pass had reported that 100 % as proof of correctness. It was found only by decoding the audio and comparing against an independent reader. The rule: a consistency check cannot detect a shared-mode error. Assert the invariant on the emitted artefact, and make at least one leg of the check independent of the pipeline that produced it. The base-model card carries the decoding warning — use proc.decode_audio_codes(..., return_stereo=False); the direct tokenizer path yields half-speed audio; “this project lost a whole corpus to that once”. The 16 August predecessor of this bug is the “same audio twice” stereo-flatten of I.16.2.

III.2.7 A checkpoint rotation that deleted the best model of its run

The DPO run on corpus v2 produced its best checkpoint at step 3216 — reward 0.4687, at the time the best model in the project. The trainer's --keep-last 2 policy then rotated it away. The evaluation directory survives, complete, which is why the number can still be quoted; the weights do not. That checkpoint no longer exists and cannot be republished. The lesson was applied afterwards: the CFG-DPO run has a sibling keep/ directory holding every checkpoint from step 212 through 4912, twenty-four of them — which is why step 4912 was still on disk to be published. The contrastive-families run rotated again, with six retained and no keep/ directory — it kept steps 3906 through 5022, the best of them survived, and the exposure was real rather than theoretical.

III.2.8 A module-cache race that cost 48 node-hours

128 ranks raced to populate the trust_remote_code module cache for a directory nobody had loaded yet. Some read a half-written source file (cannot import name MossQwen3Model); the survivors hung in the initialisation collective. Fixed by a single-process pre-warm before the parallel launch — three lines in the batch script. A related environment hazard on this filesystem: concurrent cold imports of large Python packages put every worker into uninterruptible sleep for up to 25 minutes with the GPUs idle, so launches are staggered by 45–150 s.

III.2.9 A run that trained on the wrong base, kept as a result

One preference run intended to start from the round-3 supervised checkpoint started from round 2 instead: the batch script reads its base from one variable, a different one was exported, and line 21 silently replaced it with the default. The log stated it plainly and that is where it was caught. The run is kept under the name dpo_on_sft2_WRONGBASE, because it is a valid experiment in its own right — preference tuning on SFT-2 under the round-3 prompt format. Its telemetry is the familiar shape: preference accuracy 0.709 → 0.818 while reward(chosen) falls −0.071 → −0.635 and reward(rejected) collapses to −8.34, with "healthy": false on every row. It was never run through the standard evaluation harness, so it contributes no reward or WER number to the ranking. It used 32 nodes, giving a global batch of 1,024 and only 184 optimizer steps for the same data, where an earlier run at 8 nodes got 864. The corrected run uses 8 nodes so the two are comparable.

III.2.10 Two silent evaluation bugs, and a library trap

III.2.11 A success message is not evidence — seven times in three days

Each row is a different tool lying in a different way ([TR-0829] §19.11).
ToolWhat it saidWhat was true
SchedulerARRAY DONE … rc=0Workers died in under a second on a mis-quoted path
Scheduler3-line log, rc=0, 5.5 h elapsedActually fine — indistinguishable from failure
Hub uploaduploaded in 145s plus a URL8 commits rejected
Hub upload“retrying in smaller chunks”Retrying against an unfixable per-directory cap
Process killsuccess2 processes still alive
Process pooltasks RUNNING, no error3 Python processes where 24 should be, load average 0.12
Transfer clienttransfer progressingInfinite retry on “disk quota exceeded”

The remedy that works: process tables, repository metadata, parquet footers, decoded audio, controlled probes — never the status line. It is the reason every published repository is byte-verified after upload rather than trusted (including this report's own upload, IV.3), and the reason III.2.4's “DONE … steps=0” was caught at all. The agent era added eight more ways ([P§67], III.4.3).

III.3 Every claim that was made and later withdrawn, in order

Each version of the technical report removed or corrected claims of the previous one, and each said so. This section collects all of them, so that a reader who saw any earlier version knows what changed. The agent-era retractions (III.4.2) and the burst-programme corrections (II.8–II.10, closing notes) are listed by their own sections. Source: [TR-0829] §23.3 and §18.1; [TR-0826] §18; [ACTOR] §7.2.

Claims removed or corrected between the 26 August and 29 August reports ([TR-0829] §23.3).
Claim in the 26 August draftStatus on 29 August
Training details asserted for the vocal-burst blend predictor beyond its head architecture and validation tableRemoved. The card publishes no training-set size, no label source and no split description. I.6.3 now says so explicitly.
“The [DPO] model card says why: merge_and_unload() corrupts the model.”Corrected. No such warning exists in any of the four cards. The measurement is real; the card does not carry it (III.2.2).
“Jealousy and Envy reads 0.000 everywhere because that head sits at the very bottom of its own distribution.”Corrected. The real cause is a key mismatch — Jealousy_&_Envy in the predictor repository against Jealousy_and_Envy everywhere else — so the head is never read (I.6.4).
“There is no trend” presented as the merge-weight resultScoped. True of the single general adapter at the intense band; false of the per-emotion adapters at the extreme band. Both are reported (I.12.4).
VoiceCLAP-large-v2 “scoring better on every benchmark reported” against “the small model”, quoted next to VoiceCLAP-commercialClarified. The 0.7069 vs 0.6754 comparison on that card is against voiceclap-small, not against voiceclap-commercial.
In-flight job states (“31 of 40 adapters written”, “268 adapters written”, “138,088 pairs”)Superseded. 40 emotion and 500 voice adapters are finished and published; the pair build that shipped is 368,517 rows (I.13).
The 17 per-emotion adapters described as if the set were chosenCorrected. It is where a sharding bug landed; 40 exist (I.12.5).
“All prompts drawn from the same three speakers”Weakened. They are three source corpora; the speaker-extraction function fell through to a corpus-prefix fallback (I.12.5, III.2.10).
“CFG-DPO is the best model this project has produced”Superseded on 28 August. True when written; the contrastive-families adapter at step 5022 beats it on reward and emotion percentile (I.13.4). The CFG checkpoint remains the best on burst realisation and hit rate.
Preference corpus v1 given as 236,656 pairsCorrected to 236,166 against the shards and the build log (I.10.3).
Round-3 supervised fine-tuning described as using 247,049 rowsCorrected. That is the emotion-LoRA corpus; round 3 used corpus_x (I.9.5).
“The DPO run whose rew_chosen was worst produced the best model” ([TR-0826] §13.1)Corrected on 28 August. The worst final rew_chosen (−2.3037) belongs to the corpus-v2 run, whose best checkpoint ranks fifth. Only the within-run form is supported (I.10.8, III.1.1).
Steering “is a clean negative result” at every strength ([TR-0829] §16.3, written on a grid whose smallest α was 0.5)Superseded on 28 August by the low-α re-run (α 0.05–0.30 on h20 with matched random directions): at α = 0.10 emotion percentile 0.435 → 0.584 while WER 0.166 → 0.115, random +0.014, paired over 24 prompts, dim − rand +0.135, t = 3.11; cost genuineness 3.77 → 3.15. Everything above α = 0.3 stands as written. Then contradicted in the ear by the project leader's listening test (II.3.3) and rolled back.
An earlier reading of GRPO run 4's single steps as an emotion collapseWithdrawn. The aggregate over 176 groups does not support it; per-step numbers of two groups each are far too noisy to read as a trend (I.11.2).
The “moderate contains, intense flattens” story of contained emotionOverturned by the 1,878-clip paired re-run; it was an artefact of four anchors and sparse unpaired data (I.4.9).
Round 2's conclusion that “tags lose”An artefact of the angle-bracket notation bug in the WER stripper (I.4.11).
The 8-of-83 / “never Shriek” detector finding of the first agent reportA 60-clip / 73-event artefact; on the full corpora it is 41 / 36 labels, Shriek 50 / 49 clips (II.5.2, III.4).

Two further claims were removed from the 26 August report relative to the 16 August reference report, in the other direction: the whole DramaBox / LaionBox DiT campaign and the search-agent detail were dropped from the narrative because they concern a different model. They are restored in I.16, not because they were wrong, but because the brief for this report forbids losing them.

III.4 The agent era, September 2026 — what went wrong and what followed

A result without the road to it is worthless in this report. The lab notebook ~/paper/protocol.md has 8,219 lines and 72 sections in chronological order; nothing in it is rewritten after the fact, superseded results are marked as superseded. This section extracts what is the most interesting part of it: the failures, the corrections and the retracted claims.

III.4.1 Four measurement errors, each of which had already produced a published number

DefectConsequencehow it was foundProtocol
The percentile scale of the emotion is not comparable across heads: a raw value of zero — no measurable emotion — receives the percentile 0.822 on Awe and 0.002 on Affection.“Demand the band 0.90–0.98” (translated) was a different task for each emotion; cross-model rewards stood on two scales. Within a GRPO group of the same emotion the term is almost constant near the floor — and a constant term contributes nothing to a group-normalised advantage.by recomputing where the zero point of each head lies. Fixed by an affine rescaling of the tail; verified that all 40 heads map the raw value 0 to exactly 0 and that no head loses monotonicity — every comparison of two clips of the same emotion is thereby unchanged.[P§7.1]
The evaluation texts were written to match their emotions.35.1 % of the variance of the emotion score is explained by the choice of text, median 0.972, 58.7 % of the clips at ≥ 0.95. The instrument is saturated before the model speaks.variance decomposition over 60 utterances. A non-saturating companion metric (rank within the same emotion) agrees everywhere, so the null findings are not merely a ceiling effect — but the ceiling is real.[P§18]
Whisper writes fluent text over audio that has stopped being speech, and on identical audio is 2.33 times as noisy as Parakeet.Every word-error threshold of the first five weeks carries an error bar 2.4 times too wide.both ASRs recorded on every take and the spread compared. Side finding that reverses the expectation: Whisper is the stricter instrument, the old thresholds were not too generous.[P§28], [P§30.2]
The smallest steering strength of the first grid was α = 0.5, where the output is already destroyed.A false-negative result was published: “steering works at no strength” (translated). At α = 0.5 to 2.0 a random direction really does the same damage — that was measured correctly and read wrongly.a ladder below 0.5 with a matched random direction at every strength. At α = 0.10 the true direction moves +0.149 (t 3.17) and the random direction +0.014 (t 0.48). Every page that carried the old claim has received a superseded notice.[P§23.7]
The common denominator of the four: reading a curve from its edge, using a scale without a zero point, trusting an instrument without a control, and building test material that already contains the answer. None of them was a programming error.

III.4.2 Claims that were retracted or reversed

Claimwhat became of it
“The shipped detector uses 8 of 83 labels and never emits Shriek.” (translated)Artefact of 60 clips with 73 events. Over the full corpora it is 41 and 36 labels respectively, and Shriek occurs on 50 and 49 clips respectively. The correct statement: its *effective* vocabulary is small. Corrected in four READMEs.
“No weight brings a gain without violating the guardrails.” (translated)Flipped to the opposite through a metric correction. The pre-registered primary metric was an F1 with a duration term: every additional detection pushes precision down, and a burst of the wrong length is pulled towards zero — exactly the two things a higher dose does. On the hit rate the same dose does move (+0.061 at w = 0.5 to +0.091 at w = 1.5, t 2.3–3.0). Since then the rule is: for bursts the hit rate counts first, and that Genuineness drops on a scream is the expected price and not grounds for exclusion.
“It is the weight of a single adapter that destroys a line, not the sum.” (translated; stood in the source code, backed by a real measurement)Refuted. That measurement ran on the bare model. In the production stack two adapters at 1.5 each destroy the line in 5 of 5 seeds, one at 1.5 in none. The refuted comment still stands above the correct version in the same file to this day.
bestmem — borrowing the strongest adapter of the group — brings +0.0762 (t 3.76) and costs zero GPU hours.” (translated)A clean negative under the production stack (−0.0531, t −2.50). Same adapters, same prompts, same metric, same seed rule — opposite sign, and the difference is the stack: four duration adapters there against seven here.
“The recommended burst adapters are not published.” (translated)Superseded: 105 adapters have been published since 5 September, and config.py already points to them. The documentation contradicted the configuration of the same file collection.
“The stage direction is the central lever” (translated; the premise of the entire third prompt round)−0.002 [−0.023, +0.020] over 720 paired comparisons. It does nothing. What works is the text.
“The model holds its tempo constant at around 21 characters per spoken second.” (translated)Falsified in the very next measurement. The 21 came from a study that built *every* prompt with the same constant — there was no variation of the requested rate at all. One operating point cannot establish a slope over a range that was never covered. The measured value stands, the conclusion drawn from it does not.
“The rate adapter addresses an uncontrolled quantity.” (translated)Also withdrawn, for the same reason in the corpus: the rate was computed from the printed duration and is redundant with it. The adapter is an interface gain, not a capability gain.
“No merge weight lifts a floor of zero” (translated; about shriek and frustrated_groan)The core statement remains correct — only the floor was partly not one. In DramaBox shriek is present in 88.3 % of 1,500 clips and scream in 88.6 %, against 14 lines in the entire real corpus. It was a data problem.
“No recommendation changes when one re-scores with the second detector.” (translated; written down before it was measured)Wrong. The best cell moves for 14 of 16 checkable classes. But the decision-relevant quantity is the regret — what it costs to keep the recommended cell, measured with the other instrument — and that is symmetric (median 0.083 against 0.100 in the other direction), i.e. just above one seed spread. Since then every row is called *a good setting* rather than *the best*. And: do not write down the result of a check you have not run.

III.4.3 Eight ways a green exit code lied

This list is the most transferable part of the whole protocol, and it came into being in a single day. Its common pattern has a name: a check that reports success by doing nothing.

  1. The scheduler reports COMPLETED 0:0 for a job that lost a quarter of its work. One rank died of a CUDA error after writing 1 of its 100 owed cells; the other three ran through normally. Result 301 of 400 cells — and every aggregation would silently have averaged over three quarters of the prompts. The only reliable check is to count owed against present output rows, per rank.
  2. A second instrument is worth more than a larger n. A pooled family gain of +0.0356 (t +2.67) over 45 classes looked solid. Split by whether an independently trained second detector can name the class at all: negative on the 16 checkable classes, and the entire pooled gain lives in the 29 classes for which there is no second opinion. More prompts would have narrowed the t and left the claim just as wrong.
  3. Rates measured mid-run flatter. Four numbers for the same quantity: planning 1.87 cells/min/rank · interim 2.76 · snapshot from the log 6.94 · truth over the whole run 1.61. On the strength of the interim number a job was shortened from 5:00 to 2:30 and ran into the time limit at 5,591 of 5,600 cells — nine cells short.
  4. Pairing by list position corrupts silently. 77 of 816 conditions had fewer than ten usable prompts; a cell without decodable audio does not appear at all, so position *i* in one list was a different prompt than in the other. The comparisons produced perfectly plausible t-values. Pair on an explicit key, take the intersection, and print n next to it.
  5. Shell variables that are set but not exported kill completeness checks silently. A child process sees nothing, and “nothing owed” is read as “complete”. One case died after every one of eight training jobs with a KeyError, while every job kept printing DONE — the check that was supposed to catch a missing adapter was itself missing, eight times.
  6. A grep that skips binary files makes “no match” and “never looked” indistinguishable — and log files with progress bars are exactly such files. The true error was in there; the enclosing check returned the same as for genuine absence.
  7. Every grouping raises a hit rate, so a size-matched random control is mandatory. Of +0.235 raw gain, +0.054 was pure arithmetic; only +0.182 was the grouping itself.
  8. A script can do nothing and exit with 0. A path helper took one directory level instead of three, reported “1 adapter, 0 references, 0 evaluable classes” — the correct figures would have been 20, 57 and 45 — and exited with 0. And an over-broad text substitution silently deleted three whole content blocks in a report generator; the page built, passed its syntax check and was 44 kB instead of 52 kB.

Two more from the same family came later and are expensive enough to be named. echo "… EXIT=$?" as the last line of a batch script overwrites the exit code with that of the echo — a smoke test that died of a syntax error in its first second and wrote nothing was reported as COMPLETED, ExitCode=0:0, the --dependency=afterok chain was satisfied, and the full job ran on the same broken code. And a published page is verified when its scripts have been executed against its real data — not when its files return 200. A listening page returned HTTP 200, all 140 audio files returned 200, the JSON was valid, the tables were rendered — and not a single player worked, because a stale copy of the renderer in the same script block threw an exception and thereby aborted the whole block. A human found it with the sentence “I can't find the audios here” (translated).

III.4.4 Two places where the same file contradicts itself

Where does the vocal-burst blend live in the network? One evaluation says h25, the other h33, both on the same extraction, both correctly transcribed from their artefact. They differ because one fits a shared 99-output probe and selects on test, the other a group-specific one and selects on validation — and neither place said so. For Genuineness (h12 against h23) that is a tie under its own resolution rule. For Blend it is not: the first evaluation scores h33 at 0.3529, i.e. 0.082 below its own best, far outside any noise band. Both sections now carry their probe and their selection rule; the substantive question is not resolved.

And a number hard-typed into a report generator migrated into three further documents. The saturation step of the emotion adapters stood there in three strings as −0.007 (t −0.90, 19 of 40); computed from the data with the same generator's own helper functions it is −0.0026 (t −0.40, 20 of 40). The wrong number migrated from the generator into its HTML report, from there into the narrative of a second study and its report, and from there into the protocol, where it is the stated justification for running a whole guidance ladder at w = 1.0. The *computed* numbers of the same documents were all correct, so nothing looked wrong anywhere.

That is exactly why this report is rendered and not written: every number in a table is fetched from its source at build time, and every number from prose carries its location and is checked against it at build time.

III.4.5 The balance sheet — what works and what does not

All interventions of the burst programme with their standing verdict
InterventionVerdict
adapters per class on real audio with detector labels (the bank that ships)works for a minority of classes; the baseline against which everything else is measured
the same audio relabelled with a better annotatorzero at every weight
more adapter capacity, rank 32 and 64zero — capacity is not the bottleneck; rank 64 additionally costs word error rate
prompt form: cause in the GENERAL line, longer duration, placementworks and is additive — the largest reliable lever
guidance on the burstworks, g = 4 triples the hit rate — but the prompt carries the largest part
burst-and-stop DPO adapterzero on burst realisation; if anything, ship it for Genuineness
requesting a neighbouring class as a substitutezero at family level, significant harm on the strict one
training on a second TTS corpusone large class gain; the pooled gain does not survive a second instrument
combining member adapters at inference timenegative in 8 of 8 fields, 0 of 120 class wins
training on pooled group datanegative against every baseline
borrowing the strongest adapter of a grouppositive in the lab stack, negative in the production stack — see III.4.2
the merge weightworks; the optimum moved to 1.0–2.0 as soon as Genuineness stopped capping it
generating several candidates and selectingthe largest practical lever — and the one every recommendation now starts with

Two things are worth drawing out of this table. Almost everything that changes the model fails; what works are ways of asking — prompt form, guidance, weight and more draws. And the one training intervention that produced a large gain produced it on a single class, in a way that a second instrument could confirm nowhere else.

What was measured. This section summarises; its numbers carry the n of their own sections. For completeness, the check that applied to the protocol itself: an audit over [P§22–§33] checked every claim against its file and recomputed every table from raw rows where raw rows exist. Result: 5 contradictions (the same quantity with two values), 2 cases of two different measurements presented as the same quantity, 4 plainly wrong numbers, 1 dead file reference, 4 claims without any source file and 6 phrased too confidently. Everything recomputed from raw rows reproduced exactly — the errors all sit in prose that cites another section.

III.5 Contradictions between sources, recorded rather than smoothed over

Where the brief for a report and the sources disagreed, the sources won and the disagreement was recorded. This list merges the 29 August list ([TR-0829] §23.2, eight of whose entries were themselves found during the 28 August expansion), the 26 August list ([TR-0826] §18.2), and the contradictions this report found while merging its five sources.

III.6 Chronology

Dates as stated in the sources; where a source gives only a month or relative order, the row says so. The pre-JUPITER dates come from the run notes via [TR-0829] and [REF-V1]; the JUPITER dates from the protocol and the reports.

Chronology of the voice-acting line.
Date (2026)EventWhere
before 22 JulVoice-acting v1: full FT of v1.5 on DramaBox-style data; rank-256 LoRA on 1 M samples merged; burst LoRA; burst hit-rate study (2,304 generations)I.3.2
22–23 JulVoice-acting v2: six-dataset full FT, 185,503 rows, 6 epochs, 8×A100, 14–15 h; selected on generated audioI.3.3
26 JulEmotion LoRA v1 — the rescue result (Amusement −0.51 → +2.50)I.4.1
27 JulOvernight 34-emotion scale-up; the seven-way mixture study; Got-Talent winsI.4.1
late Jul – mid Aug40-emotion sweep; character LoRAs; VoiceNet LoRAs; explicitness; 99-vector conditioning (negative); merging and burst dose studies; contained emotion; “text carries the condition”; 500 voice profiles (≈16,150 GPU-h); DiT campaign runs 5–17I.4, I.5, I.16
1–2 AugAutonomous LoRA + prompt search agent: 163 generations, ≈10 k takes, Gemma-4-31B winsI.7.4, I.16.2
16 AugReference report v1 (two model families under one name)[REF-V1]
mid AugJUPITER round begins: corpus v2 (3,147,802 rows); SFT round 1 (2,232 steps, 64 nodes); DPO corpus v1 (236,166 pairs); SFT round 2 — timing solved, directions lost; GRPO runs 1–3I.8–I.11
25 Aug, 09:54–10:43Two prompt bugs fixed seventeen minutes before SFT-3 launches; SFT-3 trains in 30 min 46 sI.9.4–I.9.5
25 Aug, 12:00 UTCZero-point percentile correction; every evaluation after this is on the new scaleI.6.7
25–26 AugDPO on corpus v2 (best step 3216, later deleted by rotation); CFG corpus (2,327,904 pairs) and CFG-DPO (step 4912); GRPO runs 4–5; emotion-LoRA rank ablation; 17-adapter selectivity grid; sweep AI.10–I.12
26 Aug, 08:51Bucket-LoRA trainer's first launch dies on the padding assert; then the zero-step DONEIII.2.3–III.2.4
26 AugPublication round 1: SFT3, CFG-DPO adapter, 40 emotion + 500 voice adapters; first technical report; VoiceNet axis adapters start at 16:51; report uploaded to the Space at 15:50 UTCIV.3
27 AugCluster downtime[TR-0829] §21
28 AugContrastive-families corpus (2,696,421 pairs) and dpo_sft3_p2 (step 5022, new best); trajectory corpus published; layer forensics and steering grid; low-α steering re-run supersedes the negative result; report expanded and re-verified; eight inconsistencies correctedI.13, I.15, II.3.3, II.4
29 AugThree-lever programme: CFG, dose-response, combination, Parakeet becomes primary WER; demo-server loader defect and stop-talking bug diagnosed; protocol audit ([P§34]); the three levers shipped into the demo server ([P§35]); final 29 Aug reportI.14, [P§27–§35]
30–31 AugConditioning datasets, organic transitions, tempo/rate tag (rate adapter changes nothing), crossfade v2, quality adapters null then a listener disagreed, rew_chosen baseline correction, burst-and-stop DPO corpus, listening page shipped with dead playersII.8 ([P§32–§46])
1–2 SepBurst adapter dose sweep under the real stack; sixty burst adapters trained on labels the detector cannot see; re-filtering (9.8 % survive); manufactured bursts; family-relaxed hit rate becomes the metric (owner decision 2 Sep); quality DPO ships at 1.5; the other levers stacked; CFG on the burst (g = 4 triples the hit rate)II.9 ([P§43–§53])
3–5 SepDramaBox scream; multimodal annotation of bursts (88 % vs 5 %); adapter capacity not the bottleneck; calm lead-in doubles dynamic contrast; real burst data with a blind second opinion; two instruments; relabelling with Gemini spans; 73,500 clips annotated whole; detector graded across generators; detector-x2 shipped with its recall floor; four labels unmapped until 4 Sep; artefacts published 5 SepII.10 ([P§54–§68])
5–8 SepScream gain at 45-class scale; combining per-class burst adapters (clean negative); consolidation and the eight ways an exit code lied; two predictors put to a human vote; voice-annotation-data-v2 re-checked; the Best-of-N reward that sees the burst ([P§71], 7 Sep)II.10, II.6
early SepBabble bug: sum of burst adapters, not brackets — commits 08caa4b / 15c9a8c; two adapters at 1.5 → WER 0.82 median in 5/5 seedsII.7
9 SepThe actor report ([ACTOR]); protocol §72; the steering listening test recorded as the project leader's statement; this report begun; Part I uploaded to the SpacePart II, IV.3

Part IV — What is published, what is open

The open items come first, and the human evaluation comes first among them, because it is the one measurement that qualifies every other number in this report and has never been made. Then the pre-agent open list, which the agent era did not close. Then everything that is public, with URLs. Then the one application that is ready now. Then the limitations, and a one-paragraph summary.

IV.1 Open questions of the agent era — the human evaluation first

IV.1.1 The human evaluations — and they rightly come first

Nobody has listened. Not to the training data, not to the generated audio, not to the clips behind any number in this report. Every hit rate, every recall, every sentence of the form “the model produces the sound” is the judgement of one model about the output of another model. That is the only part of the chain no human has checked so far, and it is the most important caveat of this report.

The apparatus for it is in place. laion/voice-pairs-genuineness-vocalburst-blend-human-eval contains 2,000 pairwise comparisons over 4,000 distinct clips in four subsets, published as eight WebDataset tars and with metadata verified byte-for-byte against the Hub:

The four subsets, 500 pairs each, 250 synthetic and 250 real each
SetPredictorDifficultySelection rulenative Δ median
AGenuineness (0–6)easyΔ ≥ 2.02.33
BGenuineness (0–6)hard1.0 ≤ Δ ≤ 1.51.04
CVocal-burst blend (0–10)easyΔ ≥ 2.03.38
DVocal-burst blend (0–10)hard1.0 ≤ Δ ≤ 1.51.21

Everything that would otherwise determine the answer is held constant: same language, same half (synthetic never against real — the Genuineness predictor separates the two halves with medians of 1.32 against 4.66 on a 0–6 scale, so a cross pair would measure “synthetic against real”), same speaker where an identity exists (1,976 of 2,000), for C and D the same primary burst label and both members actually carry a burst, durations within 4 s. The A/B order is random — the higher-rated one is on side A in 1,012 of 2,000 cases — and the annotator sees only the two audio files. The question asked is: which recording sounds more genuine (A and B) or which blends its vocal sound in more naturally (C and D). The predictor has already made the decision; what is measured is how often a human agrees.

A caveat that must precede the evaluation: the hard band (1.0–1.5) lies at or below the two models' own test error — Genuineness MAE 1.00 on 0–6, Blend MAE 2.057 on 0–10. Sets B and D ask humans for judgements about differences that the predictors themselves cannot resolve reliably. Low agreement there is the expected result and not automatically a broken model.

The listening page: the Space laion/voice-pairs-listening-demo exists — measured on 9 September unauthenticated, api/spaces/… returns 200, while api/models/… and api/datasets/… return 401. Exactly this asymmetry is an already documented trap: a link checker that resolves only in the model namespace reports every Space as a dead link, and once a working link was deleted instead of corrected because of it. Locally the Space is referenced in no file of this home directory; the local listening page is ~/stimmpaare_demo.html with 80 pairs, 20 from each subset, with playback, voting and subsequent reveal against the predictor.

Status: the dataset is finished and published, the voting is under way, the result is open. No artefact so far contains a single human vote.

IV.1.2 What the guided path changes

The entire Best-of-N grid was generated unguided, while the server on the API path works with guidance 3.0. The small guided control arm (n = 80, four classes, one form) flips no sign, but suggests that guidance already collects a large part of the same gain by itself. Measuring the guided path at scale needs a batched two-branch decoder — today's runs at batch 1 and is about a hundred times as expensive.

IV.1.3 The instruments disagree with each other

Two independently trained detectors, the same generated audio track: agreement at cell level Pearson 0.51 strict and 0.21 at family level over 768 cells, and one names the target class more than twice as often as the other (0.135 against 0.061 over 22,825 clips). Absolute rates in this programme are properties of a measurement chain, not of the model. What survives a change of instrument is comparative and paired — and even then the identity of the best operating point does not survive it.

IV.1.4 What else is open, briefly

What was measured. Human evaluation: 2,000 pairs, 4,000 clips, eight tars, verification via an explicit key — 0 Δ violations, 0 threshold violations, 0 wrongly set answer keys, 0 burst-free C/D members, 0 cross-language pairs, 0 reused clips, each over all 2,000. Publication check over all 4,000 clips: 0 violations of the allow list, 0 podcast uids, no text fields. Result: none. Nobody has voted yet.

IV.2 What was open on 29 August and is still open

The 29 August report closed with a ranked list of next steps ([TR-0829] §21.2) and two unmeasured questions ([TR-0829] §24.10). This section records, for each, whether the agent era (30 August – 9 September) closed it. Most it did not; the September work went into the burst programme and the agent (II.6–II.10).

The 29 August list, revisited on 9 September.
Item, in the order of expected value given on 29 AugustState on 9 September
Finish the VoiceNet pilot and evaluate it (2 of 17 adapters still training; burst adapter unmeasured)Closed in part: the VoiceNet set was completed and published ([P§24–§25]); the delivery-axis adapters ship in the agent's stack; the burst adapter line became the whole burst programme of II.9–II.10.
Finish the contrastive-families epoch (run stopped at step 5,126 of 10,331)Open. The second half was never run. The same is true of the CFG and corpus-v2 runs.
A human baseline and a blind comparison — “the single largest gap in the whole project”Open, and now instrumented: 2,000 stimulus pairs and a listening Space exist, no votes yet (IV.1.1). The 29 August design constraints stand: pairwise, not absolute (an absolute scale was measured to saturate; a judge given “score 0–10” rated recency — top-ranked take chosen in 3 of 108 rounds against a 20 % chance rate, favourite the 4th or 5th presented in 71 of 81 rounds; naming concrete criteria flipped the position correlation from +0.33 to −0.12); and a real-audio ceiling control (in one adapter study the best cell scored 1.969, real human broadcast commentary 1.775, and the un-adapted base 1.750 — the metric's ceiling was below real audio). One human vote on two predictors was run ([P§69], II.10) — on the ears, not on the instrument.
Best-of-N supervised training on the RL harvest (204 clips at median emotion percentile 0.641, WER 0.049)Open. Nothing has been trained on the harvest. The agent's Best-of-N is the inference-time form of the same idea.
Band-balanced supervised data (the subset is 59.5 % extreme, 6.8 % faint)Open. Not attempted.
Train something on a trajectory (7.6 M specifications published)Open. Nothing rendered from a trajectory; the organic-transition and crossfade work of II.8 ([P§36], [P§40]) is the closest relative and it is inference-time.
Run probing task A2Open.
The three DPO follow-ups (absolute log-probability before/after a preference run; words present with only the level flipped; asymmetric 70/30 mixture)Open. Still not instrumented after a fourth preference corpus. The burst-and-stop DPO corpus of [P§44] is a fifth corpus, built for a different target.
Re-run the merge-weight sweep over all 40 emotion adapters, past 2.0, and on German-only promptsOpen for the emotion adapters. The burst adapters got their own dose sweep under the real stack ([P§43]).
Fix the Jealousy_&_Envy key mismatch and re-score offlineNot recorded as done in any source read for this report; treated as open.
Complete the per-emotion adapter measurement (17 of 40 evaluated, by accident)Superseded rather than closed: the combination study of II.3.5 measured the emotion adapters' contribution inside the stack (n = 399), not per adapter.
A reward ablation rather than a longer RL runOpen. No GRPO run since run 5.
The voice-acting agent as an agentic systemClosed — this is Part II. Its specification was met point by point: identity ranked with a floor, seeds per candidate, whole-take ranking, a supervisor prompt (the Mind) that names the failure modes.
Blending two emotions (never measured)Open. The two-emotion prompts of the GRPO prompt mix (35 %) and the corpus captions remain the only in-distribution evidence.
Spectral content — a band-energy ratio to catch a quiet low-passOpen; [P§33] asked the adjacent question (does stacking dull the sound?) and answered: no — it destroys the words instead (II.8).
Total injected steering magnitude held constant across layer countsOpen; moot for production since steering was rolled back (II.3.3).
Rotate the two cleartext Hugging Face tokens; inode quota on the primary data filesystemOperational, outside this report; the quota is still the reason everything new goes to scratch.

Two further open items arrived with the agent: the directing model was chosen (gpt-5.6-luna remotely, Gemma-4-12B locally) without re-running the 31B-vs-12B comparison of I.16.2 on the new instrument; and the BON_N = 8 / BON_GUIDANCE = 3.0 operating point was set from the guided arm's n = 80 (II.6.4), which IV.1.2 names as the largest open risk.

IV.3 Everything published, with URLs

Every URL in the 29 August tables was checked on 26 August 2026 ([TR-0829] §22); the agent-era rows were checked on 9 September (this report's build). Repositories marked private returned HTTP 401 to an anonymous request and are listed because they are the authoritative location of an artefact this report cites.

IV.3.1 The model line

StageRepository
Base checkpoint (OpenMOSS)huggingface.co/OpenMOSS-Team/MOSS-TTS-Local-Transformer-v1.5
Audio tokenizer, 12 codebooks @ 12.5 fps → 48 kHzhuggingface.co/OpenMOSS-Team/MOSS-Audio-Tokenizer-v2
v1 — DramaBox full FT (internal)huggingface.co/TTS-AGI/moss-dramabox-ft — private
v1 — public merged release (base FT + rank-256 LoRA)huggingface.co/laion/moss-tts-local-transformer-4.55b-voice-acting
v2 — the six-dataset production modelhuggingface.co/laion/moss-tts-local-transformer-4.55b-voice-acting-v2
v2 rolling checkpointshuggingface.co/TTS-AGI/moss-voiceacting-ft-checkpoints — private
Round-1 SFT (full fine-tune)…-v2-sft
Round-1 SFT + DPO, full parameter…-v2-sft-dpo
Round-1 DPO LoRA (ships unmerged — III.2.2)…-v2-dpo-lora
Round-3 SFT — SFT3, the Instrument — published 26 Aughuggingface.co/laion/moss-tts-local-transformer-4.55b-voice-acting-v2-sft3
CFG-DPO adapter, step 4912 — published 26 Aughuggingface.co/laion/moss-va-sft3-dpo-lora
Contrastive-families DPO adapter, step 5022 — the DPO of the shipped stack, published 28 Aughuggingface.co/laion/moss-va-sft3-dpo-lora-p2
8 B delay-architecture sibling (24 kHz)huggingface.co/laion/moss-tts-v1.5-8b-voice-acting

IV.3.2 Adapters

AdaptersRepository
40 emotion LoRAs on SFT-3, rank 16 — published 26 Aughuggingface.co/laion/moss-va-sft3-emotion-loras
500 voice LoRAs on SFT-3, rank 16 — published 26 Aughuggingface.co/laion/moss-va-sft3-voice-loras
VoiceNet tail (delivery-axis) adapters on SFT-3 — published in round 2 ([P§24–§25])the sft3_voicenet:* set the agent ships; repository named in [P§25]
Quality adapters (genuineness_high, blend_high, esthetics_high), quality-DPO, burst adapters on SFT-3published in [P§25], [P§27], [P§68]; ids in II.10.12
40 emotion LoRAs on v2, rank 32 (v3 line)huggingface.co/TTS-AGI/moss-emotion-loras-v3
Emotion LoRAs v1 · v2 · the 40-emotion sweep · the filtered bucketsTTS-AGI/moss-emotion-loras · -v2 · -40 · moss-emotion-filtered-buckets — all private
Voice-profile LoRAs — 10 pilot voices, rank ablationhuggingface.co/TTS-AGI/moss-voice-profile-loras
Voice-profile LoRAs — 500 production voices, rank 4huggingface.co/laion/moss-voice-profile-loras-500
114 VoiceNet dimension LoRAs on v2 (57 × high/low)huggingface.co/laion/moss-voicenet-dimension-loras
64 vocal-burst LoRAs on v2huggingface.co/laion/vocal-burst-lora-adapters
Voice-acting LoRA (rank 256, 100 k + 1 M) · inline-burst training dataTTS-AGI/moss-local-transformer-voice-acting · TTS-AGI/moss-inline-vocal-bursts — both private
Character LoRAs — refined / genuine, public mirrors…-refined-public · …-genuine-public
Character LoRAs — source repositoriesTTS-AGI/moss-character-loras-genuine · -refined · moss-12-cluster-character-loras — all private
Explicitness LoRAs (public + age-gated)huggingface.co/TTS-AGI/moss-explicitness-loras
German broadcast · sports commentarylaion/moss-mediathek-emotion-lora · laion/moss-sports-commentator-lora

IV.3.3 Measurement instruments — the Ears

InstrumentRepository
VoiceCLAP-commercial — the shared 768-d audio↔text embedderhuggingface.co/laion/voiceclap-commercial
VoiceCLAP small · small-v2 · large-v2 (3,584-d)voiceclap-small · voiceclap-small-v2 · voiceclap-large-v2
Genuineness predictor (0–6)huggingface.co/laion/voiceclap-commercial-genuineness · large sibling voiceclap-large-v2-genuineness
Vocal-burst blend predictor (0–10)huggingface.co/laion/voiceclap-commercial-vocalburst-blend
Earlier-generation genuineness / blend heads (slots 97 & 98 of the 99-vector scorer)laion/genuineness · laion/vocal-burst-blend — both private
Empathic-Insight-Voice — 40 emotion + 15 attribute experts (the agent listens to the *user* with this)huggingface.co/laion/Empathic-Insight-Voice-Small
Empathic-Insight-Voice-Plus — the same plus 4 audio-quality expertshuggingface.co/laion/Empathic-Insight-Voice-Plus
BUD-E-Whisper — the frozen encoder both suites readhuggingface.co/laion/BUD-E-Whisper
VoiceNet 57-dimension regressor + classifier (114 heads)huggingface.co/laion/voicenet-dimension-predictors-commercial
…its Gemini-3.5-flash label sethuggingface.co/datasets/laion/emolia-voicenet-gemini-annotations
Vocal-burst locator · detector v2 (83 classes)laion/vocalburst-locator · laion/vocal-burst-detector-v2
Vocal-burst detector x2 (17 classes, chance 0.0588) — the agent's burst ear, shipped 5 Sep ([P§63])huggingface.co/laion/vocal-burst-detector-x2
Speaker embedding used for identityhuggingface.co/speechbrain/spkrec-ecapa-voxceleb
Word error rate — the agent's ASR earhuggingface.co/nvidia/parakeet-tdt-0.6b-v3 (Whisper-large-v3-turbo for I.8–I.13)
Sparse autoencoder over 155 M CLAP embeddings (32× expansion)huggingface.co/TTS-AGI/audio-audio-clap-maestrino-sae-32x-k5

IV.3.4 Listening pages and the human evaluation

All audio on these pages is model-generated; none of it is source recording.
PageURL
Human evaluation — 2,000 stimulus pairs / 4,000 clips, sets A–D (genuineness, vocal-burst blend); no votes yethuggingface.co/datasets/laion/voice-pairs-genuineness-vocalburst-blend-human-eval
The listening Space for those pairs (exists on the Hub; local page ~/stimmpaare_demo.html)huggingface.co/spaces/laion/voice-pairs-listening-demo
Round-3 evaluation grid — nine models × 80 prompts × 4 completions, with ASR transcriptshuggingface.co/spaces/laion/moss-va-sft3-samples
Emotion LoRAs vs baseline — 17 adapters, matched and neutral promptshuggingface.co/spaces/laion/moss-va-emotion-loras
Merge-weight ablation — 31 emotion adapters over six weights (I.12.4)huggingface.co/spaces/laion/moss-va-lora-scale
Round-2 baseline grid (timed script, no directions)huggingface.co/spaces/laion/moss-va-sft2-samples
Round-1 grid (base / SFT / SFT+DPO / ground truth)huggingface.co/spaces/laion/moss-va-sft-samples
This report, as a hosted page (the 26 August version is archived beside it at archive/technical_report_2026-08-26.html)huggingface.co/spaces/laion/moss-va-technical-report
The demo server (the actor's home)github.com/LAION-AI/Humaneness-Voice-Demo-Server

IV.3.5 Datasets

DatasetURL
Annotated voice-profile corpus — 28,212,933 rows, 5.8 TBhuggingface.co/datasets/laion/laion-voice-profiles-annotated
Voice-profile SFT — 1,200,531 rowshuggingface.co/datasets/laion/laion-voice-profiles-sft
Real-speech SFT, EN/DE — 1,947,272 rowshuggingface.co/datasets/laion/tts-realspeech-sft-en-de
Voice-profile DPO — 3,451,531 pairshuggingface.co/datasets/laion/laion-voice-profiles-dpo
Real-speech DPO, EN/DE — 3,959,192 pairshuggingface.co/datasets/laion/tts-realspeech-dpo-en-de
Trajectory corpus, public subset — 7,638,961 specifications, no audio (I.15)huggingface.co/datasets/laion/moss-va-trajectory-corpus
Trajectory corpus, full — 10,653,713 rowshuggingface.co/datasets/TTS-AGI/moss-va-trajectory-corpus-full — private
Voice-profile reference clipshuggingface.co/datasets/TTS-AGI/moss-voice-profile-references
Game-character corpus, 750,242 rows with full VoiceNet + emotion annotationhuggingface.co/datasets/scientifi-papers/chavo-annotated
Designed character takes, ≈38,400huggingface.co/datasets/laion/moss-character-voices-top3-captioned
Gemini-generated named-character voices (30 × 1,000) · additional voice-acting clips (31,306)TTS-AGI/gemini-adult-voices · TTS-AGI/additional-data-a1-plus — both private
Vocal-burst training data with verified negatives, 3,598 real utterances over 49 classes, and the 73,500-clip annotation ([P§58], [P§61], [P§62])ids in II.10.12 / [P§68]
voice-annotation-data-v2 (re-checked and extended, [P§70])id in II.10.14

IV.3.6 Documentation and code

ItemURL
The voice-acting manual — 25 chapters, 40 emotion pages, 12 recipe chaptersprojects.laion.ai/moss-voiceacting-manual/site/index.html (mirror: laion-ai.github.io/moss-voiceacting-manual)
The manual's emotions chapter…/site/emotions/index.html
Model home — inference, prompting guide, demos (default branch master)github.com/LAION-AI/laion-moss-local-1.5-voice-acting-4.55b
The 29 August report and its two companion pagestechnical-report.html · steering-and-probes.html · trajectory-corpus.html
Experiment overview · the 13-voice showcase · the reinterpretation corpusoverview.html · voices.html · reinterpretations_25.html
40-emotion evolution · prompt × LoRA study · edge-case evolutionmoss-emotion-evolution-all · moss-emotion-prompt-lora-study · moss-emotion-edgecase-evolution
VoiceNet-LoRA evolution · explicitness dose gridmoss-voicenet-lora-evolution · moss-explicitness-grid
Autonomous LoRA + prompt search agentgithub.com/LAION-AI/voice-acting-search-agent
Expressive best-of-N methodgithub.com/LAION-AI/chatterbox-voice-conversion/tree/main/expressive_bestofn
Manual source · VoiceNet taxonomyLAION-AI/moss-voiceacting-manual · LAION-AI/voicenet
Canonical pipeline documentation (docs/01–23 + LEARNINGS)github.com/LAION-AI/Voice-Acting-Pipeline-WIP — returns 404
The WikiSkills (VOCAL_BURSTS.md and siblings) — the operator's manual of the agent era~/wikiskills/; the burst recipe table is parsed from it at render time (Appendix A.8)
This report's renderer$SC/code/vareport/en_build.py and en_part*.py

IV.3.7 What the published cards say, and two defects in them

The 26 August artefacts went public as four repositories under the laion organisation, and three more followed on 28 August; the September artefacts are listed in II.10.12. All were uploaded from the login node (compute nodes have no route to the internet) and byte-verified after upload — the uploader re-reads the repository metadata with HfApi.repo_info(files_metadata=True), size-compares every file, and for the strictest of the scripts downloads and SHA-256-hashes anything the Hub reports no size for. The reason is in that script's docstring: “upload_folder has returned a success URL for a rejected upload in this project before, so nothing is trusted.” This report is uploaded the same way.

The 26 and 28 August repositories ([TR-0829] §14).
RepositoryContentsKey numbers on the card
laion/moss-tts-local-transformer-4.55b-voice-acting-v2-sft3The full SFT round-3 model — weights, remote code, processor, provenance. CC-BY-4.0, EN/DE.Trained on the 398,282-row extreme subset, 2 epochs, 712 steps on 32 nodes. WER on direction-carrying prompts 0.447 → 0.099; median duration error 0.100 → 0.080 s; clips within 0.5 s 92.8 → 100 %; burst hit rate 0.516 → 0.666.
laion/moss-va-sft3-dpo-loraThe CFG-DPO adapter — rank 64, α 128, dropout 0.05, 23 target modules including all twelve audio_lm_heads. ckpt-step4912.2,327,904 preference pairs, 20.5 % CFG pairs (20.38 % on recount); preference accuracy on those 0.562 → 0.975. Reward 0.4708, WER 0.0950, quality 0.9235, burst 0.4271, hit rate 0.772.
laion/moss-va-sft3-emotion-loras40 per-emotion adapters, one subdirectory each, with a bucket.json recording rows, steps, minutes, non-finite count and prompt hash.Rank 16, α 32, 5 epochs, cosine 1e-4 → 5e-6, 34.4 M trainable (0.83 %). All 40 with zero non-finite batches. Carries the selectivity result and the merge-weight sweep with its recommended operating point of 1.5.
laion/moss-va-sft3-voice-loras500 per-voice adapters — emolia_*, refvoice_*, mediathek_* — each trained on every clip of that voice minus the worst decile by mean of blend and genuineness.Rank 16, α 32, 5 epochs. Rows per voice: median 2,163, min 1,971, max 2,251. Steps per adapter: median 2,705. 500 of 500 with zero non-finite batches.
laion/moss-va-sft3-dpo-lora-p2 — 28 AugThe preference adapter trained on the contrastive-families corpus. Rank 64, α 128. ckpt-step5022. Supersedes moss-va-sft3-dpo-lora, which stays up as the CFG-corpus result.2,696,421 preference pairs, 13.7 % the new families. Reward 0.4757, WER 0.0977, emotion percentile 0.3541, quality 0.9208, burst 0.4180, hit rate 0.762.
datasets/laion/moss-va-trajectory-corpus — 28 AugThe public subset of the trajectory corpus. Specifications only; no audio. CC-BY-4.0.7,638,961 trajectories over 17,722,101 distinct clips and 1,214,241 speakers; 74,858 h of referenced audio; 419 MB. Three source datasets held back.

Three card statements are the honest limits of what was shipped:

Two defects in the published cards, found on 28 August and not yet fixed as far as the sources read for this report record. (1) The “loading adapters” snippet cannot work with the adapters it ships beside: all four cards tell the reader to call model.add_weighted_adapter(["dpo","voice","emo"], [1.0,1.0,1.5], "combined", combination_type="linear"); the DPO adapter is rank 64 and the others rank 16, and PEFT refuses: “All adapters must have the same r value when using combination_type linear, ties, dare_ties or dare_linear.” The project's working route — "route": "multi-active adapters, per-adapter scaling" — is what the demo server does. (2) The DPO card carries no merge_and_unload() warning; grepping all four READMEs for merge_and_unload, weight-tied or audio_embeddings returns nothing. The warning and its measurement (III.2.2) live in the reports and the training code, not on the Hub.

IV.4 The application that is ready now: dubbing under a duration budget

The capability that is solved — timing — has an obvious and immediate use. Dubbing has a hard constraint that ordinary TTS does not: the translated line must fit the time the original line occupied. A German utterance of 4.7 seconds has to become an English utterance of 4.7 seconds, including its pauses, or the picture and the sound come apart. Conventional pipelines fight this with post-hoc time-stretching, which audibly damages the voice, or by rewriting the translation until it happens to fit. The timed-script format solves it at the source. Source: [TR-0829] §20; [TR-0826] §15.

  1. Measure the source. Force-align the German utterance to get per-word timings. That yields exactly the quantities the script format prints: per-sentence spoken durations, every silence of 0.2 s or more, and every vocal burst with its length and class (from the locator and detector).
  2. Re-render the script in the target language. Same tags, same numbers, translated words. The frame budget is floor(d × 12.5) of the original duration.
  3. Clone the voice. Attach the reference recording, or a voice-profile adapter if one has been trained for that speaker.
  4. Generate best-of-N. 32 candidates cost the same wall time as 16 on this architecture (I.4.12), so the group is nearly free.
  5. Rank and pick. Three independent criteria are available, and they answer different questions (table).
Ranking criteria for a dub ([TR-0829] §20).
CriterionInstrumentAnswers
Timing fidelityProduced frame count vs. the script's own sumDoes it fit the picture? Median error 0.08 s and essentially every clip inside 0.5 s, so this is usually a tie-break rather than a filter.
Emotional matchThe 40 Empathic-Insight heads, compared clip to clip on the same headDoes the English take carry the same feeling as the German original? Safe on the corrected percentile scale precisely because it is within-head (I.6.7); cross-emotion comparison is not.
Overall closeness in the audio–text embedding spaceVoiceCLAP — cosine similarity between the 768-d embedding of the source clip and of each candidateDoes it sound like the same performance — timbre, energy, style, affect together? Because VoiceCLAP is a dual-tower audio↔text model, the same encoder also lets a director's sentence be used as the target instead of a reference clip — which is exactly the CLAP term of the agent's Best-of-N (II.6).

Add speaker similarity (ECAPA, speechbrain/spkrec-ecapa-voxceleb) as a fourth term, and apply I.5's lesson: rank on it, do not gate on it — and if you must exclude, exclude below 0.40 rather than penalising. A note on which VoiceCLAP: the audio–text contrastive models live in the laion organisation — voiceclap-commercial (768-d, CC-BY-4.0, the one the scoring stack uses), voiceclap-small-v2, and voiceclap-large-v2 (3,584-d). The related artefact in the TTS-AGI organisation, audio-audio-clap-maestrino-sae-32x-k5, is a different kind of object: a sparse autoencoder trained on 155 M CLAP embeddings drawn from six audio corpora — 512-d in, 16,384-d hidden (32× expansion), top-k 5, with 14,127 of 16,384 features alive. For a dubbing ranker you want the embedder; the SAE is what you would reach for to ask which few features a pair of clips differ on.

Two limits belong with this recommendation. First, the emotional-intensity ceiling means that if the source performance is at the top of the human range, the dub will land below it — the ranker will pick the closest of what the model produced, and the model does not produce extremes on command. Second, cross-lingual reference conditioning does not clone (mean ECAPA 0.188 against a 0.105 floor, I.4.3), so a German→English dub that must preserve identity needs a trained voice adapter for that speaker, not just a reference clip. The 500 voice adapters give per-speaker identity and the 40 emotion adapters at weight ≈1.5 give the intensity lift, stacked rather than merged.

IV.5 Limitations of everything above

IV.6 The one-paragraph summary, for anyone who read only this section

A general-purpose 4.55-billion-parameter TTS model was turned into a voice-acting model by two full fine-tunes — one on distilled director-instruction data with a rank-256 adapter merged in, one on six datasets of prepared and self-generated material, with the released checkpoint chosen by scoring generated audio rather than by validation loss — and then surrounded by roughly a thousand LoRA adapters for emotions, characters, voices, timbre dimensions and vocal bursts. A timed-script prompt format then stated, per sentence, how long it may take, where the silences are and how long each vocal burst lasts. That worked: median duration error 0.08 s, essentially every clip within 0.5 s, and adding the format also fixed a word-error collapse (0.447 → 0.099) and improved burst realisation. The same round tried to make the model perform an emotion at a requested intensity, and four objectives — supervised fine-tuning on an extreme subset, preference tuning on contrast pairs, preference tuning on symmetric instruction-conditioned pairs, and group-relative reinforcement learning — all failed to move it; the RL failure is structural (a constant term contributes nothing to a group-normalised advantage). What moved the number before the agent was data, dose and a better contrast: an adapter trained on the most intense and most genuine 8 % of the corpus (0.3572 against 0.3494); overdriving the per-emotion adapters to 1.5 (0.408 → 0.471 across 31 adapters); and a preference corpus whose contrasts pit intense against intense (0.3541, and the first preference tuning to raise intensity while lowering word error rate). All three remain far from the 0.90–0.98 asked for. Looking inside the model found emotion roughly half as linearly legible as voice quality, losing 39 % of that legibility before the acoustic decoder, and the internal “more emotional” direction running close to opposite to “more genuine” (cosine −0.746 at h20, −0.954 at the frame slot). Steering along those directions destroyed the output at α ≥ 0.5, worked on the meters at α = 0.10 — and was rolled back after the project leader listened. The system that ships in September is an actor: a directing language model writes the prompt in one constrained pass, plays SFT3 through a stack of adapters summing to 5.75–7.25 with up to 3 burst adapters under a budget of 2.00, and — when Best-of-N is switched on — generates 8 takes with guidance 3.0 and keeps the one its ears rank highest; that selection lifts the human-proxy target by +0.0348 (t 5.47, n 1,940) and would lift it +0.0454 with the burst term the ears now have. The babble bug that named the configuration was a sum of burst adapters, not a bracket. Every one of these statements is the output of a learned scorer; the 2,000 pairs that would let a human check them are published and unvoted.

Appendix — Full tables, computed at render time

Every table in this appendix except A.1 is produced from a file on disk when the report is rendered; the file is named in each caption. Column names are the keys of the source file, untranslated, so that a reader can find the number again. A.1 is transcribed from [TR-0829] §13.3 because it exists nowhere else as a file.

A.1 The 18-run ranking of 29 August (corrected percentile scale, 80 prompts × 4 completions, plateau yardstick)

Transcribed from [TR-0829] §13.3; n = 320 clips per row; the evaluation cannot resolve ±0.01 in reward.
#ModelRewardWEREmotion pctEmotion termQualityBurstBurst hit rate
1SFT-3 + DPO contrastive families, step 5022 (published)0.47570.09770.35410.92080.41800.7625
2SFT-3 + DPO contrastive families, step 39060.47440.09160.34780.92310.41060.7625
3SFT-3 + CFG-DPO, step 4912 (published)0.47080.09500.33730.12690.92350.42710.7719
4SFT-3 + CFG-DPO, step 39630.46980.09530.34200.13970.92350.39400.7000
5SFT-3 + DPO corpus v2, step 3216 (weights deleted, III.2.7)0.46870.10940.34010.13700.92110.39290.7094
6SFT-3 + DPO corpus v1, step 8640.46680.11170.35180.14570.91080.39730.6937
7SFT-3 + emotion LoRA r160.46630.12390.35720.14720.91890.40020.6844
8SFT-3 + emotion LoRA r640.46400.12980.33450.13930.92180.38670.6844
9SFT-3 + DPO corpus v2, step 45900.46330.10240.34270.13850.90620.39300.7188
10SFT-3 + emotion LoRA r16 @ 0.750.46220.10550.33900.14120.90520.38950.6969
11SFT-3, no adapter0.45840.09870.34940.13940.91270.35640.6656
12SFT-3 + emotion LoRA r16 @ 1.250.45840.11300.33150.13920.90500.39290.6844
13SFT-3 + emotion LoRA r16 @ 0.250.45790.11740.35330.13590.91360.38270.6656
14SFT-3 + emotion LoRA r16 @ 0.50.45540.13450.33650.13240.91500.37120.6656
15SFT-3 + GRPO v40.45510.12860.34680.14860.89500.38370.6937
16SFT-3 + GRPO v50.45120.12530.32750.12020.92820.35860.6469
17SFT-3 + emotion LoRA r16 @ 1.50.45090.13250.33580.12980.92600.36110.6312
18SFT-3 + emotion LoRA r320.44940.13150.32090.12920.91260.36400.6719

A.2 The shipped configuration, every constant

Read from $SC/hv_mirror/config.py (read-only).

Numeric constants parsed from float(os.environ.get(...)) / int(...).
ConstantDefault
AESTH_LORA_LAM0.0000
ALIGN_BURST_SLACK2.0000
ALIGN_EVERY_S0.5000
ALIGN_LEAD_MIN_S0.1200
ALIGN_LEAD_RAMP_S0.1000
ALIGN_LEAD_SCAN_S1.3000
ALIGN_LEAD_WORDS3
ALIGN_LOOKAHEAD_S0.5000
ALIGN_MIN_SCORE0.3500
ALIGN_QWEN_MAX_WORD_S3.0000
ALIGN_TAIL_AFTER0.6000
ALIGN_TAIL_MIN_S0.2500
ALIGN_TAIL_PAD_S0.1200
ALIGN_TAIL_RAMP_S0.1500
APP_PORT8,792
BON_BATCH8
BON_BATCH_CFG4
BON_CLAP_WEIGHT2.0000
BON_GUIDANCE3.0000
BON_N8
BON_WER_KNEE0.0000
BREATHE_MAX2
BREATHE_WANT2
BURST_LAM0.2500
BURST_LAM_BUDGET2.0000
BURST_LAM_INTENSE0.5000
BURST_LAM_MAX1.2500
BURST_MAX_ADAPTERS3
CFG_COST_FACTOR1.9300
CFG_G_MAX3.0000
CFG_G_MIN1.5000
CTX_FRAMES160
HISTORY_TURNS_LOCAL8
HISTORY_TURNS_LUNA40
HOSTED_MAX_TOKENS3,000
MAX_CPU_ADAPTERS64
MAX_SESSIONS200
MIN_FRAME_FRACTION0.5500
PROFILE_LORA_LAM1.0000
PURE_PROFILE_LAM0.5000
RETRIEVAL_EMO_BONUS0.5000
RETRIEVAL_LEVEL_PENALTY0.0500
SFT3_DPO_LAM1.0000
SFT3_EMOTION_LAM1.0000
SFT3_VN_MAX1
SFT3_VOICE_LAM1.0000
SIDON_OUT_SR48,000
SIDON_TIMEOUT120.0000
SKILLS_MIN_HIT0.1500
SPEAKER_LORA_LAM1.0000
SPEAK_BEST_OF10
SPEAK_GUIDANCE3.0000
SPLIT_MIN_WORDS100,000
STEER_ALPHA0.1000
STEER_ALPHA_CEILING0.1500
STEER_PACK_K5
STEER_REALISED_CEILING0.2500
STOP_BIAS2.0000
TAIL_FRAMES50
TAIL_TURNS2
TIMED_FRAMES_PER_WORD4.0000
TOKENS_PER_WORD6.0000
TOKEN_HEADROOM2.5000
VC_CONTEXT_S1.5000
VC_CROSSFADE_S0.0400
String constants.
ConstantDefault
AESTH_LORA``
AGENT_PICKS_MODE1
ALIGN_BACKENDqwen,mms
ALIGN_DEVICEcuda
ALIGN_DIR/mnt/nvme/moss-15-v2-assets/aligner
ALIGN_ON1
ALIGN_QWEN_DEVICEcuda:1
ALIGN_QWEN_REPOQwen/Qwen3-ForcedAligner-0.6B-hf
ASR_DEVICEcuda:1
ASR_MODELnvidia/parakeet-tdt-0.6b-v3
ASSETS/mnt/nvme/moss-15-v2-assets/loras
BON_DEVICEcuda:0
BON_ON0
BREATHE_ON1
BURST_SETrecipe
CFG_ENABLED1
CFG_STEER_BRANCHboth
CODEC_REPOOpenMOSS-Team/MOSS-Audio-Tokenizer-v2
CODE_CACHE/mnt/nvme/moss-15-v2-assets/code_cache
DEFAULT_BRAINluna
DEFAULT_PROFILEemolia_c1699
DELIVERY_LEVERadapter
EIV_DIR/mnt/nvme/empathic-insights-voice-small
EMOTION_NUANCE_ON1
ENGLISH_CUES1
GEN_MODEadapter
HOSTED_REASONINGnone
LLM_BASEhttp://127.0.0.1:8790
LLM_GPU0
LLM_MODELgemma-4-12b-it-qat
LUNA_BASEhttps://api.hyprlab.io
LUNA_KEY_FILE/home/c4r33u19/moss15v2/.hyprlab_key
LUNA_MODELgpt-5.6-luna
MEANVC2_ROOT/mnt/nvme/moss-15-v2-assets/MeanVC2
NUMBNESS_SUBTRACTIONwith_steer
PROFILE_LORAS/mnt/nvme/moss-15-v2-assets/loras/profiles
PROFILE_REFS/mnt/nvme/moss-15-v2-assets/refs2
REF3_DIR/mnt/nvme/moss-15-v2-assets/refs3
REF_VARIANTvc_sidon
RETRIEVAL_DIR/mnt/nvme/moss-15-v2-assets/retrieval
RETRIEVAL_ON1
SCORE_DEVICEcuda:1
SFT3_DPO_LORAsft3_dpo:p2
SIDON_BASEhttp://127.0.0.1:8793
SIDON_CKPTS/mnt/nvme/moss-15-v2-assets/sidon-ckpts
SIDON_DEVICEcuda:0
SIDON_GPU0
SIDON_ON0
SIDON_SRC/mnt/nvme/moss-15-v2-assets/sidon/src
SKILLS_ON1
SPEAKER_LORAspeaker:velvet-sage-baritone
STEER_ENABLED1
STEER_PACK/mnt/nvme/moss-15-v2-assets/steering/p3_vectors_server.npz
STEER_TAP_RANK/mnt/nvme/moss-15-v2-assets/steering/tap_rank.json
TIMED_SCRIPT1
TTS_GPU1
USE_ANCHOR1
USE_LORA1
USE_TAIL_CONTEXT1
VC_DEVICEcuda:1
VN_BASELINE/mnt/nvme/moss-15-v2-assets/vn_baseline.json
VN_DIR/mnt/nvme/moss-15-v2-assets/voicenet-pred
VOICECLAP_REPOlaion/voiceclap-commercial
WAV_CACHE/mnt/nvme/moss-15-v2-assets/ref_wav_cache
WHISPER_DIR/mnt/nvme/moss-15-v2-assets/bude-whisper
The permanent stack (sum 5.75; with one delivery axis up to 7.25).
AdapterKeyλ
Praeferenzadapter (DPO)sft3_dpo:p21.00
Stimmadapter (Stimmprofil)sft3_voice1.00
Qualitaet: Genuinenesssft3_quality:genuineness_high0.25
Qualitaet: Burst-Blendsft3_quality:blend_high0.50
Qualitaet: Aesthetiksft3_quality:esthetics_high0.50
Zweiter Praeferenzadapter (Quality-DPO)sft3_qdpo:quality_dpo1.50
Emotionsadaptersft3_emotion1.00

SFT3_VN_LEVELS = 0.50, 0.75, 1.00, 1.25, 1.50; SFT3_VN_MAX = 1; burst adapters ≤ 3, each ≤ 1.25, budget 2.00; MOSS_BON default 0.

A.3 Best-of-N optimisation ($SC/out/vb_opt/*.json)

Directory $SC/out/vb_opt; 1,940 cells, 15,517 candidates; arms: bestmem, cap100, cap150, ship, ship_vn, shipset, v2, w0; classes: affirmative_grunt, breathy_giggle, chuckle, deep_breath, exasperated_sigh, exhausted_groan, frustrated_groan, heavy_breathing, humming, panting, relief_sigh, scream, sharp_inhale, soft_hum, wistful_sigh, yawn; kinds: inline, solo; recall floor {'affirmative_grunt': {'strict_real': 0.7, 'family_real': 0.7667, 'strict_dramabox': 0.68}, 'breathy_giggle': {'strict_real': 0.4688, 'family_real': 0.9688, 'strict_dramabox': 0.8571}, 'chuckle': {'strict_real': 0.6415, 'family_real': 0.8491, 'strict_dramabox': 0.8}, 'deep_breath': {'strict_real': 0.1864, 'family_real': 0.7119, 'strict_dramabox': 0.6863}, 'exasperated_sigh': {'strict_real': 0.2264, 'family_real': 0.3396, 'strict_dramabox': 0.381}, 'exhausted_groan': {'strict_real': 0.4839, 'family_real': 0.871, 'strict_dramabox': 0.32}, 'frustrated_groan': {'strict_real': 0.5938, 'family_real': 0.8125, 'strict_dramabox': 0.6538}, 'heavy_breathing': {'strict_real': 0.12, 'family_real': 0.64, 'strict_dramabox': 0.28}, 'humming': {'strict_real': 0.4722, 'family_real': 0.6111, 'strict_dramabox': 0.5161}, 'panting': {'strict_real': 0.6471, 'family_real': 0.8824, 'strict_dramabox': 0.4615}, 'relief_sigh': {'strict_real': 0.0625, 'family_real': 0.3125, 'strict_dramabox': 0.4688}, 'scream': {'strict_real': 0.5946, 'family_real': 0.5946, 'strict_dramabox': 0.963}, 'sharp_inhale': {'strict_real': 0.8974, 'family_real': 0.8974, 'strict_dramabox': 0.5312}, 'soft_hum': {'strict_real': 0.3929, 'family_real': 0.5714, 'strict_dramabox': 0.28}, 'wistful_sigh': {'strict_real': 0.1429, 'family_real': 0.3929, 'strict_dramabox': 0.4}, 'yawn': {'strict_real': 0.4595, 'family_real': 0.4595, 'strict_dramabox': 0.5}}.

A.3 strict hit — the shipped reward R_old (large-v2 encoder, local window).
keymeantnci95n_upn_downrn_setsdf
delivered0.153118.72181,940[0.1371, 0.1691]2970
mean0.118325.67131,940[0.1093, 0.1273]7350
lift0.03485.47481,940[0.0223, 0.0473]287438
auc0.597658.2627725[0.5775, 0.6177]7030
auc_minus_half0.09769.5125725[0.0775, 0.1177]452242
r_within15,517[0.0591, 0.0925]0.07581,94013,576
wer0.079027.14461,940[0.0733, 0.0847]8450
genuineness2.734795.42261,940[2.6785, 2.7908]1,9400
clap0.144177.47341,940[0.1405, 0.1478]1,84595
A.3 strict hit — the candidate rewards R_new by λ.
keydelivered.meandelivered.tdelivered.ndelivered.ci95delivered.n_updelivered.n_downlift.meanlift.tlift.nlift.ci95lift.n_uplift.n_downvs_old.dvs_old.tvs_old.nvs_old.n_up
lam0_soft0.153118.72181,940[0.1371, 0.1691]29700.03485.47481,940[0.0223, 0.0473]2874380.00000.00001,9400
lam1_soft0.183020.83961,940[0.1658, 0.2002]35500.06479.46661,940[0.0513, 0.0781]3453800.02995.97231,94077
lam2_soft0.194321.62621,940[0.1767, 0.2119]37700.076011.03521,940[0.0625, 0.0895]3673580.04126.55871,940116
lam3_soft0.208222.58311,940[0.1902, 0.2263]40400.089912.60381,940[0.0760, 0.1039]3943310.05527.73831,940152
lam3_tier0.210822.75951,940[0.1927, 0.2290]40900.092512.84531,940[0.0784, 0.1066]3993260.05778.70291,940142
lam3_tierfb0.213422.93571,940[0.1952, 0.2316]41400.095113.12561,940[0.0809, 0.1093]4043210.06038.92031,940148
lam5_soft0.206222.44181,940[0.1882, 0.2242]40000.087912.40551,940[0.0740, 0.1018]3903350.05316.94991,940164
A.3 strict hit — by BON_N.
Nold.meanold.told.nold.ci95old.n_upold.n_downnew.meannew.tnew.nnew.ci95new.n_upnew.n_downnew_vs_old.dnew_vs_old.tnew_vs_old.nnew_vs_old.n_up
20.131417.13011,940[0.1164, 0.1465]25500.145918.19791,940[0.1302, 0.1616]28300.01443.13761,94054
40.139717.74371,940[0.1243, 0.1551]27100.180420.65971,940[0.1633, 0.1975]35000.04076.82831,940108
80.153118.72181,940[0.1371, 0.1691]29700.208222.58311,940[0.1902, 0.2263]40400.05527.73831,940152
A.3 strict hit — by δ.
δdelivered.meandelivered.tdelivered.ndelivered.ci95delivered.n_updelivered.n_downvs_old.dvs_old.tvs_old.nvs_old.n_upvs_old.ci95
00.206222.44181,940[0.1882, 0.2242]40000.05317.60101,940146[0.0394, 0.0668]
0.50.208222.58311,940[0.1902, 0.2263]40400.05527.73831,940152[0.0412, 0.0691]
10.207222.51251,940[0.1892, 0.2253]40200.05417.58931,940151[0.0401, 0.0681]
A.3 strict hit — every arm paired against ship.
armdelivered.ddelivered.tdelivered.ndelivered.n_updelivered.ci95mean_over_set.dmean_over_set.tmean_over_set.nmean_over_set.n_upmean_over_set.ci95wer.dwer.twer.nwer.n_upwer.ci95genuineness.d
bestmem−0.0531−2.500032015[−0.0948, −0.0115]−0.0492−4.700232043[−0.0697, −0.0287]−0.0088−1.459832059[−0.0206, 0.0030]−0.3383
cap100−0.0400−1.417820012[−0.0953, 0.0153]−0.0394−3.218920035[−0.0634, −0.0154]−0.0048−0.612120040[−0.0204, 0.0107]0.0530
cap1500.00500.152120022[−0.0594, 0.0694]0.02752.305820070[0.0041, 0.0509]−0.0083−0.998620044[−0.0245, 0.0080]−0.0631
ship_vn−0.0250−0.684916015[−0.0965, 0.0465]−0.0219−1.658916029[−0.0477, 0.0040]0.01772.556216057[0.0041, 0.0313]0.1159
shipset−0.2200−4.36401004[−0.3188, −0.1212]−0.1212−4.617410019[−0.1727, −0.0698]0.01141.055110027[−0.0098, 0.0327]0.6172
v2−0.0312−1.416432020[−0.0745, 0.0120]−0.0453−4.454732047[−0.0652, −0.0254]−0.0042−0.736832054[−0.0155, 0.0070]−0.2405
w0−0.0938−3.960932015[−0.1401, −0.0474]−0.0867−7.778632029[−0.1086, −0.0649]0.02102.235032075[0.0026, 0.0394]0.3540
A.3 strict hit — gates.
armd_wer_vs_w0.dd_wer_vs_w0.td_wer_vs_w0.nd_wer_vs_w0.n_upd_wer_vs_w0.ci95passes_paired_werabs_wer_inline_meanpasses_abs_wer_inlined_genuineness_vs_w0.dd_genuineness_vs_w0.td_genuineness_vs_w0.nd_genuineness_vs_w0.n_upd_genuineness_vs_w0.ci95bound_paired_werbound_abs_wer_inlinenote
bestmem−0.0298−3.376232065[−0.0471, −0.0125]yes0.0634yes−0.6924−11.337432099[−0.8121, −0.5727]0.10400.2500genuineness is reported, never gating
cap100−0.0239−2.127120042[−0.0459, −0.0019]yes0.0799yes−0.2756−4.437320094[−0.3973, −0.1538]0.10400.2500genuineness is reported, never gating
cap150−0.0273−2.219120042[−0.0515, −0.0032]yes0.0751yes−0.3917−5.479320083[−0.5317, −0.2516]0.10400.2500genuineness is reported, never gating
ship−0.0210−2.235032069[−0.0394, −0.0026]yes0.0697yes−0.3540−6.7041320140[−0.4576, −0.2505]0.10400.2500genuineness is reported, never gating
ship_vn−0.0093−1.162016048[−0.0250, 0.0064]yes0.0874yes−0.2273−3.032116074[−0.3742, −0.0804]0.10400.2500genuineness is reported, never gating
shipset−0.0300−1.336810020[−0.0739, 0.0140]yes0.0767yes−0.1657−2.000310046[−0.3280, −0.0033]0.10400.2500genuineness is reported, never gating
v2−0.0253−2.713932069[−0.0435, −0.0070]yes0.0675yes−0.5945−9.4376320118[−0.7180, −0.4710]0.10400.2500genuineness is reported, never gating
w00.00000.00003200[0.0000, 0.0000]yes0.0967yes0.00000.00003200[0.0000, 0.0000]0.10400.2500genuineness is reported, never gating
A.3 strict hit — per class.
classold_delivered.meanold_delivered.told_delivered.nold_delivered.ci95old_delivered.n_upold_delivered.n_downold_mean.meanold_mean.told_mean.nold_mean.ci95old_mean.n_upold_mean.n_downold_lift.meanold_lift.told_lift.nold_lift.ci95
affirmative_grunt0.07782.739790[0.0221, 0.1334]700.08896.479190[0.0620, 0.1158]370−0.0111−0.407490[−0.0646, 0.0423]
breathy_giggle0.13084.4053130[0.0726, 0.1890]1700.06447.2619130[0.0470, 0.0818]4600.06632.6159130[0.0166, 0.1161]
chuckle0.43089.8804130[0.3453, 0.5162]5600.332712.2541130[0.2795, 0.3859]9000.09813.0147130[0.0343, 0.1618]
deep_breath0.03852.2716130[0.0053, 0.0716]500.05876.0651130[0.0397, 0.0776]370−0.0202−1.2378130[−0.0522, 0.0118]
exasperated_sigh0.27276.3934110[0.1891, 0.3563]3000.15917.2864110[0.1163, 0.2019]5500.11363.5083110[0.0502, 0.1771]
exhausted_groan0.12004.5076150[0.0678, 0.1722]1800.09757.9025150[0.0733, 0.1217]5600.02250.9709150[−0.0229, 0.0679]
frustrated_groan0.14675.0606150[0.0899, 0.2035]2200.14338.9511150[0.1119, 0.1747]6800.00330.1379150[−0.0440, 0.0507]
heavy_breathing0.01541.4197130[−0.0059, 0.0366]200.00381.6436130[−0.0007, 0.0084]300.01151.1901130[−0.0075, 0.0305]
humming0.23645.8085110[0.1566, 0.3161]2600.09895.4752110[0.0635, 0.1343]4900.13753.8053110[0.0667, 0.2083]
panting0.00000.000090[0.0000, 0.0000]000.00000.000090[0.0000, 0.0000]000.00000.000090[0.0000, 0.0000]
relief_sigh0.03331.751890[−0.0040, 0.0706]300.01943.297890[0.0079, 0.0310]1200.01390.869290[−0.0174, 0.0452]
scream0.26007.2354150[0.1896, 0.3304]3900.20589.9815150[0.1654, 0.2463]8100.05421.9834150[0.0006, 0.1077]
sharp_inhale0.446210.1940130[0.3604, 0.5319]5800.378816.8604130[0.3348, 0.4229]11700.06731.7967130[−0.0061, 0.1407]
soft_hum0.10003.7859130[0.0482, 0.1518]1300.09718.8062130[0.0755, 0.1187]6400.00290.1152130[−0.0462, 0.0520]
wistful_sigh0.00000.000090[0.0000, 0.0000]000.01813.568790[0.0081, 0.0280]120−0.0181−3.568790[−0.0280, −0.0081]
yawn0.00771.0000130[−0.0074, 0.0228]100.00872.7831130[0.0026, 0.0147]80−0.0010−0.1297130[−0.0155, 0.0136]
A.3 strict hit — per arm.
armold_delivered.meanold_delivered.told_delivered.nold_delivered.ci95old_delivered.n_upold_delivered.n_downold_mean.meanold_mean.told_mean.nold_mean.ci95old_mean.n_upold_mean.n_downold_lift.meanold_lift.told_lift.nold_lift.ci95
bestmem0.13447.0370320[0.0969, 0.1718]4300.09779.8614320[0.0782, 0.1171]11200.03672.3638320[0.0063, 0.0672]
cap1000.16006.1567200[0.1091, 0.2109]3200.13698.3212200[0.1046, 0.1691]8000.02311.1303200[−0.0170, 0.0632]
cap1500.20507.1634200[0.1489, 0.2611]4100.203711.5650200[0.1692, 0.2383]11100.00130.0552200[−0.0431, 0.0456]
ship0.18758.5799320[0.1447, 0.2303]6000.146911.4224320[0.1217, 0.1721]13600.04062.5094320[0.0089, 0.0724]
ship_vn0.16885.6814160[0.1105, 0.2270]2700.13208.0089160[0.0997, 0.1643]6600.03671.6410160[−0.0071, 0.0806]
shipset0.14004.0145100[0.0716, 0.2084]1400.10256.3181100[0.0707, 0.1343]3900.03751.3389100[−0.0174, 0.0924]
v20.15627.6860320[0.1164, 0.1961]5000.101610.5387320[0.0827, 0.1205]12400.05473.3580320[0.0228, 0.0866]
w00.09385.7446320[0.0618, 0.1257]3000.06026.9187320[0.0431, 0.0772]6700.03362.7250320[0.0094, 0.0578]
A.3 family-relaxed hit — the shipped reward R_old (large-v2 encoder, local window).
keymeantnci95n_upn_downrn_setsdf
delivered0.274727.10231,940[0.2549, 0.2946]5330
mean0.239841.05421,940[0.2284, 0.2513]1,2480
lift0.03494.25561,940[0.0188, 0.0510]511715
auc0.579674.68261,226[0.5644, 0.5948]1,1900
auc_minus_half0.079610.25671,226[0.0644, 0.0948]718449
r_within15,517[0.0687, 0.1021]0.08541,94013,576
wer0.079027.14461,940[0.0733, 0.0847]8450
genuineness2.734795.42261,940[2.6785, 2.7908]1,9400
clap0.144177.47341,940[0.1405, 0.1478]1,84595
A.3 family-relaxed hit — the candidate rewards R_new by λ.
keydelivered.meandelivered.tdelivered.ndelivered.ci95delivered.n_updelivered.n_downlift.meanlift.tlift.nlift.ci95lift.n_uplift.n_downvs_old.dvs_old.tvs_old.nvs_old.n_up
lam0_soft0.274727.10231,940[0.2549, 0.2946]53300.03494.25561,940[0.0188, 0.0510]5117150.00000.00001,9400
lam1_soft0.319130.14271,940[0.2983, 0.3398]61900.07939.24391,940[0.0624, 0.0961]5976290.04436.75111,940126
lam2_soft0.339731.58331,940[0.3186, 0.3608]65900.099911.59471,940[0.0830, 0.1168]6375890.06497.84521,940196
lam3_soft0.359833.01081,940[0.3384, 0.3812]69800.120013.74091,940[0.1029, 0.1371]6765500.08519.29481,940247
lam3_tier0.363933.30691,940[0.3425, 0.3853]70600.124114.03581,940[0.1068, 0.1414]6845420.089210.49521,940230
lam3_tierfb0.366033.45531,940[0.3445, 0.3874]71000.126214.23271,940[0.1088, 0.1435]6885380.091210.52161,940238
lam5_soft0.366033.45531,940[0.3445, 0.3874]71000.126214.44191,940[0.1090, 0.1433]6885380.09129.31531,940277
A.3 family-relaxed hit — by BON_N.
Nold.meanold.told.nold.ci95old.n_upold.n_downnew.meannew.tnew.nnew.ci95new.n_upnew.n_downnew_vs_old.dnew_vs_old.tnew_vs_old.nnew_vs_old.n_up
20.254625.73761,940[0.2352, 0.2740]49400.271626.89201,940[0.2519, 0.2914]52700.01702.68981,94092
40.267026.57691,940[0.2473, 0.2867]51800.330430.93241,940[0.3095, 0.3513]64100.06347.95081,940185
80.274727.10231,940[0.2549, 0.2946]53300.359833.01081,940[0.3384, 0.3812]69800.08519.29481,940247
A.3 family-relaxed hit — by δ.
δdelivered.meandelivered.tdelivered.ndelivered.ci95delivered.n_updelivered.n_downvs_old.dvs_old.tvs_old.nvs_old.n_upvs_old.ci95
00.360833.08471,940[0.3394, 0.3822]70000.08619.47301,940246[0.0683, 0.1039]
0.50.359833.01081,940[0.3384, 0.3812]69800.08519.29481,940247[0.0671, 0.1030]
10.357732.86321,940[0.3364, 0.3791]69400.08309.00291,940247[0.0649, 0.1011]
A.3 family-relaxed hit — every arm paired against ship.
armdelivered.ddelivered.tdelivered.ndelivered.n_updelivered.ci95mean_over_set.dmean_over_set.tmean_over_set.nmean_over_set.n_upmean_over_set.ci95wer.dwer.twer.nwer.n_upwer.ci95genuineness.d
bestmem−0.0219−0.733332042[−0.0803, 0.0366]−0.0547−3.642432089[−0.0841, −0.0253]−0.0088−1.459832059[−0.0206, 0.0030]−0.3383
cap100−0.0050−0.139720025[−0.0752, 0.0652]−0.0406−2.905820049[−0.0680, −0.0132]−0.0048−0.612120040[−0.0204, 0.0107]0.0530
cap1500.05501.435920035[−0.0201, 0.1301]0.03752.552020083[0.0087, 0.0663]−0.0083−0.998620044[−0.0245, 0.0080]−0.0631
ship_vn0.02500.506816033[−0.0717, 0.1217]−0.0344−1.934916037[−0.0692, 0.0004]0.01772.556216057[0.0041, 0.0313]0.1159
shipset−0.2200−4.05371006[−0.3264, −0.1136]−0.1375−4.949410022[−0.1920, −0.0830]0.01141.055110027[−0.0098, 0.0327]0.6172
v2−0.0531−1.749832039[−0.1126, 0.0064]−0.0602−4.499332079[−0.0864, −0.0340]−0.0042−0.736832054[−0.0155, 0.0070]−0.2405
w0−0.1313−4.360432028[−0.1902, −0.0723]−0.1289−9.411232054[−0.1558, −0.1021]0.02102.235032075[0.0026, 0.0394]0.3540
A.3 family-relaxed hit — gates.
armd_wer_vs_w0.dd_wer_vs_w0.td_wer_vs_w0.nd_wer_vs_w0.n_upd_wer_vs_w0.ci95passes_paired_werabs_wer_inline_meanpasses_abs_wer_inlined_genuineness_vs_w0.dd_genuineness_vs_w0.td_genuineness_vs_w0.nd_genuineness_vs_w0.n_upd_genuineness_vs_w0.ci95bound_paired_werbound_abs_wer_inlinenote
bestmem−0.0298−3.376232065[−0.0471, −0.0125]yes0.0634yes−0.6924−11.337432099[−0.8121, −0.5727]0.10400.2500genuineness is reported, never gating
cap100−0.0239−2.127120042[−0.0459, −0.0019]yes0.0799yes−0.2756−4.437320094[−0.3973, −0.1538]0.10400.2500genuineness is reported, never gating
cap150−0.0273−2.219120042[−0.0515, −0.0032]yes0.0751yes−0.3917−5.479320083[−0.5317, −0.2516]0.10400.2500genuineness is reported, never gating
ship−0.0210−2.235032069[−0.0394, −0.0026]yes0.0697yes−0.3540−6.7041320140[−0.4576, −0.2505]0.10400.2500genuineness is reported, never gating
ship_vn−0.0093−1.162016048[−0.0250, 0.0064]yes0.0874yes−0.2273−3.032116074[−0.3742, −0.0804]0.10400.2500genuineness is reported, never gating
shipset−0.0300−1.336810020[−0.0739, 0.0140]yes0.0767yes−0.1657−2.000310046[−0.3280, −0.0033]0.10400.2500genuineness is reported, never gating
v2−0.0253−2.713932069[−0.0435, −0.0070]yes0.0675yes−0.5945−9.4376320118[−0.7180, −0.4710]0.10400.2500genuineness is reported, never gating
w00.00000.00003200[0.0000, 0.0000]yes0.0967yes0.00000.00003200[0.0000, 0.0000]0.10400.2500genuineness is reported, never gating
A.3 family-relaxed hit — per class.
classold_delivered.meanold_delivered.told_delivered.nold_delivered.ci95old_delivered.n_upold_delivered.n_downold_mean.meanold_mean.told_mean.nold_mean.ci95old_mean.n_upold_mean.n_downold_lift.meanold_lift.told_lift.nold_lift.ci95
affirmative_grunt0.17784.386790[0.0983, 0.2572]1600.19318.438490[0.1482, 0.2379]570−0.0153−0.416590[−0.0872, 0.0566]
breathy_giggle0.34628.2640130[0.2641, 0.4283]4500.295212.0136130[0.2470, 0.3434]9100.05101.5614130[−0.0130, 0.1149]
chuckle0.546212.4594130[0.4602, 0.6321]7100.429816.3246130[0.3782, 0.4814]10800.11633.2972130[0.0472, 0.1855]
deep_breath0.37698.8339130[0.2933, 0.4606]4900.343317.2175130[0.3042, 0.3823]11600.03370.9070130[−0.0391, 0.1064]
exasperated_sigh0.29096.6871110[0.2056, 0.3762]3200.17737.4476110[0.1306, 0.2239]5800.11363.5437110[0.0508, 0.1765]
exhausted_groan0.21336.3566150[0.1476, 0.2791]3200.230810.6224150[0.1882, 0.2734]820−0.0175−0.6392150[−0.0712, 0.0362]
frustrated_groan0.24676.9848150[0.1774, 0.3159]3700.226711.1723150[0.1869, 0.2664]9000.02000.7179150[−0.0346, 0.0746]
heavy_breathing0.31547.7089130[0.2352, 0.3956]4100.351915.8432130[0.3084, 0.3955]1130−0.0365−1.0216130[−0.1066, 0.0336]
humming0.27276.3934110[0.1891, 0.3563]3000.19899.2511110[0.1567, 0.2410]7400.07392.0039110[0.0016, 0.1461]
panting0.33336.670890[0.2354, 0.4313]3000.315312.610090[0.2663, 0.3643]8000.01810.379190[−0.0753, 0.1114]
relief_sigh0.20004.717090[0.1169, 0.2831]1800.15147.969490[0.1142, 0.1886]5500.04861.269790[−0.0264, 0.1236]
scream0.26007.2354150[0.1896, 0.3304]3900.20589.9815150[0.1654, 0.2463]8100.05421.9834150[0.0006, 0.1077]
sharp_inhale0.446210.1940130[0.3604, 0.5319]5800.388516.9989130[0.3437, 0.4333]11800.05771.5269130[−0.0164, 0.1318]
soft_hum0.16154.9853130[0.0980, 0.2250]2100.14049.7604130[0.1122, 0.1686]7400.02120.7264130[−0.0359, 0.0782]
wistful_sigh0.14443.876390[0.0714, 0.2175]1300.11676.268090[0.0802, 0.1531]4300.02780.912090[−0.0319, 0.0875]
yawn0.00771.0000130[−0.0074, 0.0228]100.00872.7831130[0.0026, 0.0147]80−0.0010−0.1297130[−0.0155, 0.0136]
A.3 family-relaxed hit — per arm.
armold_delivered.meanold_delivered.told_delivered.nold_delivered.ci95old_delivered.n_upold_delivered.n_downold_mean.meanold_mean.told_mean.nold_mean.ci95old_mean.n_upold_mean.n_downold_lift.meanold_lift.told_lift.nold_lift.ci95
bestmem0.281211.1726320[0.2319, 0.3306]9000.229316.6154320[0.2022, 0.2563]21000.05202.5124320[0.0114, 0.0925]
cap1000.31509.5661200[0.2505, 0.3795]6300.266213.1099200[0.2264, 0.3061]12800.04881.8763200[−0.0022, 0.0997]
cap1500.375010.9270200[0.3077, 0.4423]7500.344417.1841200[0.3051, 0.3837]14800.03061.0927200[−0.0243, 0.0856]
ship0.303111.7796320[0.2527, 0.3536]9700.284018.3346320[0.2536, 0.3143]22000.01910.9253320[−0.0214, 0.0597]
ship_vn0.35009.2529160[0.2759, 0.4241]5600.254713.0962160[0.2166, 0.2928]10900.09533.2751160[0.0383, 0.1524]
shipset0.17004.5030100[0.0960, 0.2440]1700.16887.6945100[0.1258, 0.2117]5200.00130.0377100[−0.0637, 0.0662]
v20.250010.3118320[0.2025, 0.2975]8000.223817.4007320[0.1986, 0.2490]22200.02621.2910320[−0.0136, 0.0659]
w00.17198.1368320[0.1305, 0.2133]5500.155112.8570320[0.1314, 0.1787]15900.01680.9420320[−0.0182, 0.0517]

Cross-validation: 1,940 cells, window local, level hit_strict. ranked with rank_enc, hit measured with hit_enc. Absolute hit rates are NOT comparable across cells: each hit_enc is a different instrument with a different recall. Only the paired difference against R_old within a cell is comparable.

A.3 cross-validation cells (rank encoder × hit encoder).
celln_setsn_localisation_fallbacksdiagonalold_delivered.meanold_delivered.told_delivered.nold_delivered.ci95old_delivered.n_upold_delivered.n_downold_mean.meanold_mean.told_mean.nold_mean.ci95old_mean.n_upold_mean.n_downold_lift.mean
rank=large-v2|hit=large-v21,9402yes0.153118.72181,940[0.1371, 0.1691]29700.118325.67131,940[0.1093, 0.1273]73500.0348
rank=large-v2|hit=commercial1,9402no0.124216.58441,940[0.1095, 0.1389]24100.110826.16571,940[0.1025, 0.1191]73100.0135
rank=commercial|hit=large-v21,9402no0.153118.72181,940[0.1371, 0.1691]29700.118325.67131,940[0.1093, 0.1273]73500.0348
rank=commercial|hit=commercial1,9402yes0.124216.58441,940[0.1095, 0.1389]24100.110826.16571,940[0.1025, 0.1191]73100.0135
A.3 retention, hit encoder large-v2.
λdiagonal_dcrossed_dcrossed_tratio
lam1_soft0.02990.02785.80470.9310
lam2_soft0.04120.04547.59981.1000
lam3_soft0.05520.05728.70131.0374
lam3_tier0.05770.05889.32741.0179
lam3_tierfb0.06030.06139.59621.0171
lam5_soft0.05310.06398.99181.2039
A.3 retention, hit encoder commercial.
λdiagonal_dcrossed_dcrossed_tratio
lam1_soft0.03350.02274.42100.6769
lam2_soft0.05570.03145.09830.5648
lam3_soft0.07010.03875.74910.5515
lam3_tier0.06600.03665.60460.5547
lam3_tierfb0.06910.03715.59820.5373
lam5_soft0.07370.04075.56000.5524
A.3 oracle and headroom.
windowrandomoldneworacleheadroomfallbacksold_sharenew_share
local0.11830.15310.20820.37890.260620.13350.3452
global0.11830.15310.19070.37890.260600.13350.2779
A.3 term correlations against the human-proxy target, encoder large-v2.
termrnn_setsdfci95
genuineness−0.008515,5171,94013,576[−0.0253, 0.0083]
blend−0.083915,5171,94013,576[−0.1006, −0.0672]
clap0.148215,5171,94013,576[0.1317, 0.1646]
inv_wer0.038815,5171,94013,576[0.0220, 0.0556]
burst_term0.305315,5171,94013,576[0.2900, 0.3205]
A.3 term correlations against the human-proxy target, encoder commercial.
termrnn_setsdfci95
genuineness−0.008415,5171,94013,576[−0.0253, 0.0084]
blend−0.077515,5171,94013,576[−0.0942, −0.0607]
clap0.125415,5171,94013,576[0.1088, 0.1419]
inv_wer0.032315,5171,94013,576[0.0155, 0.0491]
burst_term0.329115,5171,94013,576[0.3140, 0.3440]
A.3 split by has_voice.
has_voicen_sets
True1,164
False776
A.3 has_voice = True — delivered.
deliveredmeantnci95n_upn_down
delivered0.088510.62551,164[0.0722, 0.1048]1030
A.3 has_voice = True — mean.
meanmeantnci95n_upn_down
mean0.078117.49691,164[0.0693, 0.0868]3560
A.3 has_voice = True — lift.
liftmeantnci95n_upn_down
lift0.01041.51301,164[−0.0031, 0.0239]102253
A.3 has_voice = True — auc_minus_half.
auc_minus_halfmeantnci95n_upn_down
auc_minus_half0.07375.0069355[0.0449, 0.1026]207135
A.3 has_voice = True — r_within.
r_withinrnn_setsdfci95
r_within0.04999,3091,1648,144[0.0282, 0.0715]
A.3 has_voice = True — wer.
wermeantnci95n_upn_down
wer0.024613.91371,164[0.0212, 0.0281]2720
A.3 has_voice = True — new_vs_old.
new_vs_olddtnn_upci95
new_vs_old0.06016.73941,16491[0.0426, 0.0776]
A.3 has_voice = False — delivered.
deliveredmeantnci95n_upn_down
delivered0.250016.0728776[0.2195, 0.2805]1940
A.3 has_voice = False — mean.
meanmeantnci95n_upn_down
mean0.178619.9502776[0.1611, 0.1962]3790
A.3 has_voice = False — lift.
liftmeantnci95n_upn_down
lift0.07145.9663776[0.0479, 0.0948]185185
A.3 has_voice = False — auc_minus_half.
auc_minus_halfmeantnci95n_upn_down
auc_minus_half0.12048.4755370[0.0926, 0.1483]245107
A.3 has_voice = False — r_within.
r_withinrnn_setsdfci95
r_within0.11996,2087765,431[0.0936, 0.1460]
A.3 has_voice = False — wer.
wermeantnci95n_upn_down
wer0.160628.5536776[0.1496, 0.1716]5730
A.3 has_voice = False — new_vs_old.
new_vs_olddtnn_upci95
new_vs_old0.04774.052977661[0.0246, 0.0707]

Guided arm: 80 sets, 640 clips, guidance 3.0, classes chuckle, relief_sigh, scream, sharp_inhale.

A.3 guided minus unguided, paired.
quantitydtnci95
delivered0.12502.290080[0.0182, 0.2318]
set_mean0.01560.670080[−0.0297, 0.0610]
lift0.10941.980080[0.0013, 0.2174]
wer−0.0009−0.080080
genuineness−0.0236−0.360080
clap0.00150.250080
A.3 new reward vs old, guided and unguided.
armdtn
unguided0.11252.390080
guided0.01250.300080
A.3 shipped values as read by the optimiser.
BURST_LAMBURST_LAM_MAXBURST_LAM_BUDGETLAM_DPOLAM_QDPOLAM_EMOTIONLAM_VOICE
shipped0.25001.25002.00001.00001.50001.00001.0000

Recipes above the ceiling: 24 of 50; ceiling binding on 10 classes: breathy_giggle, chuckle, deep_breath, exhausted_groan, frustrated_groan, heavy_breathing, scream, sharp_inhale, soft_hum, yawn.

A.4 The 17-class detector's recall floor (per_class_recall.json)

Read from $SC/out/vb_seg2/clf_prod/per_class_recall.json.
ClassFamilyRecall, real95 % CIFamily recall, realn realreliableRecall, DramaBox95 % CIFamily recall, DBXn DBX
Affirmative Gruntgroan0.700[0.521, 0.833]0.76730yes0.680[0.484, 0.828]0.68025
Breathy Gigglelaugh0.469[0.309, 0.635]0.96932yes0.857[0.685, 0.943]1.00028
Chucklelaugh0.641[0.507, 0.757]0.84953yes0.800[0.652, 0.895]0.97540
Deep Breathbreath0.186[0.107, 0.304]0.71259yes0.686[0.550, 0.797]0.86351
Exasperated Sighsigh0.226[0.135, 0.355]0.34053yes0.381[0.250, 0.532]0.59542
Exhausted Groangroan0.484[0.320, 0.652]0.87131yes0.320[0.172, 0.516]0.76025
Frustrated Groangroan0.594[0.423, 0.745]0.81232yes0.654[0.462, 0.806]0.65426
Heavy Breathingbreath0.120[0.042, 0.300]0.64025no0.280[0.143, 0.476]0.92025
Humminghum0.472[0.320, 0.630]0.61136yes0.516[0.348, 0.680]0.80631
Pantingbreath0.647[0.479, 0.785]0.88234yes0.462[0.288, 0.645]0.57726
Relief Sighsigh0.062[0.017, 0.202]0.31232yes0.469[0.309, 0.635]0.75032
Screamscream0.595[0.435, 0.737]0.59537yes0.963[0.817, 0.993]0.96327
Sharp Inhalebreath0.897[0.764, 0.959]0.89739yes0.531[0.364, 0.691]0.75032
Soft Humhum0.393[0.236, 0.576]0.57128no0.280[0.143, 0.476]0.60025
Wistful Sighsigh0.143[0.057, 0.315]0.39328no0.400[0.234, 0.593]0.56025
Yawnyawn0.460[0.310, 0.616]0.46037yes0.500[0.332, 0.668]0.50030
no_burstno_burst0.980[0.968, 0.988]0.980802yes0.964[0.910, 0.986]0.964110

A.5 Encoder comparison (encoder_compare.json)

Read from $SC/out/vb_merge/encoder_compare.json.
EncoderLicenceParamsDimval_acc_meanexact acc: real/heldout_balancedexact acc: proposed23exact acc: dramabox
voiceclap-large-v2CC BY 4.07 B + LoRA3,5840.70560.4659
voiceclap-commercialCC BY 4.0110 M7680.62580.3929
voiceclap-small-v2CC BY-NC 4.0110 M7680.63980.4024

A.6 The guidance-contrast preference pairs ($SC/cfg_rows)

128 shards, 474,418 rows, counted 2026-09-09 (cached count).

By family.
familyrows
cfg_high237,209
cfg_low237,209
By language.
langrows
de247,350
en227,068

Columns: uid, family, src, voice_key, lang, text, text_bursts, caption_general, caption_script, caption_tpl_text, caption_tpl, speaker_name, spoken_dur_s, dur_s, words_json, burst_starts, burst_ends, burst_labels, in_extreme, frames, codes, rej_frames, rej_codes, ref_frames, ref_codes, has_ref, is_val, cfg_free_text, cfg_neutral, emo_strength

A.7 The combination study (~/combination_study/stats/analysis.json)

10,459 records, 31,377 clips, 60 attributes; by family: emo 40, vn 17, qual 3.

A.7 2×2×2 effects on target_z, pooled.
effectemo/Affection.n_slotsemo/Amusement.n_slotsemo/Anger.n_slotsemo/Astonishment_Surprise.n_slotsemo/Awe.n_slotsemo/Bitterness.n_slotsemo/Concentration.n_slotsemo/Confusion.n_slotsemo/Contemplation.n_slotsemo/Contempt.n_slotsemo/Contentment.n_slotsemo/Disappointment.n_slotsemo/Disgust.n_slotsemo/Distress.n_slotsemo/Doubt.n_slotsemo/Elation.n_slots
per_attr10101010101010101010101010101010
pooled
metric
A.7 2×2×2 effects on target, pooled.
effectemo/Affection.n_slotsemo/Amusement.n_slotsemo/Anger.n_slotsemo/Astonishment_Surprise.n_slotsemo/Awe.n_slotsemo/Bitterness.n_slotsemo/Concentration.n_slotsemo/Confusion.n_slotsemo/Contemplation.n_slotsemo/Contempt.n_slotsemo/Contentment.n_slotsemo/Disappointment.n_slotsemo/Disgust.n_slotsemo/Distress.n_slotsemo/Doubt.n_slotsemo/Elation.n_slots
per_attr10101010101010101010101010101010
pooled
metric
A.7 2×2×2 effects on wer_parakeet, pooled.
effectemo/Affection.n_slotsemo/Amusement.n_slotsemo/Anger.n_slotsemo/Astonishment_Surprise.n_slotsemo/Awe.n_slotsemo/Bitterness.n_slotsemo/Concentration.n_slotsemo/Confusion.n_slotsemo/Contemplation.n_slotsemo/Contempt.n_slotsemo/Contentment.n_slotsemo/Disappointment.n_slotsemo/Disgust.n_slotsemo/Distress.n_slotsemo/Doubt.n_slotsemo/Elation.n_slots
per_attr10101010101010101010101010101010
pooled
metric
A.7 2×2×2 effects on wer, pooled.
effectemo/Affection.n_slotsemo/Amusement.n_slotsemo/Anger.n_slotsemo/Astonishment_Surprise.n_slotsemo/Awe.n_slotsemo/Bitterness.n_slotsemo/Concentration.n_slotsemo/Confusion.n_slotsemo/Contemplation.n_slotsemo/Contempt.n_slotsemo/Contentment.n_slotsemo/Disappointment.n_slotsemo/Disgust.n_slotsemo/Distress.n_slotsemo/Doubt.n_slotsemo/Elation.n_slots
per_attr10101010101010101010101010101010
pooled
metric
A.7 2×2×2 effects on genuineness, pooled.
effectemo/Affection.n_slotsemo/Amusement.n_slotsemo/Anger.n_slotsemo/Astonishment_Surprise.n_slotsemo/Awe.n_slotsemo/Bitterness.n_slotsemo/Concentration.n_slotsemo/Confusion.n_slotsemo/Contemplation.n_slotsemo/Contempt.n_slotsemo/Contentment.n_slotsemo/Disappointment.n_slotsemo/Disgust.n_slotsemo/Distress.n_slotsemo/Doubt.n_slotsemo/Elation.n_slots
per_attr10101010101010101010101010101010
pooled
metric
A.7 2×2×2 effects on blend, pooled.
effectemo/Affection.n_slotsemo/Amusement.n_slotsemo/Anger.n_slotsemo/Astonishment_Surprise.n_slotsemo/Awe.n_slotsemo/Bitterness.n_slotsemo/Concentration.n_slotsemo/Confusion.n_slotsemo/Contemplation.n_slotsemo/Contempt.n_slotsemo/Contentment.n_slotsemo/Disappointment.n_slotsemo/Disgust.n_slotsemo/Distress.n_slotsemo/Doubt.n_slotsemo/Elation.n_slots
per_attr10101010101010101010101010101010
pooled
metric
A.7 2×2×2 effects on r_burst, pooled.
effectemo/Affection.n_slotsemo/Amusement.n_slotsemo/Anger.n_slotsemo/Astonishment_Surprise.n_slotsemo/Awe.n_slotsemo/Bitterness.n_slotsemo/Concentration.n_slotsemo/Confusion.n_slotsemo/Contemplation.n_slotsemo/Contempt.n_slotsemo/Contentment.n_slotsemo/Disappointment.n_slotsemo/Disgust.n_slotsemo/Distress.n_slotsemo/Doubt.n_slotsemo/Elation.n_slots
per_attr10101010101010101010101010101010
pooled
metric
A.7 2×2×2 effects on dur_err_abs_s, pooled.
effectemo/Affection.n_slotsemo/Amusement.n_slotsemo/Anger.n_slotsemo/Astonishment_Surprise.n_slotsemo/Awe.n_slotsemo/Bitterness.n_slotsemo/Concentration.n_slotsemo/Confusion.n_slotsemo/Contemplation.n_slotsemo/Contempt.n_slotsemo/Contentment.n_slotsemo/Disappointment.n_slotsemo/Disgust.n_slotsemo/Distress.n_slotsemo/Doubt.n_slotsemo/Elation.n_slots
per_attr10101010101010101010101010101010
pooled
metric
A.7 family emo, metric target_z.
effectmeantnn_upvalue
eff_L0.07662.8067399185
eff_S0.38409.4113399240
eff_C0.04961.8450399172
observed_LSC0.776214.1288399301
sum_of_parts0.51026.6668399229
excess0.26604.4634399230
int_LS0.03821.3562399195
int_LC−0.0312−1.2246399178
int_SC0.27667.5125399249
int_LSC0.03500.6161399196
cumulativity_ratio1.5214
A.7 family emo, metric target.
effectmeantnn_upvalue
eff_L0.02573.4286399185
eff_S0.114010.2492399240
eff_C0.01602.1917399172
observed_LSC0.232614.6192399301
sum_of_parts0.15577.4014399229
excess0.07694.6341399230
int_LS0.00771.0459399195
int_LC−0.0119−1.7237399178
int_SC0.08438.0480399249
int_LSC0.00660.4265399196
cumulativity_ratio1.4937
A.7 family emo, metric wer_parakeet.
effectmeantnn_upvalue
eff_L−0.0003−0.1701399175
eff_S0.02785.5364399209
eff_C0.00411.8276399192
observed_LSC0.100811.1937399280
sum_of_parts0.03154.6476399238
excess0.06939.0839399260
int_LS−0.0082−2.6191399174
int_LC0.00010.0254399195
int_SC0.078211.8010399320
int_LSC0.00160.2765399203
cumulativity_ratio3.1967
A.7 family emo, metric genuineness.
effectmeantnn_upvalue
eff_L0.02521.2317399211
eff_S−0.9316−18.918239980
eff_C0.10975.3644399261
observed_LSC−1.5935−22.730439957
sum_of_parts−0.7966−12.3340399101
excess−0.7969−15.055239988
int_LS0.08203.3292399215
int_LC0.02611.1602399207
int_SC−0.8623−21.529539954
int_LSC0.08541.8124399208
cumulativity_ratio2.0004
A.7 family emo, metric blend.
effectmeantnn_upvalue
eff_L0.06651.0808399211
eff_S0.11581.3981399218
eff_C−0.0072−0.1153399195
observed_LSC0.84406.2487399260
sum_of_parts0.17511.1089399207
excess0.66894.7264399232
int_LS0.09741.6463399211
int_LC0.05780.9733399205
int_SC0.53155.1796399241
int_LSC0.03560.2927399210
cumulativity_ratio4.8195
A.7 family emo, metric r_burst.
effectmeantnn_upvalue
eff_L−0.0113−1.3860399173
eff_S−0.0488−4.1040399164
eff_C−0.0078−1.0618399160
observed_LSC−0.1405−10.0174399110
sum_of_parts−0.0680−3.3268399153
excess−0.0725−4.5648399157
int_LS0.00230.2703399188
int_LC0.00290.3374399202
int_SC−0.0704−6.7293399129
int_LSC0.01470.9185399181
cumulativity_ratio2.0666
A.7 family vn, metric target_z.
effectmeantnn_upvalue
eff_L0.37658.6642170126
eff_S0.61436.1469170115
eff_C0.02591.132617096
observed_LSC0.98517.4312170121
sum_of_parts1.01668.0187170121
excess−0.0316−0.368017085
int_LS−0.1643−3.747917071
int_LC−0.1253−3.499017059
int_SC0.14352.019217090
int_LSC−0.2289−3.221417065
cumulativity_ratio0.9689
A.7 family vn, metric target.
effectmeantnn_upvalue
eff_L0.36928.8987170126
eff_S0.78767.0034170115
eff_C0.02971.427717096
observed_LSC1.24338.5464170121
sum_of_parts1.18668.4351170121
excess0.05670.644517085
int_LS−0.1412−3.327617071
int_LC−0.1448−4.324717059
int_SC0.20933.029217090
int_LSC−0.2668−4.295517065
cumulativity_ratio1.0478
A.7 family vn, metric wer_parakeet.
effectmeantnn_upvalue
eff_L−0.0006−0.129117072
eff_S0.466918.2196170154
eff_C0.00251.070017051
observed_LSC0.772144.3702170169
sum_of_parts0.468817.1056170150
excess0.303415.1105170155
int_LS−0.0157−2.365117069
int_LC0.01702.472017099
int_SC0.323018.9847170169
int_LSC0.04193.3073170103
cumulativity_ratio1.6471
A.7 family vn, metric genuineness.
effectmeantnn_upvalue
eff_L0.02290.696817088
eff_S−1.5492−15.676317020
eff_C0.00540.214817085
observed_LSC−2.2779−25.248517011
sum_of_parts−1.5209−14.562117020
excess−0.7570−8.606917049
int_LS−0.0235−0.548417081
int_LC0.02590.782217086
int_SC−0.7946−13.111517026
int_LSC−0.0703−1.017017081
cumulativity_ratio1.4977
A.7 family vn, metric blend.
effectmeantnn_upvalue
eff_L−0.3203−3.415817070
eff_S−3.0166−13.711317033
eff_C−0.3424−4.687617068
observed_LSC−3.4009−16.738617020
sum_of_parts−3.6793−13.992217028
excess0.27841.4753170102
int_LS0.25562.567417092
int_LC0.13011.908417092
int_SC−0.4702−4.123217072
int_LSC−0.7259−4.527017050
cumulativity_ratio0.9243
A.7 family vn, metric r_burst.
effectmeantnn_upvalue
eff_L−0.0483−3.744317060
eff_S−0.1862−6.919117038
eff_C−0.0264−2.229117051
observed_LSC−0.2826−8.817217035
sum_of_parts−0.2610−6.778117043
excess−0.0216−0.916917068
int_LS0.03772.614517091
int_LC0.03272.640517092
int_SC−0.0902−5.625517046
int_LSC0.00370.145317083
cumulativity_ratio1.0829
A.7 family qual, metric target_z.
effectmeantnn_upvalue
eff_L0.39953.26133021
eff_S0.00610.04263017
eff_C0.03200.34463014
observed_LSC0.18420.82023018
sum_of_parts0.43761.56843017
excess−0.2534−1.08843013
int_LS0.25882.05033019
int_LC−0.0605−0.62863013
int_SC−0.3335−2.24313011
int_LSC0.23631.04183018
cumulativity_ratio0.4210
A.7 family qual, metric target.
effectmeantnn_upvalue
eff_L0.42973.31703021
eff_S−0.0320−0.12683017
eff_C−0.0754−0.64203014
observed_LSC−0.0008−0.00233018
sum_of_parts0.32240.92103017
excess−0.3232−0.89763013
int_LS0.33091.98473019
int_LC−0.0632−0.37243013
int_SC−0.5490−1.96673011
int_LSC0.08390.23553018
cumulativity_ratio−0.0026
A.7 family qual, metric wer_parakeet.
effectmeantnn_upvalue
eff_L0.00020.01643012
eff_S0.02361.83553020
eff_C0.00250.4432309
observed_LSC0.22064.91773024
sum_of_parts0.02631.01353014
excess0.19444.70833026
int_LS−0.0210−1.31963013
int_LC−0.0299−2.1434309
int_SC0.21565.60913024
int_LSC−0.0594−2.06683012
cumulativity_ratio8.3979
A.7 family qual, metric genuineness.
effectmeantnn_upvalue
eff_L0.03580.36143017
eff_S0.02280.19943017
eff_C0.00540.08903015
observed_LSC−0.3813−2.07533012
sum_of_parts0.06400.32563018
excess−0.4453−2.3914308
int_LS0.11531.01103014
int_LC0.10311.50713018
int_SC−0.6215−3.5805307
int_LSC0.08450.38813016
cumulativity_ratio−5.9625
A.7 family qual, metric blend.
effectmeantnn_upvalue
eff_L0.02820.14243015
eff_S−0.3951−1.10793016
eff_C−0.3424−1.94183012
observed_LSC−0.8817−1.74733013
sum_of_parts−0.7093−1.49153014
excess−0.1724−0.32623016
int_LS0.48191.84843017
int_LC−0.0092−0.03253018
int_SC−0.7299−1.99733012
int_LSC−0.1695−0.39053017
cumulativity_ratio1.2431
A.7 family qual, metric r_burst.
effectmeantnn_upvalue
eff_L−0.0188−0.68803014
eff_S−0.0722−1.86613010
eff_C−0.0264−0.9234309
observed_LSC−0.1668−3.5421308
sum_of_parts−0.1174−1.49743010
excess−0.0494−0.66413013
int_LS0.04071.18643018
int_LC−0.0032−0.15663017
int_SC−0.0941−1.87013012
int_LSC−0.0142−0.22613012
cumulativity_ratio1.4211
A.7 the α ladder and its controls, pooled.
celltarget.meantarget.ttarget.ntarget.n_uptarget.n_attrstarget_z.meantarget_z.ttarget_z.ntarget_z.n_uptarget_z.n_attrswer_parakeet.meanwer_parakeet.twer_parakeet.nwer_parakeet.n_upwer_parakeet.n_attrswer.mean
ctl_a0 — ctl_a00.25686.017417090170.19385.249617090170.00160.490317073170.0064
ctl_rand — ctl_rand−0.0433−1.55421706717−0.0303−0.992817067170.01211.869817074170.0248
ctl_rand_LC — ctl_rand_LC0.14213.657717086170.10582.776917086170.06886.1529170120170.0757
lad_L0.5 — lad_L0.50.85606.3163170129170.946010.0239170129170.299011.7347170136170.3603
lad_L1.5 — lad_L1.51.01507.5499170135171.135511.8956170135170.306712.2239170139170.3742
lad_a0.05 — lad_a0.050.90608.0558170129170.824110.8216170129170.08225.9875170104170.1208
lad_a0.15 — lad_a0.150.71915.2005170128170.95398.6344170128170.501518.2364170158170.6396
lad_both — lad_both0.95938.7224170137170.898311.8026170137170.09796.0726170106170.1126
lad_g1.5 — lad_g1.51.00277.8445170139170.997011.2676170139170.20979.2406170131170.2801
lad_g3.0 — lad_g3.00.84646.3242170128171.041110.1409170128170.443915.1384170155170.5161
A.7 ladder against the 1/1/1 cell.
cellref
ctl_a0111
ctl_rand111
ctl_rand_LC111
lad_L0.5111
lad_L1.5111
lad_a0.05111
lad_a0.15111
lad_both111
lad_g1.5111
lad_g3.0111

A.8 The burst recipes as served (~/wikiskills/VOCAL_BURSTS.md)

31 recipes over 117 pages (177 incl. empty); 18 recipes name a weight above the shipped ceiling.
ClassAdapterλPrompt formFamily hitConsistentStrictnn cons.Source
relief_sighvb0.80GENERAL cause0.750.680.0023§51
chucklevb2.00GENERAL cause + longer duration0.730.660.3823§51
contented_sighvb1.50GENERAL cause + longer duration0.680.610.6833§51
soft_humvb2.00GENERAL cause + longer duration0.680.610.3033§51
wistful_sighvb1.00base form0.550.480.0734§52
cacklevb1.80base form0.540.470.0234§51
clears_throatvb1.25label mid-sentence0.480.410.0045§51
low_mumblevb2.00GENERAL cause + longer duration0.480.410.4545§51
childlike_gigglevb0.80base form0.470.400.0545§51
exasperated_sighvb1.25base form0.450.380.0545§52
breathy_gigglevb2.00base form0.400.330.2356§52
sharp_inhalevb2.30base form0.380.320.3857§52
resonant_humvb1.50base form0.370.300.2367§52
nervous_gigglevbr1.50GENERAL cause + longer duration0.330.270.0068§52
ahemvb2.00GENERAL cause + longer duration0.330.260.3368§52 ⚠
fearful_gaspvb1.25GENERAL cause + longer duration0.280.220.00710§52
frustrated_groanvb2.00GENERAL cause0.270.200.00811§51
deep_breathvb2.30base form0.270.200.23811§52
guffawvbr1.00base form0.270.200.00811§52
displeased_gruntvb2.00GENERAL cause + longer duration0.250.180.00912§52 ⚠
purrvb1.50base form0.250.180.00912§52
snickervbr1.25GENERAL cause + longer duration0.250.180.00912§52
whispered_mumblevbr0.25base form0.230.170.00913§52
coughvb1.00base form0.230.170.02913§52
hummingvbr0.80base form0.230.170.00913§52
exhausted_groanvb1.80base form0.220.150.221015§51
pain_moanvbr1.50GENERAL cause + longer duration0.180.120.001219§52
surprised_gaspvb1.00longer duration0.180.110.171219§51
coughingvbr0.25base form0.180.120.001219§52 ⚠
screamvb1.30longer duration0.170.100.151323§51
sniffvbr2.00GENERAL cause + longer duration0.150.080.031527§52

A.9 The coefficient table shipped with the three levers (coefficients.json)

Schema laion-voice-acting/wikiskill-coefficients/1, generated 2026-09-02 by wikiskills/code/build_wikiskills.py; primary WER instrument wer_parakeet.

A.9 global — steer.
keyemovnqual
k_by_family130
A.9 global — cfg.
keyformulag_emotiong_deliveryg_min_usefulcost_factorsteer_branch_when_bothsteer_branch_note
cfglogits = logits_uncond + g * (logits_cond - logits_uncond)3.00002.50001.00001.9300bothsteering both CFG branches rather than only the conditioned one keeps 82% of the effect and returns 0.209 of word error and 0.75 of genuineness
A.9 global — quality_dpo.
keyadapterrepoweightchosen_bynotesource
quality_dposft3_dpo:quality1504laion/moss-va-sft3-quality-dpo-lora1.5000listeningthe general-quality DPO adapter (checkpoint 1504, laion/moss-va-sft3-quality-dpo-lora) ships at merge weight 1.5. Eleven weights were swept over 220 takes and NO weight reached significance on genuineness (+0.42 at 1.0, +0.48 at 1.5, +0.58 at 1.75); the only significant effect in the whole sweep was a HARM, word error +0.148 at 2.5. The statistic could not separate the candidates, so the operating point is a listening decision, taken 2026-09-02, and it is recorded as onequality-adapter listening sweep, 280 takes, hf.co/spaces/laion/moss-quality-adapter-listening; choice by ear
A.9 global — numbness_subtraction.
keykeyalphatapsapplies_tonote
numbness_subtractionemo:Emotional_Numbness−0.1000top1emosubtracting Emotional_Numbness at alpha = -0.10 returns +0.60 of genuineness (t 9.64, 67 of 80 prompts) at no cost in emotion when the adapter carries the emotion
A.9 global — family_effects.
keyeff_L.meaneff_L.teff_L.neff_L.n_upeff_S.meaneff_S.teff_S.neff_S.n_upeff_C.meaneff_C.teff_C.neff_C.n_upint_LS.meanint_LS.tint_LS.nint_LS.n_up
emo0.07662.80673991850.38409.41133992400.04961.84503991720.03821.3562399195
vn0.37658.66421701260.61436.14691701150.02591.132617096−0.1643−3.747917071
qual0.39953.261330210.00610.042630170.03200.344630140.25882.05033019
A.9 thresholds.
regimewer_parakeet_absd_wer_parakeet_maxd_genuineness_mind_blend_mind_r_burst_mind_dur_err_abs_s_max
balanced0.15000.0500−0.6000−1.0000−0.10000.3000
high_effect0.3000−1.5000−3.0000−0.25001.0000
A.9 guardrails.
keydgendwerdburstdabsdurwernburstok_frac
safe−0.34130.1038−0.21660.07630.15000.62001.0000
strong−0.68260.2075−0.43320.15260.25000.40000.9000

60 attributes carry per-attribute coefficients; the per-attribute table is omitted here for length and is in the file.

A.10 Index of the protocol

~/paper/protocol.md: 8,219 lines, 72 numbered sections, last §72. Sections whose titles are German are the German-language entries of the protocol; this report translates their content where it uses them.

Numbered sections, in file order.
§LineTitle
116The problem this work addresses
225Prompt surface
3103Data
4146Training runs
5169Round-2 result: timing solved, emotion control lost
6187GRPO: a measured negative result
7215Round-3 result: the format collapse is fixed, emotion control is not
8295DPO corpus v2 — recovering the lost word timings
9319Published artefacts
10337Operational notes worth recording
113492026-08-25, afternoon — round 3 completed, corpus v2 built
12453The four levers, and why they failed
13529Published
14537The emotion LoRA: selection, and what the VoiceNet arm adds
15611Emotion LoRAs — results
16667Classifier-free-guidance DPO
17757Per-bucket LoRAs — 40 emotions and 500 voices
18824The four-factor study — what actually drives the emotion score
19898A measurement bug in every evaluation so far
20916Classifier-free guidance on the emotion condition
21987A vocal-burst adapter
221,018Activation forensics: where in the network each attribute lives
231,166Layer forensics — four workstreams over the stored activations
241,355VoiceNet tail adapters -- published
251,467Publication round 2 -- burst, quality and the completed VoiceNet set
261,488The multi-layer, multi-vector steering study
271,562The three-lever programme -- what the agents were told (2026-08-29)
281,638The LoRA dose-response study
291,884The combination study -- CFG x LoRA x steering vector
302,175Classifier-free guidance -- the A1 study
312,290The demo server's stop-talking investigation, and a defect in how it loads adapters
332,378Does stacking adapters dull the sound? No -- it destroys the words instead
342,436The protocol audit, and the state of the archive (2026-08-29)
322,595Three conditioning datasets: quality, speaking rate, and speed
352,922Shipping the three levers into the demo server, and the first WikiSkill draft (2026-08-29)
363,137Organic emotional transitions -- making a performance turn inside one utterance
373,242Duration control is already solved; what is NOT controlled is TEMPO
383,294The rate tag is a re-encoding of the duration, not a second control
393,355RESULT: the rate adapter changes nothing measurable
403,415Crossfade v2: the same performance turn, without the steering vectors
413,563The quality adapters: a null measurement, then a listener who disagreed
423,629rew_chosen is measured against bare SFT3, not against the run's own starting point
443,668Burst realisation and stop/no-improvisation: one DPO corpus, and a per-pair scoring harness that answers "which checkpoint"
453,831The qual1504 weight sweep cannot separate the doses
433,905The vocal-burst merge weight: a dose sweep under the real stack, and a metric that could not see the answer
464,158The listening page shipped with its audio players dead, and nobody would have known
474,195The burst adapters do not fail because of a weight: sixty of them were trained on a label the detector contradicts
484,497Re-filtering the burst buckets: 9.8 % of the training rows survive, and 34 of 70 classes have nothing to re-train on
494,755Manufacturing the missing bursts: voice conversion fails the gate, and the corpus that works scores 0.000
504,922Five more classes from manufactured data: the family-level effect replicates, the strict metric still cannot see it, and five classes are un
255,124The family-relaxed hit rate becomes the burst metric, and the table that carries it
265,169Two framings corrected by the owner
275,188The quality DPO adapter ships at 1.5
515,200The other levers, stacked: the DPO adapter is a null, the dose ladder had not ended, and two prompt forms add up
525,342The other fifty-five burst classes: nineteen recipes, nineteen classes that do not exist for this model, and a prompt result that did not ge
535,520Classifier-free guidance on the burst: it works, g = 4 triples the hit rate, and the prompt is doing most of the work
545,678DramaBox can be asked for a scream in plain English; our detector cannot be asked whether it got one
555,786A multimodal model annotates the bursts our detector cannot name: 88 % against 5 %, zero invented labels, and spans too wide to train a loca
565,891Adapter capacity is not the bottleneck: rank 32 and rank 64 on the five classes that have data and still will not fire
576,039The calm lead-in doubles DramaBox's dynamic contrast: 11.5 dB against §54's 6.1 dB, and a 24,500-prompt corpus balanced 50/50 by gender acro
586,122Real burst training data with a blind multimodal second opinion: 3,598 utterances over 49 classes, family agreement 64.8 % against exact-cla
596,296Two instruments instead of one: keeping the eval audio so a multimodal model can score it, and tuning the classifier head on the 11 classes
606,391Re-labelling the same audio with a better annotator: burst LoRAs trained on Gemini's spans against the ones that ship
616,441Does DramaBox actually produce the burst it was asked for? 73,500 clips, annotated whole
626,513Two sources, one cut policy: a burst dataset with verified negatives, a detector graded across generators, and an annotator that confabulate
636,703Shipping the burst detector, and the recall floor that makes a hit-rate table readable
646,822The scream gain at 45-class scale: a family-level effect, a dose curve that turns on at 25 %, and a recipe table merged rather than replaced
656,955Addendum to the previous section: a second detector over the same audio, and the half of the claim that does not survive it
667,030Combining the per-class burst adapters we already have: a clean negative at equal weight budget
677,100Consolidation: what the vocal-burst programme established, what it disproved, and eight ways a green exit code lied
687,485Where the work ended up, and the one convention every consumer of it needs
697,566Two predictors put to a human vote, and the scale that nearly got the question wrong
707,697Nachprüfung und Erweiterung von voice-annotation-data-v2
717,791Eine Best-of-N-Belohnung, die den Burst sieht, den sie auswählen soll — und der Stapel, der zwei Studien entgegengesetzte Vorzeichen gibt
728,057Der Voice-Acting-Agent als Schauspieler: ein Bericht entlang von Kopf, Instrument, Werkzeugen und Ohren

Provenance of the numbers

This report is rendered by $SC/code/vareport/en_build.py from one block list into HTML and Markdown, so the two cannot drift apart. The demo server and its mirror are only ever read. The 29 August report is a source and is never written.

WhatSourceHow
Backbone for Part I and the prehistory~/technical_report.html (29 Aug 2026; the uploaded report)read in full
Control that nothing was lost~/technical_report_prev.html (26 Aug) · ~/reference_report_v1.html (16 Aug)residual paragraphs merged
Part II backbone~/voice_acting_agent_report.html (9 Sep) · ~/paper/VOICE_ACTING_AGENT.mdtranslated and extended
Details no report carried/e/home/jusers/schuhmann1/jupiter/paper/protocol.md §1–§725 sections indexed
Shipped adapter stack and every ceiling/e/scratch/reformo/schuhmann1_moss/hv_mirror/config.pyread (read-only)
Best-of-N (Part II.6)/e/scratch/reformo/schuhmann1_moss/out/vb_optread
Detector recall floor/e/scratch/reformo/schuhmann1_moss/out/vb_seg2/clf_prod/per_class_recall.jsonread
Encoder comparison/e/scratch/reformo/schuhmann1_moss/out/vb_merge/encoder_compare.jsonread
Combination study 2×2×2/e/home/jusers/schuhmann1/jupiter/combination_study/stats/analysis.jsonread
Coefficient table/e/home/jusers/schuhmann1/jupiter/wikiskills/coefficients.jsonread
Burst recipes/e/home/jusers/schuhmann1/jupiter/wikiskills/VOCAL_BURSTS.mdparsed
Guidance contrast pairs/e/scratch/reformo/schuhmann1_moss/cfg_rowscounted (474,418 rows, cached)
Protocol prose citations/e/home/jusers/schuhmann1/jupiter/paper/protocol.md21 citations verified, 0 not found