A sample per class from Dataset A (burst_gemini_utterances).
The player is the shipped OGG, read back out of the WebDataset shard.
Blue = what gemini-3.8-flash heard, asked blind; amber = the burst span and class the
corpus already carried.
These annotations are a second opinion, not verified ground truth.
Every agreement number on this page is between two models — this project's burst detector and a
large multimodal model — and both are downstream of models trained on expressive speech. A shared
prior would inflate all of them and nothing measured here could see it. Whether a label is
correct is a listening question, which is what this page is for.
Gemini's spans are also wide (§55: 38.2 % clip coverage for a 1.27× lift over chance at
containing the loudest moment, against our detector's 6.5 % for 3.57×): it is better at
what and worse at where. Dataset B's segment boundaries are the weak part.