Gemini-annotated real vocal-burst training rows

A sample per class from Dataset A (burst_gemini_utterances). The player is the shipped OGG, read back out of the WebDataset shard. Blue = what gemini-3.8-flash heard, asked blind; amber = the burst span and class the corpus already carried.

These annotations are a second opinion, not verified ground truth. Every agreement number on this page is between two models — this project's burst detector and a large multimodal model — and both are downstream of models trained on expressive speech. A shared prior would inflate all of them and nothing measured here could see it. Whether a label is correct is a listening question, which is what this page is for. Gemini's spans are also wide (§55: 38.2 % clip coverage for a 1.27× lift over chance at containing the loudest moment, against our detector's 6.5 % for 3.57×): it is better at what and worse at where. Dataset B's segment boundaries are the weak part.
37.3 %
corpus class in Gemini's top-3, per row
19.0 %
corpus class is Gemini's top-1
64.8 %
same burst family
70
of 83 classes used (detector: 8)
0
invented labels
10 %
median clip covered by Gemini spans

Cards