A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation
In this tutorial, we work with MSEB, the Massive Sound Embedding Benchmark from Google Research, and approach it from the perspective of what a leaderboard number actually means: the evaluator surface. We install the package and map its three layers, then write two deliberately different encoders against the framework’s own abstract base class: one that measures loudness over time and one that measures timbre, and encode a small synthetic corpus we generate in the notebook so nothing has to be downloaded. We drive the classification, clustering, retrieval, and segmentation evaluators over those embeddings, call the metric functions directly to see what each one rewards, and finish by assembling the TaskMetadata a real submission carries. The result is a comparison in which the two encoders trade places depending on which evaluator is asked, which is the argument for a multi-task benchmark made in numbers rather than in prose. Copy CodeCopiedUse a different Browser import os import sys import json import math import traceback import subprocess import numpy as np RESULTS = {} BENCH = {} def banner(title): print(“n” + “=” * 78) print(title) print(“=” * 78) def section(name): def wrap(fn): def run(*a, **kw): banner(name) try: out = fn(*a, **kw) RESULTS[name] = out if isinstance(out, str) else “ok” return out except Exception as e: RESULTS[name] = f”SKIPPED / FAILED -> {type(e).__name__}: {e}” print(f”n[!] {name} did not complete: {type(e).__name__}: {e}”) traceback.print_exc(limit=3) return None return run return wrap banner(“0. Install MSEB and map the three layers we will use”) subprocess.run([sys.executable, “-m”, “pip”, “install”, “-q”, “mseb==0.1.0″], check=True) import mseb from mseb import types, encoder as encoder_lib, evaluator as evaluator_lib, metrics from mseb.evaluators import ( classification_evaluator, clustering_evaluator, retrieval_evaluator, segmentation_evaluator, ) print(f” mseb {mseb.__version__} | Python {sys.version.split()[0]} | numpy {np.__version__}”) print(“n MSEB is three layers, and a benchmark run walks down them:”) print(” types -> Sound, SoundEmbedding, Score, TaskMetadata: the shapes every task speaks”) print(” encoder -> MultiModalEncoder: the contract YOUR model implements”) print(” evaluators -> classification, clustering, retrieval, reranking, transcription, segmentation, …”) print(“n evaluator entry points we will drive:”) for module, cls in [(classification_evaluator, “ClassificationEvaluator”), (clustering_evaluator, “ClusteringEvaluator”), (retrieval_evaluator, “RetrievalEvaluator”), (segmentation_evaluator, “SegmentationEvaluator”)]: print(f” {module.__name__.split(‘.’)[-1]:28s} {cls}”) print(“n Everything below runs on CPU with no dataset download: we synthesise the audio.”) We install mseb and import the three layers that a benchmark run walks down. The types module holds the shapes every task speaks, Sound, SoundEmbedding, Score and TaskMetadata; the encoder module holds MultiModalEncoder, the contract our own model implements; and the evaluators package holds one module per task family. We import only the four evaluators this notebook drives, because the classification, clustering, retrieval, and segmentation modules depend on nothing heavier than NumPy and scikit-learn. In contrast, the reranking and transcription evaluators pull in Whisper and the task runner pulls in TensorFlow and apache-beam. Everything below therefore runs on a free CPU runtime with no dataset download and no accelerator. Copy CodeCopiedUse a different Browser SR = 16000 @section(“1. The type contract: Sound, SoundEmbedding, Score”) def type_contract(): t = np.arange(SR) / SR waveform = (0.5 * np.sin(2 * np.pi * 440 * t)).astype(np.float32) sound = types.Sound( waveform=waveform, context=types.SoundContextParams(id=”demo_000″, sample_rate=SR, length=len(waveform), language=”en_us”, text=”a 440 Hz tone”), ) print(f” Sound id={sound.context.id!r} {sound.waveform.shape} @ {sound.context.sample_rate} Hz” f” -> {sound.size_bytes:,} bytes”) embedding = types.SoundEmbedding( embedding=np.zeros((1, 16), dtype=np.float32), # (N, D): one utterance-level vector timestamps=np.array([[0.0, 1.0]], dtype=np.float32), # (M, 2): [start, end] in seconds context=sound.context, encoding_stats=types.EncodingStats(input_size_bytes=sound.size_bytes, embedding_size_bytes=16 * 4), ) print(f” SoundEmbedding embedding{embedding.embedding.shape} timestamps{embedding.timestamps.shape}” f” -> {embedding.size_bytes} bytes”) print(f” compression_ratio = {embedding.encoding_stats.compression_ratio:.5f}” f” ({1 / embedding.encoding_stats.compression_ratio:,.0f}x smaller than the audio)”) print(” N embeddings and M timestamps: M == N is frame-aligned, M == 1 is utterance-level.”) print(” `embedding` may also hold N strings instead of vectors – step 8 uses exactly that.”) score = types.Score(metric=”Accuracy”, description=”Overall classification accuracy”, value=0.875, min=0.0, max=1.0) print(f”n Score {score.metric}={score.value} in [{score.min}, {score.max}] :: {score.description}”) for bad, why in [(dict(metric=””, description=”d”, value=0.5, min=0.0, max=1.0), “empty metric name”), (dict(metric=”m”, description=”d”, value=0.5, min=1.0, max=0.0), “min > max”)]: try: types.Score(**bad) except Exception as e: print(f” rejected at construction ({why}): {type(e).__name__}: {e}”) return f”Sound {sound.size_bytes:,} B -> embedding {embedding.size_bytes} B” type_contract() We start with the type contract, because every other layer is expressed in it. A Sound carries a waveform, along with SoundContextParams, the identifier, sample rate, length, language, and optional transcript, which follow the audio through the whole pipeline. A SoundEmbedding carries an array of N embeddings and an array of M timestamp pairs, and the relation between N and M is the benchmark’s vocabulary: M equal to N means one vector per frame, while M equal to one means a single utterance-level vector, which is what our encoders produce. EncodingStats records the input and embedding sizes and exposes compression_ratio, here a thousandfold reduction from audio to vector. A Score is a metric name, a value and its bounds, and it validates itself at construction, rejecting an empty metric name or a minimum above its maximum, so a malformed number cannot reach a leaderboard. The embedding field also accepts N strings instead of N vectors, which is the door that step 8 walks through. Copy CodeCopiedUse a different Browser class EnergyEnvelopeEncoder(encoder_lib.MultiModalEncoder): “””Baseline: average energy in `n_bins` equal time slices. Loud/quiet, nothing about timbre.””” def __init__(self, n_bins: int = 16): super().__init__() self.n_bins = n_bins def _setup(self): self._ready = True # a real encoder loads weights here def _check_input_types(self, batch): for item in batch: if not isinstance(item, types.Sound): raise ValueError(f”{type(self).__name__} takes types.Sound, got {type(item).__name__}”) def _encode(self, batch) -> list[types.SoundEmbedding]: out = [] for sound in batch: slices = np.array_split(sound.waveform.astype(np.float32), self.n_bins) vec = np.array([[float(np.sqrt(np.mean(s ** 2) + 1e-12)) for s in slices]], dtype=np.float32) vec /= np.linalg.norm(vec) + 1e-9 out.append(types.SoundEmbedding( embedding=vec, timestamps=np.array([[0.0, sound.context.length / sound.context.sample_rate]], dtype=np.float32), context=sound.context)) return out class SpectralProfileEncoder(encoder_lib.MultiModalEncoder): “””Contender: mean log-magnitude spectrum pooled into `n_bands` bands. Describes timbre.””” def __init__(self, n_bands: int = 16, frame: int = 512): super().__init__() self.n_bands, self.frame = n_bands, frame def _setup(self): self._window = np.hanning(self.frame).astype(np.float32) def _check_input_types(self, batch): for item in batch: if not isinstance(item, types.Sound): raise ValueError(f”{type(self).__name__} takes types.Sound, got {type(item).__name__}”) def _encode(self, batch) -> list[types.SoundEmbedding]: out = [] for sound in batch: w = sound.waveform.astype(np.float32) n_frames =

