跳到主要內容
AI News HubLIVE
站內改寫6 分鐘閱讀

待翻譯:A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:A comprehensive coding tutorial on Google Research's Massive Sound Embedding Benchmark (MSEB), demonstrating how to implement custom sound encoders, drive classification, clustering, retrieval, and segmentation evaluators, and analyze multi-task benchmark performance. The post A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation appeared first on MarkTechPost.

來源MarkTechPost作者: Sana Hassan
待翻譯:A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation
報告錯誤

更正渠道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

In this tutorial, we work with MSEB, the Massive Sound Embedding Benchmark from Google Research, and approach it from the perspective of what a leaderboard number actually means: the evaluator surface. We install the package and map its three layers, then write two deliberately different encoders against the framework’s own abstract base class: one that measures loudness over time and one that measures timbre, and encode a small synthetic corpus we generate in the notebook so nothing has to be downloaded. We drive the classification, clustering, retrieval, and segmentation evaluators over those embeddings, call the metric functions directly to see what each one rewards, and finish by assembling the TaskMetadata a real submission carries. The result is a comparison in which the two encoders trade places depending on which evaluator is asked, which is the argument for a multi-task benchmark made in numbers rather than in prose. Copy CodeCopiedUse a different Browser import os import sys import json import math import traceback import subprocess import numpy as np RESULTS = {} BENCH = {} def banner(title): print("\n" + "=" * 78) print(title) print("=" * 78) def section(name): def wrap(fn): def run(*a, kw): banner(name) try: out = fn(*a, kw) RESULTS[name] = out if isinstance(out, str) else "ok" return out except Exception as e: RESULTS[name] = f"SKIPPED / FAILED -> {type(e).name}: {e}" print(f"\n[!] {name} did not complete: {type(e).name}: {e}") traceback.print_exc(limit=3) return None return run return wrap banner("0. Install MSEB and map the three layers we will use") subprocess.run([sys.executable, "-m", "pip", "install", "-q", "mseb==0.1.0"], check=True) import mseb from mseb import types, encoder as encoder_lib, evaluator as evaluator_lib, metrics from mseb.evaluators import ( classification_evaluator, clustering_evaluator, retrieval_evaluator, segmentation_evaluator, ) print(f" mseb {mseb.version} | Python {sys.version.split()[0]} | numpy {np.version}") print("\n MSEB is three layers, and a benchmark run walks down them:") print(" types -> Sound, SoundEmbedding, Score, TaskMetadata: the shapes every task speaks") print(" encoder -> MultiModalEncoder: the contract YOUR model implements") print(" evaluators -> classification, clustering, retrieval, reranking, transcription, segmentation, ...") print("\n evaluator entry points we will drive:") for module, cls in [(classification_evaluator, "ClassificationEvaluator"), (clustering_evaluator, "ClusteringEvaluator"), (retrieval_evaluator, "RetrievalEvaluator"), (segmentation_evaluator, "SegmentationEvaluator")]: print(f" {module.name.split('.')[-1]:28s} {cls}") print("\n Everything below runs on CPU with no dataset download: we synthesise the audio.") We install mseb and import the three layers that a benchmark run walks down. The types module holds the shapes every task speaks, Sound, SoundEmbedding, Score and TaskMetadata; the encoder module holds MultiModalEncoder, the contract our own model implements; and the evaluators package holds one module per task family. We import only the four evaluators this notebook drives, because the classification, clustering, retrieval, and segmentation modules depend on nothing heavier than NumPy and scikit-learn. In contrast, the reranking and transcription evaluators pull in Whisper and the task runner pulls in TensorFlow and apache-beam. Everything below therefore runs on a free CPU runtime with no dataset download and no accelerator. Copy CodeCopiedUse a different Browser SR = 16000 @section("1. The type contract: Sound, SoundEmbedding, Score") def type_contract(): t = np.arange(SR) / SR waveform = (0.5 * np.sin(2 * np.pi * 440 * t)).astype(np.float32) sound = types.Sound( waveform=waveform, context=types.SoundContextParams(id="demo_000", sample_rate=SR, length=len(waveform), language="en_us", text="a 440 Hz tone"), ) print(f" Sound id={sound.context.id!r} {sound.waveform.shape} @ {sound.context.sample_rate} Hz" f" -> {sound.size_bytes:,} bytes") embedding = types.SoundEmbedding( embedding=np.zeros((1, 16), dtype=np.float32), # (N, D): one utterance-level vector timestamps=np.array([[0.0, 1.0]], dtype=np.float32), # (M, 2): [start, end] in seconds context=sound.context, encoding_stats=types.EncodingStats(input_size_bytes=sound.size_bytes, embedding_size_bytes=16 * 4), ) print(f" SoundEmbedding embedding{embedding.embedding.shape} timestamps{embedding.timestamps.shape}" f" -> {embedding.size_bytes} bytes") print(f" compression_ratio = {embedding.encoding_stats.compression_ratio:.5f}" f" ({1 / embedding.encoding_stats.compression_ratio:,.0f}x smaller than the audio)") print(" N embeddings and M timestamps: M == N is frame-aligned, M == 1 is utterance-level.") print(" embedding may also hold N strings instead of vectors - step 8 uses exactly that.") score = types.Score(metric="Accuracy", description="Overall classification accuracy", value=0.875, min=0.0, max=1.0) print(f"\n Score {score.metric}={score.value} in [{score.min}, {score.max}] :: {score.description}") for bad, why in [(dict(metric="", description="d", value=0.5, min=0.0, max=1.0), "empty metric name"), (dict(metric="m", description="d", value=0.5, min=1.0, max=0.0), "min > max")]: try: types.Score(bad) except Exception as e: print(f" rejected at construction ({why}): {type(e).name}: {e}") return f"Sound {sound.size_bytes:,} B -> embedding {embedding.size_bytes} B" type_contract() We start with the type contract, because every other layer is expressed in it. A Sound carries a waveform, along with SoundContextParams, the identifier, sample rate, length, language, and optional transcript, which follow the audio through the whole pipeline. A SoundEmbedding carries an array of N embeddings and an array of M timestamp pairs, and the relation between N and M is the benchmark’s vocabulary: M equal to N means one vector per frame, while M equal to one means a single utterance-level vector, which is what our encoders produce. EncodingStats records the input and embedding sizes and exposes compression_ratio, here a thousandfold reduction from audio to vector. A Score is a metric name, a value and its bounds, and it validates itself at construction, rejecting an empty metric name or a minimum above its maximum, so a malformed number cannot reach a leaderboard. The embedding field also accepts N strings instead of N vectors, which is the door that step 8 walks through. Copy CodeCopiedUse a different Browser class EnergyEnvelopeEncoder(encoder_lib.MultiModalEncoder): """Baseline: average energy in n_bins equal time slices. Loud/quiet, nothing about timbre.""" def init(self, n_bins: int = 16): super().init() self.n_bins = n_bins def _setup(self): self._ready = True # a real encoder loads weights here def _check_input_types(self, batch): for item in batch: if not isinstance(item, types.Sound): raise ValueError(f"{type(self).name} takes types.Sound, got {type(item).name}") def _encode(self, batch) -> list[types.SoundEmbedding]: out = [] for sound in batch: slices = np.array_split(sound.waveform.astype(np.float32), self.n_bins) vec = np.array([[float(np.sqrt(np.mean(s 2) + 1e-12)) for s in slices]], dtype=np.float32) vec /= np.linalg.norm(vec) + 1e-9 out.append(types.SoundEmbedding( embedding=vec, timestamps=np.array([[0.0, sound.context.length / sound.context.sample_rate]], dtype=np.float32), context=sound.context)) return out class SpectralProfileEncoder(encoder_lib.MultiModalEncoder): """Contender: mean log-magnitude spectrum pooled into n_bands bands. Describes timbre.""" def init(self, n_bands: int = 16, frame: int = 512): super().init() self.n_bands, self.frame = n_bands, frame def _setup(self): self._window = np.hanning(self.frame).astype(np.float32) def _check_input_types(self, batch): for item in batch: if not isinstance(item, types.Sound): raise ValueError(f"{type(self).name} takes types.Sound, got {type(item).name}") def _encode(self, batch) -> list[types.SoundEmbedding]: out = [] for sound in batch: w = sound.waveform.astype(np.float32) n_frames = max(1, len(w) // self.frame) spectra = [np.abs(np.fft.rfft(w[i * self.frame:(i + 1) * self.frame] * self._window)) for i in range(n_frames)] mean_spectrum = np.log1p(np.mean(spectra, axis=0)) vec = np.array([[float(b.mean()) for b in np.array_split(mean_spectrum, self.n_bands)]], dtype=np.float32) vec /= np.linalg.norm(vec) + 1e-9 out.append(types.SoundEmbedding( embedding=vec, timestamps=np.array([[0.0, sound.context.length / sound.context.sample_rate]], dtype=np.float32), context=sound.context)) return out @section("2. The encoder contract: three methods, and the framework does the rest") def encoder_contract(): print(" MultiModalEncoder abstract methods a subclass must implement:") for name in sorted(encoder_lib.MultiModalEncoder.abstractmethods): print(f" {name}") print(" final (framework-owned, do not override): setup(), encode()") t = np.arange(SR) / SR fade = np.exp(-2.5 * t).astype(np.float32) # a decaying note, so the envelope is not flat sound = types.Sound(waveform=(0.5 * fade * np.sin(2 * np.pi * 440 * t)).astype(np.float32), context=types.SoundContextParams(id="demo_000", sample_rate=SR, length=SR)) for enc in (EnergyEnvelopeEncoder(), SpectralProfileEncoder()): enc.setup() emb = enc.encode([sound])[0] stats = emb.encoding_stats # attached by encode(), not by our code print(f"\n {type(enc).name:24s} -> {emb.embedding.shape} {emb.embedding.dtype}" f" output_type={enc.output_type().name}") print(f" {'':24s} EncodingStats(input={stats.input_size_bytes:,} B, " f"embedding={stats.embedding_size_bytes} B, flops={stats.flops})") print(f" {'':24s} first 6 dims: {np.round(emb.embedding[0][:6], 3)}") print("\n The envelope encoder sees the note decay; the spectral encoder sees one peak at 440 Hz.") try: EnergyEnvelopeEncoder().encode(["not a Sound"]) except ValueError as e: print(f"\n wrong input type is caught by _check_input_types: {e}") return "two encoders satisfying MultiModalEncoder" encoder_contract() We write two encoders by subclassing MultiModalEncoder, whose abstract methods are exactly three: _setup loads whatever the model needs, _check_input_types rejects anything that is not a Sound, and _encode turns a batch into SoundEmbedding objects. The framework owns setup and encode, and encode is what attaches EncodingStats to every result, so our code never fills that in by hand. EnergyEnvelopeEncoder averages energy in sixteen equal time slices and therefore describes only how loudness moves; SpectralProfileEncoder pools the mean log-magnitude spectrum into sixteen bands and therefore describes timbre. Both L2-normalise their output so a dot product is a cosine. Encoding one decaying note through each shows the difference immediately: the envelope encoder sees the decay, and the spectral encoder sees a single peak at 440 Hz. Copy CodeCopiedUse a different Browser CLASSES = ["tone", "chirp", "noise"] N_PER_CLASS = 12 def synthesize(kind: str, index: int, take: int) -> types.Sound: """One second of audio. take 0 is the document, take 1 is a noisier recording of the SAME clip. Two cues are deliberately separated: the spectrum says which class it is, and the amplitude envelope - drawn per item, independent of class - says which item it is. """ item = np.random.default_rng(1000 + CLASSES.index(kind) * 100 + index) control = 0.25 + 0.75 * item.random(8) envelope = np.interp(np.linspace(0, 7, SR), np.arange(8), control).astype(np.float32) t = np.arange(SR) / SR if kind == "tone": w = np.sin(2 * np.pi * (380 + 80 * item.random()) * t) elif kind == "chirp": f0, f1 = 200 + 50 * item.random(), 3200 + 400 * item.random() w = np.sin(2 * np.pi * (f0 * t + 0.5 * (f1 - f0) * t 2)) else: w = item.standard_normal(SR) w /= np.sqrt(np.mean(w 2)) + 1e-9 # unit RMS: the envelope is the only loudness cue take_rng = np.random.default_rng(50_00 [truncated for AI cost control]

展開要點與分析

文章情報

工程師進階

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • A comprehensive coding tutorial on Google Research's Massive Sound Embedding Benchmark (MSEB), demonstrating how to implement custom sound encoders, drive classification, clusteri…

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。