Audio Fingerprinting Basics in Rust
Audio fingerprinting lets us answer questions like:
“Does this noisy recording match anything in my catalog?”
without storing or comparing raw waveforms.
In this article we’ll:
- Load WAV files with
symphonia. - Compute magnitude spectra using
rustfft. - Build a very naive hash just to get the pipeline in place.
Loading audio with Symphonia
We start with 16-bit mono WAV files to keep things simple:
rust
use symphonia::core::{
audio::{AudioBufferRef, Signal},
codecs::CODEC_TYPE_NULL,
formats::FormatOptions,
io::MediaSourceStream,
meta::MetadataOptions,
probe::Hint,
};
use std::{fs::File, path::Path};
pub fn load_samples(path: &Path) -> anyhow::Result<Vec<f32>> {
let file = File::open(path)?;
let mss = MediaSourceStream::new(Box::new(file), Default::default());
let mut hint = Hint::new();
hint.with_extension("wav");
let probed = symphonia::default::get_probe().format(
&hint,
mss,
&FormatOptions::default(),
&MetadataOptions::default(),
)?;
let mut format = probed.format;
let track = format
.tracks()
.iter()
.find(|t| t.codec_params.codec != CODEC_TYPE_NULL)
.ok_or_else(|| anyhow::anyhow!("no audio track"))?;
let mut decoder = symphonia::default::get_codecs().make(&track.codec_params, &Default::default())?;
let mut samples = Vec::new();
loop {
let packet = match format.next_packet() {
Ok(packet) => packet,
Err(_) => break,
};
let decoded = decoder.decode(&packet)?;
if let AudioBufferRef::S16(buf) = decoded {
for frame in buf.chan(0) {
samples.push(*frame as f32 / i16::MAX as f32);
}
}
}
Ok(samples)
}From samples to spectra
We’ll use a fixed-size FFT (e.g. 4096 points). Each window produces one spectrum:
rust
use rustfft::{FftPlanner, num_complex::Complex32};
pub fn spectrogram(samples: &[f32], fft_size: usize, hop: usize) -> Vec<Vec<f32>> {
let mut planner = FftPlanner::new();
let fft = planner.plan_fft_forward(fft_size);
let mut out = Vec::new();
let mut buf = vec![Complex32::ZERO; fft_size];
let mut i = 0;
while i + fft_size <= samples.len() {
for (dst, &s) in buf.iter_mut().zip(&samples[i..i + fft_size]) {
*dst = Complex32::new(s, 0.0);
}
fft.process(&mut buf);
let magnitudes = buf
.iter()
.map(|c| c.norm())
.collect::<Vec<_>>();
out.push(magnitudes);
i += hop;
}
out
}A toy fingerprint
Real systems use clever peak picking and hashing. For now we’ll do something intentionally simple:
- Divide each spectrum into
Nbands. - For each band, record the index of the maximum bin.
- Concatenate those indices into a small vector.
rust
pub fn toy_fingerprint(spec: &[Vec<f32>], bands: usize) -> Vec<u16> {
let bins = spec[0].len();
let band_width = bins / bands;
spec.iter()
.map(|frame| {
let mut fp = 0u16;
for b in 0..bands {
let start = b * band_width;
let end = ((b + 1) * band_width).min(bins);
let (idx, _) = frame[start..end]
.iter()
.enumerate()
.max_by(|a, b| a.1.partial_cmp(b.1).unwrap())
.unwrap();
// 4 bits per band (toy!), shift and or
fp <<= 4;
fp |= (idx as u16 & 0x0F);
}
fp
})
.collect()
}This is deliberately bad
Don’t ship this to production 🙂
The goal is to have something concrete we can later replace with a better, well-researched scheme.
Sanity-checking in the terminal
Once everything compiles, we run a quick smoke test:
cargo run --release --bin toy-fp ./data/snare.wav
frames: 214
fingerprints: 214
first 10: [0x3AF2, 0x39F1, 0x39F1, 0x39E1, 0x39E1, 0x39D1, 0x39C1, 0x39C1, 0x39C1, 0x39C1]If two different takes of the same audio produce a similar sequence of toy fingerprints, we’re on the right track.
In later posts we’ll:
- Replace the toy hash with a constellation hash.
- Add a database backend for large catalogs.
- Benchmark the system on real-world noisy recordings.