--- license: mit language: en library_name: transformers tags: - audio-classification - speech-commands - cnn - audio datasets: - speech-commands metrics: - accuracy model-index: - name: tone results: - task: type: audio-classification dataset: name: Speech Commands type: speech-commands split: test metrics: - type: accuracy value: 0.9116 --- # tone A small convolutional classifier over spoken commands: one second of 16 kHz audio in, one of ten words out. Input is a 64-bin log-mel spectrogram (25 ms window, 10 ms hop), classified by two convolution layers (1 to 32 channels, 32 to 64, 3 by 3 kernels) with max pooling, adaptive pooling to 8 by 8, a 128-unit dense layer with dropout, and a 10-way head. Trained for eight epochs with Adam at learning rate 1e-3, batch size 128, seed 0. Test accuracy 0.9116 on the canonical `testing_list.txt` split of 4,074 clips. CPU training. The ten classes, in order, are yes, no, up, down, left, right, on, off, stop and go. No transformer here. The `PreTrainedModel` wrapper exists only so the weights serialize as `config.json` plus `model.safetensors` and load through `AutoModel` (including `trust_remote_code` via `auto_map`), the same arrangement as `wear` and `pole`. There is no attention and no pretraining: it classifies one-second command clips and nothing else. ## Usage ```python import numpy as np from transformers import AutoModel from modeling_tone import ToneCNN # registers the architecture model = AutoModel.from_pretrained("harpertoken/tone") model.eval() waveform = np.zeros(16000, dtype=np.float32) print(model.predict_label(waveform)) ``` `predict_label` takes 16,000 float samples at 16 kHz and returns the command word. Clips shorter than a second are zero-padded and longer ones truncated, matching training exactly. Needs `torch`, `torchaudio` for the mel transform, and `transformers`. ## Predictions ![First test clips of three words, with model predictions on their log-mel inputs](tone_grid.png) The first test clip of three words in `testing_list.txt` order, shown as the 64-bin log-mel spectrograms the model actually classifies, with the predicted word, the true word, and the file. All three are correct, which is what the selection rule produced rather than a curated set; overall test accuracy is 0.9116, so roughly one in twelve predictions elsewhere is wrong. ## Training Canonical Speech Commands v0.02 via torchaudio, restricted to the ten command words (34,472 train clips after the test split is removed, 4,074 test). Torchaudio's own loader needed torchcodec, whose native library would not load on the training machine, so waveforms were read with scipy and the mel transform uses torchaudio's pure-torch `MelSpectrogram`. An earlier hand-rolled mel filterbank trained to chance (loss stuck at ln 10), which is why the dependency is there. Per-epoch test accuracy ran 0.7239, 0.7955, 0.8441, 0.8675, 0.8893, 0.9001, 0.9043, 0.9116. The wrapper was checked for exact equivalence: identical predictions on 500 test clips in eval mode, after catching that dropout made train-mode outputs differ. No dataset is published alongside this model. Speech Commands already exists canonically and re-hosting it would add a duplicate, so the card cites the source instead. ## Limitations One-second isolated English command words recorded in ordinary rooms. Anything else, continuous speech, other languages, heavy noise, is out of scope and the scores will be meaningless. 0.9116 is an ordinary result for this setup, not a benchmark claim.