Transcript poster
Poster for MLSP 2015 Retrieving Sounds by Vocal Imitation Recognition Yichi Zhang & Zhiyao Duan AIR Lab, Department of Electrical and Computer Engineering, University of Rochester Introduction Q: How to search for a sound that matches the concept in your head? A: Current ways: through its name or other semantic labels. Q: What if you don’t remember its name, or what you are looking for simply doesn’t have a semantic meaning? A: Imitate the concept with your voice! Dog barking sound: infantile bark Synthesized sound: Classification & Retrieval: Use multi-class Support Vector Machine (SVM) to generate probability output for concept retrieval. (Limitation: close-set scenario) Majority Vote: Label 2 Label 2 Label 3 Label 1 Label 2 Label 2 Label 4 nth Imitation: Patch 1 Patch 2 Patch 3 Patch 4 Patch 5 Patch 6 Feature Extraction threat bark (a) the 1st hidden layer Est. Prob. Vector 7 Classification Sound Retrieval Table 1. Description of the VocalSketch v1.0.4 dataset A big challenge in vocal imitation recognition is feature extraction. People tend to imitate different aspects for different recordings: car horn: [ ] cat: [ ] guitar note: [ ] Even for the same recording, different people imitate differently: car horn 1: car horn 2: car horn 3: Category # classes Acoustic instruments 40 Commercial synthesizers Everyday 120 Single synthesizer 40 Automatic Feature Learning f 40 Category Feature Extraction: Use Stacked Auto-encoder (SAE) to learn features from training patches automatically. (a) autoencoder xN zM b ' y3 yN x3 (b) stacked auto-encoder xN z1M b2 z22 y2 ... z12 w2 z2K b3 y3 Hidden layer Output layer Input layer Acoustic instruments Commercial synthesizers Everyday Single synthesizer # classes Proposed MFCC Accuracy MRR Accuracy MRR 17 23.61% 0.4259 21.94% 0.3789 13 20.00% 0.3577 12.69% 0.2960 48 10.71% 0.2666 10.00% 0.2368 40 12.00% 0.2732 6.25% 0.2188 yN Acknowledgements b1 b Input layer x2 z11 ... y2 y1 w1 ... t x1 ... x3 z2 y1 ... 525 ms x2 z1 ... Patch m ... 26 ms … Patch n w' Sound Concepts Orchestral instruments playing a single note with the pitch C (in an appropriate octave chosen for each instrument) Various recordings from Apple’s Logic Pro music production suite A wide variety of acoustic events in everyday life Recordings from a single 15-parameter subtractive synthesizer playing a note with the pitch C (octave varies depending on the parameter settings) Table 2. Recording-level 10-fold cross validation results. Sigmoid w (b) the 2nd hidden layer Experimental Results Challenges x1 50 Hz Patch N Label Prob. Vector Pre-processing: Convert imitation audio into spectrogram by Constant-Q Transform (CQT), then segment it into overlapping patches. 3200 Hz … Est. Prob. Est. Prob. Est. Prob. Est. Prob. Est. Prob. Est. Prob. Vector 1 Vector 2 Vector 3 Vector 4 Vector 5 Vector 6 Proposed System Pre-processing Label 2 1st hidden 2nd hidden Output layer layer layer We thank Mark Cartwright and Bryan Pardo for generously providing us with the VocalSketch Data Set v1.0.4.