Transcript poster

Poster for MLSP 2015
Retrieving Sounds by Vocal Imitation Recognition
Yichi Zhang & Zhiyao Duan
AIR Lab, Department of Electrical and Computer Engineering, University of Rochester
Introduction
Q: How to search for a sound that matches the concept in your head?
A: Current ways: through its name or other semantic labels.
Q: What if you don’t remember its name, or what you are looking for
simply doesn’t have a semantic meaning?
A: Imitate the concept with your voice!
 Dog barking sound: infantile bark
 Synthesized sound:
Classification & Retrieval:
Use multi-class Support Vector Machine (SVM) to generate probability
output for concept retrieval. (Limitation: close-set scenario)
Majority Vote: Label 2
Label 2 Label 3 Label 1 Label 2 Label 2 Label 4
nth
Imitation: Patch 1 Patch 2 Patch 3 Patch 4 Patch 5 Patch 6
Feature Extraction
threat bark
(a) the 1st hidden layer
Est. Prob.
Vector 7
Classification
Sound Retrieval
Table 1. Description of the VocalSketch v1.0.4 dataset
A big challenge in vocal imitation recognition is feature extraction.
 People tend to imitate different aspects for different recordings:
car horn:
[
] cat:
[
] guitar note:
[
]
 Even for the same recording, different people imitate differently:
car horn 1:
car horn 2:
car horn 3:
Category
# classes
Acoustic
instruments
40
Commercial
synthesizers
Everyday
120
Single synthesizer
40
Automatic Feature Learning
f
40
Category
Feature Extraction:
Use Stacked Auto-encoder (SAE) to learn features from training
patches automatically.
(a) autoencoder
xN
zM
b
'
y3
yN
x3
(b) stacked
auto-encoder
xN
z1M
b2
z22
y2
...
z12
w2
z2K
b3
y3
Hidden
layer
Output
layer
Input
layer
Acoustic
instruments
Commercial
synthesizers
Everyday
Single synthesizer
# classes
Proposed
MFCC
Accuracy
MRR
Accuracy
MRR
17
23.61%
0.4259
21.94%
0.3789
13
20.00%
0.3577
12.69%
0.2960
48
10.71%
0.2666
10.00%
0.2368
40
12.00%
0.2732
6.25%
0.2188
yN
Acknowledgements
b1
b
Input
layer
x2
z11
...
y2
y1
w1
...
t
x1
...
x3
z2
y1
...
525 ms
x2
z1
...
Patch m
...
26 ms
…


Patch n
w'
Sound Concepts
Orchestral instruments playing a single note with the
pitch C (in an appropriate octave chosen for each
instrument)
Various recordings from Apple’s Logic Pro music
production suite
A wide variety of acoustic events in everyday life
Recordings from a single 15-parameter subtractive
synthesizer playing a note with the pitch C (octave
varies depending on the parameter settings)
Table 2. Recording-level 10-fold cross validation results.
Sigmoid
w
(b) the 2nd hidden layer
Experimental Results
Challenges
x1
50 Hz
Patch N
Label Prob. Vector
Pre-processing:
Convert imitation audio into spectrogram by Constant-Q Transform
(CQT), then segment it into overlapping patches.
3200 Hz
…
Est. Prob. Est. Prob. Est. Prob. Est. Prob. Est. Prob. Est. Prob.
Vector 1 Vector 2 Vector 3 Vector 4 Vector 5 Vector 6
Proposed System
Pre-processing
Label 2
1st hidden 2nd hidden Output
layer
layer
layer
We thank Mark Cartwright and Bryan Pardo for generously providing
us with the VocalSketch Data Set v1.0.4.