Wav2Keyword

Wav2Keyword is keyword spotting(KWS) based on Wav2Vec 2.0. This model shows state-of-the-art in Speech commands dataset V1 and V2.

Preparation

PyTorch version >= 1.5.0
Python version >= 3.6
To install fairseq and develop locally:

git clone https://github.com/pytorch/fairseq
cd fairseq
pip install --editable ./

The pretrained Wav2Vec 2.0 base model not finetuned (https://dl.fbaipublicfiles.com/fairseq/wav2vec/wav2vec_small.pt) must exist in the model directory.

Training

python downstream_kws.py [pretrained model path] [dataset path] [saving model path]

And you can benchmark this model number of samples per each class.

python downstream_kws_benchmark.py [pretrained model path] [dataset path] [saving model path]

Model Architecture

With Wav2Vec 2.0 as the backbone, speech representation, output of transformer, is transferred to the structure of the model for speech commands recognition.

Performance

Accuracy of baseline models and proposed Wav2Keyword model on Google Speech Command Datasets V1 and V2 considering their 12 shared commands.

Dataset	Accuracy (%)
Dataset V1	97.9
Dataset V2	98.5

Accuracy of baseline models and proposed Wav2Keyword model on Google Speech Command Dataset V2 with its 22 commands

Dataset	Accuracy (%)
Dataset V12	97.8

Reference

[0] https://github.com/pytorch/fairseq/tree/master/examples/wav2vec

[1] https://arxiv.org/abs/1804.03209

[2] https://paperswithcode.com/sota/keyword-spotting-on-google-speech-commands

Citation

This paper has been submitted. If accept, will add.

@ARTICLE{9427206,  
  author={Seo, Deokjin and Oh, Heung-Seon and Jung, Yuchul},  
  journal={IEEE Access},   
  title={Wav2KWS: Transfer Learning from Speech Representations for Keyword Spotting},   
  year={2021},  
  pages={1-1},  
  doi={10.1109/ACCESS.2021.3078715}
}

Name		Name	Last commit message	Last commit date
Latest commit History 15 Commits
build		build
dist		dist
docs		docs
examples		examples
fairseq.egg-info		fairseq.egg-info
fairseq		fairseq
fairseq_cli		fairseq_cli
scripts		scripts
tests		tests
.gitattributes		.gitattributes
CODE_OF_CONDUCT.md		CODE_OF_CONDUCT.md
CONTRIBUTING.md		CONTRIBUTING.md
LICENSE		LICENSE
README.md		README.md
downstream_kws.py		downstream_kws.py
downstream_kws_benchmark.py		downstream_kws_benchmark.py
hubconf.py		hubconf.py
infer.py		infer.py
infer_single.py		infer_single.py
pyproject.toml		pyproject.toml
recognize.py		recognize.py
setup.py		setup.py
train.py		train.py
trans.txt		trans.txt
wav2vec2.CPU.Dockerfile		wav2vec2.CPU.Dockerfile

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

Wav2Keyword

Preparation

Training

Model Architecture

Performance

Reference

Citation

About

Releases

Packages

Contributors 3

Languages

License

dobby-seo/Wav2Keyword

Folders and files

Latest commit

History

Repository files navigation

Wav2Keyword

Preparation

Training

Model Architecture

Performance

Reference

Citation

About

Topics

Resources

License

Code of conduct

Stars

Watchers

Forks

Releases

Packages 0

Contributors 3

Languages

Packages