GitHub - FurKan7/speaker-recognition at 30f4b0eca03c3c19525585e9005ee2f27eb574c0

Name	Name	Last commit message	Last commit date
Latest commit History 268 Commits
doc	doc
log	log
src	src
.gitattributes	.gitattributes
.gitignore	.gitignore
README.md	README.md
complete-report.pdf	complete-report.pdf
demo.avi	demo.avi
presentation.pdf	presentation.pdf

Name

Last commit message

Last commit date

About

This is a Speaker Recognition system with GUI. At first, it served as an SRT project for the course Signal Processing (2013Fall) in Tsinghua University. But we did find it pretty useful!

For more details of this project, please see:

Our presentation slides
Our complete report

Dependencies

Installation / Compilation

(Optional) Bob:

See here for instructions on bob core library installation.

Bob python bindings are available on PyPI. You may need to install bob packages in the following order:

bob.extension
bob.blitz (require blitz++)
bob.core
bob.sp
bob.ap

Note: We also have MFCC feature implemented on our own, which will be used as a fallback when bob is unavailable. But it's not so efficient as the C implementation in bob.

(Optional) GMM

Run make -C src/gmm to compile our fast gmm implementation. Require gcc >= 4.7.

It will be used as default, if successfully compiled.

Algorithms Used

Voice Activity Detection(VAD):

Long-Term Spectral Divergence (LTSD)

Feature:

Mel-Frequency Cepstral Coefficient (MFCC)
Linear Predictive Coding (LPC)

Model:

Gaussian Mixture Model (GMM)
Universal Background Model (UBM)
Continuous Restricted Boltzman Machine (CRBM)
Joint Factor Analysis (JFA)

GUI Demo

Our GUI not only has basic functionality for recording, enrollment, training and testing, but also has a visualization of real-time speaker recognition:

You should understand that real-time speaker recognition is extremely hard, because we only use corpus of about 1 second length to identify the speaker. Therefore the real-time system doesn't work very perfect. You can See our demo video (in Chinese).

Command Line Tools

usage: speaker-recognition.py [-h] -t TASK -i INPUT -m MODEL

Speaker Recognition Command Line Tool

optional arguments:
  -h, --help            show this help message and exit
  -t TASK, --task TASK  Task to do. Either "enroll" or "predict"
  -i INPUT, --input INPUT
                        Input Files(to predict) or Directories(to enroll)
  -m MODEL, --model MODEL
                        Model file to save(in enroll) or use(in predict)

Wav files in each input directory will be labeled as the basename of the directory.
Note that wildcard inputs should be *quoted*, and they will be sent to glob module.

Examples:
    Train:
    ./speaker-recognition.py -t enroll -i "/tmp/person* ./mary" -m model.out

    Predict:
    ./speaker-recognition.py -t predict -i "./*.wav" -m model.out

Languages

C++ 64.1%

Python 26.0%

MATLAB 5.5%

Makefile 4.3%

Shell 0.1%

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

About

Dependencies

Installation / Compilation

(Optional) Bob:

(Optional) GMM

Algorithms Used

GUI Demo

Command Line Tools

About

Releases

Packages

Languages

FurKan7/speaker-recognition

Folders and files

Latest commit

History

Repository files navigation

About

Dependencies

Installation / Compilation

(Optional) Bob:

(Optional) GMM

Algorithms Used

GUI Demo

Command Line Tools

About

Resources

Stars

Watchers

Forks

Releases

Packages 0

Languages

Packages