Skip to content
PreprintRTP-00002804Open AccessDOI 10.48550/arxiv.1910.11416

A study of semi-supervised speaker diarization system using gan mixture\n model

arXiv (Cornell University) · 2019 · Cornell University

0 views · 0 downloads

Abstract

We propose a new speaker diarization system based on a recently introduced\nunsupervised clustering technique namely, generative adversarial network\nmixture model (GANMM). The proposed system uses x-vectors as front-end\nrepresentation. Spectral embedding is used for dimensionality reduction\nfollowed by k-means initialization during GANMM pre-training. GANMM performs\nunsupervised speaker clustering by efficiently capturing complex data\ndistributions. Experimental results on the AMI meeting corpus show that the\nproposed semi-supervised diarization system matches or exceeds the performance\nof competitive baselines. On an evaluation set containing fifty sessions with\nvarying durations, the best achieved average diarization error rate (DER) is\n17.11%, a relative improvement of 33% over the information bottleneck baseline\nand comparable to xvector baseline.\n

Citations by source

  • openalex2

Counts differ by provider and are shown separately, never combined.