Skip to content
PreprintRTP-00002808Open AccessDOI 10.48550/arxiv.1910.11398

Speaker diarization using latent space clustering in generative\n adversarial network

arXiv (Cornell University) · 2019 · Cornell University

2 views · 0 downloads

Abstract

In this work, we propose deep latent space clustering for speaker diarization\nusing generative adversarial network (GAN) backprojection with the help of an\nencoder network. The proposed diarization system is trained jointly with GAN\nloss, latent variable recovery loss, and a clustering-specific loss. It uses\nx-vector speaker embeddings at the input, while the latent variables are\nsampled from a combination of continuous random variables and discrete one-hot\nencoded variables using the original speaker labels. We benchmark our proposed\nsystem on the AMI meeting corpus, and two child-clinician interaction corpora\n(ADOS and BOSCC) from the autism diagnosis domain. ADOS and BOSCC contain\ndiagnostic and treatment outcome sessions respectively obtained in clinical\nsettings for verbal children and adolescents with autism. Experimental results\nshow that our proposed system significantly outperform the state-of-the-art\nx-vector based diarization system on these databases. Further, we perform\nembedding fusion with x-vectors to achieve a relative DER improvement of 31%,\n36% and 49% on AMI eval, ADOS and BOSCC corpora respectively, when compared to\nthe x-vector baseline using oracle speech segmentation.\n

Citations by source

  • openalex0

Counts differ by provider and are shown separately, never combined.