Learning Domain Invariant Representations for Child-Adult Classification\n from Speech
- Rimita Lahiri
- M. Kumar
- Somer Bishop
- Shrikanth Narayanan
- MKManoj Kumar
arXiv (Cornell University) · 2019 · Cornell University
0 views · 0 downloads
Abstract
Diagnostic procedures for ASD (autism spectrum disorder) involve\nsemi-naturalistic interactions between the child and a clinician. Computational\nmethods to analyze these sessions require an end-to-end speech and language\nprocessing pipeline that go from raw audio to clinically-meaningful behavioral\nfeatures. An important component of this pipeline is the ability to\nautomatically detect who is speaking when i.e., perform child-adult speaker\nclassification. This binary classification task is often confounded due to\nvariability associated with the participants' speech and background conditions.\nFurther, scarcity of training data often restricts direct application of\nconventional deep learning methods. In this work, we address two major sources\nof variability - age of the child and data source collection location - using\ndomain adversarial learning which does not require labeled target domain data.\nWe use two methods, generative adversarial training with inverted label loss\nand gradient reversal layer to learn speaker embeddings invariant to the above\nsources of variability, and analyze different conditions under which the\nproposed techniques improve over conventional learning methods. Using a large\ncorpus of ADOS-2 (autism diagnostic observation schedule, 2nd edition)\nsessions, we demonstrate upto 13.45% and 6.44% relative improvements over\nconventional learning methods.\n
