Skip to content
PreprintRTP-00002807Open AccessDOI 10.48550/arxiv.1910.11472

Learning Domain Invariant Representations for Child-Adult Classification\n from Speech

arXiv (Cornell University) · 2019 · Cornell University

0 views · 0 downloads

Abstract

Diagnostic procedures for ASD (autism spectrum disorder) involve\nsemi-naturalistic interactions between the child and a clinician. Computational\nmethods to analyze these sessions require an end-to-end speech and language\nprocessing pipeline that go from raw audio to clinically-meaningful behavioral\nfeatures. An important component of this pipeline is the ability to\nautomatically detect who is speaking when i.e., perform child-adult speaker\nclassification. This binary classification task is often confounded due to\nvariability associated with the participants' speech and background conditions.\nFurther, scarcity of training data often restricts direct application of\nconventional deep learning methods. In this work, we address two major sources\nof variability - age of the child and data source collection location - using\ndomain adversarial learning which does not require labeled target domain data.\nWe use two methods, generative adversarial training with inverted label loss\nand gradient reversal layer to learn speaker embeddings invariant to the above\nsources of variability, and analyze different conditions under which the\nproposed techniques improve over conventional learning methods. Using a large\ncorpus of ADOS-2 (autism diagnostic observation schedule, 2nd edition)\nsessions, we demonstrate upto 13.45% and 6.44% relative improvements over\nconventional learning methods.\n

Citations by source

  • openalex0

Counts differ by provider and are shown separately, never combined.