Language Discrimination Improves Linguistic Learning in Multilingual Speech Models

Apple Machine Learning Research Version 1 original current

Imported from official source

Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. We show that strengthening the model’s ability to discriminate languages during pretraining reduces and, on some measures, closes this multilingual gap on continuous phonetic and higher-level linguistic measures, while preserving substantial cross-language sharing. Using a controlled English/French HuBERT setting, we test two interventions which strengthen language discrimination: an auxiliary language…

This version

Version
1 of 1
Recorded
October 02, 2026 15:00
Change
Initial
Content hash
de811f9f337fb9b8c00f1fc1872d0b12
All versions
Revision history

Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.