Compressing Streaming Neural Audio Encoders via Latent-Space Distillation

Imported from official source

AI Classified by Officially

Compressing Streaming Neural Audio Encoders via Latent-Space Distillation

AuthorsPrasanth Yadla‡, Mohammad Samragh Razlighi‡, Dongseong Hwang, Mingbin Xu, Yuanyuan Zhang, Chung-Cheng Chiu, Yongqiang Wang†**, Yuan Liu§**, Zhen Huang, Xiaodan Zhuang

System-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the foundation model through a tokenizer: an encoder that maps short windows of waveform onto the representation the language model reads. Because that model is sparsely activated under Instruction-Following Pruning, only a small subset of its experts occupies DRAM at any time, so the always-on tokenizer competes for the same memory, and its parameter count bears directly on power and latency. In this work we study how to compress such a tokenizer by distillation, taking as the supervision target neither the discrete token nor the output distribution but the pre-quantizer latent the model actually consumes—the last representation the two token interfaces share. We train only the student encoder to regress the teacher’s per-frame latent under a squared-error objective, with a single affine layer absorbing the teacher–student width mismatch. Because the target precedes both the quantizer and the language-model bridge, one recipe covers both token interfaces we support, and applies both to a tokenizer pretrained alone and to one jointly trained with a language model. At 2.8× compression the distilled student stays within 1.9% relative WER of its teacher on five of six teacher–student pairs without any fine-tuning, and improves on an independently trained tokenizer of identical capacity by 3.9% relative.

Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why

This is an extract. The publication continues at the source.

Read the original at the source: https://machinelearning.apple.com/research/latent-space-distillation

Officially imported this from Apple Machine Learning Research’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.

Provenance

Organization
Apple Machine Learning Research — imported from official source
Official source
https://machinelearning.apple.com/rss.xml RSS
Imported
September 24, 2026 16:00
Versions
1 recorded
Identity
latent-space-distillation

Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.