Eeshan Gunesh Dhekane

Alumni

Publications

Scaling Properties of Continuous Diffusion Spoken Language Models

Jason Ramapuram

Eeshan Gunesh Dhekane

Amitis Shidani

Dan Busbridge

Bogdan Mazoure

Zijin Gu

Russ Webb

Tatiana Likhomanenko

Navdeep Jaitly

Speech-only spoken language models (SLMs) lag behind text and text-speech models in performance, with recent discrete autoregressive (AR) SL… (see more)Ms indicating significant computational and data demands to match text models. Since discretizing continuous speech for AR creates bottlenecks, we explore whether continuous diffusion (CD) SLM is more viable. To quantify the SLMs linguistic quality, we introduce the phoneme Jensen-Shannon divergence (pJSD) metric. Our analysis reveals CD SLMs, mirroring AR behavior, exhibit scaling laws for validation loss and pJSD, and show optimal token-to-parameter ratios decreasing as compute scales. However, for the latter, loss becomes insensitive to choice of data and model sizes, showing potential for fast inference. Scaling CD SLMs to 16B parameters with tens of millions of hours of conversational data enables generation of emotive, prosodic, multi-speaker, multilingual speech, though achieving long-form coherence remains a significant challenge.

2026-04-26

arXiv (preprint)

doi.org

arxiv.org

Poly-View Contrastive Learning

Amitis Shidani

R Devon Hjelm

Jason Ramapuram

Russell Webb

Eeshan Gunesh Dhekane

Dan Busbridge

2024-01-15

ICLR.cc/2024/Conference (poster)

openreview.net

AI Policy Fellowship Publications

Mila Ventures Launchpad

AI Policy Compass

Eeshan Gunesh Dhekane

Publications

AI Policy Fellowship Publications

Mila Ventures Launchpad

AI Policy Compass

Popular keywords:

Eeshan Gunesh Dhekane

Publications