Table of Contents

Intended runsheet

Tuesday 29th September 2pm to 4pm - Area 14-3

In terms of what to expect for the Special Session, the intended runsheet is:

  • Special Session Chair welcome and introduction - Kathy Reid - 10 mins
  • Oral presentations by author(s) of each selected paper (10-13 mins presentation time each with 7-10 mins question time, total 20 mins per paper)
  • The suggested order to group similarly-themed papers is:

    • Kumar et al - Who Synthesized This? Joint Deepfake Detection and Generative Source Attribution
    • Staněk et al - Ethical and Technical Limits of Deepfake Speech Datasets

    • Nanni - Rethinking Consent Acquisition for Voice Synthesis: From Static to Dynamic Consent
    • Williams - AI Regulation and the Technical Language of Speech Synthesis

    • Bassey Edet et al - Towards Digital Preservation of Efik: TTS for a Low-Resource African Language

This runsheet allows a 10 minute buffer in the session for running over time or for equipment setup (e.g. 2 mins for each speaker to set up a laptop). If we are running head of time, the Session Chair will facilitate a discussion on what we would like to see for similar Special Sessions at future Interspeech conferences.

Speakers should prepare a 10-13 minute oral presentation and be prepared to answer 7-10 minutes of questions, with a total of 20 minutes per paper. You may wish to prepare questions for the other papers, and the Session Chair will also prepare questions.

Accepted papers

NOTE: Please see the main Interspeech site for the full conference program.

Who Synthesized This? Joint Deepfake Detection and Generative Source Attribution

Paper ID: 2442  ·  Time: 14:00–16:00  ·  Presenting author: Vishal Kumar

Abstract

We propose a few-shot, open-set framework for joint deepfake detection and generative source attribution. Addressing the rapid evolution of synthetic speech, we introduce a temporal evaluation strategy using the MLAAD v9 dataset, partitioning models by launch year and architectural family. Our framework utilises a WavLM-Large model with targeted LoRA adapters, optimised via a hierarchical metric learning objective: Conditional AAM Softmax + Center Loss. We simulate real-world forensic deployment by performing prototype-based inference on unseen 2025+ model families, establishing new class centroids with only 12 samples. Evaluated on the ASVspoof 5.0 benchmark, our framework achieves SOTA performance with 0.49% EER and 0.09 minDCF, while attaining 99.3% accuracy in tracking generated speech to its source. These results demonstrate that hierarchical representation learning and few-shot prototype adaptation provide a robust, scalable defence against the shifting landscape of audio deepfakes.

Authors

Towards Digital Preservation of Efik: TTS for a Low-Resource African Language

Paper ID: 1868  ·  Time: 14:00–16:00  ·  Presenting author: Offiong Bassey Edet

Abstract

Efik, a tonal language spoken by about 3 million second-language speakers and 1.5 million native speakers in southeastern Nigeria, remains underrepresented in speech synthesis research. We present the first documented end-to-end text-to-speech study for Efik, introducing a curated single-speaker corpus of 2,632 utterances totalling three hours and a comparative evaluation of four neural models (VITS, MMS-TTS, SpeechT5, and Orpheus-TTS) under low-resource conditions. Native speakers evaluated the systems using MOS, Nat-MOS, and A-MOS. MMS-TTS achieved the highest MOS of 3.80 ± 0.63 and produced more stable long-form speech, though tonal errors persisted. Other models showed greater tonal and prosodic inconsistencies. These results provide a reproducible baseline and highlight the need for larger corpora and tone-aware modelling of African tonal languages.

Authors


Paper ID: 2379  ·  Time: 14:00–16:00  ·  Presenting author: Matilde Nanni

Abstract

Voice synthesis raises specific ethical challenges by enabling third parties to generate a person’s voice without their involvement or control. Existing safeguards focus on consent acquisition, typically through contractual agreements, and attempt to regulate this practice via standard data-processing models of authorisation. However, voices function not only as data but also as markers of personal identity. As a result, one-time authorisation of voice data, which cannot be de-identified, leaves later uses insufficiently regulated and does not prevent identity-based harms to the original voice sources. This paper critically examines two models of contractual consent acquisition, one general and one specific, to show that neither satisfies the requirements for informed consent. As an alternative, the paper proposes dynamic consent, a type of informed consent that has been used in bioethics, as a governance solution better suited for synthetic voices.

Authors


Ethical and Technical Limits of Deepfake Speech Datasets

Paper ID: 124  ·  Time: 14:00–16:00  ·  Presenting author: Vojtěch Staněk

Abstract

Claims about the robustness and fairness of deepfake speech detectors are only as credible as the datasets used to train and evaluate those systems. We present a dataset-level audit of the deepfake speech landscape. We compile and analyse 39 deepfake speech datasets, examining key attributes including accessibility, documentation, demographic and language coverage, dataset scale, and the underlying bona fide speech sources. Our audit reveals two important takeaways. Firstly, fairness assessment is largely infeasible because most datasets lack demographic metadata, and only a few contain gender or language labels. This prevents any meaningful subgroup analysis and leaves other demographic attributes unaddressed. Secondly, we identify substantial overlap in underlying bona fide source corpora across datasets, which can undermine cross-dataset evaluation and lead to overstated generalisation claims.

Authors


AI Regulation and the Technical Language of Speech Synthesis

Paper ID: 210  ·  Time: 14:00–16:00  ·  Presenting author: Jennifer Williams

Abstract

Speech synthesis has garnered increasing interest with global efforts to regulate artificial intelligence (AI), including deepfakes. Policymakers must draft careful wording to prevent societal harm and identify responsible parties. Legal AI provisions must be accurate, especially when instructing responsible parties to act, as seen in the transparency obligations across current AI regulations worldwide. But AI and generative AI are broad terms, and analogies from image and video domains do not translate well to audio. This paper critically assesses current AI policy focussing on speech synthesis outputs overlooks portable models and complex workflows wherein speaker embeddings that encode voice identity can be developed independently and later reproduced. This paper traces the convergence of text-to-speech, voice cloning, speaker verification, and speech recognition, highlighting where legal definitions may fail to capture the technical components that enable voice cloning.

Authors


Why is this Special Session so needed?

Speech synthesis has advanced rapidly in the last five years (e.g. (X. Chen et al., 2024; Hayashi et al., 2019; Valle, Li, et al., 2020; Valle, Shih, et al., 2020). Some speech synthesis models such as VALL-E require merely seconds of speech data to produce a synthesised voice (S. Chen et al., 2024). Moreover, this technology has been widely adopted due to its new-found accessibility, creating new opportunities for personalised communication, accessibility applications and creative content generation.

However, these new capabilities have raised significant ethical, legal and social concerns surrounding consent, identity protection, copyright, misuse and authenticity (Burgess et al., 2025). Recent legal developments such as the successful court case brought by German actor Manfred Lehmann - the German voice actor for Bruce Willis - who had his voice cloned without consent by a YouTuber (Reinholz & Schmidt, 2025), and the Ensuring Likeness Voice and Image Security Act (ELVIS) legislated by the US State of Tennessee (Kirkwood, 2025), as well as efforts by industry, such as Hugging Face’s voice consent gate initiative (Mitchell & Kaffe, 2025), have centred the need to establish social, legal and ethical norms for the use of synthesised speech.

This special session addresses the critical gap between technological capabilities, ethical and legal frameworks for governing and steering synthetic speech systems and establishing social norms for the use of these systems. The field urgently needs interdisciplinary approaches to consent management, deepfake prevention, watermarketing, and other safety mechanisms and safeguarding protocols.

What will this Special Session cover?

The session will explore multiple dimensions of this challenge:

  • Technical approaches to consent verification, watermarking, and authentication in TTS systems
  • Legal frameworks for personality rights, data protection, and liability in voice synthesis
  • Ethical considerations around vulnerable populations, posthumous voice rights, and cultural sensitivities
  • Industry perspectives on implementing consent management at scale
  • User studies on public perception, trust, and acceptance of voice cloning technologies

The objectives are to: (a) establish shared understanding of current challenges and regulatory landscape; (b) present state-of-the-art technical solutions for consent and safety; (c) foster collaboration between technical, legal, and ethical experts; (d) develop recommendations for best practices and standards; and (e) identify critical research gaps requiring community attention.

Session Format

The special session will combine:

  • Keynote presentation on legal, social or ethical implications of voice cloning (30 minutes)
  • Oral presentations of peer-reviewed papers (8-10 papers, 12 minutes each)
  • Panel discussion with industry representatives, legal experts, and researchers (30 minutes)

If the topics in this Special Session are of interest to you, please submit via the general Interspeech Call for Papers process.

Visit the Interspeech 2026 Call for Papers

References

Burgess, J., Carlon, D., & Doyuran, E. B. (2025). Voice AI and authenticity: Current issues and emerging challenges. ARC Centre of Excellence for Automated Decision-Making and Society. https://apo.org.au/node/331920

Chen, S., Liu, S., Zhou, L., Liu, Y., Tan, X., Li, J., Zhao, S., Qian, Y., & Wei, F. (2024). VALL-e 2: Neural codec language models are human parity zero-shot text to speech synthesizers. https://arxiv.org/abs/2406.05370

Chen, X., Wang, X., Zhang, S., He, L., Wu, Z., Wu, X., & Meng, H. (2024). Stylespeech: Self-supervised style enhancing with vq-vae-based pre-training for expressive audiobook speech synthesis. ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 12316–12320.

Hayashi, T., Yamamoto, R., Inoue, K., Yoshimura, T., Watanabe, S., Toda, T., Takeda, K., Zhang, Y., & Tan, X. (2019). ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit. arXiv Preprint arXiv:1910.10909.

Kirkwood, J. (2025, March 24). Why Tennessee’s ELVIS Act Is the King of Artificial Intelligence Protections. Vanderbilt Law School. https://law.vanderbilt.edu/why-tennessees-elvis-act-is-the-king-of-artificial-intelligence-protections

Mitchell, M., & Kaffe, L.-A. (2025, October 31). Voice Cloning with Consent. https://huggingface.co/blog/voice-consent-gate

Reinholz, F., & Schmidt, R. (2025, October 14). Voice clones by AI in court—Dubbing artist wins at Berlin Regional Court. HÄRTING Rechtsanwälte. https://haerting.de/en/insights/voice-clones-by-ai-in-court-dubbing-artist-wins-at-berlin-regional-court/

Valle, R., Li, J., Prenger, R., & Catanzaro, B. (2020). Mellotron: Multispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens. ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6189–6193.

Valle, R., Shih, K. J., Prenger, R., & Catanzaro, B. (2020). Flowtron: An autoregressive flow-based generative network for text-to-speech synthesis. International Conference on Learning Representations. https://iclr.cc/virtual/2021/poster/3204