Speech After Gender: A Trans-Feminine Perspective on Next Steps for Speech Science and Technology

detalle del documento

IDENTIFICACIÓN

oai:arXiv.org:2407.07235

Tema

Computer Science - Sound Computer Science - Machine Learnin... Electrical Engineering and Systems...

Autor

Netzorg, Robin Cote, Alyssa Koshin, Sumi Garoute, Klo Vivienne Anumanchipalli, Gopala Krishna

Categoría

Computer Science

Año

2024

fecha de cotización

17/7/2024

Palabras clave

systems speech speaker science voice

Métrico

Resumen

As experts in voice modification, trans-feminine gender-affirming voice teachers have unique perspectives on voice that confound current understandings of speaker identity.

To demonstrate this, we present the Versatile Voice Dataset (VVD), a collection of three speakers modifying their voices along gendered axes.

The VVD illustrates that current approaches in speaker modeling, based on categorical notions of gender and a static understanding of vocal texture, fail to account for the flexibility of the vocal tract.

Utilizing publicly-available speaker embeddings, we demonstrate that gender classification systems are highly sensitive to voice modification, and speaker verification systems fail to identify voices as coming from the same speaker as voice modification becomes more drastic.

As one path towards moving beyond categorical and static notions of speaker identity, we propose modeling individual qualities of vocal texture such as pitch, resonance, and weight.

Netzorg, Robin,Cote, Alyssa,Koshin, Sumi,Garoute, Klo Vivienne,Anumanchipalli, Gopala Krishna, 2024, Speech After Gender: A Trans-Feminine Perspective on Next Steps for Speech Science and Technology