detalle del documento
IDENTIFICACIÓN

oai:arXiv.org:2407.07235

Tema
Computer Science - Sound Computer Science - Machine Learnin... Electrical Engineering and Systems...
Autor
Netzorg, Robin Cote, Alyssa Koshin, Sumi Garoute, Klo Vivienne Anumanchipalli, Gopala Krishna
Categoría

Computer Science

Año

2024

fecha de cotización

17/7/2024

Palabras clave
systems speech speaker science voice
Métrico

Resumen

As experts in voice modification, trans-feminine gender-affirming voice teachers have unique perspectives on voice that confound current understandings of speaker identity.

To demonstrate this, we present the Versatile Voice Dataset (VVD), a collection of three speakers modifying their voices along gendered axes.

The VVD illustrates that current approaches in speaker modeling, based on categorical notions of gender and a static understanding of vocal texture, fail to account for the flexibility of the vocal tract.

Utilizing publicly-available speaker embeddings, we demonstrate that gender classification systems are highly sensitive to voice modification, and speaker verification systems fail to identify voices as coming from the same speaker as voice modification becomes more drastic.

As one path towards moving beyond categorical and static notions of speaker identity, we propose modeling individual qualities of vocal texture such as pitch, resonance, and weight.

Netzorg, Robin,Cote, Alyssa,Koshin, Sumi,Garoute, Klo Vivienne,Anumanchipalli, Gopala Krishna, 2024, Speech After Gender: A Trans-Feminine Perspective on Next Steps for Speech Science and Technology

Documento

Abrir

Compartir

Fuente

Artículos recomendados por ES/IODE IA