Resumen
In this paper, we describe a phonotactic language recognition model that effectively manages long and short n-gram input sequences to learn contextual phonotactic-based vector embeddings. Our approach uses a transformer-based encoder that integrates a sliding window attention to attempt finding discriminative short and long cooccurrences of language dependent n-gram phonetic units. We then evaluate and compare the use of different phoneme recognizers (Brno and Allosaurus) and sub-unit tokenizers to help select the more discriminative n-grams. The proposed architecture is evaluated using the Kalaka-3 database that contains clean and noisy audio recordings for very similar languages (i.e. Iberian languages, e.g., Spanish, Galician, Catalan). We provide results using the Cavg and accuracy metrics used in NIST evaluations. The experimental results show that our proposed approach outperforms by 21% of relative improvement to the best system presented in the Albayzin LR competition.
| Idioma original | Inglés |
|---|---|
| Título de la publicación alojada | 2022 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022 - Proceedings |
| Editorial | Institute of Electrical and Electronics Engineers Inc. |
| Páginas | 6872-6876 |
| Número de páginas | 5 |
| ISBN (versión digital) | 9781665405409 |
| ISBN (versión impresa) | 9781665405409 |
| DOI | |
| Estado | Publicada - 2022 |
| Evento | 47th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022 - Virtual, Online, Singapur Duración: 23 may 2022 → 27 may 2022 |
Serie de la publicación
| Nombre | ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings |
|---|---|
| Volumen | 2022-May |
Conferencia
| Conferencia | 47th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022 |
|---|---|
| País/Territorio | Singapur |
| Ciudad | Virtual, Online |
| Período | 23/05/22 → 27/05/22 |
Nota bibliográfica
Publisher Copyright:© 2022 IEEE
Areas de Conocimiento del CACES
- 316A Desarrollo y análisis de software y aplicaciones
Huella
Profundice en los temas de investigación de 'PHONOTACTIC LANGUAGE RECOGNITION USING A UNIVERSAL PHONEME RECOGNIZER AND A TRANSFORMER ARCHITECTURE'. En conjunto forman una huella única.Citar esto
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver