Resumen
Embedding models turn words/documents into real-number vectors via co-occurrence data from unrelated texts. Crafting domain-specific embeddings from general corpora with limited domain vocabulary is challenging. Existing solutions retrain models on small domain datasets, overlooking potential of gathering rich in-domain texts. We exploit Named Entity Recognition and Doc2Vec for autonomous in-domain corpus creation. Our experiments compare models from general and in-domain corpora, highlighting that domain-specific training attains the best outcome.
| Idioma original | Inglés |
|---|---|
| Páginas (desde-hasta) | 491-527 |
| Número de páginas | 37 |
| Publicación | Informatica (Netherlands) |
| Volumen | 34 |
| N.º | 3 |
| DOI | |
| Estado | Publicada - 8 sept 2023 |
Nota bibliográfica
Publisher Copyright:© 2023 Vilnius University.
Areas de Conocimiento del CACES
- 316A Desarrollo y análisis de software y aplicaciones
Citar esto
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver