Abstract
Embedding models turn words/documents into real-number vectors via co-occurrence data from unrelated texts. Crafting domain-specific embeddings from general corpora with limited domain vocabulary is challenging. Existing solutions retrain models on small domain datasets, overlooking potential of gathering rich in-domain texts. We exploit Named Entity Recognition and Doc2Vec for autonomous in-domain corpus creation. Our experiments compare models from general and in-domain corpora, highlighting that domain-specific training attains the best outcome.
| Original language | English |
|---|---|
| Pages (from-to) | 491-527 |
| Number of pages | 37 |
| Journal | Informatica (Netherlands) |
| Volume | 34 |
| Issue number | 3 |
| DOIs | |
| State | Published - 8 Sep 2023 |
Bibliographical note
Publisher Copyright:© 2023 Vilnius University.
Keywords
- ad hoc corpus
- Doc2Vec
- embedding models
- Named Entity Recognition
CACES Knowledge Areas
- 316A Software and Applications Development and Analysis
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver