Refine
Year of publication
- 2021 (4) (remove)
Document Type
- Conference Proceeding (4) (remove)
Language
- English (4)
Has Fulltext
- yes (4)
Is part of the Bibliography
- no (4) (remove)
Keywords
- Computerlinguistik (2)
- Korpus <Linguistik> (2)
- Urheberrecht (2)
- corpus linguistics (2)
- Antwort (1)
- Ausrichten <Technik> (1)
- Automatische Sprachanalyse (1)
- Data Mining (1)
- Dialog (1)
- Europäische Kommission. Digital Single Market (1)
Publicationstate
Reviewstate
- Peer-Review (4)
Publisher
We describe a simple procedure for the automatic creation of word-level alignments between printed documents and their respective full-text versions. The procedure is unsupervised, uses standard, off-the-shelf components only, and reaches an F-score of 85.01 in the basic setup and up to 86.63 when using pre- and post-processing. Potential areas of application are manual database curation (incl. document triage) and biomedical expression OCR.