Refine
Document Type
- Part of a Book (1)
- Conference Proceeding (1)
Language
- English (2) (remove)
Has Fulltext
- yes (2)
Is part of the Bibliography
- yes (2)
Keywords
- Algorithmus (1)
- Annotation (1)
- Computerlinguistik (1)
- Direkte Rede (1)
- Korpus <Linguistik> (1)
- Maschinelles Lernen (1)
- Methodik (1)
- Natürliche Sprache (1)
- Redeerwähnung (1)
- Text Mining (1)
Publicationstate
- Zweitveröffentlichung (2) (remove)
Reviewstate
- Peer-Review (2)
Publisher
Corpus REDEWIEDERGABE
(2020)
This article presents the corpus REDEWIEDERGABE, a German-language historical corpus with detailed annotations for speech, thought and writing representation (ST&WR). With approximately 490,000 tokens, it is the largest resource of its kind. It can be used to answer literary and linguistic research questions and serve as training material for machine learning. This paper describes the composition of the corpus and the annotation structure, discusses some methodological decisions and gives basic statistics about the forms of ST&WR found in this corpus.
This paper describes a rule-based approach to detect direct speech without the help of any quotation markers. As datasets fictional and non-fictional texts were used. Our evaluation shows that the results appear stable throughout different datasets in the fictional domain and are comparable to the results achieved in related work.