OPUS 4 | Search

Refine

Has Fulltext

yes (3)
no (1)

4 search hits

1 to 4

Sort by

Präliminarien einer Korpusgrammatik (2014)

Bubenhofer, Noah ; Konopka, Marek ; Schneider, Roman

Der korpuslinguistische Ansatz des Projekts »Korpusgrammatik« eröffnet neue Perspektiven auf unsere Sprachwirklichkeit allgemein und grammatische Regularitäten im Besonderen. Der vorliegende Band klärt auf, wie man korpuslinguistisch nach dem Standard fragen kann, wie die Projektkorpora aufgebaut und in einer Korpusdatenbank erschlossen sind, wie man in einem automatischen Abfragesystem der Variabilität der Sprache zu Leibe rückt und sie sogar messbar macht, schließlich aber auch, wo die Grenzen quantitativer Korpusanalysen liegen. Pilotstudien deuten an, wie der Ansatz unsere grammatischen Horizonte erweitert und die Grammatikografie voranbringt.

GenitivDB 1.0 – Datenbank zur Genitivmarkierung (Release vom 01.06.2014). Elektronische Ressource (2014)

Bubenhofer, Noah ; Hansen-Morath, Sandra ; Konopka, Marek ; Schneider, Roman

Datenbasis für Untersuchungen zur grammatischen Variabilität im Standarddeutschen

GenitivDB - a corpus-generated database for German genitive classification (2014)

Schneider, Roman

We present a novel NLP resource for the explanation of linguistic phenomena, built and evaluated exploring very large annotated language corpora. For the compilation, we use the German Reference Corpus (DeReKo) with more than 5 billion word forms, which is the largest linguistic resource worldwide for the study of contemporary written German. The result is a comprehensive database of German genitive formations, enriched with a broad range of intra- und extralinguistic metadata. It can be used for the notoriously controversial classification and prediction of genitive endings (short endings, long endings, zero-marker). We also evaluate the main factors influencing the use of specific endings. To get a general idea about a factor’s influences and its side effects, we calculate chi-square-tests and visualize the residuals with an association plot. The results are evaluated against a gold standard by implementing tree-based machine learning algorithms. For the statistical analysis, we applied the supervised LMT Logistic Model Trees algorithm, using the WEKA software. We intend to use this gold standard to evaluate GenitivDB, as well as to explore methodologies for a predictive genitive model.

Hypertext, Wissensnetz und Datenbank: die Webinformationssysteme Grammis und ProGr@mm (2014)

Schneider, Roman ; Schwinn, Horst

1 to 4

Open Access

Refine

Author

Year of publication

Document Type

Language

Has Fulltext

Is part of the Bibliography

Keywords

Publicationstate

Reviewstate

Publisher

4 search hits