OPUS 4 | Search

50 search hits

1 to 10

Sort by

Year
Year
Title
Title
Author
Author

GenitivDB 2.0 – Datenbank zur Genitivmarkierung (Release vom 01.09.2015) (2015)

Bubenhofer, Noah ; Hansen-Morath, Sandra ; Konopka, Marek ; Schneider, Roman

Datenbasis für Untersuchungen zur grammatischen Variabilität im Standarddeutschen

Discovering Subtle Word Relations in Large German Corpora (2015)

Buschjäger, Sebastian ; Pfahler, Lukas ; Morik, Katharina

With an increasing amount of text data available it is possible to automatically extract a variety of information about language. One way to obtain knowledge about subtle relations and analogies between words is to observe words which are used in the same context. Recently, Mikolov et al. proposed a method to efficiently compute Euclidean word representations which seem to capture subtle relations and analogies between words in the English language. We demonstrate that this method also captures analogies in the German language. Furthermore, we show that we can transfer information extracted from large non-annotated corpora into small annotated corpora, which are then, in turn, used for training NLP systems.

Valenz im Fokus: Vorwort (2015)

Dominguez Vázquez, Maria José ; Eichinger, Ludwig M.

Die Festschrift Valenz im Fokus: Grammatische und lexikografische Studien enthält zum einen die Beiträge des internationalen Kolloquiums „Valenz im Fokus“, das am 12. Juli 2013 im Institut für Deutsche Sprache in Mannheim zu Ehren von Jacqueline Kubczak veranstaltet wurde, zum anderen weitere Beiträge von Kollegen aus der ganzen Welt, die zum einen als elektronische Publikation während des Kolloquiums präsentiert wurden, zum anderen speziell für diese Festschrift hinzukamen.

Der Tanz um das Verb (2015)

Engel, Ulrich

Ziggurat: A new data model and indexing format for large annotated text corpora (2015)

Evert, Stefan ; Hardie, Andrew

The IMS Open Corpus Workbench (CWB) software currently uses a simple tabular data model with proven limitations. We outline and justify the need for a new data model to underlie the next major version of CWB. This data model, dubbed Ziggurat, defines a series of types of data layer to represent different structures and relations within an annotated corpus; each such layer may contain variables of different types. Ziggurat will allow us to gradually extend and enhance CWB’s existing CQP-syntax for corpus queries, and also make possible more radical departures relative not only to the current version of CWB but also to other contemporary corpus-analysis software.

Sind "logische" Wörter ambig? (2015)

Frosch, Helmut

Spezialstudie: Regionale Verteilung (2015)

Fürbacher, Monica

Challenges in the Alignment, Management and Exploitation of Large and Richly Annotated Multi-Parallel Corpora (2015)

Graën, Johannes ; Clematide, Simon

The availability of large multi-parallel corpora offers an enormous wealth of material to contrastive corpus linguists, translators and language learners, if we can exploit the data properly. Necessary preparation steps include sentence and word alignment across multiple languages. Additionally, linguistic annotation such as partof- speech tagging, lemmatisation, chunking, and dependency parsing facilitate precise querying of linguistic properties and can be used to extend word alignment to sub-sentential groups. Such highly interconnected data is stored in a relational database to allow for efficient retrieval and linguistic data mining, which may include the statistics-based selection of good example sentences. The varying information needs of contrastive linguists require a flexible linguistic query language for ad hoc searches. Such queries in the format of generalised treebank query languages will be automatically translated into SQL queries.

KoGra-R - Standardisierte statistische Verfahren für korpusbasierte Häufigkeiten. Elektronische Ressource (2015)

Hansen-Morath, Sandra ; Schmitz, Hans-Christian ; Wolfer, Sascha

Des Iraks, des Irakes oder des Irak - von Sprachzweifeln und Sprachveriation (2015)

Konopka, Marek