OPUS 4 | Korpuslinguistik

Korpuslinguistik

9 search hits

1 to 9

Sort by

DeReWo: Korpusbasierte Wortformenliste. Technical Report IDS-KL-2009-02 (2009)

Perkuhn, Rainer ; Belica, Cyril ; Kupietz, Marc ; Keibel, Holger ; Hennig, Sophie

"Corpus-driven": Systematische Auswertung automatisch ermittelter sprachlicher Muster (2007)

Perkuhn, Rainer

Approaching grammar: Detecting, conceptualizing and generalizing paradigmatic variation (2011)

Keibel, Holger ; Belica, Cyril ; Kupietz, Marc ; Perkuhn, Rainer

This paper presents ongoing research which is embedded in an empirical-linguistic research program, set out to devise viable research strategies for developing an explanatory theory of grammar as a psychological and social phenomenon. As this phenomenon cannot be studied directly, the program attempts to approach it indirectly through its correlates in language corpora, which is justified by referring to the core tenets of Emergent Grammar. The guiding principle for identifying such corpus correlates of grammatical regularities is to imitate the psychological processes underlying the emergent nature of these regularities. While previous work in this program focused on syntagmatic structures, the current paper goes one step further by investigating schematic structures that involve paradigmatic variation. It introduces and explores a general strategy by which corpus correlates of such structures may be uncovered, and it further outlines how these correlates may be used to study the nature of the psychologically real schematic structures.

Putting corpora into perspective. Rethinking synchronicity in corpus linguistics (2010)

Belica, Cyril ; Keibel, Holger ; Kupietz, Marc ; Perkuhn, Rainer ; Vachková, Marie

Empirical synchronic language studies generally seek to investigate language phenomena for one point in time, even though this point in time is often not stated explicitly. Until today, surprisingly little research has addressed the implications of this time-dependency of synchronic research on the composition and analysis of data that are suitable for conducting such studies. Existing solutions and practices tend to be too general to meet the needs of all kinds of research questions. In this theoretical paper that is targeted at both corpus creators and corpus users, we propose to take a decidedly synchronic perspective on the relevant language data. Such a perspective may be realised either in terms of sampling criteria or in terms of analytical methods applied to the data. As a general approach for both realisations, we introduce and explore the FReD strategy (Frequency Relevance Decay) which models the relevance of language events from a synchronic perspective. This general strategy represents a whole family of synchronic perspectives that may be customised to meet the requirements imposed by the specific research questions and language domain under investigation.

Systematic Exploration of Collocation Profiles (2007)

Perkuhn, Rainer

The central issue in corpus-driven linguistics is the detection and description of patterns in language usage. The features that constitute the notion of a pattern can be computed to a certain extent by statistical (collocation) methods, but a crucial part of the notion may vary depending on applications and users. Thus, typically, any computed collocation cluster will have to be interpreted hermeneutically. Often it might be captured by a generalized, more abstract pattern. We present a generic process model that supports the recognition, interpretation, and expression of the patterns inside and of the relations between clusters. By this, clusters can be merged virtually according to any notion of a 'pattern', and their relations can be exploited for different applications

A brief tutorial on using collocations for uncovering and contrasting meaning potentials of lexical items (2009)

Perkuhn, Rainer ; Keibel, Holger

This introductory tutorial describes a strictly corpus-driven approach for uncovering indications for aspects of use of lexical items. These aspects include ‘(lexical) meaning’ in a very broad sense and involve different dimensions, they are established in and emerge from respective discourses. Using data-driven mathematical-statistical methods with minimal (linguistic) premises, a word’s usage spectrum is summarized as a collocation profile. Self-organizing methods are applied to visualize the complex similarity structure spanned by these profiles. These visualizations point to the typical aspects of a word’s use, and to the common and distinctive aspects of any two words.

Korpustechnologie am Institut für Deutsche Sprache (2005)

Perkuhn, Rainer ; Belica, Cyril ; al-Wadi, Doris ; Lauer, Meike ; Steyer, Kathrin ; Weiß, Christian

Web as corpus: Kooperation mit der Universität Bologna (2007)

Belica, Cyril ; Keibel, Holger ; Kupietz, Marc ; Perkuhn, Rainer

Korpuslinguistik – Das unbekannte Wesen oder Mythen über Korpora und Korpuslinguistik (2006)

Perkuhn, Rainer ; Belica, Cyril

Eine angemessene, sachgemäße Diskussion über Stärken und Schwächen, Möglichkeiten und Grenzen der Korpuslinguistik ist überschattet von vielen Mythen, die sich mittlerweile eingebürgert haben und die in vielen Diskussionen – gerade unter Linguisten – immer wieder aufkommen. An dieser Stelle möchten wir einige der verbreitetsten Mythen zusammenstellen und die Hintergründe aus dieser korpuslinguistischen Perspektive erörtern.

1 to 9

Open Access

Korpuslinguistik

Refine

Author

Year of publication

Document Type

Language

Has Fulltext

Is part of the Bibliography

Keywords

Publicationstate

Reviewstate

Publisher

9 search hits