OPUS 4 | Search

Konnektoren im gesprochenen Deutsch. Eine Untersuchung am Beispiel der kommunikativen Gattung «autobiographisches Interview» (2016)

Antonioli, Giorgio

Dieses Buch schließt eine Lücke in der Konnektorenforschung, indem es den Gebrauch von Konnektoren im gesprochenen Deutsch untersucht. Die Fragestellung bringt Elemente aus dem traditionellen grammatischen Ansatz und aus der pragmatisch basierten Forschung zur gesprochenen Sprache zusammen. In Anlehnung an die Methode der Interaktionalen Linguistik analysiert der Autor den Gebrauch der Konjunktoren «und», «aber» und der Adverbkonnektoren «also», «dann» in zwei Korpora von autobiographischen Interviews. Die Untersuchung zeigt, wie Konnektoren zur Bewältigung von verschiedenartigen kommunikativen Aufgaben zur Stiftung von Intersubjektivität und zur Gesprächsorganisation eingesetzt werden können.

Konnektoren im Gespräch: Eine Fallstudie zur verstehensdokumentatorischen Funktion von und, also und dann im Israelkorpus (2016)

Antonioli, Giorgio

Die Rolle der antizipatorischen Verstehensdokumentation erweist sich in den Interviews aus dem Israelkorpus m. E. als besonders wichtig. Es wird von der Tatsache ausgegangen, dass es sich bei den Informanten um Personen mit besonders delikaten biographischen Hintergründen handele. Die Interviewerinnen müssen demzufolge mit der starken emotionalen Belastung rechnen, der die Interviewten während der Rekonstruktion ihrer Lebensgeschichte ausgesetzt sind. Ein sehr direkter Frage-Antwort-Stil könnte wegen dieser emotionalen Belastung als unangenehm empfunden werden. Der Einsatz von Verfahren antizipatorischer Verstehensdokumentation weist stattdessen m. E. eindeutig darauf hin, wie sich die Interviewerinnen offensichtlich um Empathie bemühen und im Sinne einer intersubjektiven Inreraktionskonstitution mit den Interviewten kooperieren. Ziel dieses Beitrages ist es zu zeigen, wie solche Verfahren der antizipatorischen Verstehensdokumentation durch den systematischen Einsatz der Konnektoren und, also, dann realisiert werden können.

Im Radiostudio arbeiten: Vielgestaltige Handlungen in einem flexibel architekturierten Raum (2016)

Mondada, Lorenza ; Oloff, Florence

Dieses Kapitel befasst sich mit dem Zusammenspiel von Raum und Interaktion und konzentriert sich auf die dynamischen Organisationsformen sozialer Handlungen unter Berücksichtigung verbaler und sichtbarer Ressourcen. Durch die Untersuchung eines spezifischen Settings – professionelle Interaktionen in einem Radiostudio – werden wir empirisch beschreiben und konzeptualisieren, wie ein gebauter bzw. stark architekturierter Raum im Rahmen institutioneller Praktiken genutzt und relevant gesetzt wird. So soll zu aktuellen Überlegungen zu Interaktionsraum und -architektur, zu Raum als Ressource sowie als materiellem Umfeld beigetragen werden. Unsere ethnomethodologische und konversationsanalytische Perspektive wird von aktuellen Debatten über den sogenannten spatial turn in der interaktionalen Forschung beeinflusst (Kap. 1.1). Auf Grundlage eines in einem Radiostudio erstellten Videokorpus (Kap. 1.2) wird zunächst die Verbindung zwischen einem architektonisch und technologisch komplexen Umfeld und dem interaktionalen Handeln der Teilnehmer skizziert (Kap. 2.1, Kap. 2.2). Es folgt die detaillierte Analyse eines Einzelfalls (Kap. 3), in dem die Radiomoderatoren einen Text für den nächsten Sendeabschnitt vorbereiten. Hier werden die räumlichen Charakteristika sichtbar, die bei der Arbeit nach und nach relevant gesetzt werden (Kap. 4).

Vom vielschichtigen Planen. Textproduktions-Praxis empirisch erforscht (2016)

Perrin, Daniel

Der Beitrag diskutiert das Konzept sprachlicher Praktiken am Beispiel des Planens in kollaborativem beruflichem Schreiben. Gestützt auf eine Fallstudie aus großen Korpora natürlicher empirischer Daten, werden Praktiken herausgearbeitet, die flexibles Planen im dynamischen System der Textproduktion ermöglichen. Deutlich wird, dass die Praktiken wie auch die durch sie geprägten Schreibphasen skalieren, also ähnliche Muster bilden im Kleineren wie im Größeren. Ein solches Verständnis von Planen geht weit über den Planungsbegriff in bisherigen Modellen von Schreibprozessen hinaus. So erweist sich empirische Forschung am Arbeitsplatz als gewinnbringend auch für die theoretische Schärfung des Praktiken-Konzepts. Schreiben als Prozess der Herstellung schriftsprachlicher Äußerungen wurde früh aus sprachpsychologischem Blickwinkel erforscht und modelliert. Bedeutende Phasen und Praktiken des natürlichen Schreibens, außerhalb psychologischer Laborexperimente, sind durch die Dominanz dieser Forschungstradition lange außer Acht geblieben. Der vorliegende Beitrag entwickelt ein dynamisches und komplexes Konzept von Schreibphasen und den sie bestimmenden Praktiken beruflicher Textproduktion (Teil 1). Linguistisch basierte ethnografische Forschung (2) erschließt Schreiben jenseits des Labors als vielschichtiges Zusammenspiel situierter Praktiken im dynamischen System arbeitsteiliger Textproduktion (3). Ein Beispiel einer Analyse erklärt, wie Praktiken flexiblen Planens im Nachrichtenschreiben skalieren (4). Deutlich wird dabei der Sinn empirischer Analyse von Schreibphasen und -praktiken für Theorie und Praxis (5).

Praktiken modellieren: Dialogmodellierung als Methode der Interaktionalen Linguistik (2016)

Scharloth, Joachim

Der Beitrag plädiert dafür, die Interaktionale Linguistik stärker für modellorientierte Forschung und datengeleitete Methoden zu öffnen. Er stellt eine Methode vor, wie auf der Basis von Korpora datengeleitet Praktiken rekonstruiert und modelliert werden können. Ausgehend von einer Diskussion der tiefgreifenden Veränderungen, die die Digitalisierung für die Linguistik mit sich bringt, und einer Auseinandersetzung mit dem Modellbegriff, wird der Begriff der (Kommunikativen) Praktik in Abgrenzung zum Begriff der Kommunikativen Gattung bestimmt. Im Anschluss wird am Beispiel von Trostdialogen in OnlineForen eine korpusgeleitete Methode zur Dialogmodellierung vorgestellt. Schließlich werden die Folgen der menschlichen Interaktion mit maschinellen Dialogsystemen reflektiert.

A Discourse-structured Blog Corpus for German: Challenges of Compilation and Annotation (2016)

Suarez, Holger Grumt ; Karlova-Bourbonus, Natali ; Lobin, Henning

The present paper reports the first results of the compilation and annotation of a blog corpus for German. The main aim of the project is the representation of the blog discourse structure and relations between its elements (blog posts, comments) and participants (bloggers, commentators). The data included in the corpus were manually collected from the scientific blog portal SciLogs. The feature catalogue for the corpus annotation includes three types of information which is directly or indirectly provided in the blog or can be construed by means of statistical analysis or computational tools. At this point, only directly available information (e.g. title of the blog post, name of the blogger etc.) has been annotated. We believe, our blog corpus can be of interest for the general study of blog structure or related research questions as well as for the development of NLP methods and techniques (e.g. for authorship detection).

Compilation and Annotation of the Discourse-structured Blog Corpus for German (2016)

Grumt Suárez, Holger ; Karlova-Bourbonus, Natali ; Lobin, Henning

The present paper reports the first results of the compilation and annotation of a blog corpus for German. The main aim of the project is the representation of the blog discourse structure and relations between its elements (blog posts, comments) and participants (bloggers, commentators). The data included in the corpus were manually collected from the scientific blog portal SciLogs. The feature catalogue for the corpus annotation includes three types of information which is directly or indirectly provided in the blog or can be construed by means of statistical analysis or computational tools. At this point, only directly available information (e.g., title of the blog post, name of the blogger etc.) has been annotated. We believe, our blog corpus can be of interest for the general study of blog structure or related research questions as well as for the development of NLP methods and techniques (e.g. for authorship detection).

Editorial (2016)

Geyken, Alexander ; Kupietz, Marc

Journal for language technology and computational linguistics. Corpus linguistic software tools (2016)

With the growing availability and importance of (large) corpora in all fields of linguistics, the role of software tools is gradually moving from useful, possibly intelligent informationtechnological “helpers” towards scientific instruments that are as integral parts of the research process as data, methodology and interpretations. Both aspects are present in this special issue of JLCL on corpus linguistic software tools.

Construction and dissemination of a corpus of spoken interaction - tools and workflows in the FOLK project (2016)

Schmidt, Thomas

This paper is about the workflow for construction and dissemination of FOLK (Forschungs - und Lehrkorpus Gesprochenes Deutsch – Research and Teaching Corpus of Spoken German), a large corpus of authentic spoken interaction data, recorded on audio and video. Section 2 describes in detail the tools used in the individual steps of transcription, anonymization, orthographic normalization, lemmatization and POS tagging of the data, as well as some utilities used for corpus management. Section 3 deals with the DGD (Datenbank für Gesprochenes Deutsch - Database of Spoken German) as a tool for distributing completed data sets and making them available for qualitative and quantitative analysis. In section 4, some plans for further development are sketched.

Krill: KorAP search and analysis engine (2016)

Diewald, Nils ; Margaretha, Eliza

Integrating corpora of computer-mediated communication into the language resources landscape: Initiatives and best practices from French, German, Italian and Slovenian projects (2016)

Beißwenger, Michael ; Chanier, Thierry ; Chiari, Isabella ; Erjavec, Tomaž ; Fišer, Darja ; Herold, Axel ; Ljubešić, Nikola ; Lüngen, Harald ; Poudat, Céline ; Stemle, Egon W. ; Storrer, Angelika ; Wigham, Ciara

The paper presents best practices and results from projects in four countries dedicated to the creation of corpora of computer-mediated communication and social media interactions (CMC). Even though there are still many open issues related to building and annotating corpora of that type, there already exists a range of accessible solutions which have been tested in projects and which may serve as a starting point for a more precise discussion of how future standards for CMC corpora may (and should) be shaped like.

Reformulierungsindikatoren im gesprochenen Deutsch: Die Benutzung der Ressourcen DGD und FOLK für gesprächsanalytische Zwecke (2016)

Kaiser, Julia

Dieser Beitrag stellt nach einer kurzen allgemeinen Einführung die Datenbank für Gesprochenes Deutsch (DGD) und das Forschungs- und Lehrkorpus Gesprochenes Deutsch (FOLK) als Instrumente speziell für gesprächsanalytisches Arbeiten vor. Anhand des Beispiels sprich als Diskursmarker für Reformulierungen werden Schritt für Schritt die Ressourcen und Tools für systematische korpus- und datenbankgesteuerte Recherchen illustriert: Nutzungsmöglichkeiten der Token-, Kontext-, Metadaten- und Positionssuche werden gezeigt, jeweils in Bezug auf und im wechselseitigen Verhältnis mit qualitativen Fallanalysen, auch mit Belegannotationen nach analyserelevanten (strukturellen und funktionalen) Kategorien. Schließlich wird das heißt als weiterer Reformulierungsindikator für eine vergleichende Analyse herangezogen. Dieser Beitrag stellt eine detailliertere Ausarbeitung einer kürzeren, eher technisch-didaktischen Online-Handreichung (Kaiser/ Schmidt 2016) zu diesem Thema dar, und hat einen stärker inhaltlich-analytischen Fokus.

The Karl Eberhards Corpus of spontaneously spoken southern German in dialogues - audio and articulatory recordings (2016)

Arnold, Denis ; Tomaschek, Fabian

The current paper presents a corpus containing 35 dialogues of spontaneously spoken southern German, including half an hour of articulography for 13 of the speakers. Speakers were seated in separate recording chambers, mimicking a telephone call, and recorded on individual audio channels. The corpus provides manually corrected word boundaries and automatically aligned segment boundaries. Annotations are provided in the Praat format. In addition to audio recordings, speakers filled out a detailed questionnaire, assessing among others their audio-visual consumption habits.

Creating an extensible, levelled study corpus of Russian (2016)

Batinić, Dolores ; Birzer, Sandra ; Zinsmeister, Heike

In this paper, we present first results of training a classifier for discriminating Russian texts into different levels of difficulty. For the classification we considered both surface-oriented features adopted from readability assessments and more linguistically informed, positional features to classify texts into two levels of difficulty. This text classification is the main focus of our Levelled Study Corpus of Russian (LeStCoR), in which we aim to build a corpus adapted for language learning purposes – selecting simpler texts for beginner second language learners and more complex texts for advanced learners. The most discriminative feature in our pilot study was a lexical feature that approximates accessibility of the vocabulary by the second language learner in terms of the proportion of familiar words in the texts. The best feature setting achieved an accuracy of 0.91 on a pilot corpus of 209 texts.

The IFCASL Corpus of French and German Non-native and Native Read Speech (2016)

Trouvain, Jürgen ; Bonneau, Anne ; Colotte, Vincent ; Fauth, Camille ; Fohr, Dominique ; Jouvet, Denis ; Jügler, Jeanin ; Laprie, Yves ; Mella, Odile ; Möbius, Bernd ; Zimmerer, Frank

The IFCASL corpus is a French-German bilingual phonetic learner corpus designed, recorded and annotated in a project on individualized feedback in computer-assisted spoken language learning. The motivation for setting up this corpus was that there is no phonetically annotated and segmented corpus for this language pair of comparable of size and coverage. In contrast to most learner corpora, the IFCASL corpus incorporate data for a language pair in both directions, i.e. in our case French learners of German, and German learners of French. In addition, the corpus is complemented by two sub-corpora of native speech by the same speakers. The corpus provides spoken data by about 100 speakers with comparable productions, annotated and segmented on the word and the phone level, with more than 50% manually corrected data. The paper reports on inter-annotator agreement and the optimization of the acoustic models for forced speech-text alignment in exercises for computer-assisted pronunciation training. Example studies based on the corpus data with a phonetic focus include topics such as the realization of /h/ and glottal stop, final devoicing of obstruents, vowel quantity and quality, pitch range, and tempo.

Integrating corpora of computer-mediated communication into the language resources landscape: Initiatives and best practices from French, German, Italian and Slovenian projects (2016)

Beißwenger, Michael ; Chanier, Thierry ; Chiari, Isabella ; Erjavec, Tomaž ; Fišer, Darja ; Herold, Axel ; Ljubešić, Nikola ; Lüngen, Harald ; Poudat, Céline ; Stemle, Egon W. ; Storrer, Angelika ; Wigham, Ciara

The paper presents best practices and results from projects in four countries dedicated to the creation of corpora of computer-mediated communication and social media interactions (CMC). Even though there are still many open issues related to building and annotating corpora of that type, there already exists a range of accessible solutions which have been tested in projects and which may serve as a starting point for a more precise discussion of how future standards for CMC corpora may (and should) be shaped like.

Fremdwörter zwischen Isolation und Integration. Empirische Analysen zum Schreibusus auf der Basis von Textkorpora professioneller und informeller Schreiber (2016)

Krome, Sabine ; Roll, Bernhard

When becoming integrated into the German vocabulary, foreign words reflect paradigmatic changes regarding orthography, grammar as well as semantics. In this context,German orthography is also highly determined by orthographic codification, which continues to influence the development of spelling to the present day. This study compares digital linguistically annotated corpora containing texts written by professional as well as non-professional writers; these corpora contain several billion foreign words (of Greek, Latin and French origin, and in the second part of the study of English/American and Italian origin), studied over a period of 20 years following the German orthographic reform of 1996. The results may potentially help the official regulations to adapt to the spelling practices observed – either by describing the rules more precisely or by proposing possible spelling variants or eliminating those which are not in common use. The study may also help to support correct lexicographic codification in dictionaries.

Kuration und Exploration des Korpus "Diskurs in der Weimarer Republik" (2016)

Fankhauser, Peter

Datenbank für Gesprochenes Deutsch (DGD) (2016)

Schmidt, Thomas

Aufbau einer Korpusinfrastruktur für die Beobachtung des Schreibgebrauchs (2016)

Fischer, Peter M. ; Diewald, Nils ; Kupietz, Marc ; Witt, Andreas

TInCAP – ein interdisziplinäres Korpus zu Ambiguitätsphänomenen (2016)

Hartmann, Jutta M. ; Sauter, Corinna ; Schole, Gesa ; Wagner, Wiltrud ; Gietz, Peter ; Winkler, Susanne

Kookkurrenz aus korpuslinguistischer Sicht (2016)

Storjohann, Petra

Kookkurrenzen (zum Beispiel ‘Beziehungen pflegen’ oder ‘wirtschaftlich bankrott’) gehören zum zentralen Gegenstand jeder korpusanalytischen Studie. Als Wortverbindungen sind sie Einheiten, die unter bestimmten kontextuellen Voraussetzungen zustande kommen und die wichtige Funktionen im Syntagma, Satz oder Text aufweisen. Kookkurrenzen stellen den systematischen Zugang zur Erfassung von Bedeutung, Funktionen sowie von konventionalisierten Mustern dar. Ihre Relevanz wird auch zunehmend in kultur- und politikwissenschaftlich und in kognitiv orientierten Wissenschaftsbereichen anerkannt. Mit diesem Band wird Fachliteratur zu zentralen Bereichen und Themen zusammengefasst, bei denen korpusanalytische Verfahren zur Untersuchung typischer Wortkombinationen im Mittelpunkt stehen. Dazu zählen neben Überblicksliteratur und allgemeinen Einführungen auch interessante Einzelstudien, die mit diversen Korpusansätzen arbeiten, sowie weiterführende Links und Materialsammlungen. Dieser Band bildet insbesondere die Themenschwerpunkte ab, die gegenwärtig viel Aufmerksamkeit erhalten.

Annotating Discourse Relations in Spoken Language: A Comparison of the PDTB and CCR Frameworks (2016)

Rehbein, Ines ; Scholman, Merel ; Demberg, Vera

In discourse relation annotation, there is currently a variety of different frameworks being used, and most of them have been developed and employed mostly on written data. This raises a number of questions regarding interoperability of discourse relation annotation schemes, as well as regarding differences in discourse annotation for written vs. spoken domains. In this paper, we describe ouron annotating two spoken domains from the SPICE Ireland corpus (telephone conversations and broadcast interviews) according todifferent discourse annotation schemes, PDTB 3.0 and CCR. We show that annotations in the two schemes can largely be mappedone another, and discuss differences in operationalisations of discourse relation schemes which present a challenge to automatic mapping. We also observe systematic differences in the prevalence of implicit discourse relations in spoken data compared to written texts,find that there are also differences in the types of causal relations between the domains. Finally, we find that PDTB 3.0 addresses many shortcomings of PDTB 2.0 wrt. the annotation of spoken discourse, and suggest further extensions. The new corpus has roughly theof the CoNLL 2015 Shared Task test set, and we hence hope that it will be a valuable resource for the evaluation of automatic discourse relation labellers.

Einführung in die Benutzung der Ressourcen DGD und FOLK für gesprächsanalytische Zwecke. Handreichung: "sprich" als Reformulierungsindikator (2016)

Kaiser, Julia ; Schmidt, Thomas

Diese Handreichung stellt die Datenbank für Gesprochenes Deutsch (DGD) und speziell das Forschungs- und Lehrkorpus Gesprochenes Deutsch (FOLK) als Instrumente gesprächsanalytischer Arbeit vor. Nach einem kurzen einführenden Überblick werden anhand des Beispiels "sprich" als Diskursmarker bzw. Reformulierungsindikator Schritt für Schritt die Ressourcen und Tools für systematische korpus- und datenbankgesteuerte Recherchen und Analysen vorgestellt und illustriert.

Einführung in die Benutzung der Ressourcen DGD und FOLK für gesprächsanalytische Zwecke. Handreichung: Einfache Recherche-Anfragen als Übungsbeispiele (2016)

Kaiser, Julia ; Schmidt, Thomas

Diese Handreichung stellt die Datenbank für Gesprochenes Deutsch (DGD) und speziell das Forschungs- und Lehrkorpus Gesprochenes Deutsch (FOLK) als Instrumente gesprächsanalytischer Arbeit vor. Nach einem kurzen einführenden Überblick werden anhand vier verschiedener Beispiele Schritt für Schritt die Ressourcen und Tools für systematische korpus- und datenbankgesteuerte Recherchen und Analysen vorgestellt und illustriert.

Einführung in die Benutzung der Ressourcen DGD und FOLK für gesprächsanalytische Zwecke. Handreichung: Metapragmatische Modalisierungen. (2016)

Kaiser, Julia ; Schmidt, Thomas

Diese Handreichung stellt die Datenbank für Gesprochenes Deutsch (DGD) und speziell das Forschungs- und Lehrkorpus Gesprochenes Deutsch (FOLK) als Instrumente gesprächsanalytischer Arbeit vor. Nach einem kurzen einführenden Überblick werden anhand des Beispiels metapragmatischer Modalisierungen mit den Adverbien "sozusagen" und "gewissermaßen" und mit der Formel "in Anführungszeichen/-strichen" Schritt für Schritt die Ressourcen und Tools für systematische korpus- und datenbankgesteuerte Recherchen und Analysen vorgestellt und illustriert.

(Best) Practices for Annotating and Representing CMC and Social Media Corpora in CLARIN-D (2016)

Beißwenger, Michael ; Ehrhardt, Eric ; Herold, Axel ; Lüngen, Harald ; Storrer, Angelika

The paper reports the results of the curation project ChatCorpus2CLARIN. The goal of the project was to develop a workflow and resources for the integration of an existing chat corpus into the CLARIN-D research infrastructure for language resources and tools in the Humanities and the Social Sciences (http://clarin-d.de). The paper presents an overview of the resources and practices developed in the project, describes the added value of the resource after its integration and discusses, as an outlook, to what extent these practices can be considered best practices which may be useful for the annotation and representation of other CMC and social media corpora.

Das Dortmunder Chat-Korpus in CLARIN-D: Modellierung und Mehrwerte (2016)

Beißwenger, Michael ; Herold, Axel ; Lüngen, Harald ; Storrer, Angelika

Integrating corpora of computer-mediated communication in CLARIN-D: Results from the curation project ChatCorpus2CLARIN (2016)

Lüngen, Harald ; Beißwenger, Michael ; Ehrhardt, Eric ; Herold, Axel ; Storrer, Angelika

We introduce our pipeline to integrate CMC and SM corpora into the CLARIN-D corpus infrastructure. The pipeline was developed by transforming an existing CMC corpus, the Dortmund Chat Corpus, into a resource conforming to current technical and legal standards. We describe how the resource has been prepared and restructured in terms of TEI encoding, linguistic annotations, and anonymisation. The output is a CLARIN-conformant resource integrated in the CLARIN-D research infrastructure.

Converting and Representing Social Media Corpora into TEI: Schema and best practices from CLARIN-D (2016)

Beißwenger, Michael ; Ehrhardt, Eric ; Herold, Axel ; Lüngen, Harald ; Storrer, Angelika

The paper presents results from a curation project within CLARIN-D, in which an existing lMWord corpus of German chat communication has been integrated into the DEREKO and DWDS corpus infrastructures of the CLARIN-D centres at the Institute for the German Language (IDS, Mannheim) and at the Berlin-Brandenburg Academy of Sciences (BBAW, Berlin). The focus is on the solutions developed for converting and representing the corpus in a TEI format.

4th Workshop on Challenges in the Management of Large Corpora. (May 28th 2016, Portorož; part of the LREC-2016 workshop structure) / LREC 2016, CMLC-4. (2016)

Topical Diversification Over Time In The Royal Society Corpus (2016)

Fankhauser, Peter ; Knappen, Jörg ; Teich, Elke

Good practices in the compilation of FOLK, the research and teaching corpus of spoken German (2016)

Schmidt, Thomas

The paper presents practices in the compilation of FOLK, the Research and Teaching Corpus of Spoken German, a large collection of spontaneous verbal interaction from diverse discourse domains. After introducing the aims and organisational circumstances of the construction of FOLK, the general idea discussed is that good practices cannot be developed without considering methodological, technological and organisational aspects on equal footing. Starting from this idea, this paper inspects more closely some actual practices in FOLK, namely the handling of legal (especially privacy protection) issues, the decisions taken for the transcription and annotation workflow, and the question of how to best disseminate a corpus like FOLK. The final section sketches some possible future improvements for practices in FOLK.

Automatic Classification by Topic Domain for Meta Data Generation, Web Corpus Evaluation, and Corpus Comparison (2016)

Schäfer, Roland ; Bildhauer, Felix

In this paper, we describe preliminary results from an ongoing experiment wherein we classify two large unstructured text corpora—a web corpus and a newspaper corpus—by topic domain (or subject area). Our primary goal is to develop a method that allows for the reliable annotation of large crawled web corpora with meta data required by many corpus linguists. We are especially interested in designing an annotation scheme whose categories are both intuitively interpretable by linguists and firmly rooted in the distribution of lexical material in the documents. Since we use data from a web corpus and a more traditional corpus, we also contribute to the important field of corpus comparison and corpus evaluation. Technically, we use (unsupervised) topic modeling to automatically induce topic distributions over gold standard corpora that were manually annotated for 13 coarse-grained topic domains. In a second step, we apply supervised machine learning to learn the manually annotated topic domains using the previously induced topics as features. We achieve around 70% accuracy in 10-fold cross validations. An analysis of the errors clearly indicates, however, that a revised classification scheme and larger gold standard corpora will likely lead to a substantial increase in accuracy.

DRuKoLA – towards contrastive German-Romanian research based on comparable corpora (2016)

Cosma, Ruxandra ; Cristea, Dan ; Kupietz, Marc ; Tufiş, Dan ; Witt, Andreas

This paper introduces the recently started DRuKoLA-project that aims at providing mechanisms to flexibly draw virtual comparable corpora from the German Reference Corpus DeReKo and the Reference Corpus of Contemporary Romanian Language CoRoLa in order to use these virtual corpora as empirical basis for contrastive linguistic research.

Names in competition: A corpus-based quantitative investigation into the use of colonial place names (2016)

Engelberg, Stefan

Referentially equivalent toponyms occur very often in colonial and postcolonial contexts. These names are in competition, and this competition is reflected in language use and in changing frequencies of use in large corpora. The main theoretical and methodological assumption of this paper is that corpus frequencies of referentially equivalent toponyms change according to particular patterns, and that the Google Ngram Corpora and Google Ngram Viewers can be used to detect these patterns. The aims of this paper are twofold: firstly, a corpus-linguistic method for investigations into the use of names will be presented, applied, and critically evaluated; secondly, it will be shown that the correlation between patterns of frequency changes and patterns of socio-historical colonial and postcolonial events gives rise to cross-linguistic generalizations, for example, that an increase in public interest in a place strongly promotes one of the referenlially equivalent names, or that in renaming scenarios colonial toponyms in relation to new toponyms remain in stronger use in the language of the former colonial power than in languages of other colonial powers.

Collocation(s) in German minds (2016)

Perkuhn, Rainer

German research on collocation(s) focuses on many different aspects. A comprehensive documentation would be impossible in this short report. Accepting that we cannot do justice to all the contributions to this area, we just pick out some influential comerstones. This selection does not claim to be representative or balanced, but it follows the idea to constitute the backbone of the story we want to tell: Our ‘German’ view of the still ongoing evolution of a notion of ‘collocation’ Although our own work concerns the theoretical background of and the empirical rationale for collocations, lexicography occupies a large space. Some of the recent publications ( Wahrig 2008, Häcki Buhofer et al. 2014) represent a turn to the empirical legitimation for the selection of typical expressions. Nevertheless, linking the empirical evidence to the needs of an abstract lexicographic description (or a didactic format) is still an open issue.

FOLK-Gold ― A gold standard for part-of-speech-tagging of spoken German (2016)

Westpfahl, Swantje ; Schmidt, Thomas

In this paper, we present a GOLD standard of part-of-speech tagged transcripts of spoken German. The GOLD standard data consists of four annotation layers – transcription (modified orthography), normalization (standard orthography), lemmatization and POS tags – all of which have undergone careful manual quality control. It comes with guidelines for the manual POS annotation of transcripts of German spoken data and an extended version of the STTS (Stuttgart Tübingen Tagset) which accounts for phenomena typically found in spontaneous spoken German. The GOLD standard was developed on the basis of the Research and Teaching Corpus of Spoken German, FOLK, and is, to our knowledge, the first such dataset based on a wide variety of spontaneous and authentic interaction types. It can be used as a basis for further development of language technology and corpus linguistic applications for German spoken language.

User, who art thou? User profiling for oral corpus platforms (2016)

Fandrych, Christian ; Frick, Elena ; Hedeland, Hanna ; Iliash, Anna ; Jettka, Daniel ; Meißner, Cordula ; Schmidt, Thomas ; Wallner, Franziska ; Weigert, Kathrin ; Westpfahl, Swantje

This contribution presents the background, design and results of a study of users of three oral corpus platforms in Germany. Roughly 5.000 registered users of the Database for Spoken German (DGD), the GeWiss corpus and the corpora of the Hamburg Centre for Language Corpora (HZSK) were asked to participate in a user survey. This quantitative approach was complemented by qualitative interviews with selected users. We briefly introduce the corpus resources involved in the study in section 2. Section 3 describes the methods employed in the user studies. Section 4 summarizes results of the studies focusing on selected key topics. Section 5 attempts a generalization of these results to larger contexts.

Constructing a Corpus (2016)

Kupietz, Marc

Corpus Query Lingua Franca (CQLF) (2016)

Bański, Piotr ; Frick, Elena ; Witt, Andreas

The present paper describes Corpus Query Lingua Franca (ISO CQLF), a specification designed at ISO Technical Committee 37 Subcommittee 4 “Language resource management” for the purpose of facilitating the comparison of properties of corpus query languages. We overview the motivation for this endeavour and present its aims and its general architecture. CQLF is intended as a multi-part specification; here, we concentrate on the basic metamodel that provides a frame that the other parts fit in.

KorAP architecture – diving in the deep sea of corpus data (2016)

Diewald, Nils ; Hanl, Michael ; Margaretha, Eliza ; Bingel, Joachim ; Kupietz, Marc ; Bański, Piotr ; Witt, Andreas

KorAP is a corpus search and analysis platform, developed at the Institute for the German Language (IDS). It supports very large corpora with multiple annotation layers, multiple query languages, and complex licensing scenarios. KorAP’s design aims to be scalable, flexible, and sustainable to serve the German Reference Corpus DEREKO for at least the next decade. To meet these requirements, we have adopted a highly modular microservice-based architecture. This paper outlines our approach: An architecture consisting of small components that are easy to extend, replace, and maintain. The components include a search backend, a user and corpus license management system, and a web-based user frontend. We also describe a general corpus query protocol used by all microservices for internal communications. KorAP is open source, licensed under BSD-2, and available on GitHub.

Dmitrij Dobrovol'skij. 2013. Besedy o nemeckom slove. Studien zur deutschen Lexik. (=Studia Philologica). Moskau: Jazyki Slavjansko Kul'tury [Rezension] (2016)

Volodina, Anna

Der Aufmerksamkeitswettkampf um ADHS. Eine linguistische Konfliktanalyse im Bereich medizinisch-psychologischer Faktizitätsherstellung (2016)

Schnedermann, Theresa

Analyzing lexical change in diachronic corpora (2016)

Koplenig, Alexander

This thesis consists of the following three papers that all have been published in international peer-reviewed journals: Chapter 3: Koplenig, Alexander (2015c). The Impact of Lacking Metadata for the Measurement of Cultural and Linguistic Change Using the Google Ngram Data Sets—Reconstructing the Composition of the German Corpus in Times of WWII. Published in: Digital Scholarship in the Humanities. Oxford: Oxford University Press. [doi:10.1093/llc/fqv037] Chapter 4: Koplenig, Alexander (2015b). Why the quantitative analysis of dia-chronic corpora that does not consider the temporal aspect of time-series can lead to wrong conclusions. Published in: Digital Scholarship in the Humanities. Oxford: Oxford University Press. [doi:10.1093/llc/fqv030] Chapter 5: Koplenig, Alexander (2015a). Using the parameters of the Zipf–Mandelbrot law to measure diachronic lexical, syntactical and stylistic changes – a large-scale corpus analysis. Published in: Corpus Linguistics and Linguistic Theory. Berlin/Boston: de Gruyter. [doi:10.1515/cllt-2014-0049] Chapter 1 introduces the topic by describing and discussing several basic concepts relevant to the statistical analysis of corpus linguistic data. Chapter 2 presents a method to analyze diachronic corpus data and a summary of the three publications. Chapters 3 to 5 each represent one of the three publications. All papers are printed in this thesis with the permission of the publishers.

Konflikt, Sprache, korpuslinguistische Methodik (2016)

Perkuhn, Rainer ; Belica, Cyril

Open Access

Refine

Author

Year of publication

Document Type

Language

Has Fulltext

Is part of the Bibliography

Keywords

Publicationstate

Reviewstate

Publisher

47 search hits