OPUS 4 | Search

The International Comparable Corpus: Challenges in building multilingual spoken and written comparable corpora (2021)

Čermáková, Anna ; Jantunen, Jarmo ; Jauhiainen, Tommi ; Kirk, John ; Křen, Michal ; Kupietz, Marc ; Uí Dhonnchadha, Elaine

This paper reports on the efforts of twelve national teams in building the International Comparable Corpus (ICC; https://korpus.cz/icc) that will contain highly comparable datasets of spoken, written and electronic registers. The languages currently covered are Czech, Finnish, French, German, Irish, Italian, Norwegian, Polish, Slovak, Swedish and, more recently, Chinese, as well as English, which is considered to be the pivot language. The goal of the project is to provide much-needed data for contrastive corpus-based linguistics. The ICC corpus is committed to the idea of re-using existing multilingual resources as much as possible and the design is modelled, with various adjustments, on the International Corpus of English (ICE). As such, ICC will contain approximately the same balance of forty percent of written language and 60 percent of spoken language distributed across 27 different text types and contexts. A number of issues encountered by the project teams are discussed, ranging from copyright and data sustainability to technical advances in data distribution.

Rule talk: Instructing proper play with impersonal deontic statements (2021)

Zinken, Jörg ; Kaiser, Julia ; Weidner, Matylda ; Mondada, Lorenza ; Rossi, Giovanni ; Sorjonen, Marja-Leena

The present paper explores how rules are enforced and talked about in everyday life. Drawing on a corpus of board game recordings across European languages, we identify a sequential and praxeological context for rule talk. After a game rule is breached, a participant enforces proper play and then formulates a rule with an impersonal deontic statement (e.g. “It’s not allowed to do this”). Impersonal deontic statements express what may or may not be done without tying the obligation to a particular individual. Our analysis shows that such statements are used as part of multi-unit and multi-modal turns where rule talk is accomplished through both grammatical and embodied means. Impersonal deontic statements serve multiple interactional goals: they account for having changed another’s behavior in the moment and at the same time impart knowledge for the future. We refer to this complex action as an “instruction.” The results of this study advance our understanding of rules and rule-following in everyday life, and of how resources of language and the body are combined to enforce and formulate rules.

Anglizismen in der Coronakrise (2021)

Zifonun, Gisela

Anglizismen in der Coronakrise (2021)

Zifonun, Gisela

Zwischenruf zu „Herdenimmunität“ (2021)

Zifonun, Gisela

Zwischenruf zu „Warum eine Maske für Mund und Nase Mund-Nasen-Maske heißt“ (2021)

Zifonun, Gisela

Zwischenruf zu „Neue Normalität“ (2021)

Zifonun, Gisela

Zwischenruf zu „Soziale Distanz“ (2021)

Zifonun, Gisela

Eine Linguistin denkt nach über den Genderstern (2021)

Zifonun, Gisela

Das Deutsche als europäische Sprache: Ein Porträt (2021)

Zifonun, Gisela

Das Deutsche ist eine der am besten erforschten Sprachen der Welt; weniger bekannt ist, welche Gemeinsamkeiten es mit den europäischen Nachbarsprachen teilt und wo seine Besonderheiten liegen. Die insgesamt acht Kapitel des Buches stellen prägnant und anhand von anschaulichen Beispielen Wortschatz und Grammatik des Deutschen vor. Dabei verhilft ein Vergleich mit den Optionen etwa im Englischen, Französischen, Polnischen, Ungarischen oder anderen europäischen Sprachen zu einem verschärften Blick. Ausgangspunkt ist dabei ein kurzer Abriss der Facetten von Sprache allgemein sowie die Herleitung der grundlegenden Sprachfunktionen aus einer handlungsbezogenen Perspektive. Die folgenden Kapitel stehen unter Motti wie: „Das Verb – Zeiten, Modi, Szenarios und Inszenierungen“, „Der nominale Bereich – die vielerlei Arten, Gegenstände zu konstruieren“ oder „Der Text – wenn wir kohärent und dabei narrativ oder argumentativ werden“. Das letzte Kapitel trägt den Titel: „Das Deutsche – auf dem Weg zu einem Sprachporträt“. Das Buch soll Sprachinteressierten auch ohne linguistische Fachkenntnisse einen neuen Zugang zu unserer Muttersprache erschließen und die Sensibilität für die sprachliche Verbundenheit auf unserem Kontinent trotz aller Vielfalt stärken. - Grammatik anschaulich und konkret - Innovativer Blick auf das Deutsche im Kreis europäischer Sprachen - Kurzweilige Einführung für Sprachinteressierte auch ohne linguistische Fachkenntnisse

Vorwort (2021)

Ziegler, Evelyn ; Marten, Heiko F.

Linguistic Landscapes in deutschsprachigen Kontexten (2021)

Ziegler, Evelyn ; Marten, Heiko F.

This chapter starts out by giving a brief overview of the main priorities of international and German studies in the area of linguistic landscape research. The contributions to this volume are then embedded in current debates and developments in the field. Finally, we outline important desiderata of linguistic landscape research that focus on German and address challenges of knowledge transfer and application as well as possible contributions to international lines of research.

Einleitung (2021)

Wöllstein, Angelika

Mit dem zweiten Band werden vier neue „Bausteine“ zu einer korpuslinguistisch fundierten Grammatik des Deutschen vorgelegt. Sie behandeln die Bereiche Determination, syntaktische Funktionen der Nominalphrase und Attribution. Dem Fachpublikum werden zugleich die analysierten Sprachdaten und vertiefende Zusatzuntersuchungen zugänglich gemacht.

Nutzungsstatistiken von linguistischen Online-Ressourcen: Potentiale und Limitationen (2021)

Wolfer, Sascha ; Michaelis, Frank ; Müller-Spitzer, Carolin

Dictionary usage research views dictionaries primarily as tools for solving linguistic problems. A large proportion of dictionary use now takes place online and can thus be easily monitored using tracking technologies. Using the data gathered through tracking usage data, we hope to optimize user experiences of dictionaries and other linguistic resources. Usage statistics are also used for external evaluation of linguistic resources. In this paper, we pursue the following three questions from a quantitative perspective: (1) What new insights can we gain from collecting and analysing usage data? (2) What limitations of the data and/or the collection process do we need to be aware of? (3) How can these insights and limitations inform the development and evaluation of linguistic resources?

cOWIDplus Analyse: Wie sehr schränkt die Corona-Krise das Vokabular deutschsprachiger Online-Presse ein? (2021)

Wolfer, Sascha ; Koplenig, Alexander ; Michaelis, Frank ; Müller-Spitzer, Carolin

cOWIDplus Analyse ist eine kontinuierlich aktualisierte Ressource zu der Frage, ob und wie stark sich der Wortschatz ausgewählter deutscher Online-Pressemeldungen während der Corona-Pandemie systematisch einschränkt und ob bzw. wann sich das Vokabular nach der Krise wieder ausweitet. In diesem Artikel erläutern die Autor*innen die hinter der Ressource stehende Forschungsfrage, die zugrunde gelegten Daten, die Methode sowie die bisherigen Ergebnisse.

Toilettenpapier im April, Mutationen im Dezember: Einflüsse der Corona-Pandemie auf die deutsche Sprache (2021)

Wolfer, Sascha

Am 24. Februar 2020 wurde in der Schweiz die erste Infektion mit dem Coronavirus nachgewiesen. Zu diesem Zeitpunkt konnte wohl noch niemand ahnen, welche tiefgreifenden Konsequenzen die Corona-Pandemie für die Gesellschaft haben wird. Aus heutiger Perspektive überrascht es uns nicht mehr, dass das Pandemiegeschehen auch starke Auswirkungen auf die Sprache hatte und noch immer hat, denn Sprachgebrauch passt sich stets gesellschaftlichen Veränderungen an. Am Leibniz-Institut für Deutsche Sprache in Mannheim dokumentieren und erforschen wir die ungewöhnlich starken und kurzfristigen Wirkungen der Pandemie auf die deutsche Sprache und fassen unsere Ergebnisse unter anderem in zahlreichen Beiträgen zusammen.

Implicitly abusive language – What does it actually look like and why are we not getting there? (2021)

Wiegand, Michael ; Ruppenhofer, Josef ; Eder, Elisabeth

Abusive language detection is an emerging field in natural language processing which has received a large amount of attention recently. Still the success of automatic detection is limited. Particularly, the detection of implicitly abusive language, i.e. abusive language that is not conveyed by abusive words (e.g. dumbass or scum), is not working well. In this position paper, we explain why existing datasets make learning implicit abuse difficult and what needs to be changed in the design of such datasets. Arguing for a divide-and-conquer strategy, we present a list of subtypes of implicitly abusive language and formulate research tasks and questions for future research.

Exploiting emojis for abusive language detection (2021)

Wiegand, Michael ; Ruppenhofer, Josef

We propose to use abusive emojis, such as the “middle finger” or “face vomiting”, as a proxy for learning a lexicon of abusive words. Since it represents extralinguistic information, a single emoji can co-occur with different forms of explicitly abusive utterances. We show that our approach generates a lexicon that offers the same performance in cross-domain classification of abusive microposts as the most advanced lexicon induction method. Such an approach, in contrast, is dependent on manually annotated seed words and expensive lexical resources for bootstrapping (e.g. WordNet). We demonstrate that the same emojis can also be effectively used in languages other than English. Finally, we also show that emojis can be exploited for classifying mentions of ambiguous words, such as “fuck” and “bitch”, into generally abusive and just profane usages.

Implicitly abusive comparisons – a new dataset and linguistic analysis (2021)

Wiegand, Michael ; Geulig, Maja ; Ruppenhofer, Josef

We examine the task of detecting implicitly abusive comparisons (e.g. “Your hair looks like you have been electrocuted”). Implicitly abusive comparisons are abusive comparisons in which abusive words (e.g. “dumbass” or “scum”) are absent. We detail the process of creating a novel dataset for this task via crowdsourcing that includes several measures to obtain a sufficiently representative and unbiased set of comparisons. We also present classification experiments that include a range of linguistic features that help us better understand the mechanisms underlying abusive comparisons.

Verbundprojekt CLARIAH-DE – Eine nachhaltige Forschungsinfrastruktur für die Geistes-, Kultur- und Sozialwissenschaften (2021)

Werthmann, Antonina ; Witt, Andreas ; Bopp, Jutta

Das vom BMBF geförderte Verbundprojekt CLARIAH-DE, an dem über 25 Partnerinstitutionen mitwirken, unter ihnen auch das IDS, hat zum Ziel, mit der Entwicklung einer Forschungsinfrastruktur zahlreiche Angebote zur Verfügung zu stellen, die die Bedingungen der Forschungsarbeit mit digitalen Werkzeugen, Diensten sowie umfangreichen Datenbeständen im Bereich der geisteswissenschaftlichen Forschung und benachbarter Disziplinen verbessern. Die in CLARIAH-DE entwickelte Infrastruktur bietet den Forschenden Unterstützung bei der Analyse und Aufbereitung von Sprachdaten für linguistische Untersuchungen in unterschiedlichsten Anwendungskontexten und leistet somit einen Beitrag zur Entwicklung der NFDI.

Open Access

Refine

Author

Year of publication

Document Type

Language

Has Fulltext

Is part of the Bibliography

Keywords

Publicationstate

Reviewstate

Publisher

356 search hits