OPUS 4 | Computerlinguistik

Lessons learned from joining forces across disparate disciplines in the NFDI: the conference on research on text analytics. Presented at the 1st Conference on Research Data Infrastructure (CoRDI), Karlsruhe, 12. – 14. September 2023 (2023)

Krieger, Ulrich ; Trippel, Thorsten

This contribution summarizes the lessons learned from the organization of a joint conference on text analytics research by the Business, Economic, and Related Data (BERD@NFDI) and Text+ consortia within the National Research Data Infrastructure (NFDI) in Germany. The collaboration aimed to identify common ground and foster interdisciplinary dialogue between scholars in the humanities and in the business domain. The lessons learned include the importance of presenting research questions using textual data to establish common ground, similarities in methodology for processing textual data between the consortia, similarities in research data management, and the need for regular interconsortial discussions on textual analysis methods and data. The collaboration proved valuable for interdisciplinary dialogue within the NFDI, and further collaboration between the consortia is planned.

Data for my research: Where can I get it, where do I take it, what can I do with it, how do I use it in my resume? Looking at the research data infrastructures CLARIN in Europe and Text+ in Germany. Presented at the CLEOPATRA final public workshop, Hannover, 2023-05-15 (2023)

Trippel, Thorsten

"Reproducibility crisis" and "empirical turn" are only two keywords when it comes to providing reasons for research data management. Research data is omnipresent and with the more and more automatic data processing procedures, they become even more important. However, just because new methods require data and produce data, this does not mean that data are easily accessible, reusable or even make a difference in the CV of a researcher, even if a large portion of research goes into data creation, acquisition, preparation, and analysis. In this talk I will present where we find data in the research process, where we may find appropriate support for data management and advocate for a procedure for including it in research publications and resumes. This presentation relies on work within the BMBF-funded project CLARIN-D. It also builds on work within the German National Research Data Infrastructure (NFDI) consortium Text+, DFG project number 460033370.

KoMuX - Der Kompositamuster-Explorer (2023)

Brunner, Annelen ; Hein, Katrin

KoMuX, der Kompositamuster-Explorer, (www.owid.de/plus/komux) ist eine Webanwendung, die es ermöglicht, mehr als 50.000 nominale Komposita des Deutschen gezielt nach abstrakten oder lexikalisch-teilspezifizierten Mustern zu durchsuchen. Unterschiedliche Visualisierungen helfen dabei, Strukturen und Zusammenhänge innerhalb der Ergebnismenge zu erfassen.

Projektvorstellung – Sprachanfragen. Empirisch gestützte Erforschung von Zweifelsfällen (2023)

Lang, Christian ; Tu, Ngoc Duyen Tanja ; Schneider, Roman ; Volodina, Anna

"Das im Januar 2022 gestartete Projekt "Sprachanfragen" (https://www.ids-mannheim.de/gra/projekte2/sprachanfragen/) verfolgt erstmalig das Ziel, Sprachanfragedaten zu erfassen, aufzubereiten und ein wissenschaftsöffentliches Monitorkorpus aus ihnen zu erstellen. Dazukommend wird eine Rechercheschnittstelle entwickelt, mit der die Sprachanfragen systematisch wissenschaftlich analysierbar gemacht werden. Das Poster gibt einen Überblick über das Projekt, zeigt erste Ergebnisse und bietet einen Ausblick auf Überlegungen zur Konzeption eines Chatbots zur automatisierten Beantwortung von Sprachanfragen." Ein Beitrag zur 9. Tagung des Verbands "Digital Humanities im deutschsprachigen Raum" - DHd 2023 Open Humanities Open Culture.

Korpora modular, verteilt, vernetzt in Text+ (2023)

Leinen, Peter ; Trippel, Thorsten ; Weimer, Lukas ; Witt, Andreas

Als Teil der NFDI vernetzt Text+ ortsverteilt verschiedenste Daten und Dienste für die geisteswissenschaftliche Forschung und stellt sie der wissenschaftlichen Gemeinschaft FAIR zur Verfügung. In diesem Beitrag beschreiben wir die Umsetzung beispielhaft im Bereich der Text+ Datendomäne Sammlungen anhand von Korpora, die in verschiedenen Disziplinen Verwendung finden. Die Infrastruktur ist auf Erweiterbarkeit ausgelegt, so dass auch weitere Ressourcen über Text+ verfügbar gemacht werden können. Enthalten ist auch ein Ausblick auf weitere zu erwartende Entwicklungen. Ein Beitrag zur 9. Tagung des Verbands "Digital Humanities im deutschsprachigen Raum" - DHd 2023 Open Humanities Open Culture.

CLARIAH-DE work package 5 - community engagement: outreach/dissemination and liaison (2021)

Walker, Nathalie ; Werthmann, Antonina ; Trippel, Thorsten ; Buddenbohm, Stefan ; Weimer, Lukas ; Friedrichs, Sonja

This poster summarizes the results of the CLARIAH-DE Work Package 5 - Community Engagement: Outreach/Dissemination and Liaison. Work package 5 engages with the community through dissemination activities, outreach and liaison. The work package set itself the following sub goals: - Combining the existing dissemination and outreach activities of CLARIN-D and DARIAH-DE in a meaningful way and elaborating on them. In some cases this meant continuity, in other cases a new appearance for resources. - Providing a web portal as a gateway to the CLARIAH-DE project. - Creating a common identity and corporate identity and maintaining the established level of trust users already put into CLARIN-D and DARIAH-DE. - Providing a social media presence as well as a physical presence at workshops, conferences and other meetings in the Digital Humanities.

Shallow context analysis for German idiom detection (2021)

Amin, Miriam ; Fankhauser, Peter ; Kupietz, Marc ; Schneider, Roman

In order to differentiate between figurative and literal usage of verb-noun combinations for the shared task on the disambiguation of German Verbal Idioms issued for KONVENS 2021, we apply and extend an approach originally developed for detecting idioms in a dataset consisting of random ngram samples. The classification is done by implementing a rather shallow, statistics-based pipeline without intensive preprocessing and examinations on the morphosyntactic and semantic level. We describe the overall approach, the differences between the original dataset and the dataset of the KONVENS task, provide experimental classification results, and analyse the individual contributions of our feature sets.

CLARIAH-DE work package 3: skills training and promotion of junior researchers (2021)

Annisius, Marie ; Bock, Sina ; Gradl, Tobias ; Schopf, Juliane ; Stegmeier, Jörn ; Werthmann, Antonina

This poster summarizes the results of the CLARIAH-DE Work Package 3: Skills Training and Promotion of Junior Researchers. For a research field that is characterised by rapid technical development, CLARIAH-DE has to include the promotion of data literacy necessary for the efficient use of this digital research infrastructure as part of its objective. To develop, consolidate and refine a common programme in this area, work package 3 set itself the following sub goals: - Consolidation of the activities from the previous projects into a joint service - Cataloguing and reflecting on the methods and tools used in the research field, with the aim of identifying remaining gaps - Skills training of, individual support for and the promotion of junior researchers

Semantische Suche mit Word Embeddings für ein mehrsprachiges Wörterbuchportal (2022)

Tu, Ngoc Duyen Tanja ; Meyer, Peter

Das Lehnwortportal Deutsch (LWPD) ist ein Online-Informationssystem zu Entlehnungen von Wörtern aus dem Deutschen in andere Sprachen. Es beruht auf einer wachsenden Zahl von lexikographischen Ressourcen zu verschiedenen Sprachen und bietet eine einfache ressourcenübergreifende Suchfunktion an. Das Poster präsentiert eine derzeit in Entwicklung befindliche onomasiologische Suchfunktion für das LWPD.

Der CLARIAH-DE Tutorial Finder. Eine Suchumgebung für Lehr- und Schulungsmaterialien in den Digital Humanities (2022)

Werthmann, Antonina ; Gradl, Tobias

Um eine bessere Erreichbarkeit und Zugänglichkeit zu bestehenden sowie neuen Angeboten von Lehr- und Schulungsmaterialien im Bereich der Digital Humanities zu ermöglichen, sollten diese in einem zentralen Verzeichnis zur Verfügung gestellt werden. Im Rahmen des CLARIAH-DE Projekts wurde – zunächst für die Umsetzung eines Projektmeilensteins – eine Lösung gesucht, die eine übergreifende Suche in frei zugänglichen und nachnutzbaren Lehr- und Schulungsmaterialien zu Forschungsmethoden, Verfahren sowie Werkzeugen im Bereich der Digital Humanities in unterschiedlichen Plattformen und Repositorien bietet.

Open Access

Computerlinguistik

Refine

Author

Year of publication

Document Type

Language

Has Fulltext

Is part of the Bibliography

Keywords

Publicationstate

Reviewstate

Publisher

17 search hits