OPUS 4 | Search

(More) common ground for processing spoken language corpora? (2014)

A study on gaps and syntactic boundaries in spoken interaction (2018)

We present a study on gaps in spoken language interaction as a potential candidate for syntactic boundaries. On the basis of an online annotation experiment, we can show that there is an effect of gap duration and gap type on its likelihood of being a syntactic boundary. We discuss the potential of these findings for an automation of the segmentation process.

An exchange format for multimodal annotations (2009)

Schmidt, Thomas ; Duncan, Susan ; Ehmer, Oliver ; Hoyt, Jeffrey ; Kipp, Michael ; Loehr, Dan ; Magnusson, Magnus ; Rose, Travis ; Sloetjes, Han

The paper presents the results of a joint effort of a group of multimodality researchers and tool developers to improve the interoperability between several tools used for the annotation and analysis of multimodality. Each of the tools has specific strengths so that a variety of different tools, working on the same data, can be desirable for project work. However this usually requires tedious conversion between formats. We propose a common exchange format for multimodal annotation, based on the annotation graph (AG) formalism, which is supported by import and export routines in the respective tools. In the current version of this format the common denominator information can be reliably exchanged between the tools, and additional information can be stored in a standardized way.

An exchange format for multimodal annotations (2008)

Schmidt, Thomas ; Duncan, Susan ; Ehmer, Oliver ; Hoyt, Jeffrey ; Kipp, Michael ; Loehr, Dan ; Magnusson, Magnus ; Rose, Travis ; Sloetjes, Han

This paper presents the results of a joint effort of a group of multimodality researchers and tool developers to improve the interoperability between several tools used for the annotation and analysis of multimodality. Each of the tools has specific strengths so that a variety of differ-ent tools, working on the same data, can be desirable for project work. However this usually re-quires tedious conversion between formats. We propose a common exchange format for multi-modal annotation, based on the annotation graph (AG) formalism, which is supported by import and export routines in the respective tools. In the current version of this format the common de-nominator information can be reliably exchanged between the tools, and additional information can be stored in a standardized way.

CLARIN Web Services for TEI-annotated Transcripts of Spoken Language (2020)

Fisseni, Bernhard ; Schmidt, Thomas

We present web services which implement a workflow for transcripts of spoken language following the TEI guidelines, in particular ISO 24624:2016 “Language resource management – Transcription of spoken language”. The web services are available at our website and will be available via the CLARIN infrastructure, including the Virtual Language Observatory and WebLicht.

Connecting resources: Which issues have to be solved to integrate CMC corpora from heterogeneous sources and for different languages? (2017)

Beißwenger, Michael ; Wigham, Ciara ; Etienne, Carole ; Fišer, Darja ; Grumt Suárez, Holger ; Herzberg, Laura ; Hinrichs, Erhard ; Horsmann, Tobias ; Karlova-Bourbonus, Natali ; Lemnitzer, Lothar ; Longhi, Julien ; Lüngen, Harald ; Ho-Dac, Lydia-Mai ; Parisse, Christophe ; Poudat, Céline ; Schmidt, Thomas ; Stemle, Egon W. ; Storrer, Angelika ; Zesch, Torsten

The paper reports on the results of a scientific colloquium dedicated to the creation of standards and best practices which are needed to facilitate the integration of language resources for CMC stemming from different origins and the linguistic analysis of CMC phenomena in different languages and genres. The key issue to be solved is that of interoperability – with respect to the structural representation of CMC genres, linguistic annotations metadata, and anonymization/pseudonymization schemas. The objective of the paper is to convince more projects to partake in a discussion about standards for CMC corpora and for the creation of a CMC corpus infrastructure across languages and genres. In view of the broad range of corpus projects which are currently underway all over Europe, there is a great window of opportunity for the creation of standards in a bottom-up approach.

Introduction: putting practices in spoken corpora into focus (2014)

Ruhi, Şükriye ; Haugh, Michael ; Schmidt, Thomas ; Wörner, Kai

Multilingual corpora at the Hamburg centre for language corpora (2014)

Hedeland, Hanna ; Lehmberg, Timm ; Schmidt, Thomas ; Wörner, Kai

Querying Interaction Structure: Approaches to Overlap in Spoken Language Corpora (2022)

Frick, Elena ; Helmer, Henrike ; Schmidt, Thomas

In this paper, we address two problems in indexing and querying spoken language corpora with overlapping speaker contributions. First, we look into how token distance and token precedence can be measured when multiple primary data streams are available and when transcriptions happen to be tokenized, but are not synchronized with the sound at the level of individual tokens. We propose and experiment with a speaker based search mode that enables any speaker’s transcription tier to be the basic tokenization layer whereby the contributions of other speakers are mapped to this given tier. Secondly, we address two distinct methods of how speaker overlaps can be captured in the TEI based ISO Standard for Spoken Language Transcriptions (ISO 24624:2016) and how they can be queried by MTAS – an open source Lucene-based search engine for querying text with multilevel annotations. We illustrate the problems, introduce possible solutions and discuss their benefits and drawbacks.

Recent Initiatives towards New Standards for Language Resources (2015)

Herzog, Gottfried ; Heid, Ulrich ; Trippel, Thorsten ; Bański, Piotr ; Romary, Laurent ; Schmidt, Thomas ; Witt, Andreas ; Eckart, Kerstin

Reconstruction of separable particle verbs in a corpus of spoken German (2018)

Batinić, Dolores ; Schmidt, Thomas

We present a method for detecting and reconstructing separated particle verbs in a corpus of spoken German by following an approach suggested for written language. Our study shows that the method can be applied successfully to spoken language, compares different ways of dealing with structures that are specific to spoken language corpora, analyses some remaining problems, and discusses ways of optimising precision or recall for the method. The outlook sketches some possibilities for further work in related areas.

Technological and methodological challenges in creating, annotating and sharing a learner corpus of spoken German (2012)

Hedeland, Hanna ; Schmidt, Thomas

This article discusses questions concerning the creation, annotation and sharing of spoken language corpora. We use the Hamburg Map Task Corpus (HAMATAC), a small corpus in which advanced learners of German were recorded solving a map task, as an example to illustrate our main points. We first give an overview of the corpus creation and annotation process including recording, metadata documentation, transcription and semi-automatic annotation of the data. We then discuss the manual annotation of disfluencies as an example case in which many of the typical and challenging problems for data reuse – in particular the reliability of interpretative annotations – are revealed.

The Kicktionary : Combining corpus linguistics and lexical semantics for a multilingual football dictionary (2008)

Schmidt, Thomas

This paper presents the Kicktionary, a multilingual (English - German - French) electronic lexical resource of the language of football. In the Kicktionary, methods from corpus linguistics and two approaches to lexical semantics - the theory of frame semantics and the concept of semantic relations - are combined to construct a lexical resource in which the user can explore relationships between lexical units in various ways. This paper explains the theoretical background of the Kicktionary, sketches the data and methods which were used in its construction, and describes how the resulting resource is presented to users via a set of hyperlinked webpages.

The Kicktionary – A Multilingual Lexical Resource of Football Language (2009)

Schmidt, Thomas

Tools for multimodal annotation (2017)

Cassidy, Steve ; Schmidt, Thomas

Researchers interested in the sounds of speech or the physical gestures of Speakers make use of audio and video recordings in their work. Annotating these recordings presents a different set of requirements to the annotation of text. Special purpose tools have been developed to display video and audio Signals and to allow the creation of time-aligned annotations. This chapter reviews the most widely used of these tools for both manual and automatic generation of annotations on multimodal data.

Using Automatic Speech Recognition in Spoken Corpus Curation (2020)

Gorisch, Jan ; Gref, Michael ; Schmidt, Thomas

The newest generation of speech technology caused a huge increase of audio-visual data nowadays being enhanced with orthographic transcripts such as in automatic subtitling in online platforms. Research data centers and archives contain a range of new and historical data, which are currently only partially transcribed and therefore only partially accessible for systematic querying. Automatic Speech Recognition (ASR) is one option of making that data accessible. This paper tests the usability of a state-of-the-art ASR-System on a historical (from the 1960s), but regionally balanced corpus of spoken German, and a relatively new corpus (from 2012) recorded in a narrow area. We observed a regional bias of the ASR-System with higher recognition scores for the north of Germany vs. lower scores for the south. A detailed analysis of the narrow region data revealed – despite relatively high ASR-confidence – some specific word errors due to a lack of regional adaptation. These findings need to be considered in decisions on further data processing and the curation of corpora, e.g. correcting transcripts or transcribing from scratch. Such geography-dependent analyses can also have the potential for ASR-development to make targeted data selection for training/adaptation and to increase the sensitivity towards varieties of pluricentric languages.

Open Access

Refine

Author

Year of publication

Document Type

Language

Has Fulltext

Is part of the Bibliography

Keywords

Publicationstate

Reviewstate

Publisher

16 search hits