OPUS 4 | Search

32 search hits

11 to 20

Sort by

Relevancy
Year
Year
Title
Title
Author
Author

Making CONCUR work (2005)

Hilbert, Mirco ; Schonefeld, Oliver ; Witt, Andreas

The SGML feature CONCUR allowed for a document to be simultaneously marked up in multiple conflicting hierarchical tagsets but validated and interpreted in one tagset at a time. Alas, CONCUR was rarely implemented, and XML does not address the problem of conflicting hierarchies at all. The MuLaX document syntax is a non-XML syntax that enables multiply-encoded hierarchies by distinguishing different “layers” in the hierarchy by adding a layer ID as a prefix to the element names. The IDs tie all the elements in a single hierarchy together in an “annotation layer”. Extraction of a single annotation layer results in a well-formed XML document, and each annotation layer may be associated with an XML schema. The MuLaX processing model works on the nodes of one annotation layer at a time through Xpath-like navigation. CONCUR lives!

Linguistically Annotated Corpora: Quality Assurance, Reusability and Sustainability (2008)

Zinsmeister, Heike ; Witt, Andreas ; Kübler, Sandra ; Hinrichs, Erhard

Declarations of Relations, Differences and Transformations between Theory-specific Treebanks: A New Methodology (2003)

Sasaki, Felix ; Witt, Andreas ; Metzing, Dieter

This paper deals with the problem of how to interrelate theory-specific treebanks and how to transform one treebank format to another. Currently, two approaches to achieve these goals can be differentiated. The first creates a mapping algorithm between treebank formats. Categories of a source format are transformed into a target format via a given set of general or language-specific mapping rules. The second relates treebanks via a transformation to a general model of linguistic categories, for example based on the EAGLES recommendations for syntactic annotations of corpora, or relying on the HPSG framework. This paper proposes a new methodology as a solution for these desiderata.

Co-reference in Japanese Task-oriented Dialogues: A Contribution to the Development of Language-specific and Language-general Annotation Schemes and Resources (2004)

Sasaki, Felix ; Witt, Andreas

This paper describes a corpus of Japanese task-oriented dialogues, i.e. its data, annotations, analysis methodology and preliminary results for the modeling of co-referential phenomena. Current corpus based approaches to co-reference concentrate on textual data from English or other European languages. Hence, the emerging language-general models of co-reference miss input from dialogue data of non-European languages. We aim to fill this gap and contribute to a model of co-reference on various language-specific and language-general levels.

SusTEInability of linguistic resources through feature structures (2009)

Witt, Andreas ; Rehm, Georg ; Hinrichs, Erhard ; Lehmberg, Timm ; Stegmann, Jens

This article shows that the TEI tag set for feature structures can be adopted to represent a heterogeneous set of linguistic corpora. The majority of corpora is annotated using markup languages that are based on the Annotation Graph framework, the upcoming Linguistic Annotation Format ISO standard, or according to tag sets defined by or based upon the TEI guidelines. A unified representation comprises the separation of conceptually different annotation layers contained in the original corpus data (e.g. syntax, phonology, and semantics) into multiple XML files. These annotation layers are linked to each other implicitly by the identical textual content of all files. A suitable data structure for the representation of these annotations is a multi-rooted tree that again can be represented by the TEI and ISO tag set for feature structures. The mapping process and representational issues are discussed as well as the advantages and drawbacks associated with the use of the TEI tag set for feature structures as a storage and exchange format for linguistically annotated data.

Guidance through the standards jungle for linguistic resources (2012)

Stührenberg, Maik ; Werthmann, Antonina ; Witt, Andreas

Research today is often performed in collaborated projects composed of project partners with different backgrounds and from different institutions and countries. Standards can be a crucial tool to help harmonizing these differences and to create sustainable resources. However, choosing a standard depends on having enough information to evaluate and compare different annotation and metadata formats. In this paper we present ongoing work on an interactive, collaborative website that collects information on standards in the ﬁeld of linguistics as a means to guide interested researchers.

Corpus Masking: Legally Bypassing Licensing Restrictions for the Free Distribution of Text Collections (2007)

Rehm, Georg ; Witt, Andreas ; Zinsmeister, Heike ; Dellert, Johannes

Multi-Dimensional Markup: N-way relations as a generalisation over possible relations between annotation layers (2008)

Lüngen, Harald ; Witt, Andreas

Multidimensional markup and heterogeneous linguistic resources (2006)

Stührenberg, Maik ; Witt, Andreas ; Goecke, Daniela ; Metzing, Dieter ; Schonefeld, Oliver

The paper discusses two topics: firstly an approach of using multiple layers of annotation is sketched out. Regarding the XML representation this approach is similar to standoff annotation. A second topic is the use of heterogeneous linguistic resources (e.g., XML annotated documents, taggers, lexical nets) as a source for semiautomatic multi-dimensional markup to resolve typical linguistic issues, dealing with anaphora resolution as a case study.

Meaning and interpretation of concurrent markup (2002)

Witt, Andreas

11 to 20

Person(s)
Title
Subject
Abstract
Fulltext
Year(s)

Open Access

Refine

Author

Year of publication

Document Type

Language

Has Fulltext

Is part of the Bibliography

Keywords

Publicationstate

Reviewstate

Publisher

32 search hits