Modelling Linguistic Data Structures
- Linguistic corpora have been annotated by means of SGML-based markup languages for almost 20 years. We can, very roughly, differentiate between three distinct evolutionary stages of markup technologies. (1)Originally, single SGML tree-based document instances were deemed sufficient for the representation of linguistic structures. (2) Linguists began to realize that alternatives and extensions to the traditional model are needed. Formalisms such as, for example, NITE were proposed: the NITE Object Model (NOM) consists of multi-rooted trees. (3) We are now on the threshold of the third evolutionary stage: even NITE's very flexible approach is not suited for all linguistic purposes. As some structures, such as these, cannot be modeled by multi-rooted trees, an even more flexible approach is needed in order to provide a generic annotation format that is able to represent genuinely arbitrary linguistic data structures.
Author: | Kai Wörner, Andreas WittORCiDGND, Georg Rehm, Stefanie Dipper |
---|---|
URN: | urn:nbn:de:bsz:mh39-45173 |
Parent Title (English): | Proceedings of Extreme Markup Languages 2006 |
Publisher: | Extreme Markup Languages Conference |
Place of publication: | Montreal |
Document Type: | Conference Proceeding |
Language: | English |
Year of first Publication: | 2006 |
Date of Publication (online): | 2015/12/22 |
Publicationstate: | Veröffentlichungsversion |
Tag: | Markup Languages; Modeling; Trees/Graphs |
Page Number: | 13 |
DDC classes: | 400 Sprache / 410 Linguistik |
Open Access?: | ja |
Linguistics-Classification: | Korpuslinguistik |
Licence (German): | ![]() |