TY - CPAPER U1 - Konferenzveröffentlichung A1 - Faaß, Gertrud A1 - Heid, Ulrich A1 - Taljard, Elsabe A1 - Prinsloo, Danie ED - De Pauw, Guy ED - de Schryver, Gilles-Maurice ED - Levin, Lori T1 - Part-of-Speech tagging of Northern Sotho: Disambiguating polysemous function words T2 - Proceedings of the First Workshop on Language Technologies for African Languages N2 - A major obstacle to part-of-speech (=POS) tagging of Northern Sotho (Bantu, S 32) are ambiguous function words. Many are highly polysemous and very frequent in texts, and their local context is not always distinctive. With certain taggers, this issue leads to comparatively poor results (between 88 and 92 % accuracy), especially when sizeable tagsets (over 100 tags) are used. We use the RF-tagger (Schmid and Laws,2008), which is particularly designed for the annotation of fine-grained tagsets (e.g. including agreement information), and we restructure the 141 tags of the tagset proposed by Taljard et al. (2008) in a way to fit the RF tagger. This leads to over 94 % accuracy. Error analysis in addition shows which types of phenomena cause trouble in the POS-tagging of Northern Sotho. KW - Nordsotho KW - Polysemie KW - Funktionswort KW - Methodologie KW - Bantusprachen Y1 - 2009 U6 - https://nbn-resolving.org/urn:nbn:de:bsz:mh39-118813 UN - https://nbn-resolving.org/urn:nbn:de:bsz:mh39-118813 UR - https://aclanthology.org/volumes/W09-07/ SN - 1-932432-25-6 SB - 1-932432-25-6 SP - 38 EP - 45 PB - Association for Computational Linguistics CY - Stroudsburg ER -