OPUS 4 | 430 Deutsch

430 Deutsch

430 Deutsch (130)
431 Schriftsysteme und Phonologie des Deutschen (1)
432 Etymologie des Deutschen (20)
433 Deutsche Wörterbücher (51)
435 Deutsche Grammatik (111)
437 Varianten des Deutschen (121)
438 Gebrauch des Standard-Deutsch (27)
439 Andere germanische Sprachen (40)

Refine

Has Fulltext

yes (3)

3 search hits

1 to 3

Sort by

Das neue "Gesetz zur Angleichung des Urheberrechts an die aktuellen Erfordernisse der Wissensgesellschaft" und seine Auswirkungen für Digital Humanities (2018)

Kamocki, Pawel ; Ketzan, Erik ; Wildgans, Julia ; Witt, Andreas

New exceptions for Text and Data Mining and their possible impact on the CLARIN infrastructure (2018)

Kamocki, Pawel ; Ketzan, Erik ; Wildgans, Julia ; Witt, Andreas

The proposed paper discusses new exceptions for Text and Data Mining that have recently been adopted in some EU Member States, and probably will soon be adopted also at the EU level. These exceptions are of great significance for language scientists, as they exempt those who compile corpora from the obligation to obtain authorisation from rightholders. However, corpora compiled on the basis of such exceptions cannot be freely shared, which in a long run may have serious consequences for Open Science and the functioning of research infrastructure such as CLARIN ERIC.

Mining corpora of computer-mediated communication: analysis of linguistic features in Wikipedia talk pages using machine learning methods (2014)

Beißwenger, Michael ; Lüngen, Harald ; Margaretha, Eliza ; Pölitz, Christian

Machine learning methods offer a great potential to automatically investigate large amounts of data in the humanities. Our contribution to the workshop reports about ongoing work in the BMBF project KobRA (http://www.kobra.tu-dortmund.de) where we apply machine learning methods to the analysis of big corpora in language-focused research of computer-mediated communication (CMC). At the workshop, we will discuss first results from training a Support Vector Machine (SVM) for the classification of selected linguistic features in talk pages of the German Wikipedia corpus in DeReKo provided by the IDS Mannheim. We will investigate different representations of the data to integrate complex syntactic and semantic information for the SVM. The results shall foster both corpus-based research of CMC and the annotation of linguistic features in CMC corpora.

1 to 3

Open Access

430 Deutsch

Refine

Author

Year of publication

Document Type

Language

Has Fulltext

Is part of the Bibliography

Keywords

Publicationstate

Reviewstate

Publisher

3 search hits