OPUS 4 | Search

Refine

Has Fulltext

yes (3)

3 search hits

1 to 3

Sort by

The linguistic construal of disciplinarity: A data-mining approach using register features (2015)

Teich, Elke ; Degaetano-Ortlieb, Stefania ; Fankhauser, Peter ; Kermes, Hannah ; Lapshinova-Koltunski, Ekaterina

We analyze the linguistic evolution of selected scientific disciplines over a 30-year time span (1970s to 2000s). Our focus is on four highly specialized disciplines at the boundaries of computer science that emerged during that time: computational linguistics, bioinformatics, digital construction, and microelectronics. Our analysis is driven by the question whether these disciplines develop a distinctive language use—both individually and collectively—over the given time period. The data set is the English Scientific Text Corpus (scitex), which includes texts from the 1970s/1980s and early 2000s. Our theoretical basis is register theory. In terms of methods, we combine corpus-based methods of feature extraction (various aggregated features [part-of-speech based], n-grams, lexico-grammatical patterns) and automatic text classification. The results of our research are directly relevant to the study of linguistic variation and languages for specific purposes (LSP) and have implications for various natural language processing (NLP) tasks, for example, authorship attribution, text mining, or training NLP tools.

Combining macro- and microanalysis for exploring the construal of scientific disciplinarity (2014)

Fankhauser, Peter ; Kermes, Hannah ; Teich, Elke

Data Mining with Shallow vs. Linguistic Features to Study Diversification of Scientific Registers (2014)

Degaetano-Ortlieb, Stefania ; Fankhauser, Peter ; Kermes, Hannah ; Lapshinova-Koltunski, Ekaterina ; Ordan, Noam ; Teich, Elke

We present a methodology to analyze the linguistic evolution of scientific registers with data mining techniques, comparing the insights gained from shallow vs. linguistic features. The focus is on selected scientific disciplines at the boundaries to computer science (computational linguistics, bioinformatics, digital construction, microelectronics). The data basis is the English Scientific Text Corpus (SCITEX) which covers a time range of roughly thirty years (1970/80s to early 2000s) (Degaetano-Ortlieb et al., 2013; Teich and Fankhauser, 2010). In particular, we investigate the diversification of scientific registers over time. Our theoretical basis is Systemic Functional Linguistics (SFL) and its specific incarnation of register theory (Halliday and Hasan, 1985). In terms of methods, we combine corpus-based methods of feature extraction and data mining techniques.

1 to 3

Open Access

Refine

Author

Year of publication

Document Type

Language

Has Fulltext

Is part of the Bibliography

Keywords

Publicationstate

Reviewstate

Publisher

3 search hits