TY - JOUR U1 - Zeitschriftenartikel, wissenschaftlich - begutachtet (reviewed) A1 - Müller, Mark-Christoph A1 - Reitz, Florian A1 - Roy, Nicolas T1 - Data sets for author name disambiguation: an empirical analysis and a new resource JF - Scientometrics N2 - Data sets of publication meta data with manually disambiguated author names play an important role in current author name disambiguation (AND) research. We review the most important data sets used so far, and compare their respective advantages and shortcomings. From the results of this review, we derive a set of general requirements to future AND data sets. These include both trivial requirements, like absence of errors and preservation of author order, and more substantial ones, like full disambiguation and adequate representation of publications with a small number of authors and highly variable author names. On the basis of these requirements, we create and make publicly available a new AND data set, SCAD-zbMATH. Both the quantitative analysis of this data set and the results of our initial AND experiments with a naive baseline algorithm show the SCAD-zbMATH data set to be considerably different from existing ones. We consider it a useful new resource that will challenge the state of the art in AND and benefit the AND research community. KW - author name disambiguation KW - author name homography KW - author name variability KW - data sets KW - digital libraries KW - Empirische Forschung KW - Datensatz KW - Metadaten KW - Autor KW - Veröffentlichung KW - Quantitative Analyse KW - Homographie KW - Elektronische Bibliothek KW - SCAD-zbMATH Y1 - 2017 UN - https://nbn-resolving.org/urn:nbn:de:bsz:mh39-110871 SN - 1588-2861 SS - 1588-2861 U6 - https://doi.org/10.1007/s11192-017-2363-5 DO - https://doi.org/10.1007/s11192-017-2363-5 VL - 111 IS - 3 SP - 1467 EP - 1500 PB - Springer Nature CY - Berlin ER -