2025.1.1

>ČASOPIS PRO MODERNÍ FILOLOGII 2025 (107) 1

Pasti dat: srovnatelnost dat jazykových korpusů

Data Traps: Comparability of Language Corpus Data

Markus Giger — Jana Kocková

 

 FULL TEXT   

 ABSTRACT (en)

Despite the apparent unambiguity of data provided by corpora, the data reflect different composition of the corpora, different conceptions of the synchronic period of a given language, different linguistic traditions, different orthography and other factors. We focus on the most common reasons affecting the comparability of data in parallel corpora, such as unequal lemmatization, tagging and tokenization, and illustrate them with examples from Czech, German and Russian. For example, when comparing Russian and Czech verb forms and lemmas, the data provided by the corpora are not comparable, because in Russian, unlike in Czech, the reflexive and non-reflexive forms are assigned to different lemmas and the verb lemma includes participles, whereas the corresponding Czech forms are tagged as adjectives, in accordance with Czech philological tradition. The differing approaches to tokenization are also reflected in the overall size of the corpus, indirectly affecting the comparability of relative frequencies.

 KEYWORDS (cz)

korpusy, komparativní lingvistika, tagování, srovnatelnost dat, vyváženost korpusů

 KEYWORDS (en)

corpora, comparative linguistics, tagging, data comparability, corpus balance

 DOI

https://doi.org/10.14712/23366591.2025.1.1

 SOURCES

Aranea Corpora (2020): Jazykovedný ústav Ľudovíta Štúra SAV. http://unesco.uniba.sk/.

BNC: The British National Corpus, version 2 (BNC World) (2001). http://www.natcorp. ox.ac.uk/.

DeReKo: Deutsches Referenzkorpus Mannheim IDS (2022): Leibniz-Institut für Deutsche Sprache. https://cosmas2.ids-mannheim.de/.

Grac v.18: General Regionally Annotated Corpus of Ukrainian, (2017). uacorpus.org.

InterCorp v16: Institut českého národního korpusu FF UK (2022). Praha. http://www. korpus.cz.

Kernkorpus DWDS: Kernkorpus des Digitalen Wörterbuchs der deutschen Sprache, 20. Jahrhundert. (2018): www.dwds.de.

MASC: Manually Annotated Sub-Corpus, The Open American National Corpus. https://anc. org/data/masc/.

NKRJa: The Russian National Corpus (ruscorpora.ru). 2003—2023. ruscorpora.ru.

SNK: Slovenský národný korpus — Bratislava: Jazykovedný ústav Ľ. Štúra SAV. (2022). https://korpus.sk.

SYN v2020: Institut českého národního korpusu FF UK (2020). Prague. http://www.korpus.cz.

 LITERATURE

Bláha, O. (2020): Typologické aspekty ruského a českého vidového systému. Slavia, 79, s. 245–255.

Dąbrowska, A. (2013): ‘National Corpus of Polish’ and ‘Great Dictionary of Polish’: two leading projects of present-day Polish lexicography. Konferenční příspěvek. http://efnil.nytud. hu/documents/conference-publications/ budapest-2012/15-EFNIL-Budapest-Dabrowska-Final.pdf.

Geyken, A. (2007): The DWDS Corpus: a Reference Corpus for the German Language of the 20th Century. In: C. Fellbaum (ed.), Collocations and Idioms: Linguistic, lexicographic, and computational aspects. London: Continuum Press, s. 23–41.

Giger, M. (2010): Příčestí minulé činné na -(v)ší v dnešních českých publicistických textech. Korpus — Gramatika — Axiologie,1, 2, s. 3–23.

Giger, M. (2015): Subjektová rezultativa v češtině ve srovnání s ruštinou. Časopis pro moderní filologii, 97, s. 146–156.

Giger, M. (2020): Několik poznámek k příčestí přítomnému činnému v ruštině a češtině. In: J. Bílková J. — I. Kolářová — M. Vondráček (eds.), Lingvistika — Korpus — Empirie. Praha: Ústav pro jazyk český, s. 9–16.

Giger, M. (2021): C. Participia a predikativa. In: F. Štícha a kol., Velká akademická gramatika spisovné češtiny. II. Morfologie: Morfologické kategorie / Flexe. Část 1. Praha: Academia, s. 331–352.

Giger, M. — Kocková, J. (2024): Grenzüberschreitungen an der Peripherie: Aspektuelle Funktionen von Aktivpartizipien und Verbalsubstantiven im Tschechischen. Zeitschrift für Slawistik, 69, s. 1–26.

Kocková, J. — Sytar, H. (2024): The Most Frequent Lemmas in the Ukrainian and Czech Corpus as a Resource for Foreign Language Learning and Teaching. Зборник Матице српске за славистику, 105, s. 369–382.

Křen, M. (2012): Diachronní srovnání synchronních korpusů. PhD. dis., FF UK, Praha.

MČ 1. (1986): Mluvnice češtiny 1. Fonetika, fonologie, morfonologie a morfemika, tvoření slov. Pod red. M. Dokulila a kol. Praha: Academia.

Štícha, F. (2008): Uzuálnost, funkčnost a systémovost jako kritéria gramatičnosti. K jednomu typu morfologické derivace (udělajíc — udělající). Slovo a slovesnost, 69, s. 176–191.

VAGSČ 1. (2018): F. Štícha et al. Velká akademická gramatika spisovné češtiny I. Morfologie: druhy slov, tvoření slov. Praha: Academia.

Siepmann, D. — Bürgel, C. — Diwersy, S. (2015): The Corpus de référence du français contemporain (CRFC) as the first genre diverse mega-corpus of French. International Journal of Lexicography, 30, s. 63–84. Stefanowitsch, A. (2020): Corpus linguistics: A guide to the methodology. Berlin: Language Science Press (Language Sciences 7).

Гловинская, М. Я. (2010): Потенциальные глагольные формы. In: Л. П. Крысин, (ред.), Современный русский язык. Система — норма — узус. Москва: Языки славянских культур, с. 171–199.

Изотов, А. И. (1993): Чешские атрибутивные причастия на фоне русских. Москва: МГУ, Филологический Факультет.

Черткова, М. Ю. — Чанг, П.-Ч. (1998): Эволюция двувидовых глаголов в современном русском языке. Russian Linguistics, 22, s. 13–34.

Úvod > 2025.1.1