Corpus of Slovak legislative documents
Corpus of Slovak legislative documents
Author(s): Radovan GarabíkSubject(s): Language and Literature Studies, Applied Linguistics, Computational linguistics, Western Slavic Languages
Published by: Jazykovedný ústav Ľudovíta Štúra Slovenskej akadémie vied
Keywords: corpus; Slovak language; body of law; legislation
Summary/Abstract: The article describes the construction of the corpus of Slovak legislative documents. By analyzing several statistical values of the source metadata and documents, we efficiently improve corpus quality. We describe the methods used to clean up small variations in metadata, length based discrimination of document and examine the effectiveness of several strategies of deduplication. The corpus is a part of a comparable corpus of legislative documents of seven languages, created in the Multilingual Resources for CEF.AT in the Legal Domain (MARCELL) project.
Journal: Jazykovedný časopis
- Issue Year: 73/2022
- Issue No: 2
- Page Range: 175-189
- Page Count: 15
- Language: English