Corpus of Slovak legislative documents Cover Image

Corpus of Slovak legislative documents
Corpus of Slovak legislative documents

Author(s): Radovan Garabík
Subject(s): Language and Literature Studies, Applied Linguistics, Computational linguistics, Western Slavic Languages
Published by: Jazykovedný ústav Ľudovíta Štúra Slovenskej akadémie vied
Keywords: corpus; Slovak language; body of law; legislation

Summary/Abstract: The article describes the construction of the corpus of Slovak legislative documents. By analyzing several statistical values of the source metadata and documents, we efficiently improve corpus quality. We describe the methods used to clean up small variations in metadata, length based discrimination of document and examine the effectiveness of several strategies of deduplication. The corpus is a part of a comparable corpus of legislative documents of seven languages, created in the Multilingual Resources for CEF.AT in the Legal Domain (MARCELL) project.

  • Issue Year: 73/2022
  • Issue No: 2
  • Page Range: 175-189
  • Page Count: 15
  • Language: English