Out-of-Domain Evaluation of Finnish Dependency Parsing

dc.contributor.authorKanerva Jenna
dc.contributor.authorGinter Filip
dc.contributor.organizationfi=data-analytiikka|en=Data-analytiikka|
dc.contributor.organization-code1.2.246.10.2458963.20.68940835793
dc.converis.publication-id176213812
dc.converis.urlhttps://research.utu.fi/converis/portal/Publication/176213812
dc.date.accessioned2022-10-28T13:36:07Z
dc.date.available2022-10-28T13:36:07Z
dc.description.abstract<p>The prevailing practice in the academia is to evaluate the model performance on in-domain evaluation data typically set aside from the training corpus. However, in many real world applications the data on which the model is applied may very substantially differ from the characteristics of the training data. In this paper, we focus on Finnish out-of-domain parsing by introducing a novel UD Finnish-OOD out-of-domain treebank including five very distinct data sources (web documents, clinical, online discussions, tweets, and poetry), and a total of 19,382 syntactic words in 2,122 sentences released under the Universal Dependencies framework. Together with the new treebank, we present extensive out-of-domain parsing evaluation utilizing the available section-level information from three different Finnish UD treebanks (TDT, PUD, OOD). Compared to the previously existing treebanks, the new Finnish-OOD is shown include sections more challenging for the general parser, creating an interesting evaluation setting and yielding valuable information for those applying the parser outside of its training domain.<br></p>
dc.format.pagerange1114
dc.format.pagerange1124
dc.identifier.isbn979-10-95546-72-6
dc.identifier.jour-issn2522-2686
dc.identifier.olddbid183023
dc.identifier.oldhandle10024/166117
dc.identifier.urihttps://www.utupub.fi/handle/11111/58167
dc.identifier.urlhttp://www.lrec-conf.org/proceedings/lrec2022/pdf/2022.lrec-1.120.pdf
dc.identifier.urnURN:NBN:fi-fe2022091258698
dc.language.isoen
dc.okm.affiliatedauthorKanerva, Jenna
dc.okm.affiliatedauthorGinter, Filip
dc.okm.discipline113 Computer and information sciencesen_GB
dc.okm.discipline113 Tietojenkäsittely ja informaatiotieteetfi_FI
dc.okm.internationalcopublicationnot an international co-publication
dc.okm.internationalityInternational publication
dc.okm.typeA4 Conference Article
dc.publisher.countryFranceen_GB
dc.publisher.countryRanskafi_FI
dc.publisher.country-codeFR
dc.publisher.placeParis
dc.relation.conferenceInternational Conference on Language Resources and Evaluation
dc.relation.ispartofjournalLREC Proceedings
dc.relation.ispartofseriesLREC Proceedings
dc.source.identifierhttps://www.utupub.fi/handle/10024/166117
dc.titleOut-of-Domain Evaluation of Finnish Dependency Parsing
dc.title.bookProceedings of the 13th Conference on Language Resources and Evaluation (LREC 2022)
dc.year.issued2022

Tiedostot

Näytetään 1 - 1 / 1
Ladataan...
Name:
2022.lrec-1.120.pdf
Size:
258.61 KB
Format:
Adobe Portable Document Format