DaNE: A named entity resource for danish

Publikation: Bidrag til bog/antologi/rapportKonferencebidrag i proceedingsForskningfagfællebedømt

  • Rasmus Hvingelby
  • Amalie Brogaard Pauli
  • Maria Barrett
  • Christina Rosted
  • Lasse Malm Lidegaard
  • Søgaard, Anders

We present a named entity annotation for the Danish Universal Dependencies treebank using the CoNLL-2003 annotation scheme: DaNE. It is the largest publicly available, Danish named entity gold annotation. We evaluate the quality of our annotations intrinsically by double annotating the entire treebank and extrinsically by comparing our annotations to a recently released named entity annotation of the validation and test sections of the Danish Universal Dependencies treebank. We benchmark the new resource by training and evaluating competitive architectures for supervised named entity recognition (NER), including FLAIR, monolingual (Danish) BERT and multilingual BERT. We explore cross-lingual transfer in multilingual BERT from five related languages in zero-shot and direct transfer setups, and we show that even with our modestly-sized training set, we improve Danish NER over a recent cross-lingual approach, as well as over zero-shot transfer from five related languages. Using multilingual BERT, we achieve higher performance by fine-tuning on both DaNE and a larger Bokmål (Norwegian) training set compared to only using DaNE. However, the highest performance is achieved by using a Danish BERT fine-tuned on DaNE. Our dataset enables improvements and applicability for Danish NER beyond cross-lingual methods. We employ a thorough error analysis of the predictions of the best models for seen and unseen entities, as well as their robustness on un-capitalized text. The annotated dataset and all the trained models are made publicly available.

OriginalsprogEngelsk
TitelLREC 2020 - 12th International Conference on Language Resources and Evaluation, Conference Proceedings
RedaktørerNicoletta Calzolari, Frederic Bechet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Helene Mazo, Asuncion Moreno, Jan Odijk, Stelios Piperidis
ForlagEuropean Language Resources Association (ELRA)
Publikationsdato2020
Sider4597-4604
ISBN (Elektronisk)9791095546344
StatusUdgivet - 2020
Begivenhed12th International Conference on Language Resources and Evaluation, LREC 2020 - Marseille, Frankrig
Varighed: 11 maj 202016 maj 2020

Konference

Konference12th International Conference on Language Resources and Evaluation, LREC 2020
LandFrankrig
ByMarseille
Periode11/05/202016/05/2020
SponsorAmazon AWS, Bertin, Lenovo, Ontotex, Vecsys, Vocapia

ID: 258327332