Is the Best Better? Bayesian Statistical Model Comparison for Natural Language Processing

16/11/2020

Is the Best Better? Bayesian Statistical Model Comparison for Natural Language Processing

Piotr Szymański, Kyle Gorman

Keywords: natural models, bayesian technique, k-fold cross-validation, english taggers

Abstract: Recent work raises concerns about the use of standard splits to compare natural language processing models. We propose a Bayesian statistical model comparison technique which uses k-fold cross-validation across multiple data sets to estimate the likelihood that one model will outperform the other, or that the two will produce practically equivalent results. We use this technique to rank six English part-of-speech taggers across two data sets and three evaluation metrics.

EMNLP

This is an embedded video. Talk and the respective paper are published at EMNLP 2020 virtual conference. If you are one of the authors of the paper and want to manage your upload, see the question "My papertalk has been externally embedded..." in the FAQ section.

Comments

Post Comment

no comments yet