Classifier Evaluation
Real world tasks are often unbalanced. The best baseline for comparison is nonuniform. A classifier should be evaluated on skill gain instead of raw accuracy. Looking at the distribution of
Real world tasks are often unbalanced. The best baseline for comparison is nonuniform. A classifier should be evaluated on skill gain instead of raw accuracy. Looking at the distribution of classes in a benchmark dataset will reveal imbalance. We can show this using datasets lmsys/toxic-chat SetFit/sst5and ehovy/race. ToxicChat is an
Sep 22