<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title><![CDATA[ Alias - pckt ]]></title>
        <link><![CDATA[ https://alias.pckt.blog ]]></link>
        <description><![CDATA[  ]]></description>
        <language>en</language>
        <pubDate>Sun, 27 Sep 2026 16:20:58 +0000</pubDate>

                    <item>
                <title>Classifier Evaluation</title>
                <link>https://alias.pckt.blog/classifier-evaluation-dwug23r</link>
                <description><![CDATA[Real world tasks are often unbalanced. The best baseline for comparison is nonuniform. A classifier should be evaluated on skill gain instead of raw accuracy. Looking at the distribution of classes in a benchmark dataset will reveal imbalance. We can show this using datasets lmsys/toxic-chat SetFit/sst5and ehovy/race. ToxicChat is an example of extreme imbalance. More than 90% of the examples are “benign” where the task is to identify “toxic” text.]]></description>
                <author>Alias</author>
                <guid isPermaLink="false">classifier-evaluation-dwug23r</guid>
                <pubDate>Tue, 22 Sep 2026 20:04:54 +0000</pubDate>
                            </item>
            </channel>
</rss>