The problem with a single label
Severity ratings, policy calls, and many engineering judgments admit more than one defensible answer. Collapsing them to a majority vote throws away information your model could learn from — and can make eval scores look cleaner than they are.