Back to news
arXiv cs.LG · 2026-08-12 00:00 UTC
research

Procedural Fairness Failures in RLHF from Preference Averaging

arXiv:2608.10126v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single reward model, assuming preference homogeneity. When preferences are heterogeneous, this aggregation induces a procedural fairness failure where majority preference groups dominate reward learning while minority preferences are systematically under-represented. This work defines procedural fairness in alignment as preserving distinct preference signals during reward modeling and shows that standard RLHF violates this via preference averaging. Prefe

Why it matters

Preference averaging can cause procedural unfairness, signaling decision-makers to redesign RLHF aggregation to avoid systematically disadvantaging minority or context-specific preferences.

Read the original story

Published to Cognify News · Week 33, 2026