We prove that training a Bradley-Terry (BT) reward model on pairwise preferences containing an Arrow/Condorcet 3-cycle incurs an irreducible information-geometric cost. For any Arrow/Condorcet 3-cycle (a, b, c) with contextuality fraction CF = a+b+c-2 > 0, the KL divergence between the true preference distribution and its BT maximum-likelihood projection satisfies KL >= CF * ln 2, where CF is the social-choice contextual fraction (the minimum fraction of preference data not rationalisable by any mixture of transitive rankings). For the symmetric Condorcet triplet the KL gap is 3 (ln 2 - H (c) ) ; the efficiency ratio KL/ (CF ln 2) attains its unique global minimum over all Arrow 3-cycles at the golden point c* = phi/2, with rho* = 3 log2 phi = 2. 0827. . . (proved via 2 phi² - phi³ = 1). We give a sandwich bound between sharp fixed-mass (lower) and fixed-logit-curl (upper) envelopes, and prove the cost cannot be hidden by embedding the cycle in a larger tournament. The all-n positive theorem is the triangle-LP bound KL >= rho* * CFₜri * ln 2 (CFₜri the triangle-LP contextual fraction) ; for n= CFSC ln 2 is proved (companion structural paper) for n=8 the full-polytope KL bound is open, the uniform sigmoid-tilt/tied-block membership route being refuted at large n by the Paley-family obstruction of the companion spectral paper. On the empirical side we identify a scalar-rating trap: any RLHF dataset whose pairwise win matrix is extracted from a single scalar rating per response is forced into the linear ordering polytope and cannot exhibit Condorcet cycles by construction; observing cyclicity requires multi-criterion or cross-aspect comparisons. A Pythagorean identity gives an exact decomposition KL = Iₘem - Iₚred (a Still-style nostalgia term; thermodynamic interpretation left as future work).
Vladimir Riabov (Tue,) studied this question.