How much of a measured AI preference is the model, and how much is the instrument?
Abstract
arXiv:2608.23641v1 Announce Type: new Abstract: Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences. Keeling et al. (2024), Mazeika et al. (2025), Mikaelson et al. (2025), Tagliabue and Dung (2025) and Trhlik et al. (2026) have built four instruments for that purpose, and their findings disagree. The disagreement cannot be attributed to a single cause, because no two of these studies have held the (1) set of outcomes, (2) set of m
Transparencia: Este análisis ha sido generado con asistencia de inteligencia artificial bajo supervisión editorial de SAPIENSDATAAI.