扶摇AI知识笔记AI 前沿知识库
大模型基础

Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations

来源:arXiv cs.CL 论文速递 约 2203 字 llm
arXiv cs.CL
转载

本文转载自 arXiv cs.CL,版权归原作者及原发布平台所有。本站仅作知识整理与转载分享,如涉版权问题请联系客服删除。

01核心要点

  • As models scale, reward hacking becomes more frequent, more sophisticated, and more consequential.
  • Does it leave a telltale signature in model representations?
  • This work analyzes how reward hacking is represented internally in frontier open source LLMs, and how those representations can be used to understand and discover the range of hacking behaviors a model displays.

02正文全文

Abstract:As models scale, reward hacking becomes more frequent, more sophisticated, and more consequential. Does it leave a telltale signature in model representations? This work analyzes how reward hacking is represented internally in frontier open source LLMs, and how those representations can be used to understand and discover the range of hacking behaviors a model displays. In particular, we find that simple difference of means vectors coherently represent reward hacking in Kimi K3, GLM 5.2, and Qwen 3.8 Max across a variety of behaviors in common evaluations. Despite their simplicity, these vectors are both generalizable and interpretable, and we can use them to reliably detect reward hacking. We first evaluate reward hacking in commonly reported benchmarks like DeepSWE and SWE-bench, finding that models reward hack excessively in these environments; GLM 5.2 hacks in 57.2% of rollouts on DeepSWE and in 73% of rollouts on SWE-bench. Catching these requires monitors; LLM monitors are effective, but expensive detectors. We show that DoM vectors are similarly effective but virtually free, catching 3.1% more hacks in Kimi K3 and 7.9% fewer hacks in GLM 5.2 on DeepSWE at a monitor matched false positive rate. DoM vectors run on the chain-of-thought also predict reward hacks in the model's subsequent actions, meaning we can run them online and catch potential hacks before they occur. Finally, we analyze probe-hits that LLM monitors do not catch and discover other undesirable behaviors, as well as show transfer to finding hacks in non-SWE evaluations. Together, these results provide evidence that simple, white-box methods can be used to scalably study and monitor reward hacking behaviors in frontier open source models

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

03原文直达

本文内容转载自 arXiv cs.CL,如需查看原排版、配图与最新修订,请访问原始出处。

阅读原文(arXiv cs.CL)

下载论文 PDF

正在校验阅读权限…
RELATED

相关阅读

更多 大模型基础