扶摇AI知识笔记AI 前沿知识库
智能体应用

Exponential Hardness of Off-Policy Evaluation under History-Dependent Logging

来源:arXiv cs.LG 论文速递 约 1688 字 memory
arXiv cs.LG
转载

本文转载自 arXiv cs.LG,版权归原作者及原发布平台所有。本站仅作知识整理与转载分享,如涉版权问题请联系客服删除。

01核心要点

  • Can a logged dataset visit every hidden state frequently and still be exponentially uninformative about a target policy's value?
  • We show that it can when the logger depends on history.
  • For every horizon $H \ge 3$, we construct two POMDPs with at most two latent states per stage, three actions, and a common logger with three memory states.

02正文全文

Abstract:Can a logged dataset visit every hidden state frequently and still be exponentially uninformative about a target policy's value? We show that it can when the logger depends on history. For every horizon $H \ge 3$, we construct two POMDPs with at most two latent states per stage, three actions, and a common logger with three memory states. Action coverage, belief coverage, and two behavior-marginal outcome-revealing conditions all have constants independent of $H$. Nevertheless, evaluating a known deterministic target policy to accuracy $1/8$ requires $\Theta((3/2)^H \log(1/\delta))$ logged episodes at confidence $1-\delta$, for $0 < \delta \le 1/4$, even when both candidate models are known. The mechanism is simple: a reset erases the unknown transition that determines the target value. We characterize the resulting statistical experiment exactly and obtain a matching optimal estimator. A directed two-lane gridworld realizes the construction, and trajectory simulations agree with its finite-sample prediction. The result establishes intractability for the history-dependent-logging, model-based case posed by Zhang and Jiang (2025, arXiv:2503.01134), under their behavior-marginal definition of revealing.

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

03原文直达

本文内容转载自 arXiv cs.LG,如需查看原排版、配图与最新修订,请访问原始出处。

阅读原文(arXiv cs.LG)

下载论文 PDF

正在校验阅读权限…
RELATED

相关阅读

更多 智能体应用