Heuristic Search Value Iteration for POMDPs

Trey Smith and Reid Simmons

Conference Paper, Proceedings of 20th Conference on Uncertainty in Artificial Intelligence (UAI '04), pp. 520 - 527, July, 2004

View Publication

Abstract

We present a novel POMDP planning algorithm called heuristic search value iteration (HSVI). HSVI is an anytime algorithm that returns a policy and a provable bound on its regret with respect to the optimal policy. HSVI gets its power by combining two well-known techniques: attention-focusing search heuristics and piecewise linear convex representations of the value function. HSVI's soundness and convergence have been proven. On some benchmark problems from the literature, HSVI displays speedups of greater than 100 with respect to other state-of-the-art POMDP value iteration algorithms. We also apply HSVI to a new rover exploration problem 10 times larger than most POMDP problems in the literature.

BibTeX

@conference{Smith-2004-16921,
author = {Trey Smith and Reid Simmons},
title = {Heuristic Search Value Iteration for POMDPs},
booktitle = {Proceedings of 20th Conference on Uncertainty in Artificial Intelligence (UAI '04)},
year = {2004},
month = {July},
pages = {520 - 527},
keywords = {planning, probabilistic, POMDP, heuristic search},
}

Copyright notice: This material is presented to ensure timely dissemination of scholarly and technical work. Copyright and all rights therein are retained by authors or by other copyright holders. All persons copying this information are expected to adhere to the terms and constraints invoked by each author's copyright. These works may not be reposted without the explicit permission of the copyright holder.