Search

Navigator: RI | Publications | Policy Search in Reproducing Kernel Hilbert Space

Graphics enhanced version of this site

Policy Search in Reproducing Kernel Hilbert Space
J. Bagnell and J. Schneider
tech. report CMU-RI-TR-03-45, Robotics Institute, Carnegie Mellon University, November, 2003.

Jump to: Download | Abstract | Notes | Text Reference | BibTeX Reference


Download [Help]

Adobe portable document format (pdf) [351 KB]
Compressed postscript (ps.gz) [269 KB]

Copyright notice: This material is presented to ensure timely dissemination of scholarly and technical work. Copyright and all rights therein are retained by authors or by other copyright holders. All persons copying this information are expected to adhere to the terms and constraints invoked by each author's copyright. These works may not be reposted without the explicit permission of the copyright holder.


Abstract

Much recent work in reinforcement learning and stochastic optimal control has focused on algorithms that search directly through a space of policies rather than building approximate value functions. Policy search has numerous advantages: it does not rely on the Markov assumption, domain knowledge may be encoded in a policy, the policy may require less representational power than a value-function approximation, and stable and convergent algorithms are well-understood. In contrast with value-function methods, however, existing approaches to policy search have heretofore focused entirely on parametric approaches. This places fundamental limits on the kind of policies that can be represented. In this work, we show how policy search (with or without the additional guidance of value-functions) in a Reproducing Kernel Hilbert Space gives a simple and rigorous extension of the technique to non-parametric settings.

In particular, we investigate a new class of algorithms which generalize \textsc{Reinforce}-style likelihood ratio methods to yield both online and batch techniques that perform gradient search in a function space of policies. Further, we describe the computational tools that allow efficient implementation. Finally, we apply our new techniques towards interesting reinforcement learning problems.


Notes

Associated lab/group: Reliable Autonomous Systems Lab
Associated project: Federation of Intelligent Robotic Explorers Project


Text Reference

J. Bagnell and J. Schneider, Policy Search in Reproducing Kernel Hilbert Space, tech. report CMU-RI-TR-03-45, Robotics Institute, Carnegie Mellon University, November, 2003.


BibTeX Reference

@techreport{Bagnell_2003_4538,
   author = "James (Drew) Bagnell and Jeff Schneider",
   title = "Policy Search in Reproducing Kernel Hilbert Space",
   institution = "Robotics Institute, Carnegie Mellon University",
   month = "November",
   year = "2003",
   number = "CMU-RI-TR-03-45",
   address = "Pittsburgh, PA"
}


The Robotics Institute is part of the School of Computer Science, Carnegie Mellon University.
For updates and comments, please see these instructions.
This page maintained by robotwebmaster@ri.cmu.edu