Showing posts with label NIPS. Show all posts
Showing posts with label NIPS. Show all posts
A Hybridized Bayesian Parametric-Nonparametric Approach to the Pure Exploration Problem
Information-driven approaches to reinforcement learning (RL) and bandit problems largely rely on optimizing with respect to an expectation on calculated Kullback-Leibler (KL) divergence values. Although KL divergence may provide bounds on problem domain models, bounds on the expected KL divergence itself are absent from information-driven approaches. As such, we focus our investigation on the pure exploration problem, a key component of RL and bandit problems, where the objective is to efficiently gain knowledge about the problem domain. For this task, we develop an algorithm using a Poisson exposure process Cox Gaussian process (Pep-CGP), a hybridized Bayesian parametric-nonparametric L\'{e}vy process, and theoretically derive a bound for the Pep-CGP expectation on KL divergence. Our algorithm, Real-time Adaptive Prediction of Time-varying and Obscure Rewards (RAPTOR), is validated on 4 real-world datasets, wherein baseline pure exploration approaches are outperformed by RAPTOR.
Collaborative Goal and Policy Learning from Human Operators of Construction Co-Robots
Human operators of real-world co-robots, such as excavator, require extensive experience to skillfully handle these complicated machines in uncertain safety-critical environments. We consider the problem of human-robot collaborative learning and task execution, where efficient human-robot interaction is critical to safely and efficiently accomplish complex tasks in uncertain environments. Our collaborative learning algorithm enables a construction co-robot to learn latent task subgoals from the demonstrations of skilled human operators which can then be used to guide novice human operators in completing complex tasks under uncertainty. The effectiveness our algorithm is demonstrated through experimentation on a scaled model of an excavator with guided and unguided human operators. Our results demonstrate that when the co-robot’s inferred subgoals are communicated back to the novice human operator, task performance significantly improves.
| Paper | Bibtex |
Human Aware UAS Path Planning in Urban Environments using Nonstationary MDPs
A growing concern with deploying Unmanned Aerial Vehicles (UAVs) in urban environments is potential violation of human privacy, and the backlash this could entail. Therefore, there is a need for UAV path planning algorithms that minimize the likelihood of invading human privacy. Such algorithms would be useful for pipeline and agricultural survey, wildfire monitoring and other missions where surveillance of humans should be avoided. We formulate the problem of human-aware path planning as a nonstationary Markov Decision Process, and provide a novel model-based reinforcement learning solution that leverages Gaussian process clustering. Our algorithm is flexible enough to accommodate changes in human population densities, and is real-time computable, as opposed to competing approaches employing Bayesian nonparametrics. The approach is validated experimentally on a large-scale long duration experiment with both simulated and real UAVs.
Authors & Details:
2013,
2014,
A. Axelrod,
Autonomy,
C. Crick,
Conference,
G. Chowdhary,
H. Kingravi,
ICRA,
ML,
NIPS,
Nonstationarity,
R. Allamraju,
R. Grande,
RL,
Robotics,
ROS,
Sparse GPs,
W. Sheng,
Workshop
Subscribe to:
Posts (Atom)