Showing posts with label Informative Events. Show all posts
Showing posts with label Informative Events. Show all posts

The Explore-Exploit Dilemma in Nonstationary Decision Making Under Uncertainty

Chapter 2 of Handling Uncertainty and Networked Structure in Robot Control

It is often assumed that autonomous systems are operating in environments that may be described by a stationary (time-invariant) environment. However, real-world environments are often nonstationary (time-varying), where the underlying phenomena changes in time, so stationary approximations of the nonstationary environment may quickly lose relevance. Here, two approaches are presented and applied in the context of reinforcement learning in nonstationary environments. In Section 2, the first approach leverages reinforcement learning in the presence of a changing reward-model. In particular, a functional termed the Fog-of-War is used to drive exploration which results in the timely discovery of new models in nonstationary environments. In Section 3, the Fog-of-War functional is adapted in real-time to reflect the heterogeneous information content of a real-world environment; this is critically important for the use of the approach in Section 2 in real world environments.

Exploitation by Informed Exploration between Isolated Operatives for Information-theoretic Data Harvesting

We consider the problem of ferrying data between nodes of a sparsely distributed sensing network of Unattended Ground Sensors (UGS) with endurance-constrained Unmanned Aerial Systems (UAS). The sensing domain wherein the sparsely distributed UGS network is deployed is assumed to be highly nonstationary (time-varying) and noisy. This makes the dataferrying problem very complicated as the expected value-of-information at a sensing location can rapidly change. To address this issue, we present a new data ferrying algorithm termed Exploitation by Informed Exploration between Isolated Operatives (EIEIO), and show that with several reasonable assumptions and a model on the predicted accumulation of value-of-information, the problem can be simplified to a mathematical linear program. To solve the linear program, the UAS learns to anticipate regions in the sensing domain that have the highest degree of change. The degree of change, is learned using a novel implementation of a Cox Process called the Cox-Gaussian Process (CGP). Our approach does not require a priori knowledge of the sensing domain model to arrive at an optimal UAS allocation strategy.

Given knowledge of the informatic content of all sampling locations, a closed path may be formed so the data-ferrying agent visits the most informative subset of data source locations.

Paper    

Adaptive Algorithms for Autonomous Data-Ferrying in Nonstationary Environments

Unattended ground sensors (UGS) in long-term distributed sensing deployments benefit greatly from the incorporation of unmanned aerial systems (UAS). For instance, the mobility of data-ferrying UAS may be leveraged to reduce the cost of communication between UGS, as well as extend the effective coverage and endurance of the distributed UGS network. Since the UAS are also limited in endurance, a UAS may only ferry data between a subset of the UGS during each sortie. This is particularly problematic for extended operations in nonstationary spatio-temporal domains, as the model obtained from the set of UGS may rapidly lose relevance. Moreover, the informativeness of, or the Value-of-Information (VoI) available at, each UGS may not be equal. Our approach, termed Exploitation by Informed Exploration between Isolated Operatives (EIEIO), learns a generative spatio-temporal model for the arrival of VoI at each UGS. Through EIEIO, we anticipate and prioritize the subset of UGS with the highest VoI for each data ferrying sortie. Furthermore, a lower bound on the requisite sampling time for homogeneous Poisson processes is leveraged to provide a bound on how many times the UAS must visit each UGS in order to learn a spatio-temporal VoI model.


Paper