Model-Free Reinforcement Learning for Lexicographic Omega-Regular Objectives
Ernst Moritz Hahn, Mateo Perez, Sven Schewe, Fabio Somenzi, Ashutosh Trivedi, Dominik Wojtczak.
FM 2021 : 142-159
Abstract
We study the problem of finding optimal strategies in Markov decision processes with lexicographic omega-regular objectives, which are ordered collections of ordinary omega-regular objectives. The goal is to compute strategies that maximise the probability of satisfaction of the first omega-regular objective; subject to that, the strategy should also maximise the probability of satisfaction of the second omega-regular objective; then the third and so forth. For instance, one may want to guarantee critical requirements first, functional ones second and only then focus on the non-functional ones. We show how to harness the classic off-the-shelf model-free reinforcement learning techniques to solve this problem and evaluate their performance on four case studies.