For a Markov process with reward , the value starting at time in state is the supremum of expected rewards over admissible stopping times. In finite discrete time it obeys and for the transition operator , wherever expectations are well-defined. Statewise absolute integrability of every remaining-horizon reward makes these values finite.
Articles by others on the same topic
There are currently no matching articles.