Solution (source code)

= Solution

Backward induction makes the <Snell envelope> integrable and adapted: $|U_t|\le|Y_t|+\mathbb E(|U_{t+1}|\mid\mathcal F_t)$. Its definition gives $U_t\ge Y_t$ and $U_t\ge\mathbb E(U_{t+1}\mid\mathcal F_t)$, so it is a <supermartingale> dominating the reward.

For a <stopping time> $\tau$ taking values in $\{0,\ldots,T\}$, expand its stopped value as
$$
U_\tau=U_0+\sum_{t=0}^{T-1}\mathbf1_{\{\tau>t\}}(U_{t+1}-U_t).
$$
The indicators are $\mathcal F_t$-measurable. Taking <conditional expectations> in each summand makes its expectation nonpositive by the <supermartingale> property. Since $\mathcal F_0$ is trivial, $U_0$ is deterministic and
$$
\boxed{\mathbb E Y_\tau\le\mathbb E U_\tau\le U_0.}
$$
This proves the finite-horizon <optional sampling theorem> directly in the instance needed here, without assuming nonnegative rewards.