Past exam of the mathematics course of the University of Cambridge 2018 iii Paper 207 2 c Solution Created 2026-10-03 Updated 2026-10-05
For the plug-in Neyman allocation, the estimated Bernoulli distribution variances are and , soThis uses the target ratio directly as the allocation odds; a procedure that actively corrects previous allocation imbalances would need an additional rule.
For the randomized play-the-winner rule RPW, return the drawn ball and add one ball of the same treatment after a success or of the opposite treatment after a failure. Treatment 0 has three successes and two failures, and treatment 1 has one success and three failures. The urn therefore has treatment-0 balls and treatment-1 balls. HenceFor dynamic programming with only the final patient left, there is no future value from learning. Independent uniform prior distributions and Beta-binomial conjugacy give posteriors and . Their posterior predictive probabilities of success are and . The optimal terminal action therefore givesThis is optimal for expected successes with no imposed lower bound on the randomization probabilities.