Fоr а given MDP, there is а finite vаlue оf N such that the estimate оf the optimal value function computed by the policy iteration algorithm after N iterations is the same as the estimate after N+1 iterations, to arbitrary precision.