Minimize KL Divergence ↓ Remove the entropy term independent of θ ↓ Minimize Cross-Entropy ↓ Minimize NLL on training data