1. To ensure both efficient density estimation and sampling, van den Oord et al. [2017] proposed an approach called Probability Density Distillation which trains the flow f as normal and then uses this as a teacher network to train a tractable student network g.