Active Online Learning (AOL)
AOL acquires informative online training data by balancing scene and action difficulty. It pairs harder scenes with relatively easier actions, and easier scenes with harder actions. This complementary selection adapts simulator interaction to the model’s evolving prediction errors, prioritizing trajectories that are both novel and learnable.
