SourceEfficient Dynamics Modeling in Interactive Environments with Koopman Theory
While integrating their Koopman dynamics model into TD-MPC for model-based planning, the authors observed that vanilla TD-MPC suffers from training instability, performance degradation, and high variance when using a 20-step planning horizon. This forced them to reduce the vanilla TD-MPC baseline to a 5-step horizon for a fair comparison. The observation was made across four DeepMind Control Suite environments (quadruped run, quadruped walk, cheetah run, acrobot swingup) with 5 random seeds each.