IC-1249TD-MPC exhibits training instability and performance degradation when the planning horizon is set to 20 time steps

Arnab Kumar Mondal, Siba Smarak Panigrahi, Sai Rajeswar, Kaleem Siddiqi, Siamak Ravanbakhsh

SourceEfficient Dynamics Modeling in Interactive Environments with Koopman Theory

While integrating their Koopman dynamics model into TD-MPC for model-based planning, the authors observed that vanilla TD-MPC suffers from training instability, performance degradation, and high variance when using a 20-step planning horizon. This forced them to reduce the vanilla TD-MPC baseline to a 5-step horizon for a fair comparison. The observation was made across four DeepMind Control Suite environments (quadruped run, quadruped walk, cheetah run, acrobot swingup) with 5 random seeds each.

Evidence
observational
Caveat
The observation is made in passing to justify the experimental setup rather than as a systematic study of TD-MPC's limitations; no specific error or variance numbers are reported for the unstable runs in the text.
Model
TD-MPC
Concepts
Failure mode
Datasets
DeepMind Control Suite / DM Control [eval]
Methods
TD-MPC [compared-to]
Extraction
automatic-extraction