
SUNTA: Hierarchical Video Prediction with Surprise-based Chunking
Hierarchical state-space models have struggled with how to segment long sequences into meaningful chunks for prediction. SUNTA reframes chunking as a surprise-driven problem, using prediction errors rather than fixed intervals or similarity metrics to identify where the model needs longer context. The approach tackles two critical training obstacles: hierarchical collapse and the absence of surprise signals during inference. This work matters because sequence segmentation directly affects how well models handle long-horizon reasoning across video, time-series, and language tasks, making it relevant to anyone building or deploying systems that must maintain coherence over extended contexts.58























