CausalChapter: Improving Long-Video Chaptering with Interventional Dependency Modeling
Published in EMNLP 2026 Main (Accepted), 2026
This work reframes long-video understanding as temporally grounded, high-level narrative chapter generation. It introduces local causal dependency selection and context-sensitive semantic enhancement, improving CIDEr by 21.52 points with the same Qwen backbone.
