CausalChapter: Improving Long-Video Chaptering with Interventional Dependency Modeling

Published in EMNLP 2026 Main (Accepted), 2026

This work reframes long-video understanding as temporally grounded, high-level narrative chapter generation. It introduces local causal dependency selection and context-sensitive semantic enhancement, improving CIDEr by 21.52 points with the same Qwen backbone.

Read the paper