Abstract
Existing long-form video recommendation systems primarily rely on rating prediction or click-through rate estimation. However, the former is constrained by data sparsity, while the latter fails to capture actual viewing experiences. The accumulation of mid-playback abandonment behaviors undermines platform stickiness and commercial value. To address this issue, this paper seeks to improve viewing engagement. Grounded in Expectation-Confirmation Theory, this paper proposes the Long-Form Video Viewing Engagement Prediction (LVVEP) method. Specifically, LVVEP estimates user expectations from storyline semantics encoded by a pre-trained BERT model and refined via contrastive learning, weighted by historical engagement levels. Perceived experience is dynamically constructed using a GRU-based encoder enhanced with cross-attention and a neural tensor kernel, enabling the model to capture evolving preferences and fine-grained semantic interactions. The model parameters are optimized by jointly combining prediction loss with contrastive loss, achieving more accurate user viewing engagement predictions. Experiments conducted on real-world long-form video viewing records demonstrate that LVVEP outperforms baseline models, providing novel methodological contributions and empirical evidence to research on long-form video recommendation. The findings provide practical implications for optimizing platform management, improving operational efficiency, and enhancing the quality of information services in long-form video platforms.
Author supplied keywords
Cite
CITATION STYLE
Chen, Y., & Zhang, J. (2025). Modeling Viewing Engagement in Long-Form Video Through the Lens of Expectation-Confirmation Theory. Applied Sciences (Switzerland), 15(20). https://doi.org/10.3390/app152011252
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.