Repetition Rate (RR)
RR is the share of outputs that contain at least one repeated 5-gram. It captures whether repetition occurs.
1National University of Singapore 2Griffith University 3University of New South Wales
Video Large Language Models (VideoLLMs) are usually judged by what they predict. VideoSTF looks at how they generate, and it surfaces output repetition, a generation failure in which the decoder collapses into self-reinforcing loops of repeated phrases or sentences.
VideoSTF formalizes repetition with three complementary n-gram-based metrics, ships a standardized testbed of 10,000 diverse videos, and provides a library of controlled temporal stressors. With these components, it runs three tests on 10 VideoLLMs. Pervasive testing measures repetition on unperturbed videos, temporal stress testing measures how temporal transformations amplify it, and adversarial exploitation turns the stressors into a black-box attack.
VideoSTF has three components, namely repetition metrics, a standardized video testbed and a temporal stressor library, which feed three evaluation protocols.
RR is the share of outputs that contain at least one repeated 5-gram. It captures whether repetition occurs.
RI averages Rep-n, the share of duplicated unigrams in each output. It captures how much of an output is duplicated.
IE is the normalized unigram entropy of each output. Lower values mean lower lexical diversity and stronger repetition.
The testbed samples videos from four public datasets. The clips last up to 180 seconds and cover everyday scenarios, with comedy, lifestyle and sports among the most frequent categories.


Each stressor changes the temporal structure of the sampled frames while keeping their content. Add, Delete and Replace act locally, while Reverse and Shuffle change the global order. Pick one to see what it does to a 16-frame clip.
Pervasive testing measures repetition on unperturbed videos, temporal stress testing applies the stressors, and adversarial exploitation searches for a transformation that makes a benign video repeat.






This test measures repetition on unperturbed videos at 8, 16, 24 and 32 sampled frames. Output repetition appears in all 10 models, regardless of the base LLM, and it stays stable across these frame counts.
The heatmap shows the repetition rate (%) on original videos and under eight stressor variants. Temporal transformations raise the repetition rate in most cases.
The attacker only queries the model and applies stressors to the frames of videos that originally produce normal outputs. It stops once repetition appears or after 30 queries. Attack success rate (ASR) reaches 98%, and the average number of queries (AQ) never exceeds 15.8.
Pick a video, a frame count, a model and a temporal stressor to see what the model wrote. With a stressor selected, the transformed input and its output appear next to the original. The videos include cases whose original output is normal and cases whose original output is already repetitive.
@inproceedings{cao2026videostf,
title={VideoSTF: Stress-Testing Output Repetition in Video Large Language Models},
author={Cao, Yuxin and Song, Wei and Xu, Shangzhi and Xue, Jingling and Dong, Jin Song},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
year={2026}
}