VideoSTF

VideoSTF: Stress-Testing Output Repetition in Video Large Language Models

Yuxin Cao1 Wei Song2 Shangzhi Xu3 Jingling Xue3 Jin Song Dong1

1National University of Singapore 2Griffith University 3University of New South Wales

NeurIPS 2026

Output Repetition Example

A normal output that describes a video coherently next to a repetitive output that loops over the same phrases.
A normal output describes the video coherently. A repetitive output collapses into loops of the same phrases or sentences.

Overview

Video Large Language Models (VideoLLMs) are usually judged by what they predict. VideoSTF looks at how they generate, and it surfaces output repetition, a generation failure in which the decoder collapses into self-reinforcing loops of repeated phrases or sentences.

VideoSTF formalizes repetition with three complementary n-gram-based metrics, ships a standardized testbed of 10,000 diverse videos, and provides a library of controlled temporal stressors. With these components, it runs three tests on 10 VideoLLMs. Pervasive testing measures repetition on unperturbed videos, temporal stress testing measures how temporal transformations amplify it, and adversarial exploitation turns the stressors into a black-box attack.

10
VideoLLMs tested
10,000
testbed videos
91%
highest repetition rate
98%
highest attack success rate

The VideoSTF Framework

VideoSTF has three components, namely repetition metrics, a standardized video testbed and a temporal stressor library, which feed three evaluation protocols.

Metric

Repetition Rate (RR)

RR is the share of outputs that contain at least one repeated 5-gram. It captures whether repetition occurs.

Metric

Repetition Intensity (RI)

RI averages Rep-n, the share of duplicated unigrams in each output. It captures how much of an output is duplicated.

Metric

Information Entropy (IE)

IE is the normalized unigram entropy of each output. Lower values mean lower lexical diversity and stronger repetition.

Testbed

10,000 Videos from Public Datasets

The testbed samples videos from four public datasets. The clips last up to 180 seconds and cover everyday scenarios, with comedy, lifestyle and sports among the most frequent categories.

Histogram of testbed video durations, most clips are under 30 seconds and a tail reaches 180 seconds.
Video duration
Sunburst of testbed categories, led by comedy, lifestyle and sports.
Video categories
Stressor Library

Five Temporal Transformations

Each stressor changes the temporal structure of the sampled frames while keeping their content. Add, Delete and Replace act locally, while Reverse and Shuffle change the global order. Pick one to see what it does to a 16-frame clip.

Evaluation Protocols

Three Tests

Pervasive testing measures repetition on unperturbed videos, temporal stress testing applies the stressors, and adversarial exploitation searches for a transformation that makes a benign video repeat.

Pervasive testing feeds a video and the prompt to the VideoLLM and scores the output.
Radar charts of repetition rate, repetition intensity and information entropy for the ten VideoLLMs at 8, 16, 24 and 32 frames.
Temporal stress testing applies a temporal transformation to the video before the VideoLLM describes it.
Bar chart of the repetition rate of LLaVA-Video-7B-Qwen2-Video-Only on original videos and under each stressor at 8, 16, 24 and 32 frames.
Adversarial exploitation changes the transformation parameters until the output repeats.
Bar chart of attack success rate and average queries for LLaVA-Video-7B-Qwen2-Video-Only at 16 frames.

Main Results

Test 1

Pervasive Testing

This test measures repetition on unperturbed videos at 8, 16, 24 and 32 sampled frames. Output repetition appears in all 10 models, regardless of the base LLM, and it stays stable across these frame counts.

Test 2

Temporal Stress Testing

The heatmap shows the repetition rate (%) on original videos and under eight stressor variants. Temporal transformations raise the repetition rate in most cases.

Test 3

Adversarial Exploitation

The attacker only queries the model and applies stressors to the frames of videos that originally produce normal outputs. It stops once repetition appears or after 30 queries. Attack success rate (ASR) reaches 98%, and the average number of queries (AQ) never exceeds 15.8.

Examples

Output Repetition across VideoLLMs

Output Explorer

Pick a video, a frame count, a model and a temporal stressor to see what the model wrote. With a stressor selected, the transformed input and its output appear next to the original. The videos include cases whose original output is normal and cases whose original output is already repetitive.

Prompt Please describe this video in detail.

Citation

@inproceedings{cao2026videostf,
  title={VideoSTF: Stress-Testing Output Repetition in Video Large Language Models},
  author={Cao, Yuxin and Song, Wei and Xu, Shangzhi and Xue, Jingling and Dong, Jin Song},
  booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
  year={2026}
}