SPLICE Benchmark Reveals VLMs Struggle with Temporal Reasoning, Rely on Language Priors

SPLICE Benchmark Results
The SPLICE benchmark tests temporal, causal, spatial, contextual, and common sense reasoning by having models reconstruct the correct sequence of shuffled video clips. The research, co-authored by the source poster, was published at EMNLP 2025.
Model Performance Details
Tested models included Gemini Flash (1.5 and 2.0), Qwen2-VL (7B and 72B), InternVL2.5, and LLaVA-OneVision. Gemini 2.0 Flash scored 51% on the vision-only task, while human performance was 85%. Open-source models struggled significantly:
- LLaVA-OneVision-72B scored barely above random guessing in vision-only setting
- InternVL2.5-78B performed similarly poorly
- Qwen2-VL-72B reached only around 30% on vision-only
- Qwen2-VL-7B performed on par with the 72B variant, suggesting scaling the language model doesn't help when the bottleneck is in the vision encoder
Language Prior Dependency
When human-written text annotations describing clip content were added, model performance jumped significantly while human performance remained unchanged. This indicates models rely on language priors to compensate for weak visual understanding. Notably, Qwen2-VL-72B outperformed Gemini on text-only reasoning.
Visual Shortcut Behavior
Models demonstrated problematic reasoning patterns. When first and last video clips looked visually similar (like opening and closing a printer door), models predicted those clips were adjacent 57% of the time, compared to 2.5% for humans and 27% random chance. This suggests models are pattern matching on visual similarity rather than reasoning about events.
Testing Limitations and Future Work
The research didn't test Claude (which doesn't support video input) or OpenAI models (which couldn't handle multi-video input reliably at testing time). The dataset is public, and the poster notes newer models like Gemini 3 Flash and Qwen3-VL (with native 256K interleaved context, enhanced spatial-temporal modeling, and MoE variants up to 235B) should be tested on SPLICE to see if language prior issues persist. Preliminary testing suggests the language prior problem remains, though statistical significance hasn't been established across all experimental samples.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Claude Code v2.1.98 adds Vertex AI wizard, security fixes, and subprocess sandboxing
Claude Code v2.1.98 introduces an interactive Google Vertex AI setup wizard, adds subprocess sandboxing with PID namespace isolation on Linux, and fixes multiple security vulnerabilities including Bash permission bypasses and arbitrary code execution risks.

OpenClaw Client Adds Cost Tracking and Per-Agent Spending Limits
New release adds spending caps per agent, live usage UI with circular progress bar, sub-agent management, skill toggling, and per-agent model selection.

Atlassian Enables Default Data Collection for AI Training
Atlassian has enabled default data collection across its products to train AI models, according to a source posted on Hacker News with 312 points and 75 comments.

Xiaomi MiMo-V2-Pro AI Model Available Free on OpenRouter for 7 Days
Xiaomi's MiMo-V2-Pro AI model is available with free API access on OpenRouter for 7 days. The model features a 1 million token context window and benchmarks show it competing with Claude Opus 4.6 and approaching GPT-5.2 performance.