Long-form film understanding from subtitles

CineSubBench

Evaluating LLMs on long-form narrative and cultural understanding from multilingual movie subtitles

Mir Tafseer Nayeem1, Susmoy Chakraborty2, Davood Rafiei1

1 University of Alberta   2 Independent Researcher

CineSubBench overview: 1,012 films with six parallel subtitle languages are evaluated on narrative generation, genre, age suitability, national ratings, and subtitle-grounded safety.
Benchmark overview The same films connect long-form subtitle input to narrative, cultural, and evidence-grounded tasks.

A shared test of story, culture, and evidence

Movie subtitles are long, ordered records of dialogue, but they do not explicitly identify speakers, scenes, or events. CineSubBench uses this demanding subtitle-only input to ask whether a model can reconstruct a story, make culturally situated predictions, and point to the lines behind a safety judgment.

The benchmark holds the film set constant across seven tasks, six languages, and ten national rating systems. Its 1,012 films provide 6,072 subtitle tracks and 8.13 million timestamped entries. This matched design makes task and language differences easier to examine without changing the underlying films.

Core finding

A strong overall score does not guarantee complete plot reconstruction, reliable cross-language behavior, culturally calibrated ratings, or correctly categorized subtitle evidence.

Performance across nine evaluated models

The paper's composite ranking combines distinct capabilities. Use the task columns to see where models differ; a medal marks the original composite rank, even when a column is sorted.

All-language benchmark results from Table 1 of the paper
1Gemini 3.8 Flash78.043.543.3484.9381.4277.8777.8
2GPT-5.6 Sol74.703.523.4081.5578.7565.4274.5
3DeepSeek V4 Pro 1.6T65.523.252.9381.6379.3358.0348.3
4Gemini 3.5 Flash Lite64.703.252.9180.9272.1764.6542.8
5Claude Haiku 4.557.502.912.4471.4572.7554.3237.2
6Qwen3 235B53.092.802.3068.2859.2547.8829.6
7Mistral 4 119B50.672.442.0066.8354.5047.3731.7
8GPT-5 Nano46.672.501.9072.4360.7543.879.6
9Llama 4 Scout 17B41.582.031.6766.9250.2534.1212.7

Narrative, genre, age, and country results average the six language settings. Safety evidence is evaluated on English subtitles only. Scores use different scales and are not interchangeable; see the paper for metric definitions and the composite formula.

Seven tasks over the same films

Narrative

Plot, synopsis, key message

Recover the premise, event sequence, and central theme from a long subtitle track.

Classification

Multi-label genre

Identify the film's genres using the subtitle sequence rather than metadata or video.

Audience guidance

Age suitability and country ratings

Estimate an appropriate viewing age and ratings in ten distinct national systems.

Grounding

Language safety

Count safety-relevant language and cite the exact subtitle entries that support each category.

Subtitle languages: English, Arabic, Indonesian, Persian, Romanian, and Vietnamese.

Rating systems: Australia, Brazil, France, Germany, Netherlands, Singapore, South Korea, Sweden, United Kingdom, and United States.

What the benchmark reveals

RQ1: Does one rank capture every capability?

No. Gemini 3.8 Flash leads the composite score at 78.04, while GPT-5.6 Sol has the best synopsis score (3.40/5). DeepSeek V4 Pro is third overall, yet slightly exceeds Sol on genre F1 and age accuracy within one year. The leaderboard is a starting point, not a substitute for task-level evaluation.

RQ2: Does understanding transfer across subtitle languages?

For all nine models, the five non-English languages have lower average narrative scores than English in point estimate. Persian is the lowest narrative setting for every evaluated model. Models recover the plot premise more reliably than a full event-complete synopsis; the paper's evaluator annotations flag a missing major event in 79.7% of synopses and an invented one in 71.1%.

Heatmaps of narrative scores across six subtitle languages. Persian is the lowest overall narrative setting for each of the nine models; plot scores are higher than synopsis scores.
Language robustness The same films produce different narrative performance when the subtitle language changes.

RQ3: Are age and country ratings calibrated?

Near-match age accuracy can hide different directions of error. In the ten national systems, raw accuracy must be compared with each country's label distribution: France has 31.6% mean model exact match against a 72.0% majority baseline, while Brazil shows meaningful gains over a much lower baseline.

Age prediction errors vary in magnitude and direction across models. Country-rating exact-match heatmap shows sharp differences among ten national systems and their majority baselines.
Cultural calibration Similar age accuracy can mask opposing errors; country-rating results need local baselines.

RQ4: Can safety judgments cite the right evidence?

Finding a relevant subtitle line and assigning it the correct safety category are different abilities. The leading model reaches 86.8% category-agnostic evidence F1 but 77.8% strict F1. Strong profanity is much easier to ground than mild obscenity, exposing a failure mode hidden by a single safety score.

Language-safety chart comparing category-agnostic and strict subtitle evidence F1 across nine models, plus evidence precision and recall for four lexical categories.
Evidence grounding Correct lines are not always assigned the correct category; mild obscenity remains particularly difficult.

Matched coverage, with explicit limits

The dataset brings together film metadata, audience-guidance labels, and subtitle tracks. Construction filters check timing, coverage, encoding, language, and content, followed by manual quality review. A typical subtitle input is roughly 40,000 to 45,000 prompt tokens; some exceed 100,000.