Plot, synopsis, key message
Recover the premise, event sequence, and central theme from a long subtitle track.
Long-form film understanding from subtitles
Evaluating LLMs on long-form narrative and cultural understanding from multilingual movie subtitles
1 University of Alberta 2 Independent Researcher
Movie subtitles are long, ordered records of dialogue, but they do not explicitly identify speakers, scenes, or events. CineSubBench uses this demanding subtitle-only input to ask whether a model can reconstruct a story, make culturally situated predictions, and point to the lines behind a safety judgment.
The benchmark holds the film set constant across seven tasks, six languages, and ten national rating systems. Its 1,012 films provide 6,072 subtitle tracks and 8.13 million timestamped entries. This matched design makes task and language differences easier to examine without changing the underlying films.
A strong overall score does not guarantee complete plot reconstruction, reliable cross-language behavior, culturally calibrated ratings, or correctly categorized subtitle evidence.
The paper's composite ranking combines distinct capabilities. Use the task columns to see where models differ; a medal marks the original composite rank, even when a column is sorted.
| 1 | Gemini 3.8 Flash | 78.04 | 3.54 | 3.34 | 84.93 | 81.42 | 77.87 | 77.8 |
|---|---|---|---|---|---|---|---|---|
| 2 | GPT-5.6 Sol | 74.70 | 3.52 | 3.40 | 81.55 | 78.75 | 65.42 | 74.5 |
| 3 | DeepSeek V4 Pro 1.6T | 65.52 | 3.25 | 2.93 | 81.63 | 79.33 | 58.03 | 48.3 |
| 4 | Gemini 3.5 Flash Lite | 64.70 | 3.25 | 2.91 | 80.92 | 72.17 | 64.65 | 42.8 |
| 5 | Claude Haiku 4.5 | 57.50 | 2.91 | 2.44 | 71.45 | 72.75 | 54.32 | 37.2 |
| 6 | Qwen3 235B | 53.09 | 2.80 | 2.30 | 68.28 | 59.25 | 47.88 | 29.6 |
| 7 | Mistral 4 119B | 50.67 | 2.44 | 2.00 | 66.83 | 54.50 | 47.37 | 31.7 |
| 8 | GPT-5 Nano | 46.67 | 2.50 | 1.90 | 72.43 | 60.75 | 43.87 | 9.6 |
| 9 | Llama 4 Scout 17B | 41.58 | 2.03 | 1.67 | 66.92 | 50.25 | 34.12 | 12.7 |
Narrative, genre, age, and country results average the six language settings. Safety evidence is evaluated on English subtitles only. Scores use different scales and are not interchangeable; see the paper for metric definitions and the composite formula.
Recover the premise, event sequence, and central theme from a long subtitle track.
Identify the film's genres using the subtitle sequence rather than metadata or video.
Estimate an appropriate viewing age and ratings in ten distinct national systems.
Count safety-relevant language and cite the exact subtitle entries that support each category.
Subtitle languages: English, Arabic, Indonesian, Persian, Romanian, and Vietnamese.
Rating systems: Australia, Brazil, France, Germany, Netherlands, Singapore, South Korea, Sweden, United Kingdom, and United States.
No. Gemini 3.8 Flash leads the composite score at 78.04, while GPT-5.6 Sol has the best synopsis score (3.40/5). DeepSeek V4 Pro is third overall, yet slightly exceeds Sol on genre F1 and age accuracy within one year. The leaderboard is a starting point, not a substitute for task-level evaluation.
For all nine models, the five non-English languages have lower average narrative scores than English in point estimate. Persian is the lowest narrative setting for every evaluated model. Models recover the plot premise more reliably than a full event-complete synopsis; the paper's evaluator annotations flag a missing major event in 79.7% of synopses and an invented one in 71.1%.
Near-match age accuracy can hide different directions of error. In the ten national systems, raw accuracy must be compared with each country's label distribution: France has 31.6% mean model exact match against a 72.0% majority baseline, while Brazil shows meaningful gains over a much lower baseline.
Finding a relevant subtitle line and assigning it the correct safety category are different abilities. The leading model reaches 86.8% category-agnostic evidence F1 but 77.8% strict F1. Strong profanity is much easier to ground than mild obscenity, exposing a failure mode hidden by a single safety score.
The dataset brings together film metadata, audience-guidance labels, and subtitle tracks. Construction filters check timing, coverage, encoding, language, and content, followed by manual quality review. A typical subtitle input is roughly 40,000 to 45,000 prompt tokens; some exceed 100,000.