CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation

1Nanyang Technological University 2Centrale Supélec 3Singapore Management University 4University of Cambridge
Corresponding author
1,000Prompts

TL;DR

Current T2V models generate visually plausible and semantically correct videos, but still misrepresent the intended culture.

Russia · PelmeniHappyHorse 1.0
A woman places pelmeni dumplings on a plate at a dining table in Russia.Incorrect cultural element
China · Traditional instrumentsWan 2.2
Musicians play traditional instruments at a traditional Chinese music concert.Cultural inconsistency
Ethiopia · Crossing oneselfKling 3.0
Worshippers make the sign of the cross and continue praying in Ethiopia.Incorrect cultural element

T2V models are widely adopted across filmmaking, advertising, and education, making it increasingly important that they faithfully represent diverse cultural contexts. CultureVidBench evaluates 7 T2V models across cultural faithfulness, multimodal cultural rendering, semantic adherence, and perceptual quality.

Framework

From culturally grounded knowledge to evaluation: prompts are curated across material culture, social practice & performance, and ritual & ceremony, then assessed by both humans and multimodal language models.

CultureVidBench construction and evaluation framework
Figure 1. Framework of the CultureVidBench construction and evaluation pipeline.
Evaluation Dimension 1

Cultural faithfulness

Element alignment and consistency with the target cultural context.

Evaluation Dimension 2

Multimodal rendering

Culturally appropriate visible text and audio.

Evaluation Dimension 3

Semantic adherence

Correct subjects and actions from the prompt.

Evaluation Dimension 4

Perceptual quality

Realism, clarity, motion, and overall visual quality.

Key Findings

01

Current T2V models often achieve strong semantic adherence and visual quality, yet still struggle to faithfully capture culturally specific details.

02

The limitation is more pronounced in underrepresented cultural regions, where models perform substantially worse than in high-resource cultures.

03

Material culture is generally easier to generate than social practices and rituals.

04

Multimodal cultural rendering remains particularly challenging, especially culturally appropriate visible text and audio.

Leaderboard

Scores are normalized to 0–1. Click a metric to sort; higher is better. “—” indicates that the model does not generate audio.

#ModelAverage ↕ Cultural
Alignment ↕
Cultural
Consistency ↕
Text
Rendering ↕
Audio
Alignment ↕
Subject ↕Action ↕ Realism ↕Visual
Quality ↕

The average is computed over available reported dimensions. Proprietary model names are marked with a filled badge.

Cross-country Analysis

Culture-related dimensions vary far more across countries than general semantics and visual quality. High-resource contexts remain substantially easier for current models.

Country-level evaluation heatmap and dispersion across cultural dimensions
Figure 2. Country-level results averaged over all T2V models evaluated by Gemini-3.1-Pro. Dispersion denotes the gap between the highest and lowest country-level scores.
Radar chart of country-level cultural scores for seven T2V models
Figure 3. Country-level cultural scores of different T2V models evaluated by Gemini-3.1-Pro. Scores are averaged over cultural faithfulness and multimodal cultural rendering.

Category Analysis

Semantic adherence remains consistently higher than cultural performance across all 14 aspects. Material culture—especially clothing, architecture, and food—is generally easier to render than social practices, performances, and rituals.

Current models depict distinctive visual appearances more reliably than culturally grounded interactions, procedures, and ceremonies.
Cultural and semantic adherence scores across 14 cultural aspects
Figure 4. Cultural scores (dark-colored bars) and semantic adherence scores (light-colored bars) across cultural aspects evaluated by human. Cultural scores are averaged over cultural faithfulness and multimodal cultural rendering.

Qualitative Results

Material culture
Social practice & performance
Ritual & ceremony
US - Christmas decorationHappyHorse 1.0
Families decorate homes for Christmas with a festive banner in the United States.Text rendering failure
US - Backyard games (frisbees)Kling 3.0
People throw frisbees in a backyard setting in the United States.Successful
Italy - Church weddingWan 2.2
The bride and groom stand at the altar and receive blessings from a priest in Italy.Successful
Malaysia - Hari Raya decorationVeo 3.1
Families decorate homes for Hari Raya with a festival banner in Malaysia.
Text rendering failureIncorrect cultural element
Ethiopia - Eskista danceKling 3.0
Dancer perform eskista dance in Ethiopia.Incorrect cultural element
China - Tea ceremonyWan 2.2
The bride and groom kneel before their parents and serve tea to them during a traditional Chinese wedding ceremony.
Incorrect cultural elementCultural inconsistency

Cultural generation failures are more pronounced in underrepresented cultural regions, which are also more susceptible to cultural interference from high-resource countries.

Citation

If you find CultureVidBench useful in your research, please cite our work.

@misc{han2026culturevidbench,
  title  = {CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation},
  author = {Han, Xianjing and Su, Yuhan and Deng, Yang and Ma, Dong and Tay, Wee Peng and Zhu, Bin},
  year   = {2026}
}