CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation
- 🇺🇸United States
- 🇮🇳India
- 🇮🇹Italy
- 🇨🇳China
- 🇧🇷Brazil
- 🇳🇱Netherlands
- 🇳🇿New Zealand
- 🇯🇵Japan
- 🇳🇬Nigeria
- 🇷🇺Russia
- 🇲🇾Malaysia
- 🇪🇹Ethiopia
- Confucian
- West and South Asian
- African-Islamic
- Orthodox Europe
- Catholic Europe
- Protestant Europe
- English-speaking
- Latin America
- Food
- Clothing
- Architecture
- Decorations
- Wedding
- Funeral
- Religious practice
TL;DR
Current T2V models generate visually plausible and semantically correct videos, but still misrepresent the intended culture.
T2V models are widely adopted across filmmaking, advertising, and education, making it increasingly important that they faithfully represent diverse cultural contexts. CultureVidBench evaluates 7 T2V models across cultural faithfulness, multimodal cultural rendering, semantic adherence, and perceptual quality.
Framework
From culturally grounded knowledge to evaluation: prompts are curated across material culture, social practice & performance, and ritual & ceremony, then assessed by both humans and multimodal language models.
Cultural faithfulness
Element alignment and consistency with the target cultural context.
Multimodal rendering
Culturally appropriate visible text and audio.
Semantic adherence
Correct subjects and actions from the prompt.
Perceptual quality
Realism, clarity, motion, and overall visual quality.
Key Findings
Current T2V models often achieve strong semantic adherence and visual quality, yet still struggle to faithfully capture culturally specific details.
The limitation is more pronounced in underrepresented cultural regions, where models perform substantially worse than in high-resource cultures.
Material culture is generally easier to generate than social practices and rituals.
Multimodal cultural rendering remains particularly challenging, especially culturally appropriate visible text and audio.
Leaderboard
Scores are normalized to 0–1. Click a metric to sort; higher is better. “—” indicates that the model does not generate audio.
| # | Model | Average ↕ | Cultural Alignment ↕ | Cultural Consistency ↕ |
Text Rendering ↕ | Audio Alignment ↕ |
Subject ↕ | Action ↕ | Realism ↕ | Visual Quality ↕ |
|---|
The average is computed over available reported dimensions. Proprietary model names are marked with a filled badge.
Cross-country Analysis
Culture-related dimensions vary far more across countries than general semantics and visual quality. High-resource contexts remain substantially easier for current models.
Category Analysis
Semantic adherence remains consistently higher than cultural performance across all 14 aspects. Material culture—especially clothing, architecture, and food—is generally easier to render than social practices, performances, and rituals.
Qualitative Results
Cultural generation failures are more pronounced in underrepresented cultural regions, which are also more susceptible to cultural interference from high-resource countries.
Citation
If you find CultureVidBench useful in your research, please cite our work.
@misc{han2026culturevidbench,
title = {CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation},
author = {Han, Xianjing and Su, Yuhan and Deng, Yang and Ma, Dong and Tay, Wee Peng and Zhu, Bin},
year = {2026}
}