Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC
Abstract
Social media platforms enable large-scale cross-lingual communication, but translating user-generated content (UGC) remains challenging due to its informal style, cultural references, and interaction-based expressions. While recent LLMs have improved translation quality, existing benchmarks and metrics often fail to capture whether translations convey intended meaning and cultural resonance in real-world settings. In this work, we introduce CULTURE-MT, a benchmark for social media translation that focuses on both CULtural Transmission and UGC-specific emotion REsonance. CULTURE-MT consists of 1,002 UGC notes across 14 domains, categorized into four types based on culture-loaded symbol and linguistic style features. We also construct UGC-oriented training data to fine-tune Qwen3-8B and Qwen3-32B as baselines. We propose cultural effectiveness as a new evaluation criterion, focusing on expression accuracy and cultural adaptability. Testing 15 models, including the baselines, we find that traditional metrics fail to capture cultural effectiveness. We also observe that cultural effectiveness on base LLMs correlates with model size. Our work provides a comprehensive evaluation system for UGC translation models and will offers an open evaluation platform to advance research in this area. We release the CULTURE-MT benchmark and provide an online leaderboard where submitted translation results can be evaluated by our trained JUDGER.
Lay Summary
When people chat, joke, or share memes online, the words they use are deeply tied to their culture — slang, humor, and local references that machines often struggle to translate meaningfully. Current AI translation tools have improved significantly, but we lack reliable ways to test whether a translation truly conveys the original feeling and cultural meaning, not just the words. We built CULTURE-MT, a new benchmark — a standardized test — for evaluating how well AI systems translate social media content across languages. It contains over 1,000 real social media posts across 14 topics, ranging from internet humor to culturally specific expressions. We also introduced a new quality measure called "cultural effectiveness," which checks whether a translation feels natural and culturally resonant to readers in the target language, not just technically accurate. We tested 15 AI translation models and found that widely used automatic evaluation methods miss these cultural nuances entirely. Our benchmark and an open leaderboard will help researchers build AI translation tools that truly speak the language — and the culture — of their audience.