Learning Cultural Vectors for Cross-Cultural Generation

Published in The Fourtieth Annual Conference on Neural Information Processing Systems (NeurIPS), 2026

We study cultural vectors as controllable representations of cultural knowledge in text-to-image diffusion models. First, we show that they can be learned effectively from synthetic data and that inference-time scaling controls the trade-off between cultural alignment and exaggeration. Second, we analyze their properties, including where cultural information is encoded in the U-Net and how cultural vectors behave in activation space. Third, we study their compositionality for multi-cultural behavior and cross-cultural generation, and show that naive composition introduces interference between cultural vectors. We then explore two complementary directions for improving composition, learning more independent cultural representations and using culturally aware merging.