scieee AI-readable full text Open interactive document viewer

Mitigating Cultural Underrepresentation in Image Classification: A Comparative Study of Data Augmentation Techniques

Petrocco, Enzo Ubaldo; Oneto, Luca; Sgorbissa, Antonio

Abstract

In recent years, robots have been increasingly expected to exhibit cultural competence—that is, the ability to adapt to different cultural contexts—not only for social and ethical reasons but also to improve system robustness. This need is particularly evident given that robots' perception relies heavily on Machine Learning (ML) models. Because datasets cannot be assumed to be culturally balanced, the risk of underrepresentation bias in ML can directly undermine the robustness of autonomous systems such as social robots. We compare oversampling (OS), random transformations (RT), and diffusion model (DM)-generated images as strategies for mitigating cultural underrepresentation in two social robotics classification tasks. Our results show that while OS and RT provide moderate improvements, DM-generated images consistently yield the greatest reduction in Cultural InCompetence (CIC), a metric that quantifies disparities in model performance across cultures.

Full text

Mitigating Cultural Underrepresentation in Image Classification: A Comparative Study of Data Augmentation Techniques 1st Enzo Ubaldo Petrocco DIBRIS University of Genova Genova, Italy [email protected] 2nd Luca Oneto DIBRIS University of Genova Genova, Italy [email protected] 3rd Antonio Sgorbissa DIBRIS University of Genova Genova, Italy [email protected] Abstract—In recent years, robots have been increasingly expected to exhibit cultural competence—that is, the ability to adapt to different cultural contexts—not only for social and ethical reasons but also to improve system robustness. This need is particularly evident given that robots’ perception relies heavily on Machine Learning (ML) models. Because datasets cannot be assumed to be culturally balanced, the risk of underrepresentation bias in ML can directly undermine the robustness of autonomous systems such as social robots. We compare oversampling (OS), random transformations (RT), and diffusion model (DM)-generated images as strategies for mitigating cultural underrepresentation in two social robotics classification tasks. Our results show that while OS and RT provide moderate improvements, DM-generated images consistently yield the greatest reduction in Cultural InCompetence (CIC). Index Terms—Social Robotics, Machine Learning, Cultural Competence I. INTRODUCTION Cultural competence—the ability of a system to adapt to different cultural contexts—has become a key requirement in social robotics [1], [2]. Culture influences human perception, preferences, and design, which in turn shape human–robot interaction [3], [4]. While frameworks have been proposed for embedding cultural competence in robots [5], [6], perception modules remain vulnerable to cultural bias. Since machine learning (ML) constitutes the state of the art in computer vision, achieving culturally competent social robots requires culturally competent ML models. Although cultural competence has been studied in ML [7]–[9], the challenge of culturally imbalanced data [10] remains underexplored. Prior work shows that ML models trained predominantly on data from one culture often perform poorly on others, thereby exhibiting cultural incompetence due to underrepresentation [11], [12]. In this work, we compare oversampling (OS), random transformations (RT) [13], and diffusion model (DM)-generated images [14] as strategies to mitigate cultural underrepresentation in ML models. II. METHODOLOGY Retrieving information from the environment is essential for autonomous robots. Computer vision, typically implemented via machine learning (ML), extracts visual information from images or videos, which is especially relevant in social robotics due to complex tasks. For instance, a home-assistance robot may infer whether an elderly person needs support by detecting a lamp’s state (on/off) or assess navigation safety by recognizing carpets (present/absent). Both are binary classification problems. We collected LAMP (Chinese, French, Turkish) and CARPET (Indian, Japanese, Scandinavian) class and cultural balanced datasets. Formally [15], image classification consists of learning a function f:X → Y from a dataset Dn={(xi, yi)}n i=1, using algorithm AHwith hyperparameters H, to minimize a loss ℓ(y, ˆy).Dnis split into learning (L), validation (V), and test (T) sets; parameters are estimated from L, hyperparameters selected on V, and performance evaluated on T. In practice, fully optimizing all hyperparameters is computationally prohibitive. Therefore, some are tuned via grid search, while others are set to standard values (e.g., the fine-tuning learning rate, typically fixed at 10−5for pretrained networks). Prior work [11] demonstrated that cultural incompetence arises when some cultures are underrepresented. In the culturally competent ML framework, sets can be partitioned by culture C∈ C as LC,VC,TC, with underrepresentation simulated by subsampling minority cultures at pu. The Cultural InCompetence (CIC) metric [11] quantifies disparities: CIC =ˆ EC|ERRC−minC′ERRC′|, where ERRCis the error on TC. Lower CIC indicates higher cultural competence. To mitigate underrepresentation, data augmentation (DA) is commonly used. This includes OS and RT such as flips, rotations, noise, brightness adjustment, and zoom [13], applied offline or online to expand the dataset and reduce overfitting. More advanced approaches use DMs [14], which generate images via a forward-noising and reverse-denoising process. Due to DMs’ computational cost, we combine RTs with DMgenerated images in our experiments. While DMs have been used for data augmentation, their application to counterbalance underrepresentation in small image datasets (6000 images) remains underexplored. 2025 I-RIM Conference October 17-19, Rome, Italy ISBN: 9788894580570 10.5281/zenodo.17629694 79 Fig. 1: Top: samples from the CARPET dataset (left to right: Indian, Japanese, and Scandinavian cultures without carpets). Bottom: DM-generated images with Scandinavian majority, also without carpets. III. EXPERIMENTS We present the results obtained using the methodology described in Section II. The backbone model is ResNet50V2 from the Keras library1, pretrained on ImageNet. A GlobalAveragePooling2D layer, a Dropout layer (0.4), and a Dense layer were added as the head. Data were split into 70%/20%/10%. Model parameters were optimized using Adam with Binary Crossentropy. Learning rate and epochs hyperparameters were tuned via grid search. To mitigate overfitting, ReduceLROnPlateau and EarlyStopping were applied. Transfer learning was employed by initially freezing all layers except the head, then unfreezing them for fine-tuning with lr= 10−5. We adopt a residual U-Net DM enhanced with a transformer blocks. The DM is first pretrained on large-scale dataset2 and then fine-tuned on the datasets of interest. Training uses AdamW with weight decay and exponential moving average (EMA) to stabilize optimization and sampling. Sampling is performed via reverse diffusion, progressively denoising Gaussian noise into realistic images using a cosine noise schedule. Batches of synthetic images are generated with 15–50 diffusion steps. Generation quality is tracked via the Kernel Inception Distance (KID), which compares feature distributions of real and generated images. Examples of generated images using Scandinavian cultures as majority culture and without carpets can be found in Figure 1. Experiments were conducted with no augmentation (baseline) and with RT, OS, or DM applied individually. Table I reports ERR and CIC values for LAMP and CARPET, for each model trained on a majority culture and its variations. Overall, RT, OS, and DM improve ERR or CIC relative to the baseline. The effect is most pronounced for CIC under DM, where CIC decreases substantially even when ERR remains stable or improves only slightly. Results also indicate that our combination of online data augmentation, transfer learning and inclusion of the transform blocks in the original residual U-Net architecture let the DM to effectively learn the dataset distribution, as we can see also from Figure 1. 1https://keras.io/ 2https://www.tensorflow.org/datasets/catalog/places365 small AMaj. LAMP ERRLAMP CICLAMP Maj. CARPET ERRCARPET CICCARPET Baseline Chin. 19.7% 7.9% Ind. 20.8% 6.2% Fren. 19.6% 10.3% Jap. 23.8% 12.6% Turk. 22.2% 13.5% Scan. 21.9% 2.9% RT Chin. 18.0% 9.3% Ind. 20.5% 6.6% Fren. 18.1% 9.0% Jap. 22.8% 9.3% Turk. 21.3% 11.4% Scan. 20.4% 4.0% OS Chin. 16.6% 5.1% Ind. 21.1% 6.8% Fren. 18.1% 7.8% Jap. 23.3% 11.5% Turk. 19.5% 10.8% Scan. 20.6% 3.3% DM Chin. 16.7% 2.4% Ind. 18.3% 1.9% Fren. 19.8% 3.4% Jap. 24.3% 3.2% Turk. 19.5% 2.8% Scan. 19.0% 1.3% TABLE I: ERR and CIC values for LAMP and CARPET, for each majority culture and applied variation. IV. CONCLUSION We presented a comparative study of data augmentation techniques for mitigating cultural underrepresentation in social robotics image classification. Our results show that DMgenerated images provide the largest improvements in cultural competence, as measured by the CIC metric, substantially outperforming OS and RT. These findings highlight generative augmentation as a promising approach for culturally competent ML, relevant to social robotics since it does not impact inference time and so real-time processing constraints. REFERENCES [1] F. Hegel, C. Muhl, B. Wrede, M. Hielscher-Fastabend, and Others, “Understanding social robots,” in ACHI, 2009. [2] B. Bruno, N. Chong, H. Kamide, S. Kanoria, and Others, “Paving the way for culturally competent robots: a position paper,” in IEEE ROMAN, 2017. [3] G. Trovato, J. Ham, K. Hashimoto, H. Ishii, and Others, “Investigating the effect of relative cultural distance on the acceptance of robots,” in ICSR, 2015. [4] P. L. P. Rau, Y. Li, and D. Li, “Effects of communication style and culture on ability to accept recommendations from robots,” Comput. Hum. Behav., vol. 25, no. 2, pp. 587–595, 2009. [5] C. Faucher and F. C. Ditzinger, Advances in culturally-aware intelligent systems and in cross-cultural psychological studies. Springer, 2018. [6] C. Papadopoulos, N. Castro, A. Nigath, R. Davidson, N. Faulkes, and Others, “The caresses randomised controlled trial: Exploring the healthrelated impact of culturally competent artificial intelligence embedded into socially assistive robots and tested in older adult care homes,” Int. J. Soc. Robot., vol. 14, pp. 245–256, 2022. [7] A. Gjaci, C. Recchiuto, and A. Sgorbissa, “Towards culture-aware cospeech gestures for social robots,” Int. J. Soc. Robot., vol. 14, no. 6, pp. 1493–1506, 2022. [8] H. Quan, S. Li, and J. Hu, “Product innovation design based on deep learning and kansei engineering,” Appl. Sci., vol. 8, no. 12, p. 2397, 2018. [9] L. Zhou, X. Sun, G. Mu, J. Wu, and Others, “A tool to facilitate the cross-cultural design process using deep learning,” IEEE Trans. Hum. -Mach. Syst., vol. 52, no. 3, pp. 445–457, 2022. [10] S. Shankar, Y. Halpern, E. Breck, J. Atwood, and Others, “No classification without representation: Assessing geodiversity issues in open data sets for the developing world,” in NIPS, 2017. [11] E. Petrocco, A. Sgorbissa, and L. Oneto, “Culture-competent machine learning in social robotics,” in IRIM, 2023. [12] S. Zhioua and R. Binkyt˙ e, “Shedding light on underrepresentation and sampling bias in machine learning,” arXiv:2306.05068, 2023. [13] C. Shorten and T. M. Khoshgoftaar, “A survey on image data augmentation for deep learning,” J. Big Data, vol. 6, no. 1, pp. 1–48, 2019. [14] F.-A. Croitoru, V. Hondru, R. T. Ionescu, and M. Shah, “Diffusion models in vision: A survey,” IEEE Trans Pattern Anal Mach Intell., vol. 45, no. 9, pp. 10 850–10 869, 2023. [15] C. C. Aggarwal, Neural networks and deep learning. Springer, 2023. 80