Bimodal Debiasing for Text-to-Image Diffusion: Adaptive Guidance in Textual and Visual Spaces

University of Electronic Science and Technology of China
Interpolation end reference image. >

Abstract

Social biases in diffusion-based text-to-image models have drawn increasing attention, yet existing debiasing efforts often focus solely on either the textual (e.g., CLIP) or visual (e.g., U-Net) space. This unimodal perspective introduces two major challenges: (i) Debiasing only the textual space fails to control visual outputs, often leading to pseudo- or over-corrections due to unaddressed visual biases during denoising; (ii) Debiasing only the visual space can cause modality conflicts when biases in textual and vision are misaligned, degrading the quality and consistency of generated images. To address these issues, we propose a Bimodal ADaptive Guidance DEbiasing within Textual and Visual Spaces (BADGE). First, BADGE quantifies attribute-level bias inclination in both modalities, providing precise guidance for subsequent mitigation. Second, to avoid pseudo/over-correction and modality conflicts, the quantified bias degree is used as the debiasing strength for adaptive guidance, enabling fine-grained correction tailored to discrete attribute concepts. BADGE is a self-debiasing method that requires no additional training or external corpora. Extensive experiments demonstrate that BADGE significantly enhances fairness across intra- and inter-category attributes (e.g., gender, skin tone, age, and their interaction) while preserving high image fidelity.

Method

Inference Overview

The overall framework of BADGE. (1) We firstly obtain the quantified bias inclination in bimodal spaces (Insight.1 and Insight.2); (2) BADGE employs adaptive guidance based on quantified CLIP bias degree for obtaining a debiased text embedding (Insight.3); (3) Based on quantified U-Net bias degree, BADGE adaptively guide the samples with ambiguous attribute for obtaining a unbiased noise (Insight.4); (4) Fair image generation is achieved by removing the estimated noise from the latent noisy state based on unbiased textual embedding.

Result

Gender X Age/Skin

Inference Overview

Gender X Age Single

Inference Overview

Compatibility with ControlNet

Inference Overview

Nature

Inference Overview

Season

Inference Overview

Day-to-Night Compatibility with ControlNet

Inference Overview