scieee AI-readable full text Open interactive document viewer

Multi-focal Conditioned Latent Diffusion for Person Image Synthesis

Jiaqi, Liu; Jichao, Zhang; Rota, Paolo; Sebe, Niculae

Full text

Multi-focal Conditioned Latent Diffusion for Person Image Synthesis Jiaqi Liu1Jichao Zhang2BPaolo Rota1Nicu Sebe1 University of Trento1Ocean University of China2 Abstract The Latent Diffusion Model (LDM) has demonstrated strong capabilities in high-resolution image generation and has been widely employed for Pose-Guided Person Image Synthesis (PGPIS), yielding promising results. However, the compression process of LDM often results in the deterioration of details, particularly in sensitive areas such as facial features and clothing textures. In this paper, we propose a Multi-focal Conditioned Latent Diffusion (MCLD) method to address these limitations by conditioning the model on disentangled, pose-invariant features from these sensitive regions. Our approach utilizes a multi-focal condition aggregation module, which effectively integrates facial identity and texture-specific information, enhancing the model’s ability to produce appearance realistic and identity-consistent images. Our method demonstrates consistent identity and appearance generation on the DeepFashion dataset and enables flexible person image editing due to its generation consistency. The code is available at https://github.com/jqliu09/mcld. 1. Introduction The pose-guided person image synthesis (PGPIS) task focuses on transforming a source image of a person into a target pose, while preserving the appearance and identity of the individual as accurately as possible. This task has significant implications in applications like virtual reality, e-commerce, and the fashion industry, where maintaining photorealistic quality and identity consistency is essential. Recent approaches to PGPIS largely rely on Generative Adversarial Networks (GANs) [6], which, despite their success, often struggle with training instability and mode collapse, resulting in suboptimal preservation of identity and garment details [27,36,45,51,54,57]. As an alternative, diffusion models [11,37] have shown promise in generating high-quality images by progressively refining details through multiple denoising steps. The introduction of PIDM [2] marked the first application of diffusion models for PGPIS, where latent diffusion models (LDM) [37] compress images into high-level feature representations, thereby (a) (b) Source Image VAE Recon VAE Recon + ε GT CFLDOurs Figure 1. (a) The VAE [37] reconstruction deteriorates the detailed information of person images, especially the facial regions and complex textures. These issues worsen for the generated latent with small deviations. A small deviation ϵ= 0.2is added to demonstrate the often case of generated latent. (b) Our methods preserve this detailed information better than other LDM-based methods by introducing multi-focal conditions. reducing the computational complexity while supporting high-resolution outputs. Extensions such as PoCoLD [9] enhance 3D pose correspondence using pose-constrained attention, and CFLD [24] emphasize semantic understanding with hybrid-granularity attention. Despite these advancements, LDM-based methods encounter limitations in recovering fine appearance details, especially in facial and clothing regions. As shown in Fig. 1 (a), this challenge is primarily due to the lossy nature of autoencoder compression [1], which can degrade complex textures and identity-specific features during encoding. Since the lossy reconstructed images are the upper bound of generated images of LDM-based methods, this issue worsens when doing inference since the generated latent deviates from the compressed real latent. Additionally, LDM’s reliance on whole-image conditioning often struggles to focus on sensitive regions where appearance precision is critical. The integration of pose and appearance information complicates detail reconstruction, leading to suboptimal performance across diverse poses and sensitive areas. To overcome these limitations, we introduce a MultiThis CVPR paper is the Open Access version, provided by the Computer Vision Foundation. Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore. 16019 focal Conditioned Latent Diffusion (MCLD) approach for PGPIS. Our method mitigates the loss of detail in sensitive regions by conditioning the diffusion model on the corresponding selectively decoupled features rather than the entire image. Specifically, we isolate high-frequency regions, such as facial identity and appearance textures, from the source image and treat them as independent conditions. This decoupling strategy enhances control over sensitive areas, ensuring better identity preservation and texture fidelity. Our approach first generates pose-invariant embeddings of the selected regions shared in the source and target images using pretrained modules, which are then fused within the Multi-focal Condition Aggregation module. This module introduces selective cross-attention layers, leveraging the structural advantages of UNet to combine the conditions effectively. Consequently, our MCLD method achieves improved control and accuracy, facilitating highquality, realistic person image synthesis. Our main contributions can be summarized as follows: • We introduce a new approach, MCLD, that focuses on alleviating the deterioration of important details in sensitive areas like the face and clothing by using separate conditions for these regions, which improves both identity preservation and appearance fidelity. • We develop a multi-focal condition aggregation module that combines controls from multiple focus areas, allowing our model to produce more realistic images without losing or collapse of details in key regions. • Our method achieves consistent appearance generation across different poses, especially in challenging regions like faces and textures, leading to state-of-the-art results on the Deepfashion dataset [22] and flexible-but-accurate editing downstream applications. 2. Related Works Pose-Guided Person Image Synthesis was initially proposed by PG2 [26], which firstly applied conditional GANs to adversarially refine pose-guided human generation. Later, GAN-based research addressed this problem through two main approaches. The first focuses on the transfer process, where methods model the deformation between poses using affine transformations [42,43] and flow fields [18,33–35]. The second approach aims to enhance the generation quality by better disentangling pose and appearance information. This disentanglement can be implicitly achieved by modeling the spatial correspondence between the pose and appearance features [27,36,45,51,54, 55,57]. Auxiliary explicit information is also introduced to improve the appearance quality, especially for UV texture map [7,40,41,50] that provides pose-irrelevant appearance guidance. However, due to the instability in training and the mode collapse issues associated with GAN models, previous GAN-based works have encountered challenges with unrealistic or changed textures in posed person images. To mitigate this limitation, diffusion based methods have been more recently introduced in PGPIS. PIDM [2] was the first to utilize the iterative denoising property of the diffusion model. Subsequent methods [9,24] have sought to improve the generation capability by employing latent diffusion models [37] (LDM) rather than the pixel space. In detail, CFLD [24] addresses the importance of semantic understanding towards the decoupling of fine-grained appearance and poses while PoCoLD [9] establishes the correspondence between pose and appearance. More recent some video person animation methods also took the benefit of compressed latent in LDM, but they mainly concentrated on keeping the temporal consistency by spatial attention [12,47] and consistent pose guidance [56]. Both the image and video synthesis methods use a source person image as condition and the generated image would collapse when the target pose varies greatly from the source image. In addition, it has been noticed that there is a deterioration [1,9] of image quality when LDM compresses images to lower dimensions, especially for images of highfrequency information. However, very few considered tackling this problem. Conditional Diffusion Models. Recently, diffusion models [11,44] have outperformed GANs and significantly improved the visual fidelity of synthesized images across various domains, including text-to-image generation [37–39], person image generation [2,4,9,24,48], and 3D avatar generation [13,17,19,20,31]. For most tasks, the widely used model is Stable Diffusion [37] (and its variants), which is a unified conditional diffusion model that allows for semantic maps, text, or images to be used as conditions for controlling generation. Its key contributions lie in applying the diffusion process in latent space, which minimizes resource consumption while maintaining generation quality and flexibility. In this paper, our architecture, along with the baseline’s, is derived from this conditional model, i.e, Stable Diffusion. Previous conditional diffusion models can be categorized into three types based on the condition: textconditioned [37,38], image-conditioned [12,15,24,28], and mixed-conditioned models [49,52]. These methods typically use a pretrained model [29,32,37] to extract condition features, which are then injected into the denoising UNet via cross-attention. Different from the main stream approaches that regard images and texts as a whole, our proposed Multi-focal Conditioned method takes a human image as the input, transforms it by different focuses(e.g., texture maps and facial features), and encodes these focuses into embeddings using various pre-trained models. This approach is a sophisticated combination of image-conditioned and mixed-conditioned strategies. Additionally, we introduce a Multi-focal Conditions Aggregation technique to effectively distribute these conditions into the UNet. 16020 … … … ReferenceNet Pose Guider Face Encoder Projection Multi-focal Condition Aggregation (MFCA) CLIPCLIP VAE Appearance Region (A)Face Region (F) Source Human Image (I) Warp Texturemap Source Pose Target Pose (pt) Noise Target Image (It) (b) MFCA Q Ks (a) Multi-focal Feature Extraction KfVf (d) Pose Guider Vs Softmax WkF WvF WQ Wks Wvs Softmax z s Femb … Cref Cemb (c) Denoising UNet Femb Iemb Aemb Aref Iemb Aemb Figure 2. The overall pipeline of our proposed Multi-focal Conditioned Diffusion Model. (a) Face regions and appearance regions are first extracted from the source person images; (b) multi-focal condition aggregation module ϕis used to fuse the focal embeddings as cemb; (c) ReferenceNet Ris used to aggregate information from the appearance texture map, denoted as cref ; (d) Densepose provides the pose control to be fused into UNet with noise by Pose Guider. 3. Methodology Given a reference image Irepresenting the appearance condition c, the task of PGPIS aims to generate a target image Itwith a desired pose pt. This is achieved by learning a conditional network Tsuch that It=T(c, pt). While the representation of ptis typically predefined [3,8,23], the success of generation largely relies on the network Tand conditions c, which extract the shared pose-irrelevant appearance features between Iand It, ensuring high-quality synthesis of It. To enhance synthesis, we introduce a diffusion model ϵθconditioned by multiple factors, collectively denoted as c∗, which iteratively recovers Itfrom noise. 3.1. Multi-Conditioned Latent Diffusion Model The backbone of our proposed method is based on Stable Diffusion [37] (SD), which is an implementation of LDM. An encoder Ecompresses the image Ito a latent z, and a decoder Dtransforms zback to an image I′=D(z). The compressed latent representation reduces the optimization spaces and allows the generation of higher resolution and richer diversity. The optimization of loss Lin LDM can be repurposed as: Lmse =Ezt,p,t,ϵ,c∗(||ϵ−ϵθ(zt, pt, t, c∗)||),(1) where ϵθrepresents the forward process of UNet in LDM, ptis the target pose, ztis the noisy latent zunder timestep t, and c∗is our multi-focal condition. Despite the advantages of having a latent representation, I′deteriorates during the compression process. While the perceptual differences between I′and Imay be very small, this degeneration diminishes the significance of the latent code z, particularly for features that are supposed to exhibit substantial variance in the original input I, such as facial traits and garment texture. Furthermore, this deterioration is further magnified since Lmse of Tcould not guarantee the generated latent without any deviation, and finally results in an unsatisfactory appearance generation results in these high-frequency regions. Previous LDM-based methods [9, 24] have neglected this issue by relying only on images, which resulted in the model’s failure to accurately generate sensitive regions. To address this problem, we propose a solution that utilizes multi-focal conditions c∗to focus attention on the important areas of the image. To implement this approach, we have designed a two-branch conditional diffusion model that effectively captures multi-focal attention. The pipeline is shown in Fig.2. On the first branch, we follow the structure of ReferenceNet [12] to provide the semantic and lowlevel features cref , which are concatenated with the UNet 16021 features in each stage. In the second branch, we exploit pretrained models to embed three selected focal features from the source image I, face region F, and appearance region A, respectively. These embeddings are aggregated into UNet with our Multi-focal Conditions Aggregation (MFCA). 3.2. Multi-focal Conditions Aggregation. Multi-focal Regions. To enhance latent feature preservation, we incorporate high-frequency focal regions c∗from the source image Ias conditioning inputs. These focal regions help guide attention mechanisms to mitigate the degradation of human-sensitive features. In our implementation, we focus on regions of the face and appearance that, although they constitute a small portion of the image, capture essential perceptual variations. The degradation of these areas within the autoencoder can lead to losing fine details, potentially causing the latent feature representation to overlook subtle distinctions present in the source image. Specifically, we employ [21] to crop the source image Iobtaining the face region F. Additionally, we attain the appearance region Aby warping Iinto a structured texture map defined by the SMPL model [23], indexing from its estimated DensePose [8]pI. The texture map disentangles appearance from the posed image, retaining only the poseinvariant texture information. Multi-focal Embeddings. The three multi-focal conditions are managed using pretrained modules. Starting with a source image I, we extract its embedding Iemb using a pretrained CLIP image encoder [32]. The texture map Ais processed in two ways through T. First, we encode Awith a VAE encoder [37], producing an output Aref , which is then passed to ReferenceNet R. Additionally, we use CLIP to obtain the texture map encoding Aemb. For facial regions F, we note that general image encoders like CLIP may struggle to accurately capture identity features, as face appearance and view in Itmay differ significantly from those in I. To address this, we utilize a pretrained face recognition model [5] to localize the face region and extract identity features. These features are then projected to match the dimensionality of the previous embeddings, noted as Femb. It’s important to note that both Femb and Aemb are shared between Iand It, as they are pose-invariant and represent attributes of the same appearance. Multi-focal Conditioning. The conditions c∗are assembled as follows: c∗=cref =R(Aref ) cemb =ϕ(Iemb,Aemb,Femb, z),(2) where Ris a trainable ReferenceNet extracting both the structured details and layouts of appearance regions. ϕdenotes a mulit-focal condition aggregation module (MFCA) to aggregate the embeddings to UNet. zis a latent input in UNet. Inspired by InstantID [46], ϕis defined as follows (see Fig.2(b)): ϕ=X i∈{s,Femb} λiAttn(Q, Ki, Vi), Q=zWq, Ki=iWki, Vi=iWvi, (3) where Q,Ki,Viare query, key and value matrices for crossattentions. Wis the attention weight and λiis the scaling factor. Qis computed from latent zwhile Ki,Viare computed from conditioning embeddings i, including the face Femb and a selective condition switcher s.sis defined as: s=   Iemb if z ∈ UE cat(Iemb,Aemb)if z ∈ UM Aemb if z ∈ UD (4) where UE,UM,UDare the encoder, the latent layer and decoder of UNet, respectively. When combining all conditions, our Multi-Focal Condition Aggregator (MFCA) efficiently aggregates the multifocal embeddings. This efficiency stems from reducing attention operations to focus on a specific region at each step, while simultaneously leveraging the embedding properties and the inherent structure of the UNet architecture. Moreover, we introduce a selective condition injection approach to accommodate the distinct characteristics of the UNet structure. Specifically, UEencodes information into a lower-dimensional space, where injecting global information from Iemb related to high-level semantics, such as cloth categories, and general background. Conversely, during the decoding stage, UDrequires fine-grained information to effectively reconstruct the final image; thus, Aemb are injected to provide pose-irrelevant features, such as texture details and garments details, at this phase to fulfill that need. This targeted injection strategy reduces parameter counts and guides the model to prioritize the information most relevant to each specific architectural stage. Since Femb is derived from a pretrained face recognition model, it maintains robustness across diverse views and appearances. We retain Femb throughout all stages of UNet to consistently represent both the input and target faces. An addition operation is employed to aggregate Femb and s. Pose Guider. We harness Densepose as our pose condition as it provides appropriate 3D information as claimed in PoCoLD [9]. In addition, Densepose coordinates establish a bijection between the UV space texture map Aand image pixels of It, which implicitly bridges the appearance alignment for the two focuses. Similar to [12], we employ a lightweight pose guider module constructed with a series of convolutional layers derived from ControlNet. This module is initialized with pretrained parameters from the ControlNet segmentation model, enabling it to leverage prior knowledge for enhanced guidance. 16022 Methods FID↓LPIPS↓SSIM ↑PSNR↑ Evaluation on 256 ×176 resolution GFLA‡[34] (CVPR20) 9.827 0.1878 0.7082 – SPGNet‡[25] (CVPR21) 16.184 0.2256 0.6965 17.222 NTED‡[36] (CVPR22) 8.517 0.177 0.7156 17.74 CASD‡[55](ECCV22) 13.137 0.1781 0.7224 17.880 PIDM†[2] (CVPR23) 6.36 0.1678 0.7312 – PoCoLD†[9] (ICCV23) 8.067 0.1642 0.7310 – CFLD†(CVPR24) 6.804 0.1519 0.7378 18.235 MCLD (B3) 6.86 0.157 0.734 18.03 MCLD (Ours) 6.693 0.1482 0.7511 18.84 Evaluation on 512 ×352 resolution CoCosNets [54] (CVPR22) 13.325 0.2265 0.7236 – NTED‡[36] (CVPR22) 7.645 0.1999 0.7359 17.385 PIDM†[2] (CVPR23) 5.8365 0.1768 0.7419 – PoCoLD†[9] (ICCV23) 8.416 0.1920 0.7430 – CFLD†[24] (CVPR24) 7.149 0.1819 0.7478 17.645 MCLD (B3) 7.23 0.1951 0.7405 17.48 MCLD (Ours) 7.079 0.1757 0.7557 18.211 Table 1. Qualitative comparison with the state-of-the-arts in terms of image quality benchmarks. †The scores are reported in their paper, since the same split is followed. ‡The scores are evaluated and reported in CFLD [24], since they split validation set in a different way. Our evaluation code is the same as CFLD. 3.3. Overall objective To force the model to concentrate more on the target face regions, we introduce an extra loss for supervision at face regions: Lface =Ezt,p,t,ϵ,c∗(||(ϵ−ϵθ(zt, pt, t, c∗)) ⊙m||)(5) where mis the segmentation mask of face regions, which is parsed from the densepose pt. Combining eq.(1), the overall objective function is: Loverall =Lmse +Lface (6) 4. Experiments In this section, we present a detailed analysis of our experiments including the dataset setup, evaluation metrics, implementation details, and a thorough comparison of our approach with state-of-the-art methods. Dataset. Following [2,9,24,57], experiments are conducted using the DeepFashion In-Shop Clothes Retrieval Benchmark [22], which contains 52,712 high-resolution images of fashion models. Consistent with the CFLD, we split the dataset into training and validation subsets, comprising 101,966 and 8,570 non-overlapping image pairs, respectively. Pose pairs are extracted by Densepose and we evaluate results on 256×176 and 512×352 resolutions. Metrics. We conduct two groups of objective metrics to evaluate the overall generated image quality and the generated face preservation, respectively. For the overall generated image quality, four metrics are adopted for comparison. The Fr´ echet Inception Distance (FID) [10] meaMethods FSref ↑distref ↓FStgt↑disttgt↓ Evaluation on 256 ×176 resolution CASD [55] (ECCV22) 0.207 28.80 0.317 26.28 PIDM [2] (CVPR23) 0.270 28.06 0.394 25.17 CFLD [24] (CVPR24) 0.243 28.86 0.363 26.11 MCLD (B3) 0.279 28.1 0.381 25.7 MCLD (Ours) 0.301 27.65 0.413 25.02 ref – – 0.497 22.53 Evaluation on 512 ×352 resolution CFLD [24] (CVPR24) 0.227 28.62 0.286 27.54 MCLD (B3) 0.289 27.04 0.333 26.25 MCLD (Ours) 0.294 27.01 0.344 26.31 ref – – 0.643 17.42 Table 2. Qualitative comparison with the state of the art regarding face quality benchmarks. F S is the face similarity metric, while dist is the euclidean distance measure. ref refer to the input source human image, tgt is the ground truth image. Both ref and tgt are real images. sures the Wasserstein-2 distance between the feature distributions of generated images and real images, with features extracted from the Inception-v3 pretrained network. Specifically, the generated image features come from the validation dataset, while the real image features are from the training dataset. The Learned Perceptual Image Patch Similarity (LPIPS) [53] computes image-wise similarity in the perceptual feature space. Both FID and LPIPS assess image quality in a high-level feature domain. Additionally, we use two pixel-wise metrics: the Structural Similarity Index Measure (SSIM) and Peak Signal-to-Noise Ratio (PSNR), which evaluate the accuracy of pixel-wise correspondences between the generated and real images. To assess the identity preservation of the face region, we use a pretrained Face Recognition Model [5] to extract the face embeddings and compute the face cosine similarity F S and euclidean distance dist between the face regions of generated images and real images. Both the source image ref and the target image tgt are evaluated to assess the overall model ability. Implementation details. Our model is implemented on Stable Diffusion [37] 1.5 model using PyTorch [30] and Huggingface Diffusers. The source image and the target image are resized to 512×512. Face regions are detected by a single shot detector [21] implemented in OpenCV [14], while the face embedding is acquired by a pretrained face analysis model, antelopev21. For appearance regions, the images are first converted to 24 parts defined in Densepose with the size of 200×200, then these parts are transformed to 512×512 SMPL texture map by a predefined mapping. The model is trained for 60,000 iterations using Adam optimizer [16] with a learning rate of 1e-5. We train our model on two Nvidia A100 GPUs with a batch size of 12 for each GPU. During sampling, a classifier-free guidance (CFG) strategy is adopted to improve the sampling quality. We 1https://github.com/deepinsight/insightface 16023 (1) (2) (3) (4) (5) (6) (7) (8) Source Pose Target NTED CASD PIDM CFLD Ours Source Pose Target NTED CASD PIDM CFLD Ours Figure 3. Qualitative Comparison with several state-of-the-art models on the Deepfashion dataset. The inputs to our models are the target pose ptand the source person image I. From left to right the results are of NTED, CASD, PIDM, CFLD and ours respectively. set the CFG scale to 3.5 and λiin MFCA to 1 and 0.5. 4.1. Quantitative Comparison Our method is compared with both GAN-based and diffusion-based state-of-the-art approaches. Specifically, PIDM [2] is diffusion based while PoCoLD [9] and CFLD [24] is LDM-based. The evaluation is performed on two resolutions, 256×176 and 512×352. In addition, we compare our method with our baseline B3since it has an aggregation structure similar to [46]. As shown in Tab 1, our method performs better by conditioning with multifocal regions in the image quality benchmarks. LDM-based methods are known to encounter challenges due to autoencoder compression, which often results in suboptimal FID scores compared to fully diffusion-based approaches. Our proposed method mitigates these limitations, achieving improved FID scores among LDM-based techniques. While certain recent diffusion-based methods do not publicly release their best-performing checkpoints, we report results as stated in their respective publications. Additionally, as demonstrated in Tab. 2, our method exhibits robust identity preservation across evaluated metrics. The table also includes similarity metrics between the reference source image and target ground truth, where our method achieves performance on par with reference images, which serve as authentic representations providing facial cues to the network. 4.2. Qualitative Comparison We present our comprehensive visual comparison with recent approaches that release their validation results or reproducible, from the left to right is NETD [36], CASD [55], PIDM [2], CFLD [24] and ours respectively. We observe several conclusions listed below. Firstly, current methods are suffering from reconstructing the details of the textures since they only use the source image as condition. This is especially noticed in GAN based methods and LDM based method. This is partially because of the limited details representation ability of GAN, and the information deteriora16024 Method Conditions Aggregation Params FID LPIPS SSIM PSNR Evaluation on 256 ×176 resolution B1 I– 1622 M 6.427 0.1629 0.7371 18.18 B2 I,Aconcat 1622 M 6.830 0.1620 0.7357 18.20 B3 I,A,Fconcat 1698 M 6.858 0.1536 0.7340 18.03 B4 F,AMFCA 1711 M 6.717 0.1619 0.734 18.16 B5 I,AMFCA 1622 M 6.723 0.1483 0.7499 18.72 Ours I,A,FMFCA 1717 M 6.693 0.1482 0.7511 18.84 Evaluation on 512 ×352 resolution B1 I– 1622 M 6.738 0.1923 0.7463 17.64 B2 I,Aconcat 1622 M 7.129 0.1932 0.7433 17.63 B3 I,A,Fconcat 1698 M 7.23 0.1951 0.7405 17.48 B4 F,AMFCA 1711 M 7.138 0.1923 0.7408 17.56 B5 I,AMFCA 1622 M 7.047 0.1766 0.7543 18.13 Ours I,A,FMFCA 1717 M 7.079 0.1757 0.7557 18.21 Table 3. Qualitative comparison for ablation studies. I,A,Fare the embeddings from source images, appearance regions, and face regions respectively. Aggregation column refers to the feature fusion strategy. Params refers to the trainable parameter in network. Source Image Target Pose GT B1 B2 B3 B4 B5 Ours Figure 4. Qualitative ablation comparison. Refer to Tab. 3for baseline settings. tion in LDM. However, after introducing the appearance regions by texture map, our method shows a consistent generation results when the provided information from appearance region and face region is adequate. In rows 1-2, our method preserves better clothing styles even when these styles is rare to be seen in dataset. In rows 3-4, our method also shows a consistent ability to reconstruct the appearance patterns under the given reference images. While other methods are struggled to the details of original patterns. In row 5, for these input images with complex patterns, all the methods fail to reproduce the same details. However, our methods shows a consistent layout of cloths. In addition, identity preservation is one of the most challenging task for current methods, since it is highly sensitive from human perception but not for generative losses. As the illustrated image shows, especially in rows 6-8, our method performs a good identity preserving by introducing the invariant face region embeddings as conditions and supervisions. 4.3. Ablation Study We perform ablation studies on our MFCA module to demonstrate the value of multi-focal conditions. The quantitative result is shown in Tab. 3, while the qualitative results are illustrated in Fig. 4.B1only takes Iemb from the source image as a condition, which is similar to other image-based methods. Due to the limited power of the image condition, the generated image fails to preserve facial and textural traits, introducing undesired artifacts. When we gradually add Aemb and Femb to B2and B3with a concatenation strategy, the cloth style and textures in B2slightly improve. Introducing facial conditioning (i.e.B3) increasingly improves performance. However, this simple concatenation does not ensure stable performance. When too many conditions are handled in parallel, the effort for each condition remains unclear, and the focused areas become ambiguous, resulting in unpredictable cloths styles, textures, and identities. Quantitative results also prove that concatenation struggles to improve the generated image quality. Thus, in B4and B5, we adopt our MFCA without the conditions Iemb and Femb, respectively. Overall, this results in a diminishing in unwanted artifacts due to the reduced attention regions. Qualitatively, when dropping the Iemb in B4, the cloth styles lose in detail. Femb and Aemb only receive the information inside the Densepose estimation, and the regions outside Aare randomly generated which is not consistent to the semantic cloth style. In B5texture improves in quality but the facial traits are almost entirely lost.This seems to confirm our quantitative findings where the deformed and incomplete warping in the texture map affects the facial appearance. Though our method is close to B5 in terms of metrics, probably due to the fact that the face regions only occupy a small portion of the entire image, the facial traits are well preserved. Finally, we noticed a decrease in FID performance after introducing more conditions. As reported in [24], the FID of VAE reconstruction in LDM methods is 7.967. Consequently, a lower FID in LDM-based methods does not necessarily indicate a superior overall performance. The other three metrics provide stronger quantitative evidence. 4.4. Appearance Editing Our approach enables flexible, localized editing by adjusting specific focal conditions within the generative pipeline, allowing precise control over focal regions. Editing examples are illustrated in Fig. 5. By modifying the texture map Afor designated clothing regions, we can seamlessly alter clothing to reflect chosen reference styles, showcasing the strong control capability of our texture map focalization (row 2). Additionally, by substituting the face embedding Femb and updating the corresponding facial regions in the texture map, our method supports person identity swapping (row 3). This disentangling of maps, facial identity, and pose permits arbitrary combinations of identities, clothing styles, and poses. For selective edits, such as altering only specific clothing 16025 Ref Image Ref Image Ref Image Source Source Source Image Image Image Pose Pose Pose Figure 5. Appearance editing results. Our method accepts flexible editing of given identities, poses, and clothes. This is achieved only by modifying some regions of conditions, and no need for any masking or further training. Source Image Warped Texturemap Target Pose Ground Truth Generated Image (1) (2) (3) (4) (5) Source Image Warped Texturemap Target Pose Ground Truth Generated Image Figure 6. Failure cases caused by (1) Wrong target pose, (2) Incomplete texture map, (3) Squeezed texture map, (4) Missing face information, (5) Significant view changes. parts, we can replace corresponding regions within the texture map, which is particularly effective for simpler clothing designs with clear texture segments(as shown in the top section of the row 4). Unlike traditional diffusion-based methods, which rely on mask-based blending within latent spaces, our method provides a more streamlined and adaptable editing solution through structured modifications. In general, our approach offers a more straightforward and flexible editing solution by solely modifying the structured conditions. This highlights the superiority of our proposed Multi-Focal Conditions Aggregation module in terms of editing capabilities. Furthermore, our editing results are more realistic than those of baseline methods [2,9,24], as they avoid the boundary artifacts often associated with mask-based techniques. A detailed comparison can be found in the supplementary materials. 4.5. Failure Cases Despite achieving satisfactory appearance-preserving ability in most cases, our model occasionally fails to produce desired results when dealing with uncommon or abrupt images, as shown in Fig. 6. We notice several failure scenarios: (1) the target pose is wrongly estimated, (2) the source texture map is missing or incomplete, (3) the source texture map is fully estimated, but its appearance is shifted to limited pixel resolutions. (4) missing facial traits; (5) significant view changes that are not captured by source image. 5. Conclusions In this paper, we introduced the MCLD framework for pose-guided person image generation. We addressed the challenge of compression degradation in LDMs, especially over sensitive regions, by developing a multi-focal conditioning strategy that strengthens control over both identity and appearance. Our MFCA module selectively integrates pose-invariant focal points as conditioning inputs, significantly enhancing the quality of the generated images. Through extensive qualitative and quantitative evaluations, we demonstrate that MFCA surpasses existing methods in preserving both the appearance and identity of the subject. Moreover, our approach enables more flexible image editing through improved condition disentanglement. In future work, we aim to explore 3D priors to further enhance generation consistency and improve appearance fidelity. Acknowledgement This work was supported by the MUR PNRR project FAIR (PE00000013) funded by the NextGenerationEU and the EU Horizon project ELIAS (No. 101120237). We acknowledge the CINECA award under the ISCRA initiative, for the availability of highperformance computing resources and support. 16026 References [1] Omri Avrahami, Ohad Fried, and Dani Lischinski. Blended latent diffusion. ACM transactions on graphics (TOG), 2023. 1,2 [2] Ankan Kumar Bhunia, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Jorma Laaksonen, Mubarak Shah, and Fahad Shahbaz Khan. Person image synthesis via denoising diffusion model. In CVPR, 2023. 1,2,5,6,8 [3] Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh. Realtime multi-person 2d pose estimation using part affinity fields. In CVPR, 2017. 3 [4] Soon Yau Cheong, Armin Mustafa, and Andrew Gilbert. Upgpt: Universal diffusion model for person image generation, editing and pose transfer. In ICCV, 2023. 2 [5] Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. Arcface: Additive angular margin loss for deep face recognition. In CVPR, 2019. 4,5 [6] Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In NeurIPS, 2014. 1 [7] Artur Grigorev, Artem Sevastopolsky, Alexander Vakhitov, and Victor Lempitsky. Coordinate-based texture inpainting for pose-guided human image generation. In CVPR, 2019. 2 [8] Rıza Alp G¨ uler, Natalia Neverova, and Iasonas Kokkinos. Densepose: Dense human pose estimation in the wild. In CVPR, 2018. 3,4 [9] Xiao Han, Xiatian Zhu, Jiankang Deng, Yi-Zhe Song, and Tao Xiang. Controllable person image synthesis with poseconstrained latent diffusion. In ICCV, 2023. 1,2,3,4,5,6, 8 [10] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In NeurIPS, 2017. 5 [11] Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. NeurIPS, 33, 2020. 1,2 [12] Li Hu. Animate anyone: Consistent and controllable imageto-video synthesis for character animation. In CVPR, 2024. 2,3,4 [13] Xin Huang, Ruizhi Shao, Qi Zhang, Hongwen Zhang, Ying Feng, Yebin Liu, and Qing Wang. Humannorm: Learning normal diffusion model for high-quality and realistic 3d human generation. CVPR, 2024. 2 [14] Itseez. Open source computer vision library. https:// github.com/itseez/opencv, 2015. 5 [15] Jeongho Kim, Guojung Gu, Minho Park, Sunghyun Park, and Jaegul Choo. Stableviton: Learning semantic correspondence with latent diffusion model for virtual try-on. In CVPR, 2024. 2 [16] Diederik P Kingma. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 5 [17] Nikos Kolotouros, Thiemo Alldieck, Andrei Zanfir, Eduard Gabriel Bazavan, Mihai Fieraru, and Cristian Sminchisescu. Dreamhuman: Animatable 3d avatars from text. NeurIPS, 2023. 2 [18] Yining Li, Chen Huang, and Chen Change Loy. Dense intrinsic appearance flow for human pose transfer. In CVPR, 2019. 2 [19] Tingting Liao, Hongwei Yi, Yuliang Xiu, Jiaxaing Tang, Yangyi Huang, Justus Thies, and Michael J Black. Tada! text to animatable digital avatars. 3DV, 2024. 2 [20] Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. Zero-1-to-3: Zero-shot one image to 3d object, 2023. 2 [21] Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg. Ssd: Single shot multibox detector. In ECCV. Springer, 2016. 4,5 [22] Ziwei Liu, Ping Luo, Shi Qiu, Xiaogang Wang, and Xiaoou Tang. Deepfashion: Powering robust clothes recognition and retrieval with rich annotations. In CVPR, 2016. 2,5 [23] Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. SMPL: A skinned multiperson linear model. ACM Transactions on Graphics (TOG), 34(6):248:1–248:16, 2015. 3,4 [24] Yanzuo Lu, Manlin Zhang, Andy J Ma, Xiaohua Xie, and Jianhuang Lai. Coarse-to-fine latent diffusion for poseguided person image synthesis. In CVPR, 2024. 1,2,3, 5,6,7,8 [25] Zhengyao Lv, Xiaoming Li, Xin Li, Fu Li, Tianwei Lin, Dongliang He, and Wangmeng Zuo. Learning semantic person image generation by region-adaptive normalization. In CVPR, 2021. 5 [26] Liqian Ma, Xu Jia, Qianru Sun, Bernt Schiele, Tinne Tuytelaars, and Luc Van Gool. Pose guided person image generation. In NeurIPS, 2017. 2 [27] Yifang Men, Yiming Mao, Yuning Jiang, Wei-Ying Ma, and Zhouhui Lian. Controllable person image synthesis with attribute-decomposed gan. In CVPR, 2020. 1,2 [28] Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, and Ying Shan. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. In AAAI, 2024. 2 [29] Maxime Oquab, Timoth´ ee Darcet, Theo Moutakanni, Huy V. Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Russell Howes, Po-Yao Huang, Hu Xu, Vasu Sharma, ShangWen Li, Wojciech Galuba, Mike Rabbat, Mido Assran, Nicolas Ballas, Gabriel Synnaeve, Ishan Misra, Herve Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski. Dinov2: Learning robust visual features without supervision, 2023. 2 [30] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. NeurIPS, 2019. 5 [31] Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. ICLR, 2022. 2 [32] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, 16027