Logo image
StructFuse: Harmonizing multiple structural cues for diffusion-driven image compositing
Journal article   Open access   Peer reviewed

StructFuse: Harmonizing multiple structural cues for diffusion-driven image compositing

Waqas Ahmed, Dean Diepeveen and Ferdous Sohel
Pattern recognition, Vol.180, 114609
2026
pdf
Published4.84 MBDownloadView
Published (Version of Record) Open Access CC BY V4.0

Abstract

Image compositing Image harmonization Multi-modal structural fusion Synthetic data generation Diffusion model
Compositing visually coherent foreground and background elements remains a pivotal challenge in generative image editing. Prominent diffusion-based approaches either rely on rigid image-level inputs, which limit scene diversity, or focus on localized mask conditioning (e.g., edges, depth), while struggling to effectively combine multiple structural cues. To address these limitations, we propose an advanced diffusion-based framework that achieves precise foreground–background integration through the use of structured visual cues. Our approach introduces three key innovations: (1) a learnable adaptive gating strategy that dynamically fuses multiple structural cues (edge, contour, and depth) while effectively capturing richer and balanced conditional representations; (2) a bidirectional feature modulation method that injects multi-resolution control signals across the diffusion U-Net encoder–decoder hierarchy, significantly enhancing compositional quality; and (3) a mask-routed cross-attention mechanism that isolates text influence to user-specified regions, eliminating semantic bleeding. Additionally, we propose a novel pipeline for constructing a composite image dataset tailored to the image compositing task. Experimental results demonstrate that our framework outperforms state-of-the-art baseline methods across benchmark metrics, including Harmony Score, FID, SSIM, CLIP Score, and MUSIQ, reflecting strong compositional integrity and visual coherence. This approach offers a scalable and extensible solution for compositional image generation, with broad applications in synthetic dataset generation, creative content design, and augmented reality.

Details

Metrics

1 Record Views
Logo image