Abstract:Facial attribute editing refers to the task of modifying facial images by altering specific target facial attributes. Existing approaches typically rely on the combination of generative adversarial networks and encoder–decoder architectures to tackle this problem. However, current methods may exhibit limited accuracy when dealing with certain attributes. The primary objective of this research is to enhance facial image modification based on user-specified target facial attributes, such as hair color, beard removal, or gender transformation. During the editing process, it is crucial to selectively modify only the regions relevant to the target attributes while preserving the details of other unrelated facial attributes. This ensures that the editing results appear more natural and realistic. This study introduces a novel approach called MAGAN (Combining GRU Structure and Additive Attention with AGU—Adaptive Gated Units). Moreover, a discriminative attention mechanism is introduced to automatically identify key regions in the input images that are relevant to facial attributes. This mechanism concentrates attention on these regions, enhancing the model's ability to accurately capture and analyze subtle facial attribute features. The method incorporates external attention within the convolutional layers of the encoder–decoder architecture, facilitating the modeling of linear complexity across image regions and implicitly considering correlations among all data samples. By employing discriminative attention in the discriminator, the model achieves more precise attribute editing. To evaluate the effectiveness of MAGAN, experiments were conducted on the CelebA dataset. The average precision of facial attribute generation in images edited by our model stands at 91.83%. PSNR and SSIM for reconstructed images are 32.52 and 0.957, respectively. In comparison with existing methodologies (AttGAN, STGAN, MUGAN), noteworthy enhancements have been achieved in the domain of facial attribute manipulation.

MAKIMA: Tuning-free Multi-Attribute Open-domain Video Editing via Mask-Guided Attention Modulation

Motion Guided Token Compression for Efficient Masked Video Modeling

FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing

MaskINT: Video Editing via Interpolative Non-autoregressive Masked Transformers

EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing

Temporally Consistent Object Editing in Videos using Extended Attention

Blended Latent Diffusion under Attention Control for Real-World Video Editing

MagicStick: Controllable Video Editing via Control Handle Transformations

UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing

Attention-guided Temporally Coherent Video Object Matting

Vision-guided and Mask-enhanced Adaptive Denoising for Prompt-based Image Editing

Dual-Path Temporal Map Optimization for Make-up Temporal Video Grounding

Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion Models

Video-P2P: Video Editing with Cross-attention Control

VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking

Video Editing with Temporal, Spatial and Appearance Consistency

Multi-Attention Infused Integrated Facial Attribute Editing Model: Enhancing the Robustness of Facial Attribute Manipulation

Text-Guided Video Masked Autoencoder

Enhancing Text-to-Image Editing via Hybrid Mask-Informed Fusion

Perceptual Attributes Optimization For Multivideo Summarization

MaTe3D: Mask-guided Text-based 3D-aware Portrait Editing