Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Eldesokey, Abdelrahman, Cvejic, Aleksandar, Ghanem, Bernard, Wonka, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PartEdit: Fine-Grained Image Editing using Pre-Trained Diffusion Models
by: Cvejic, Aleksandar, et al.
Published: (2025)
by: Cvejic, Aleksandar, et al.
Published: (2025)
NearID: Identity Representation Learning via Near-identity Distractors
by: Cvejic, Aleksandar, et al.
Published: (2026)
by: Cvejic, Aleksandar, et al.
Published: (2026)
EditCLIP: Representation Learning for Image Editing
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
Build-A-Scene: Interactive 3D Layout Control for Diffusion-Based Image Generation
by: Eldesokey, Abdelrahman, et al.
Published: (2024)
by: Eldesokey, Abdelrahman, et al.
Published: (2024)
LatentMan: Generating Consistent Animated Characters using Image Diffusion Models
by: Eldesokey, Abdelrahman, et al.
Published: (2023)
by: Eldesokey, Abdelrahman, et al.
Published: (2023)
Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation
by: Eldesokey, Abdelrahman, et al.
Published: (2026)
by: Eldesokey, Abdelrahman, et al.
Published: (2026)
ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language Models
by: Gong, Bingchen, et al.
Published: (2024)
by: Gong, Bingchen, et al.
Published: (2024)
AvatarMMC: 3D Head Avatar Generation and Editing with Multi-Modal Conditioning
by: Para, Wamiq Reyaz, et al.
Published: (2024)
by: Para, Wamiq Reyaz, et al.
Published: (2024)
Zero-Shot Video Semantic Segmentation based on Pre-Trained Diffusion Models
by: Wang, Qian, et al.
Published: (2024)
by: Wang, Qian, et al.
Published: (2024)
SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy
by: Elsharkawi, Ismael, et al.
Published: (2026)
by: Elsharkawi, Ismael, et al.
Published: (2026)
CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models
by: Alzahrani, Reem, et al.
Published: (2026)
by: Alzahrani, Reem, et al.
Published: (2026)
Vivid-ZOO: Multi-View Video Generation with Diffusion Model
by: Li, Bing, et al.
Published: (2024)
by: Li, Bing, et al.
Published: (2024)
Out-of-Distribution Segmentation via Wasserstein-Based Evidential Uncertainty
by: Brosch, Arnold, et al.
Published: (2025)
by: Brosch, Arnold, et al.
Published: (2025)
PlaceIt3D: Language-Guided Object Placement in Real 3D Scenes
by: Abdelreheem, Ahmed, et al.
Published: (2025)
by: Abdelreheem, Ahmed, et al.
Published: (2025)
RESP: Reference-guided Sequential Prompting for Visual Glitch Detection in Video Games
by: Yu, Yakun, et al.
Published: (2026)
by: Yu, Yakun, et al.
Published: (2026)
Generative Human Geometry Distribution
by: Tang, Xiangjun, et al.
Published: (2025)
by: Tang, Xiangjun, et al.
Published: (2025)
SHIC: Shape-Image Correspondences with no Keypoint Supervision
by: Shtedritski, Aleksandar, et al.
Published: (2024)
by: Shtedritski, Aleksandar, et al.
Published: (2024)
LASPA: Latent Spatial Alignment for Fast Training-free Single Image Editing
by: Alharbi, Yazeed, et al.
Published: (2024)
by: Alharbi, Yazeed, et al.
Published: (2024)
Fine-Tuning Visual Autoregressive Models for Subject-Driven Generation
by: Chung, Jiwoo, et al.
Published: (2025)
by: Chung, Jiwoo, et al.
Published: (2025)
TempGlitch: Evaluating Vision-Language Models for Temporal Glitch Detection in Gameplay Videos
by: Yu, Yakun, et al.
Published: (2026)
by: Yu, Yakun, et al.
Published: (2026)
Human Geometry Distribution for 3D Animation Generation
by: Tang, Xiangjun, et al.
Published: (2025)
by: Tang, Xiangjun, et al.
Published: (2025)
Back to 3D: Few-Shot 3D Keypoint Detection with Back-Projected 2D Features
by: Wimmer, Thomas, et al.
Published: (2023)
by: Wimmer, Thomas, et al.
Published: (2023)
MindTuner: Cross-Subject Visual Decoding with Visual Fingerprint and Semantic Correction
by: Gong, Zixuan, et al.
Published: (2024)
by: Gong, Zixuan, et al.
Published: (2024)
EasyV2V: A High-quality Instruction-based Video Editing Framework
by: Mai, Jinjie, et al.
Published: (2025)
by: Mai, Jinjie, et al.
Published: (2025)
efunc: An Efficient Function Representation without Neural Networks
by: Zhang, Biao, et al.
Published: (2025)
by: Zhang, Biao, et al.
Published: (2025)
LaGeM: A Large Geometry Model for 3D Representation Learning and Diffusion
by: Zhang, Biao, et al.
Published: (2024)
by: Zhang, Biao, et al.
Published: (2024)
MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement
by: She, Dong, et al.
Published: (2025)
by: She, Dong, et al.
Published: (2025)
End-to-End Temporal Action Detection with 1B Parameters Across 1000 Frames
by: Liu, Shuming, et al.
Published: (2023)
by: Liu, Shuming, et al.
Published: (2023)
Dissolving Is Amplifying: Towards Fine-Grained Anomaly Detection
by: Shi, Jian, et al.
Published: (2023)
by: Shi, Jian, et al.
Published: (2023)
Video Self-Stitching Graph Network for Temporal Action Localization
by: Zhao, Chen, et al.
Published: (2020)
by: Zhao, Chen, et al.
Published: (2020)
ImmersePro: End-to-End Stereo Video Synthesis Via Implicit Disparity Learning
by: Shi, Jian, et al.
Published: (2024)
by: Shi, Jian, et al.
Published: (2024)
Zero-Shot Anomaly Detection in Battery Thermal Images Using Visual Question Answering with Prior Knowledge
by: Astrid, Marcella, et al.
Published: (2025)
by: Astrid, Marcella, et al.
Published: (2025)
Generative Timelines for Instructed Visual Assembly
by: Pardo, Alejandro, et al.
Published: (2024)
by: Pardo, Alejandro, et al.
Published: (2024)
SSR-Encoder: Encoding Selective Subject Representation for Subject-Driven Generation
by: Zhang, Yuxuan, et al.
Published: (2023)
by: Zhang, Yuxuan, et al.
Published: (2023)
HierRelTriple: Guiding Indoor Layout Generation with Hierarchical Relationship Triplet Losses
by: Sun, Kaifan, et al.
Published: (2025)
by: Sun, Kaifan, et al.
Published: (2025)
No Mesh, No Problem: Estimating Coral Volume and Surface from Sparse Multi-View Images
by: Farchione, Diego Eustachio, et al.
Published: (2025)
by: Farchione, Diego Eustachio, et al.
Published: (2025)
PatchRefiner: Leveraging Synthetic Data for Real-Domain High-Resolution Monocular Metric Depth Estimation
by: Li, Zhenyu, et al.
Published: (2024)
by: Li, Zhenyu, et al.
Published: (2024)
Stencil: Subject-Driven Generation with Context Guidance
by: Chen, Gordon, et al.
Published: (2025)
by: Chen, Gordon, et al.
Published: (2025)
DreamVAR: Taming Reinforced Visual Autoregressive Model for High-Fidelity Subject-Driven Image Generation
by: Jiang, Xin, et al.
Published: (2026)
by: Jiang, Xin, et al.
Published: (2026)
iFlame: Interleaving Full and Linear Attention for Efficient Mesh Generation
by: Wang, Hanxiao, et al.
Published: (2025)
by: Wang, Hanxiao, et al.
Published: (2025)
Similar Items
-
PartEdit: Fine-Grained Image Editing using Pre-Trained Diffusion Models
by: Cvejic, Aleksandar, et al.
Published: (2025) -
NearID: Identity Representation Learning via Near-identity Distractors
by: Cvejic, Aleksandar, et al.
Published: (2026) -
EditCLIP: Representation Learning for Image Editing
by: Wang, Qian, et al.
Published: (2025) -
Build-A-Scene: Interactive 3D Layout Control for Diffusion-Based Image Generation
by: Eldesokey, Abdelrahman, et al.
Published: (2024) -
LatentMan: Generating Consistent Animated Characters using Image Diffusion Models
by: Eldesokey, Abdelrahman, et al.
Published: (2023)