Saved in:
| Main Authors: | Burgess, James, Abdal, Rameen, Stoddart, Dan, Tulyakov, Sergey, Yeung-Levy, Serena, Wang, Kuan-Chieh Jackson |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.09475 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Personalization Turing Test
by: Abdal, Rameen, et al.
Published: (2026)
by: Abdal, Rameen, et al.
Published: (2026)
Tuning-free Visual Effect Transfer across Videos
by: Jones, Maxwell, et al.
Published: (2026)
by: Jones, Maxwell, et al.
Published: (2026)
Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Models
by: Burgess, James, et al.
Published: (2023)
by: Burgess, James, et al.
Published: (2023)
Dynamic Concepts Personalization from Single Videos
by: Abdal, Rameen, et al.
Published: (2025)
by: Abdal, Rameen, et al.
Published: (2025)
Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA
by: Abdal, Rameen, et al.
Published: (2025)
by: Abdal, Rameen, et al.
Published: (2025)
Motion Diffusion-Guided 3D Global HMR from a Dynamic Camera
by: Heo, Jaewoo, et al.
Published: (2024)
by: Heo, Jaewoo, et al.
Published: (2024)
Improving the Diffusability of Autoencoders
by: Skorokhodov, Ivan, et al.
Published: (2025)
by: Skorokhodov, Ivan, et al.
Published: (2025)
Helix4D: Complex 4D Mesh Generation
by: Yenphraphai, Jiraphon, et al.
Published: (2026)
by: Yenphraphai, Jiraphon, et al.
Published: (2026)
Interpreting the Weight Space of Customized Diffusion Models
by: Dravid, Amil, et al.
Published: (2024)
by: Dravid, Amil, et al.
Published: (2024)
NearID: Identity Representation Learning via Near-identity Distractors
by: Cvejic, Aleksandar, et al.
Published: (2026)
by: Cvejic, Aleksandar, et al.
Published: (2026)
The Impact of Image Resolution on Biomedical Multimodal Large Language Models
by: Chen, Liangyu, et al.
Published: (2025)
by: Chen, Liangyu, et al.
Published: (2025)
MyVLM: Personalizing VLMs for User-Specific Queries
by: Alaluf, Yuval, et al.
Published: (2024)
by: Alaluf, Yuval, et al.
Published: (2024)
Ask, Pose, Unite: Scaling Data Acquisition for Close Interactions with Vision Language Models
by: Bravo-Sánchez, Laura, et al.
Published: (2024)
by: Bravo-Sánchez, Laura, et al.
Published: (2024)
Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
by: Tang, Yuqi, et al.
Published: (2026)
by: Tang, Yuqi, et al.
Published: (2026)
ComposeMe: Attribute-Specific Image Prompts for Controllable Human Image Generation
by: Qian, Guocheng Gordon, et al.
Published: (2025)
by: Qian, Guocheng Gordon, et al.
Published: (2025)
Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
by: Endo, Mark, et al.
Published: (2025)
by: Endo, Mark, et al.
Published: (2025)
MoA: Mixture-of-Attention for Subject-Context Disentanglement in Personalized Image Generation
by: Wang, Kuan-Chieh, et al.
Published: (2024)
by: Wang, Kuan-Chieh, et al.
Published: (2024)
Improving Artifact Robustness for CT Deep Learning Models Without Labeled Artifact Images via Domain Adaptation
by: Cheung, Justin, et al.
Published: (2025)
by: Cheung, Justin, et al.
Published: (2025)
Foundation Models Secretly Understand Neural Network Weights: Enhancing Hypernetwork Architectures with Foundation Models
by: Gu, Jeffrey, et al.
Published: (2025)
by: Gu, Jeffrey, et al.
Published: (2025)
Synthesizing Artifact Dataset for Pixel-level Detection
by: Menn, Dennis, et al.
Published: (2025)
by: Menn, Dennis, et al.
Published: (2025)
Canvas-to-Image: Compositional Image Generation with Multimodal Controls
by: Dalva, Yusuf, et al.
Published: (2025)
by: Dalva, Yusuf, et al.
Published: (2025)
Continuous Perception Matters: Diagnosing Temporal Integration Failures in Multimodal Models
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
by: Aklilu, Josiah, et al.
Published: (2024)
by: Aklilu, Josiah, et al.
Published: (2024)
Multi-Human Mesh Recovery with Transformers
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration
by: Endo, Mark, et al.
Published: (2024)
by: Endo, Mark, et al.
Published: (2024)
Omni-ID: Holistic Identity Representation Designed for Generative Tasks
by: Qian, Guocheng, et al.
Published: (2024)
by: Qian, Guocheng, et al.
Published: (2024)
μ-Bench: A Vision-Language Benchmark for Microscopy Understanding
by: Lozano, Alejandro, et al.
Published: (2024)
by: Lozano, Alejandro, et al.
Published: (2024)
Detecting Human Artifacts from Text-to-Image Models
by: Wang, Kaihong, et al.
Published: (2024)
by: Wang, Kaihong, et al.
Published: (2024)
Simple Visual Artifact Detection in Sora-Generated Videos
by: Sugiyama, Misora, et al.
Published: (2025)
by: Sugiyama, Misora, et al.
Published: (2025)
SynArtifact: Classifying and Alleviating Artifacts in Synthetic Images via Vision-Language Model
by: Cao, Bin, et al.
Published: (2024)
by: Cao, Bin, et al.
Published: (2024)
Improving Image Clustering with Artifacts Attenuation via Inference-Time Attention Engineering
by: Nakamura, Kazumoto, et al.
Published: (2024)
by: Nakamura, Kazumoto, et al.
Published: (2024)
Diffusion-HPC: Synthetic Data Generation for Human Mesh Recovery in Challenging Domains
by: Weng, Zhenzhen, et al.
Published: (2023)
by: Weng, Zhenzhen, et al.
Published: (2023)
DiffusionQC: Artifact Detection in Histopathology via Diffusion Model
by: Wang, Zhenzhen, et al.
Published: (2026)
by: Wang, Zhenzhen, et al.
Published: (2026)
Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual Illusions
by: Sun, Xiaoxiao, et al.
Published: (2026)
by: Sun, Xiaoxiao, et al.
Published: (2026)
ArtifactWorld: Scaling 3D Gaussian Splatting Artifact Restoration via Video Generation Models
by: Wang, Xinliang, et al.
Published: (2026)
by: Wang, Xinliang, et al.
Published: (2026)
Automated Motion Artifact Check for MRI (AutoMAC-MRI): An Interpretable Framework for Motion Artifact Detection and Severity Assessment
by: Jerald, Antony, et al.
Published: (2025)
by: Jerald, Antony, et al.
Published: (2025)
See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis
by: Park, Jaehyun, et al.
Published: (2026)
by: Park, Jaehyun, et al.
Published: (2026)
Zero-Shot Artifact2Artifact: Self-incentive artifact removal for photoacoustic imaging without any data
by: Li, Shuang, et al.
Published: (2024)
by: Li, Shuang, et al.
Published: (2024)
DeforHMR: Vision Transformer with Deformable Cross-Attention for 3D Human Mesh Recovery
by: Heo, Jaewoo, et al.
Published: (2024)
by: Heo, Jaewoo, et al.
Published: (2024)
Counterfactual Explanations for Face Forgery Detection via Adversarial Removal of Artifacts
by: Li, Yang, et al.
Published: (2024)
by: Li, Yang, et al.
Published: (2024)
Similar Items
-
Visual Personalization Turing Test
by: Abdal, Rameen, et al.
Published: (2026) -
Tuning-free Visual Effect Transfer across Videos
by: Jones, Maxwell, et al.
Published: (2026) -
Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Models
by: Burgess, James, et al.
Published: (2023) -
Dynamic Concepts Personalization from Single Videos
by: Abdal, Rameen, et al.
Published: (2025) -
Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA
by: Abdal, Rameen, et al.
Published: (2025)