$\textit{Revelio}$: Interpreting and leveraging semantic information in diffusion models
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Dahye, Thomas, Xavier, Ghadiyaram, Deepti |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Concept Steerers: Leveraging K-Sparse Autoencoders for Test-Time Controllable Generations
by: Kim, Dahye, et al.
Published: (2025)
by: Kim, Dahye, et al.
Published: (2025)
DDiT: Dynamic Patch Scheduling for Efficient Diffusion Transformers
by: Kim, Dahye, et al.
Published: (2026)
by: Kim, Dahye, et al.
Published: (2026)
What's in a Latent? Leveraging Diffusion Latent Space for Domain Generalization
by: Thomas, Xavier, et al.
Published: (2025)
by: Thomas, Xavier, et al.
Published: (2025)
Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance
by: Qiu, Jason, et al.
Published: (2026)
by: Qiu, Jason, et al.
Published: (2026)
A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning
by: Chen, Tianle, et al.
Published: (2026)
by: Chen, Tianle, et al.
Published: (2026)
Swift Sampling: Selecting Temporal Surprises via Taylor Series
by: Kim, Dahye, et al.
Published: (2026)
by: Kim, Dahye, et al.
Published: (2026)
Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
by: Thomas, Xavier, et al.
Published: (2025)
by: Thomas, Xavier, et al.
Published: (2025)
Progressive Prompt Detailing for Improved Alignment in Text-to-Image Generative Models
by: Saichandran, Ketan Suhaas, et al.
Published: (2025)
by: Saichandran, Ketan Suhaas, et al.
Published: (2025)
Some Modalities are More Equal Than Others: Decoding and Architecting Multimodal Integration in MLLMs
by: Chen, Tianle, et al.
Published: (2025)
by: Chen, Tianle, et al.
Published: (2025)
Improving Physical Object State Representation in Text-to-Image Generative Systems
by: Chen, Tianle, et al.
Published: (2025)
by: Chen, Tianle, et al.
Published: (2025)
FuTCR: Future-Targeted Contrast and Repulsion for Continual Panoptic Segmentation
by: Ikechukwu, Nicholas, et al.
Published: (2026)
by: Ikechukwu, Nicholas, et al.
Published: (2026)
FAGER: Factually Grounded Evaluation and Refinement of Text-to-Image Models
by: Lim, Youngsun, et al.
Published: (2026)
by: Lim, Youngsun, et al.
Published: (2026)
Seeing Isn't Orienting: A Cognitively Grounded Benchmark Reveals Systematic Orientation Failures in MLLMs Supplementary
by: Tasnim, Nazia, et al.
Published: (2026)
by: Tasnim, Nazia, et al.
Published: (2026)
Seeing Isn't Orienting: A Cognitively Grounded Benchmark Reveals Systematic Orientation Failures in MLLMs
by: Tasnim, Nazia, et al.
Published: (2025)
by: Tasnim, Nazia, et al.
Published: (2025)
GeoDE: a Geographically Diverse Evaluation Dataset for Object Recognition
by: Ramaswamy, Vikram V., et al.
Published: (2023)
by: Ramaswamy, Vikram V., et al.
Published: (2023)
SLIM: Semantic-based Low-bitrate Image compression for Machines by leveraging diffusion
by: Lee, Hyeonjin, et al.
Published: (2025)
by: Lee, Hyeonjin, et al.
Published: (2025)
Multimodal generative semantic communication based on latent diffusion model
by: Fu, Weiqi, et al.
Published: (2024)
by: Fu, Weiqi, et al.
Published: (2024)
FindMeIfYouCan: Bringing Open Set metrics to $\textit{near} $, $ \textit{far} $ and $\textit{farther}$ Out-of-Distribution Object Detection
by: Montoya, Daniel, et al.
Published: (2025)
by: Montoya, Daniel, et al.
Published: (2025)
\textit{FocaLogic}: Logic-Based Interpretation of Visual Model Decisions
by: Zhao, Chenchen, et al.
Published: (2026)
by: Zhao, Chenchen, et al.
Published: (2026)
Enriching thermal point clouds of buildings using semantic 3D building models
by: Zhu, Jingwei, et al.
Published: (2024)
by: Zhu, Jingwei, et al.
Published: (2024)
Radiometric fingerprinting of object surfaces using mobile laser scanning and semantic 3D road space models
by: Schwab, Benedikt, et al.
Published: (2026)
by: Schwab, Benedikt, et al.
Published: (2026)
Revelio: A Real-World Screen-Camera Communication System with Visually Imperceptible Data Embedding
by: Nishar, Abbaas Alif Mohamed, et al.
Published: (2025)
by: Nishar, Abbaas Alif Mohamed, et al.
Published: (2025)
Novel class discovery meets foundation models for 3D semantic segmentation
by: Riz, Luigi, et al.
Published: (2023)
by: Riz, Luigi, et al.
Published: (2023)
Topological SLAM in colonoscopies leveraging deep features and topological priors
by: Morlana, Javier, et al.
Published: (2024)
by: Morlana, Javier, et al.
Published: (2024)
IPNET:Influential Prototypical Networks for Few Shot Learning
by: Chowdhury, Ranjana Roy, et al.
Published: (2022)
by: Chowdhury, Ranjana Roy, et al.
Published: (2022)
A Novel Approach to for Multimodal Emotion Recognition : Multimodal semantic information fusion
by: Dai, Wei, et al.
Published: (2025)
by: Dai, Wei, et al.
Published: (2025)
Enhancing the vision-language foundation model with key semantic knowledge-emphasized report refinement
by: Huang, Weijian, et al.
Published: (2024)
by: Huang, Weijian, et al.
Published: (2024)
Hybridnet for depth estimation and semantic segmentation
by: Sánchez-Escobedo, Dalila, et al.
Published: (2024)
by: Sánchez-Escobedo, Dalila, et al.
Published: (2024)
\textit{4DSurf}: High-Fidelity Dynamic Scene Surface Reconstruction
by: Wu, Renjie, et al.
Published: (2026)
by: Wu, Renjie, et al.
Published: (2026)
UDHF2-Net: Uncertainty-diffusion-model-based High-Frequency TransFormer Network for Remotely Sensed Imagery Interpretation
by: Zhang, Pengfei, et al.
Published: (2024)
by: Zhang, Pengfei, et al.
Published: (2024)
An empirical study for the early detection of Mpox from skin lesion images using pretrained CNN models leveraging XAI technique
by: Rahim, Mohammad Asifur, et al.
Published: (2025)
by: Rahim, Mohammad Asifur, et al.
Published: (2025)
Fine color guidance in diffusion models and its application to image compression at extremely low bitrates
by: Bordin, Tom, et al.
Published: (2024)
by: Bordin, Tom, et al.
Published: (2024)
CLIP-VQDiffusion : Langauge Free Training of Text To Image generation using CLIP and vector quantized diffusion model
by: Han, Seungdae, et al.
Published: (2024)
by: Han, Seungdae, et al.
Published: (2024)
Dynamic semantic VSLAM with known and unknown objects
by: Gu, Sanghyoup, et al.
Published: (2024)
by: Gu, Sanghyoup, et al.
Published: (2024)
Shift and matching queries for video semantic segmentation
by: Mizuno, Tsubasa, et al.
Published: (2024)
by: Mizuno, Tsubasa, et al.
Published: (2024)
$\textit{A Contrario}$ Paradigm for YOLO-based Infrared Small Target Detection
by: Ciocarlan, Alina, et al.
Published: (2024)
by: Ciocarlan, Alina, et al.
Published: (2024)
GaussExplorer: 3D Gaussian Splatting for Embodied Exploration and Reasoning
by: Yu-Ji, Kim, et al.
Published: (2026)
by: Yu-Ji, Kim, et al.
Published: (2026)
Interpreting vision transformers via residual replacement model
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy
by: Ging, Simon, et al.
Published: (2024)
by: Ging, Simon, et al.
Published: (2024)
Deep-NFA: a Deep $\textit{a contrario}$ Framework for Small Object Detection
by: Ciocarlan, Alina, et al.
Published: (2023)
by: Ciocarlan, Alina, et al.
Published: (2023)
Similar Items
-
Concept Steerers: Leveraging K-Sparse Autoencoders for Test-Time Controllable Generations
by: Kim, Dahye, et al.
Published: (2025) -
DDiT: Dynamic Patch Scheduling for Efficient Diffusion Transformers
by: Kim, Dahye, et al.
Published: (2026) -
What's in a Latent? Leveraging Diffusion Latent Space for Domain Generalization
by: Thomas, Xavier, et al.
Published: (2025) -
Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance
by: Qiu, Jason, et al.
Published: (2026) -
A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning
by: Chen, Tianle, et al.
Published: (2026)