RespoDiff: Dual-Module Bottleneck Transformation for Responsible & Faithful T2I Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sreelatha, Silpa Vadakkeeveetil, Nag, Sauradip, Awais, Muhammad, Belongie, Serge, Dutta, Anjan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RAIGen: Rare Attribute Identification in Text-to-Image Generative Models
von: Sreelatha, Silpa Vadakkeeveetil, et al.
Veröffentlicht: (2026)
von: Sreelatha, Silpa Vadakkeeveetil, et al.
Veröffentlicht: (2026)
DeNetDM: Debiasing by Network Depth Modulation
von: Sreelatha, Silpa Vadakkeeveetil, et al.
Veröffentlicht: (2024)
von: Sreelatha, Silpa Vadakkeeveetil, et al.
Veröffentlicht: (2024)
OmniCount: Multi-label Object Counting with Semantic-Geometric Priors
von: Mondal, Anindya, et al.
Veröffentlicht: (2024)
von: Mondal, Anindya, et al.
Veröffentlicht: (2024)
CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance
von: Mondal, Anindya, et al.
Veröffentlicht: (2025)
von: Mondal, Anindya, et al.
Veröffentlicht: (2025)
Actor-agnostic Multi-label Action Recognition with Multi-modal Query
von: Mondal, Anindya, et al.
Veröffentlicht: (2023)
von: Mondal, Anindya, et al.
Veröffentlicht: (2023)
Articulate That Object Part (ATOP): 3D Part Articulation via Text and Motion Personalization
von: Vora, Aditya, et al.
Veröffentlicht: (2025)
von: Vora, Aditya, et al.
Veröffentlicht: (2025)
In-2-4D: Inbetweening from Two Single-View Images to 4D Generation
von: Nag, Sauradip, et al.
Veröffentlicht: (2025)
von: Nag, Sauradip, et al.
Veröffentlicht: (2025)
MMEarth-Bench: Global Model Adaptation via Multimodal Test-Time Training
von: Gordon, Lucia, et al.
Veröffentlicht: (2026)
von: Gordon, Lucia, et al.
Veröffentlicht: (2026)
FaithDiff: Unleashing Diffusion Priors for Faithful Image Super-resolution
von: Chen, Junyang, et al.
Veröffentlicht: (2024)
von: Chen, Junyang, et al.
Veröffentlicht: (2024)
ASIA: Adaptive 3D Segmentation using Few Image Annotations
von: Perla, Sai Raj Kishore, et al.
Veröffentlicht: (2025)
von: Perla, Sai Raj Kishore, et al.
Veröffentlicht: (2025)
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning
von: Deria, Ankan, et al.
Veröffentlicht: (2025)
von: Deria, Ankan, et al.
Veröffentlicht: (2025)
PhysConvex: Physics-Informed 3D Dynamic Convex Radiance Fields for Reconstruction and Simulation
von: Wang, Dan, et al.
Veröffentlicht: (2026)
von: Wang, Dan, et al.
Veröffentlicht: (2026)
HiddenObjects: Scalable Diffusion-Distilled Spatial Priors for Object Placement
von: Schouten, Marco, et al.
Veröffentlicht: (2026)
von: Schouten, Marco, et al.
Veröffentlicht: (2026)
Labeled Data Selection for Category Discovery
von: Zhao, Bingchen, et al.
Veröffentlicht: (2024)
von: Zhao, Bingchen, et al.
Veröffentlicht: (2024)
SMITE: Segment Me In TimE
von: Alimohammadi, Amirhossein, et al.
Veröffentlicht: (2024)
von: Alimohammadi, Amirhossein, et al.
Veröffentlicht: (2024)
TaleDiffusion: Multi-Character Story Generation with Dialogue Rendering
von: Banerjee, Ayan, et al.
Veröffentlicht: (2025)
von: Banerjee, Ayan, et al.
Veröffentlicht: (2025)
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation
von: Yu, Hong-Tao, et al.
Veröffentlicht: (2025)
von: Yu, Hong-Tao, et al.
Veröffentlicht: (2025)
POEM: Precise Object-level Editing via MLLM control
von: Schouten, Marco, et al.
Veröffentlicht: (2025)
von: Schouten, Marco, et al.
Veröffentlicht: (2025)
Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model
von: An, Zhaochong, et al.
Veröffentlicht: (2025)
von: An, Zhaochong, et al.
Veröffentlicht: (2025)
Towards Faithful Multimodal Concept Bottleneck Models
von: Moreau, Pierre, et al.
Veröffentlicht: (2026)
von: Moreau, Pierre, et al.
Veröffentlicht: (2026)
Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
von: Bader, Jessica, et al.
Veröffentlicht: (2025)
von: Bader, Jessica, et al.
Veröffentlicht: (2025)
Unlearning-based Neural Interpretations
von: Choi, Ching Lam, et al.
Veröffentlicht: (2024)
von: Choi, Ching Lam, et al.
Veröffentlicht: (2024)
DC-ViT: Modulating Spatial and Channel Interactions for Multi-Channel Images
von: Marikkar, Umar, et al.
Veröffentlicht: (2026)
von: Marikkar, Umar, et al.
Veröffentlicht: (2026)
DeltaDiff: Reality-Driven Diffusion with AnchorResiduals for Faithful SR
von: Yang, Chao, et al.
Veröffentlicht: (2025)
von: Yang, Chao, et al.
Veröffentlicht: (2025)
Advances in 4D Representation: Geometry, Motion, and Interaction
von: Zhao, Mingrui, et al.
Veröffentlicht: (2025)
von: Zhao, Mingrui, et al.
Veröffentlicht: (2025)
Noise-Coded Illumination for Forensic and Photometric Video Analysis
von: Michael, Peter F., et al.
Veröffentlicht: (2025)
von: Michael, Peter F., et al.
Veröffentlicht: (2025)
Familiarity-Based Open-Set Recognition Under Adversarial Attacks
von: Enevoldsen, Philip, et al.
Veröffentlicht: (2023)
von: Enevoldsen, Philip, et al.
Veröffentlicht: (2023)
CLIPDraw++: Text-to-Sketch Synthesis with Simple Primitives
von: Mathur, Nityanand, et al.
Veröffentlicht: (2023)
von: Mathur, Nityanand, et al.
Veröffentlicht: (2023)
Learning Conditional Invariances through Non-Commutativity
von: Chaudhuri, Abhra, et al.
Veröffentlicht: (2024)
von: Chaudhuri, Abhra, et al.
Veröffentlicht: (2024)
VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward
von: An, Zhaochong, et al.
Veröffentlicht: (2026)
von: An, Zhaochong, et al.
Veröffentlicht: (2026)
On the Faithfulness of Vision Transformer Explanations
von: Wu, Junyi, et al.
Veröffentlicht: (2024)
von: Wu, Junyi, et al.
Veröffentlicht: (2024)
DualDiff+: Dual-Branch Diffusion for High-Fidelity Video Generation with Reward Guidance
von: Yang, Zhao, et al.
Veröffentlicht: (2025)
von: Yang, Zhao, et al.
Veröffentlicht: (2025)
SuperF: Neural Implicit Fields for Multi-Image Super-Resolution
von: Jyhne, Sander Riisøen, et al.
Veröffentlicht: (2025)
von: Jyhne, Sander Riisøen, et al.
Veröffentlicht: (2025)
Taxonomy-Aware Evaluation of Vision-Language Models
von: Snæbjarnarson, Vésteinn, et al.
Veröffentlicht: (2025)
von: Snæbjarnarson, Vésteinn, et al.
Veröffentlicht: (2025)
Revisiting the Perception-Distortion Trade-off with Spatial-Semantic Guided Super-Resolution
von: Wang, Dan, et al.
Veröffentlicht: (2026)
von: Wang, Dan, et al.
Veröffentlicht: (2026)
Domain Adaptation Without the Compute Burden for Efficient Whole Slide Image Analysis
von: Marikkar, Umar, et al.
Veröffentlicht: (2026)
von: Marikkar, Umar, et al.
Veröffentlicht: (2026)
DiffFormer: a Differential Spatial-Spectral Transformer for Hyperspectral Image Classification
von: Ahmad, Muhammad, et al.
Veröffentlicht: (2024)
von: Ahmad, Muhammad, et al.
Veröffentlicht: (2024)
Better Language Models Exhibit Higher Visual Alignment
von: Ruthardt, Jona, et al.
Veröffentlicht: (2024)
von: Ruthardt, Jona, et al.
Veröffentlicht: (2024)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
von: Pach, Mateusz, et al.
Veröffentlicht: (2025)
von: Pach, Mateusz, et al.
Veröffentlicht: (2025)
The Latent Color Subspace: Emergent Order in High-Dimensional Chaos
von: Pach, Mateusz, et al.
Veröffentlicht: (2026)
von: Pach, Mateusz, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
RAIGen: Rare Attribute Identification in Text-to-Image Generative Models
von: Sreelatha, Silpa Vadakkeeveetil, et al.
Veröffentlicht: (2026) -
DeNetDM: Debiasing by Network Depth Modulation
von: Sreelatha, Silpa Vadakkeeveetil, et al.
Veröffentlicht: (2024) -
OmniCount: Multi-label Object Counting with Semantic-Geometric Priors
von: Mondal, Anindya, et al.
Veröffentlicht: (2024) -
CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance
von: Mondal, Anindya, et al.
Veröffentlicht: (2025) -
Actor-agnostic Multi-label Action Recognition with Multi-modal Query
von: Mondal, Anindya, et al.
Veröffentlicht: (2023)