Do-Undo Bench: Reversibility for Action Understanding in Image Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mahajan, Shweta, Kadambi, Shreya, Le, Hoang, Yasarla, Rajeev, Bhattacharyya, Apratim, Hayat, Munawar, Porikli, Fatih |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DDIL: Diversity Enhancing Diffusion Distillation With Imitation Learning
von: Garrepalli, Risheek, et al.
Veröffentlicht: (2024)
von: Garrepalli, Risheek, et al.
Veröffentlicht: (2024)
Attention Guided Alignment in Efficient Vision-Language Models
von: Mahajan, Shweta, et al.
Veröffentlicht: (2025)
von: Mahajan, Shweta, et al.
Veröffentlicht: (2025)
RoCA: Robust Cross-Domain End-to-End Autonomous Driving
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2025)
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2025)
Distilling Multi-modal Large Language Models for Autonomous Driving
von: Hegde, Deepti, et al.
Veröffentlicht: (2025)
von: Hegde, Deepti, et al.
Veröffentlicht: (2025)
Resolving the Identity Crisis in Text-to-Image Generation
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
MultiHuman-Testbench: Benchmarking Image Generation for Multiple Humans
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing
von: Kadambi, Shreya, et al.
Veröffentlicht: (2025)
von: Kadambi, Shreya, et al.
Veröffentlicht: (2025)
DySS: Dynamic Queries and State-Space Learning for Efficient 3D Object Detection from Multi-Camera Videos
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2025)
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2025)
Generative Scenario Rollouts for End-to-End Autonomous Driving
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2026)
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2026)
ToSA: Token Selective Attention for Efficient Vision Transformers
von: Singh, Manish Kumar, et al.
Veröffentlicht: (2024)
von: Singh, Manish Kumar, et al.
Veröffentlicht: (2024)
SubZero: Composing Subject, Style, and Action via Zero-Shot Personalization
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
MAMo: Leveraging Memory and Attention for Monocular Video Depth Estimation
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2023)
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2023)
FouRA: Fourier Low Rank Adaptation
von: Borse, Shubhankar, et al.
Veröffentlicht: (2024)
von: Borse, Shubhankar, et al.
Veröffentlicht: (2024)
HexaGen3D: StableDiffusion is just one step away from Fast and Diverse Text-to-3D Generation
von: Mercier, Antoine, et al.
Veröffentlicht: (2024)
von: Mercier, Antoine, et al.
Veröffentlicht: (2024)
FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
von: Cho, Janghoon, et al.
Veröffentlicht: (2025)
von: Cho, Janghoon, et al.
Veröffentlicht: (2025)
Personalized OVSS: Understanding Personal Concept in Open-Vocabulary Semantic Segmentation
von: Park, Sunghyun, et al.
Veröffentlicht: (2025)
von: Park, Sunghyun, et al.
Veröffentlicht: (2025)
CustomKD: Customizing Large Vision Foundation for Edge Model Improvement via Knowledge Distillation
von: Lee, Jungsoo, et al.
Veröffentlicht: (2025)
von: Lee, Jungsoo, et al.
Veröffentlicht: (2025)
OCAI: Improving Optical Flow Estimation by Occlusion and Consistency Aware Interpolation
von: Jeong, Jisoo, et al.
Veröffentlicht: (2024)
von: Jeong, Jisoo, et al.
Veröffentlicht: (2024)
Generalized Contrastive Learning for Universal Multimodal Retrieval
von: Lee, Jungsoo, et al.
Veröffentlicht: (2025)
von: Lee, Jungsoo, et al.
Veröffentlicht: (2025)
ConsNoTrainLoRA: Data-driven Weight Initialization of Low-rank Adapters using Constraints
von: Das, Debasmit, et al.
Veröffentlicht: (2025)
von: Das, Debasmit, et al.
Veröffentlicht: (2025)
PosSAM: Panoptic Open-vocabulary Segment Anything
von: VS, Vibashan, et al.
Veröffentlicht: (2024)
von: VS, Vibashan, et al.
Veröffentlicht: (2024)
FutureDepth: Learning to Predict the Future Improves Video Depth Estimation
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2024)
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2024)
Ar2Can: An Architect and an Artist Leveraging a Canvas for Multi-Human Generation
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
Memory-Efficient Fine-Tuning Diffusion Transformers via Dynamic Patch Sampling and Block Skipping
von: Park, Sunghyun, et al.
Veröffentlicht: (2026)
von: Park, Sunghyun, et al.
Veröffentlicht: (2026)
Erasing Undesirable Influence in Diffusion Models
von: Wu, Jing, et al.
Veröffentlicht: (2024)
von: Wu, Jing, et al.
Veröffentlicht: (2024)
Segmentation-Free Guidance for Text-to-Image Diffusion Models
von: Azarian, Kambiz, et al.
Veröffentlicht: (2024)
von: Azarian, Kambiz, et al.
Veröffentlicht: (2024)
ForeSea: AI Forensic Search with Multi-modal Queries for Video Surveillance
von: Park, Hyojin, et al.
Veröffentlicht: (2026)
von: Park, Hyojin, et al.
Veröffentlicht: (2026)
Conditional Distribution Modelling for Few-Shot Image Synthesis with Diffusion Models
von: Gupta, Parul, et al.
Veröffentlicht: (2024)
von: Gupta, Parul, et al.
Veröffentlicht: (2024)
MAPLE: Latent Multi-Agent Play for End-to-End Autonomous Driving
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2026)
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2026)
Neural Mesh Fusion: Unsupervised 3D Planar Surface Understanding
von: Zanjani, Farhad G., et al.
Veröffentlicht: (2024)
von: Zanjani, Farhad G., et al.
Veröffentlicht: (2024)
1M-Deepfakes Detection Challenge
von: Cai, Zhixi, et al.
Veröffentlicht: (2024)
von: Cai, Zhixi, et al.
Veröffentlicht: (2024)
Tripartite Weight-Space Ensemble for Few-Shot Class-Incremental Learning
von: Lee, Juntae, et al.
Veröffentlicht: (2025)
von: Lee, Juntae, et al.
Veröffentlicht: (2025)
Sort-free Gaussian Splatting via Weighted Sum Rendering
von: Hou, Qiqi, et al.
Veröffentlicht: (2024)
von: Hou, Qiqi, et al.
Veröffentlicht: (2024)
HyperNet Fields: Efficiently Training Hypernetworks without Ground Truth by Learning Weight Trajectories
von: Hedlin, Eric, et al.
Veröffentlicht: (2024)
von: Hedlin, Eric, et al.
Veröffentlicht: (2024)
Visual Concept-driven Image Generation with Text-to-Image Diffusion Model
von: Rahman, Tanzila, et al.
Veröffentlicht: (2024)
von: Rahman, Tanzila, et al.
Veröffentlicht: (2024)
AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset
von: Cai, Zhixi, et al.
Veröffentlicht: (2023)
von: Cai, Zhixi, et al.
Veröffentlicht: (2023)
EdgeRelight360: Text-Conditioned 360-Degree HDR Image Generation for Real-Time On-Device Video Portrait Relighting
von: Lin, Min-Hui, et al.
Veröffentlicht: (2024)
von: Lin, Min-Hui, et al.
Veröffentlicht: (2024)
H3O: Hyper-Efficient 3D Occupancy Prediction with Heterogeneous Supervision
von: Shi, Yunxiao, et al.
Veröffentlicht: (2025)
von: Shi, Yunxiao, et al.
Veröffentlicht: (2025)
LoRA-X: Bridging Foundation Models with Training-Free Cross-Model Adaptation
von: Farhadzadeh, Farzad, et al.
Veröffentlicht: (2025)
von: Farhadzadeh, Farzad, et al.
Veröffentlicht: (2025)
Improving Optical Flow and Stereo Depth Estimation by Leveraging Uncertainty-Based Learning Difficulties
von: Jeong, Jisoo, et al.
Veröffentlicht: (2025)
von: Jeong, Jisoo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DDIL: Diversity Enhancing Diffusion Distillation With Imitation Learning
von: Garrepalli, Risheek, et al.
Veröffentlicht: (2024) -
Attention Guided Alignment in Efficient Vision-Language Models
von: Mahajan, Shweta, et al.
Veröffentlicht: (2025) -
RoCA: Robust Cross-Domain End-to-End Autonomous Driving
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2025) -
Distilling Multi-modal Large Language Models for Autonomous Driving
von: Hegde, Deepti, et al.
Veröffentlicht: (2025) -
Resolving the Identity Crisis in Text-to-Image Generation
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)