Grounding Stylistic Domain Generalization with Quantitative Domain Shift Measures and Synthetic Scene Images
Fuente:
arXiv
Salvato in:
| Autori principali: | Luo, Yiran, Feinglass, Joshua, Gokhale, Tejas, Lee, Kuan-Cheng, Baral, Chitta, Yang, Yezhou |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models
di: Chatterjee, Agneet, et al.
Pubblicazione: (2024)
di: Chatterjee, Agneet, et al.
Pubblicazione: (2024)
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
di: Patel, Maitreya, et al.
Pubblicazione: (2023)
di: Patel, Maitreya, et al.
Pubblicazione: (2023)
On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
di: Chatterjee, Agneet, et al.
Pubblicazione: (2024)
di: Chatterjee, Agneet, et al.
Pubblicazione: (2024)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
TROPE: TRaining-Free Object-Part Enhancement for Seamlessly Improving Fine-Grained Zero-Shot Image Captioning
di: Feinglass, Joshua, et al.
Pubblicazione: (2024)
di: Feinglass, Joshua, et al.
Pubblicazione: (2024)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
di: Yilmaz, Nilay, et al.
Pubblicazione: (2025)
di: Yilmaz, Nilay, et al.
Pubblicazione: (2025)
ActionCOMET: A Zero-shot Approach to Learn Image-specific Commonsense Concepts about Actions
di: Sampat, Shailaja Keyur, et al.
Pubblicazione: (2024)
di: Sampat, Shailaja Keyur, et al.
Pubblicazione: (2024)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs
di: Vani, Sameep, et al.
Pubblicazione: (2025)
di: Vani, Sameep, et al.
Pubblicazione: (2025)
AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
di: Malaviya, Vatsal, et al.
Pubblicazione: (2025)
di: Malaviya, Vatsal, et al.
Pubblicazione: (2025)
Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts
di: Saxon, Michael, et al.
Pubblicazione: (2024)
di: Saxon, Michael, et al.
Pubblicazione: (2024)
RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions
di: Pathiraja, Bimsara, et al.
Pubblicazione: (2025)
di: Pathiraja, Bimsara, et al.
Pubblicazione: (2025)
Dual Caption Preference Optimization for Diffusion Models
di: Saeidi, Amir, et al.
Pubblicazione: (2025)
di: Saeidi, Amir, et al.
Pubblicazione: (2025)
`Eyes of a Hawk and Ears of a Fox': Part Prototype Network for Generalized Zero-Shot Learning
di: Feinglass, Joshua, et al.
Pubblicazione: (2024)
di: Feinglass, Joshua, et al.
Pubblicazione: (2024)
Help Me Identify: Is an LLM+VQA System All We Need to Identify Visual Concepts?
di: Sampat, Shailaja Keyur, et al.
Pubblicazione: (2024)
di: Sampat, Shailaja Keyur, et al.
Pubblicazione: (2024)
Getting it Right: Improving Spatial Consistency in Text-to-Image Models
di: Chatterjee, Agneet, et al.
Pubblicazione: (2024)
di: Chatterjee, Agneet, et al.
Pubblicazione: (2024)
The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs
di: Anvekar, Tejas, et al.
Pubblicazione: (2025)
di: Anvekar, Tejas, et al.
Pubblicazione: (2025)
Chimera: Compositional Image Generation using Part-based Concepting
di: Singh, Shivam, et al.
Pubblicazione: (2025)
di: Singh, Shivam, et al.
Pubblicazione: (2025)
Improving Shift Invariance in Convolutional Neural Networks with Translation Invariant Polyphase Sampling
di: Saha, Sourajit, et al.
Pubblicazione: (2024)
di: Saha, Sourajit, et al.
Pubblicazione: (2024)
Investigating VLM Hallucination from a Cognitive Psychology Perspective: A First Step Toward Interpretation with Intriguing Observations
di: Liu, Xiangrui, et al.
Pubblicazione: (2025)
di: Liu, Xiangrui, et al.
Pubblicazione: (2025)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
di: Fallah, Forouzan, et al.
Pubblicazione: (2025)
di: Fallah, Forouzan, et al.
Pubblicazione: (2025)
EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven Alignment
di: Kusumba, Abhiram, et al.
Pubblicazione: (2025)
di: Kusumba, Abhiram, et al.
Pubblicazione: (2025)
Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation
di: Chatterjee, Agneet, et al.
Pubblicazione: (2025)
di: Chatterjee, Agneet, et al.
Pubblicazione: (2025)
Latent Diffusion Unlearning: Protecting Against Unauthorized Personalization Through Trajectory Shifted Perturbations
di: Devulapally, Naresh Kumar, et al.
Pubblicazione: (2025)
di: Devulapally, Naresh Kumar, et al.
Pubblicazione: (2025)
Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model
di: Cheng, Sheng, et al.
Pubblicazione: (2024)
di: Cheng, Sheng, et al.
Pubblicazione: (2024)
VL-GLUE: A Suite of Fundamental yet Challenging Visuo-Linguistic Reasoning Tasks
di: Sampat, Shailaja Keyur, et al.
Pubblicazione: (2024)
di: Sampat, Shailaja Keyur, et al.
Pubblicazione: (2024)
Uncertainty-Aware ControlNet: Bridging Domain Gaps with Synthetic Image Generation
di: Niemeijer, Joshua, et al.
Pubblicazione: (2025)
di: Niemeijer, Joshua, et al.
Pubblicazione: (2025)
GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
di: Siingh, Shikhhar, et al.
Pubblicazione: (2025)
di: Siingh, Shikhhar, et al.
Pubblicazione: (2025)
DomainVerse: A Benchmark Towards Real-World Distribution Shifts For Tuning-Free Adaptive Domain Generalization
di: Hou, Feng, et al.
Pubblicazione: (2024)
di: Hou, Feng, et al.
Pubblicazione: (2024)
Zero Shot Domain Adaptive Semantic Segmentation by Synthetic Data Generation and Progressive Adaptation
di: Luo, Jun, et al.
Pubblicazione: (2025)
di: Luo, Jun, et al.
Pubblicazione: (2025)
Pseudo Dataset Generation for Out-of-Domain Multi-Camera View Recommendation
di: Lee, Kuan-Ying, et al.
Pubblicazione: (2024)
di: Lee, Kuan-Ying, et al.
Pubblicazione: (2024)
Stylecodes: Encoding Stylistic Information For Image Generation
di: Rowles, Ciara
Pubblicazione: (2024)
di: Rowles, Ciara
Pubblicazione: (2024)
Soft Segmented Randomization: Enhancing Domain Generalization in SAR-ATR for Synthetic-to-Measured
di: Kim, Minjun, et al.
Pubblicazione: (2024)
di: Kim, Minjun, et al.
Pubblicazione: (2024)
Federated Learning with Domain Shift Eraser
di: Wang, Zheng, et al.
Pubblicazione: (2025)
di: Wang, Zheng, et al.
Pubblicazione: (2025)
Generalized Category Discovery under Domain Shift: A Frequency Domain Perspective
di: Feng, Wei, et al.
Pubblicazione: (2025)
di: Feng, Wei, et al.
Pubblicazione: (2025)
GEN3D: Generating Domain-Free 3D Scenes from a Single Image
di: Zhang, Yuxin, et al.
Pubblicazione: (2025)
di: Zhang, Yuxin, et al.
Pubblicazione: (2025)
Attributed Synthetic Data Generation for Zero-shot Domain-specific Image Classification
di: Wang, Shijian, et al.
Pubblicazione: (2025)
di: Wang, Shijian, et al.
Pubblicazione: (2025)
Weighted Risk Invariance: Domain Generalization under Invariant Feature Shift
di: Wong, Gina, et al.
Pubblicazione: (2024)
di: Wong, Gina, et al.
Pubblicazione: (2024)
Let Synthetic Data Shine: Domain Reassembly and Soft-Fusion for Single Domain Generalization
di: Li, Hao, et al.
Pubblicazione: (2025)
di: Li, Hao, et al.
Pubblicazione: (2025)
Photovoltaic Defect Image Generator with Boundary Alignment Smoothing Constraint for Domain Shift Mitigation
di: Li, Dongying, et al.
Pubblicazione: (2025)
di: Li, Dongying, et al.
Pubblicazione: (2025)
Documenti analoghi
-
REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models
di: Chatterjee, Agneet, et al.
Pubblicazione: (2024) -
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
di: Patel, Maitreya, et al.
Pubblicazione: (2023) -
On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
di: Chatterjee, Agneet, et al.
Pubblicazione: (2024) -
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
di: Patel, Maitreya, et al.
Pubblicazione: (2024) -
TROPE: TRaining-Free Object-Part Enhancement for Seamlessly Improving Fine-Grained Zero-Shot Image Captioning
di: Feinglass, Joshua, et al.
Pubblicazione: (2024)