What Drives Compositional Generalization? The Importance of Continuous Training Objectives in Visual Generative Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Farid, Karim, Sahay, Rajat, Alnaggar, Yumna Ali, Schrodi, Simon, Fischer, Volker, Schmid, Cordelia, Brox, Thomas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When and How Does CLIP Enable Domain and Compositional Generalization?
von: Kempf, Elias, et al.
Veröffentlicht: (2025)
von: Kempf, Elias, et al.
Veröffentlicht: (2025)
Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models
von: Schrodi, Simon, et al.
Veröffentlicht: (2024)
von: Schrodi, Simon, et al.
Veröffentlicht: (2024)
Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems
von: Hoffmann, David T., et al.
Veröffentlicht: (2023)
von: Hoffmann, David T., et al.
Veröffentlicht: (2023)
Concept Bottleneck Models Without Predefined Concepts
von: Schrodi, Simon, et al.
Veröffentlicht: (2024)
von: Schrodi, Simon, et al.
Veröffentlicht: (2024)
Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models
von: Mousakhan, Arian, et al.
Veröffentlicht: (2025)
von: Mousakhan, Arian, et al.
Veröffentlicht: (2025)
ComposeAnything: Composite Object Priors for Text-to-Image Generation
von: Khan, Zeeshan, et al.
Veröffentlicht: (2025)
von: Khan, Zeeshan, et al.
Veröffentlicht: (2025)
BrickNet: Graph-Backed Generative Brick Assembly
von: Kulits, Peter, et al.
Veröffentlicht: (2026)
von: Kulits, Peter, et al.
Veröffentlicht: (2026)
A Generative Approach for Wikipedia-Scale Visual Entity Recognition
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
Grounded Video Caption Generation
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2024)
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2024)
LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
von: Kim, Jae Myung, et al.
Veröffentlicht: (2025)
von: Kim, Jae Myung, et al.
Veröffentlicht: (2025)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
von: Schrodi, Simon, et al.
Veröffentlicht: (2025)
von: Schrodi, Simon, et al.
Veröffentlicht: (2025)
Large-scale Pre-training for Grounded Video Caption Generation
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2025)
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2025)
MoPEFT: A Mixture-of-PEFTs for the Segment Anything Model
von: Sahay, Rajat, et al.
Veröffentlicht: (2024)
von: Sahay, Rajat, et al.
Veröffentlicht: (2024)
What Are You Doing? A Closer Look at Controllable Human Video Generation
von: Bugliarello, Emanuele, et al.
Veröffentlicht: (2025)
von: Bugliarello, Emanuele, et al.
Veröffentlicht: (2025)
StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization
von: Gaur, Gopalji, et al.
Veröffentlicht: (2025)
von: Gaur, Gopalji, et al.
Veröffentlicht: (2025)
RAVEN: Rethinking Adversarial Video Generation with Efficient Tri-plane Networks
von: Ghosh, Partha, et al.
Veröffentlicht: (2024)
von: Ghosh, Partha, et al.
Veröffentlicht: (2024)
SUGAR: Pre-training 3D Visual Representations for Robotics
von: Chen, Shizhe, et al.
Veröffentlicht: (2024)
von: Chen, Shizhe, et al.
Veröffentlicht: (2024)
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
Is Mamba Capable of In-Context Learning?
von: Grazzi, Riccardo, et al.
Veröffentlicht: (2024)
von: Grazzi, Riccardo, et al.
Veröffentlicht: (2024)
RECODE: Reasoning Through Code Generation for Visual Question Answering
von: Shen, Junhong, et al.
Veröffentlicht: (2025)
von: Shen, Junhong, et al.
Veröffentlicht: (2025)
Time-, Memory- and Parameter-Efficient Visual Adaptation
von: Mercea, Otniel-Bogdan, et al.
Veröffentlicht: (2024)
von: Mercea, Otniel-Bogdan, et al.
Veröffentlicht: (2024)
MetricNet: Recovering Metric Scale in Generative Navigation Policies
von: Nayak, Abhijeet, et al.
Veröffentlicht: (2025)
von: Nayak, Abhijeet, et al.
Veröffentlicht: (2025)
DataDream: Few-shot Guided Dataset Generation
von: Kim, Jae Myung, et al.
Veröffentlicht: (2024)
von: Kim, Jae Myung, et al.
Veröffentlicht: (2024)
Memory-Modular Classification: Learning to Generalize with Memory Replacement
von: Kang, Dahyun, et al.
Veröffentlicht: (2025)
von: Kang, Dahyun, et al.
Veröffentlicht: (2025)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
von: Bousselham, Walid, et al.
Veröffentlicht: (2025)
von: Bousselham, Walid, et al.
Veröffentlicht: (2025)
Learning text-to-video retrieval from image captioning
von: Ventura, Lucas, et al.
Veröffentlicht: (2024)
von: Ventura, Lucas, et al.
Veröffentlicht: (2024)
Online 3D Scene Reconstruction Using Neural Object Priors
von: Chabal, Thomas, et al.
Veröffentlicht: (2025)
von: Chabal, Thomas, et al.
Veröffentlicht: (2025)
FOM-Nav: Frontier-Object Maps for Object Goal Navigation
von: Chabal, Thomas, et al.
Veröffentlicht: (2025)
von: Chabal, Thomas, et al.
Veröffentlicht: (2025)
SEED4D: A Synthetic Ego--Exo Dynamic 4D Data Generator, Driving Dataset and Benchmark
von: Kästingschäfer, Marius, et al.
Veröffentlicht: (2024)
von: Kästingschäfer, Marius, et al.
Veröffentlicht: (2024)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
von: Chen, Shizhe, et al.
Veröffentlicht: (2026)
von: Chen, Shizhe, et al.
Veröffentlicht: (2026)
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
von: Garcia, Ricardo, et al.
Veröffentlicht: (2024)
von: Garcia, Ricardo, et al.
Veröffentlicht: (2024)
Nepotistically Trained Generative-AI Models Collapse
von: Bohacek, Matyas, et al.
Veröffentlicht: (2023)
von: Bohacek, Matyas, et al.
Veröffentlicht: (2023)
Simple LLM Baselines are Competitive for Model Diffing
von: Kempf, Elias, et al.
Veröffentlicht: (2026)
von: Kempf, Elias, et al.
Veröffentlicht: (2026)
Dense Optical Tracking: Connecting the Dots
von: Moing, Guillaume Le, et al.
Veröffentlicht: (2023)
von: Moing, Guillaume Le, et al.
Veröffentlicht: (2023)
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
von: Fiastre, Gabriel, et al.
Veröffentlicht: (2025)
von: Fiastre, Gabriel, et al.
Veröffentlicht: (2025)
HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching
von: Chen, Zerui, et al.
Veröffentlicht: (2026)
von: Chen, Zerui, et al.
Veröffentlicht: (2026)
Visual Lexicon: Rich Image Features in Language Space
von: Wang, XuDong, et al.
Veröffentlicht: (2024)
von: Wang, XuDong, et al.
Veröffentlicht: (2024)
Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy
von: Ging, Simon, et al.
Veröffentlicht: (2024)
von: Ging, Simon, et al.
Veröffentlicht: (2024)
Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
von: Ventura, Lucas, et al.
Veröffentlicht: (2025)
von: Ventura, Lucas, et al.
Veröffentlicht: (2025)
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
von: Wysoczańska, Monika, et al.
Veröffentlicht: (2025)
von: Wysoczańska, Monika, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
When and How Does CLIP Enable Domain and Compositional Generalization?
von: Kempf, Elias, et al.
Veröffentlicht: (2025) -
Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models
von: Schrodi, Simon, et al.
Veröffentlicht: (2024) -
Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems
von: Hoffmann, David T., et al.
Veröffentlicht: (2023) -
Concept Bottleneck Models Without Predefined Concepts
von: Schrodi, Simon, et al.
Veröffentlicht: (2024) -
Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models
von: Mousakhan, Arian, et al.
Veröffentlicht: (2025)