Weak-to-Strong Compositional Learning from Generative Models for Language-based Object Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Kwanyong, Saito, Kuniaki, Kim, Donghyun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2025)
by: Saito, Kuniaki, et al.
Published: (2025)
Talk in Pieces, See in Whole: Disentangling and Hierarchical Aggregating Representations for Language-based Object Detection
by: An, Sojung, et al.
Published: (2025)
by: An, Sojung, et al.
Published: (2025)
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models
by: Yi, Hao, et al.
Published: (2024)
by: Yi, Hao, et al.
Published: (2024)
Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP
by: Abbasi, Reza, et al.
Published: (2024)
by: Abbasi, Reza, et al.
Published: (2024)
LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation
by: Kim, Kibum, et al.
Published: (2023)
by: Kim, Kibum, et al.
Published: (2023)
Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
by: Park, Jihwan, et al.
Published: (2025)
by: Park, Jihwan, et al.
Published: (2025)
Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models
by: Robicheaux, Peter, et al.
Published: (2025)
by: Robicheaux, Peter, et al.
Published: (2025)
Language-Pretraining-Induced Bias: A Strong Foundation for General Vision Tasks
by: Luo, Yaxin, et al.
Published: (2026)
by: Luo, Yaxin, et al.
Published: (2026)
Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
by: Park, Dongmin, et al.
Published: (2024)
by: Park, Dongmin, et al.
Published: (2024)
Realistic Evaluation of Model Merging for Compositional Generalization
by: Tam, Derek, et al.
Published: (2024)
by: Tam, Derek, et al.
Published: (2024)
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
by: Zhou, Yiyang, et al.
Published: (2023)
by: Zhou, Yiyang, et al.
Published: (2023)
TroL: Traversal of Layers for Large Language and Vision Models
by: Lee, Byung-Kwan, et al.
Published: (2024)
by: Lee, Byung-Kwan, et al.
Published: (2024)
ERM++: An Improved Baseline for Domain Generalization
by: Teterwak, Piotr, et al.
Published: (2023)
by: Teterwak, Piotr, et al.
Published: (2023)
Is Large-Scale Pretraining the Secret to Good Domain Generalization?
by: Teterwak, Piotr, et al.
Published: (2024)
by: Teterwak, Piotr, et al.
Published: (2024)
Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference
by: Miranda, Imanol, et al.
Published: (2026)
by: Miranda, Imanol, et al.
Published: (2026)
Semantic Prompt Learning for Weakly-Supervised Semantic Segmentation
by: Lin, Ci-Siang, et al.
Published: (2024)
by: Lin, Ci-Siang, et al.
Published: (2024)
RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction
by: Park, Jonggwon, et al.
Published: (2025)
by: Park, Jonggwon, et al.
Published: (2025)
RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Chest X-ray with Zero-Shot Multi-Task Capability
by: Park, Jonggwon, et al.
Published: (2025)
by: Park, Jonggwon, et al.
Published: (2025)
TLDR: Token-Level Detective Reward Model for Large Vision Language Models
by: Fu, Deqing, et al.
Published: (2024)
by: Fu, Deqing, et al.
Published: (2024)
Verbalized Machine Learning: Revisiting Machine Learning with Language Models
by: Xiao, Tim Z., et al.
Published: (2024)
by: Xiao, Tim Z., et al.
Published: (2024)
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
by: Motamed, Saman, et al.
Published: (2023)
by: Motamed, Saman, et al.
Published: (2023)
Adding simple structure at inference improves Vision-Language Compositionality
by: Miranda, Imanol, et al.
Published: (2025)
by: Miranda, Imanol, et al.
Published: (2025)
Enhancing Clinical Efficiency through LLM: Discharge Note Generation for Cardiac Patients
by: Jung, HyoJe, et al.
Published: (2024)
by: Jung, HyoJe, et al.
Published: (2024)
Federated Learning from Vision-Language Foundation Models: Theoretical Analysis and Method
by: Pan, Bikang, et al.
Published: (2024)
by: Pan, Bikang, et al.
Published: (2024)
Residual-based Language Models are Free Boosters for Biomedical Imaging
by: Lai, Zhixin, et al.
Published: (2024)
by: Lai, Zhixin, et al.
Published: (2024)
Improving Multimodal Large Language Models Using Continual Learning
by: Srivastava, Shikhar, et al.
Published: (2024)
by: Srivastava, Shikhar, et al.
Published: (2024)
LatentExplainer: Explaining Latent Representations in Deep Generative Models with Multimodal Large Language Models
by: Zhu, Mengdan, et al.
Published: (2024)
by: Zhu, Mengdan, et al.
Published: (2024)
BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval
by: Miranda, Imanol, et al.
Published: (2024)
by: Miranda, Imanol, et al.
Published: (2024)
Weighted Multi-Prompt Learning with Description-free Large Language Model Distillation
by: Lee, Sua, et al.
Published: (2025)
by: Lee, Sua, et al.
Published: (2025)
Exploring the Limits of Zero Shot Vision Language Models for Hate Meme Detection: The Vulnerabilities and their Interpretations
by: Rizwan, Naquee, et al.
Published: (2024)
by: Rizwan, Naquee, et al.
Published: (2024)
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models
by: Dumpala, Sri Harsha, et al.
Published: (2024)
by: Dumpala, Sri Harsha, et al.
Published: (2024)
Learning Self-Correction in Vision-Language Models via Rollout Augmentation
by: Ding, Yi, et al.
Published: (2026)
by: Ding, Yi, et al.
Published: (2026)
Weakly Semi-supervised Tool Detection in Minimally Invasive Surgery Videos
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
Efficient and Versatile Robust Fine-Tuning of Zero-shot Models
by: Kim, Sungyeon, et al.
Published: (2024)
by: Kim, Sungyeon, et al.
Published: (2024)
Mixture of Experts Made Personalized: Federated Prompt Learning for Vision-Language Models
by: Luo, Jun, et al.
Published: (2024)
by: Luo, Jun, et al.
Published: (2024)
Troika: Multi-Path Cross-Modal Traction for Compositional Zero-Shot Learning
by: Huang, Siteng, et al.
Published: (2023)
by: Huang, Siteng, et al.
Published: (2023)
MoTe: Learning Motion-Text Diffusion Model for Multiple Generation Tasks
by: Wu, Yiming, et al.
Published: (2024)
by: Wu, Yiming, et al.
Published: (2024)
Are We Truly Forgetting? A Critical Re-examination of Machine Unlearning Evaluation Protocols
by: Kim, Yongwoo, et al.
Published: (2025)
by: Kim, Yongwoo, et al.
Published: (2025)
Erase at the Core: Representation Unlearning for Machine Unlearning
by: Lee, Jaewon, et al.
Published: (2026)
by: Lee, Jaewon, et al.
Published: (2026)
Weak-to-Strong Diffusion with Reflection
by: Bai, Lichen, et al.
Published: (2025)
by: Bai, Lichen, et al.
Published: (2025)
Similar Items
-
CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2025) -
Talk in Pieces, See in Whole: Disentangling and Hierarchical Aggregating Representations for Language-based Object Detection
by: An, Sojung, et al.
Published: (2025) -
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models
by: Yi, Hao, et al.
Published: (2024) -
Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP
by: Abbasi, Reza, et al.
Published: (2024) -
LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation
by: Kim, Kibum, et al.
Published: (2023)