CoGen: Learning from Feedback with Coupled Comprehension and Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gul, Mustafa Omer, Artzi, Yoav |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Retrospective Learning from Interactions
von: Chen, Zizhao, et al.
Veröffentlicht: (2024)
von: Chen, Zizhao, et al.
Veröffentlicht: (2024)
Talk Less, Interact Better: Evaluating In-context Conversational Adaptation in Multimodal LLMs
von: Hua, Yilun, et al.
Veröffentlicht: (2024)
von: Hua, Yilun, et al.
Veröffentlicht: (2024)
Knot So Simple: A Minimalistic Environment for Spatial Reasoning
von: Chen, Zizhao, et al.
Veröffentlicht: (2025)
von: Chen, Zizhao, et al.
Veröffentlicht: (2025)
SurGen: Text-Guided Diffusion Model for Surgical Video Generation
von: Cho, Joseph, et al.
Veröffentlicht: (2024)
von: Cho, Joseph, et al.
Veröffentlicht: (2024)
R2Gen-Mamba: A Selective State Space Model for Radiology Report Generation
von: Sun, Yongheng, et al.
Veröffentlicht: (2024)
von: Sun, Yongheng, et al.
Veröffentlicht: (2024)
HistGen: Histopathology Report Generation via Local-Global Feature Encoding and Cross-modal Context Interaction
von: Guo, Zhengrui, et al.
Veröffentlicht: (2024)
von: Guo, Zhengrui, et al.
Veröffentlicht: (2024)
DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation
von: Li, Baiqi, et al.
Veröffentlicht: (2024)
von: Li, Baiqi, et al.
Veröffentlicht: (2024)
No Mean Feat: Simple, Strong Baselines for Context Compression
von: Feldman, Yair, et al.
Veröffentlicht: (2025)
von: Feldman, Yair, et al.
Veröffentlicht: (2025)
A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
von: Wang, Andrew Z., et al.
Veröffentlicht: (2025)
von: Wang, Andrew Z., et al.
Veröffentlicht: (2025)
Learning to Steer: Input-dependent Steering for Multimodal LLMs
von: Parekh, Jayneel, et al.
Veröffentlicht: (2025)
von: Parekh, Jayneel, et al.
Veröffentlicht: (2025)
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
von: Tong, Chengzhuo, et al.
Veröffentlicht: (2025)
von: Tong, Chengzhuo, et al.
Veröffentlicht: (2025)
Feedback-Driven Vision-Language Alignment with Minimal Human Supervision
von: Giannone, Giorgio, et al.
Veröffentlicht: (2025)
von: Giannone, Giorgio, et al.
Veröffentlicht: (2025)
Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback
von: Zhou, Yucheng, et al.
Veröffentlicht: (2025)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2025)
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
Vision Mamba: A Comprehensive Survey and Taxonomy
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
MedCLM: Learning to Localize and Reason via a CoT-Curriculum in Medical Vision-Language Models
von: Kim, Soo Yong, et al.
Veröffentlicht: (2025)
von: Kim, Soo Yong, et al.
Veröffentlicht: (2025)
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
von: Ryoo, Michael S., et al.
Veröffentlicht: (2024)
von: Ryoo, Michael S., et al.
Veröffentlicht: (2024)
Mamba in Vision: A Comprehensive Survey of Techniques and Applications
von: Rahman, Md Maklachur, et al.
Veröffentlicht: (2024)
von: Rahman, Md Maklachur, et al.
Veröffentlicht: (2024)
Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback
von: Xiao, Wenyi, et al.
Veröffentlicht: (2024)
von: Xiao, Wenyi, et al.
Veröffentlicht: (2024)
A Concept-Based Explainability Framework for Large Multimodal Models
von: Parekh, Jayneel, et al.
Veröffentlicht: (2024)
von: Parekh, Jayneel, et al.
Veröffentlicht: (2024)
INS-MMBench: A Comprehensive Benchmark for Evaluating LVLMs' Performance in Insurance
von: Lin, Chenwei, et al.
Veröffentlicht: (2024)
von: Lin, Chenwei, et al.
Veröffentlicht: (2024)
Generative AI for Character Animation: A Comprehensive Survey of Techniques, Applications, and Future Directions
von: Abootorabi, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
von: Abootorabi, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
Zero-Shot Refinement of Buildings' Segmentation Models using SAM
von: Mayladan, Ali, et al.
Veröffentlicht: (2023)
von: Mayladan, Ali, et al.
Veröffentlicht: (2023)
MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs
von: Khayatan, Pegah, et al.
Veröffentlicht: (2026)
von: Khayatan, Pegah, et al.
Veröffentlicht: (2026)
Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
GenURL: A General Framework for Unsupervised Representation Learning
von: Li, Siyuan, et al.
Veröffentlicht: (2021)
von: Li, Siyuan, et al.
Veröffentlicht: (2021)
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
von: Luo, Yaxin, et al.
Veröffentlicht: (2025)
von: Luo, Yaxin, et al.
Veröffentlicht: (2025)
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
von: Li, Jialu, et al.
Veröffentlicht: (2024)
von: Li, Jialu, et al.
Veröffentlicht: (2024)
A Surprising Failure? Multimodal LLMs and the NLVR Challenge
von: Wu, Anne, et al.
Veröffentlicht: (2024)
von: Wu, Anne, et al.
Veröffentlicht: (2024)
Post-training for Efficient Communication via Convention Formation
von: Hua, Yilun, et al.
Veröffentlicht: (2025)
von: Hua, Yilun, et al.
Veröffentlicht: (2025)
ECCV 2024 W-CODA: 1st Workshop on Multimodal Perception and Comprehension of Corner Cases in Autonomous Driving
von: Chen, Kai, et al.
Veröffentlicht: (2025)
von: Chen, Kai, et al.
Veröffentlicht: (2025)
ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
von: Liang, Yijun, et al.
Veröffentlicht: (2025)
von: Liang, Yijun, et al.
Veröffentlicht: (2025)
GenOL: Generating Diverse Examples for Name-only Online Learning
von: Seo, Minhyuk, et al.
Veröffentlicht: (2024)
von: Seo, Minhyuk, et al.
Veröffentlicht: (2024)
CausalChaos! Dataset for Comprehensive Causal Action Question Answering Over Longer Causal Chains Grounded in Dynamic Visual Scenes
von: Parmar, Paritosh, et al.
Veröffentlicht: (2024)
von: Parmar, Paritosh, et al.
Veröffentlicht: (2024)
Salsa as a Nonverbal Embodied Language -- The CoMPAS3D Dataset and Benchmarks
von: Burkanova, Bermet, et al.
Veröffentlicht: (2025)
von: Burkanova, Bermet, et al.
Veröffentlicht: (2025)
A Comparative Study of Machine Unlearning Techniques for Image and Text Classification Models
von: Safa, Omar M., et al.
Veröffentlicht: (2024)
von: Safa, Omar M., et al.
Veröffentlicht: (2024)
Analyzing the Roles of Language and Vision in Learning from Limited Data
von: Chen, Allison, et al.
Veröffentlicht: (2024)
von: Chen, Allison, et al.
Veröffentlicht: (2024)
MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective
von: Huang, Hailang, et al.
Veröffentlicht: (2024)
von: Huang, Hailang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Retrospective Learning from Interactions
von: Chen, Zizhao, et al.
Veröffentlicht: (2024) -
Talk Less, Interact Better: Evaluating In-context Conversational Adaptation in Multimodal LLMs
von: Hua, Yilun, et al.
Veröffentlicht: (2024) -
Knot So Simple: A Minimalistic Environment for Spatial Reasoning
von: Chen, Zizhao, et al.
Veröffentlicht: (2025) -
SurGen: Text-Guided Diffusion Model for Surgical Video Generation
von: Cho, Joseph, et al.
Veröffentlicht: (2024) -
R2Gen-Mamba: A Selective State Space Model for Radiology Report Generation
von: Sun, Yongheng, et al.
Veröffentlicht: (2024)