Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Song, Yingjin, Paperno, Denis, Gatt, Albert |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Image Captioning to Visual Storytelling
di: Passadakis, Admitos, et al.
Pubblicazione: (2025)
di: Passadakis, Admitos, et al.
Pubblicazione: (2025)
Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?
di: Song, Yingjin, et al.
Pubblicazione: (2025)
di: Song, Yingjin, et al.
Pubblicazione: (2025)
Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication
di: Merlo, Filippo, et al.
Pubblicazione: (2025)
di: Merlo, Filippo, et al.
Pubblicazione: (2025)
Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding
di: Ignatev, Daniil, et al.
Pubblicazione: (2025)
di: Ignatev, Daniil, et al.
Pubblicazione: (2025)
Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization
di: Deng, Haolin, et al.
Pubblicazione: (2026)
di: Deng, Haolin, et al.
Pubblicazione: (2026)
Relevance-aware Multi-context Contrastive Decoding for Retrieval-augmented Visual Question Answering
di: Kim, Jongha, et al.
Pubblicazione: (2026)
di: Kim, Jongha, et al.
Pubblicazione: (2026)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
di: Du, Yifan, et al.
Pubblicazione: (2023)
di: Du, Yifan, et al.
Pubblicazione: (2023)
Visual In-Context Learning for Large Vision-Language Models
di: Zhou, Yucheng, et al.
Pubblicazione: (2024)
di: Zhou, Yucheng, et al.
Pubblicazione: (2024)
VAGUE: Visual Contexts Clarify Ambiguous Expressions
di: Nam, Heejeong, et al.
Pubblicazione: (2024)
di: Nam, Heejeong, et al.
Pubblicazione: (2024)
VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
di: Gado, Mohamed, et al.
Pubblicazione: (2025)
di: Gado, Mohamed, et al.
Pubblicazione: (2025)
Retrieving Counterfactuals Improves Visual In-Context Learning
di: Xiong, Guangzhi, et al.
Pubblicazione: (2026)
di: Xiong, Guangzhi, et al.
Pubblicazione: (2026)
Generating Storytelling Images with Rich Chains-of-Reasoning
di: Song, Xiujie, et al.
Pubblicazione: (2025)
di: Song, Xiujie, et al.
Pubblicazione: (2025)
Customized Visual Storytelling with Unified Multimodal LLMs
di: Li, Wei-Hua, et al.
Pubblicazione: (2026)
di: Li, Wei-Hua, et al.
Pubblicazione: (2026)
LongVILA: Scaling Long-Context Visual Language Models for Long Videos
di: Chen, Yukang, et al.
Pubblicazione: (2024)
di: Chen, Yukang, et al.
Pubblicazione: (2024)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
di: Song, Wei, et al.
Pubblicazione: (2025)
di: Song, Wei, et al.
Pubblicazione: (2025)
Exploring the Design Space of Visual Context Representation in Video MLLMs
di: Du, Yifan, et al.
Pubblicazione: (2024)
di: Du, Yifan, et al.
Pubblicazione: (2024)
CommVQA: Situating Visual Question Answering in Communicative Contexts
di: Naik, Nandita Shankar, et al.
Pubblicazione: (2024)
di: Naik, Nandita Shankar, et al.
Pubblicazione: (2024)
Learning To See But Forgetting To Follow: Visual Instruction Tuning Makes LLMs More Prone To Jailbreak Attacks
di: Pantazopoulos, Georgios, et al.
Pubblicazione: (2024)
di: Pantazopoulos, Georgios, et al.
Pubblicazione: (2024)
Less is More: High-value Data Selection for Visual Instruction Tuning
di: Liu, Zikang, et al.
Pubblicazione: (2024)
di: Liu, Zikang, et al.
Pubblicazione: (2024)
Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning
di: Xu, Zhiyang, et al.
Pubblicazione: (2024)
di: Xu, Zhiyang, et al.
Pubblicazione: (2024)
Refining Skewed Perceptions in Vision-Language Contrastive Models through Visual Representations
di: Dai, Haocheng, et al.
Pubblicazione: (2024)
di: Dai, Haocheng, et al.
Pubblicazione: (2024)
Enhancing Micro Gesture Recognition for Emotion Understanding via Context-aware Visual-Text Contrastive Learning
di: Li, Deng, et al.
Pubblicazione: (2024)
di: Li, Deng, et al.
Pubblicazione: (2024)
Learning to Instruct for Visual Instruction Tuning
di: Zhou, Zhihan, et al.
Pubblicazione: (2025)
di: Zhou, Zhihan, et al.
Pubblicazione: (2025)
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
di: Zhang, Yanzhe, et al.
Pubblicazione: (2023)
di: Zhang, Yanzhe, et al.
Pubblicazione: (2023)
Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning
di: Wei, Xinyu, et al.
Pubblicazione: (2025)
di: Wei, Xinyu, et al.
Pubblicazione: (2025)
EmoGist: Efficient In-Context Learning for Visual Emotion Understanding
di: Seoh, Ronald, et al.
Pubblicazione: (2025)
di: Seoh, Ronald, et al.
Pubblicazione: (2025)
Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
di: Zhou, Yucheng, et al.
Pubblicazione: (2024)
di: Zhou, Yucheng, et al.
Pubblicazione: (2024)
CODIS: Benchmarking Context-Dependent Visual Comprehension for Multimodal Large Language Models
di: Luo, Fuwen, et al.
Pubblicazione: (2024)
di: Luo, Fuwen, et al.
Pubblicazione: (2024)
VERA: Identifying and Leveraging Visual Evidence Retrieval Heads in Long-Context Understanding
di: Pei, Rongcan, et al.
Pubblicazione: (2026)
di: Pei, Rongcan, et al.
Pubblicazione: (2026)
CLEAR: Character Unlearning in Textual and Visual Modalities
di: Dontsov, Alexey, et al.
Pubblicazione: (2024)
di: Dontsov, Alexey, et al.
Pubblicazione: (2024)
Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment
di: Xiao, Xin, et al.
Pubblicazione: (2024)
di: Xiao, Xin, et al.
Pubblicazione: (2024)
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
di: Guo, Pinxue, et al.
Pubblicazione: (2025)
di: Guo, Pinxue, et al.
Pubblicazione: (2025)
Little Data, Big Impact: Privacy-Aware Visual Language Models via Minimal Tuning
di: Samson, Laurens, et al.
Pubblicazione: (2024)
di: Samson, Laurens, et al.
Pubblicazione: (2024)
Visual Reasoning at Urban Intersections: FineTuning GPT-4o for Traffic Conflict Detection
di: Masri, Sari, et al.
Pubblicazione: (2025)
di: Masri, Sari, et al.
Pubblicazione: (2025)
Openstory++: A Large-scale Dataset and Benchmark for Instance-aware Open-domain Visual Storytelling
di: Ye, Zilyu, et al.
Pubblicazione: (2024)
di: Ye, Zilyu, et al.
Pubblicazione: (2024)
Is Your Image a Good Storyteller?
di: Song, Xiujie, et al.
Pubblicazione: (2024)
di: Song, Xiujie, et al.
Pubblicazione: (2024)
Learning Speaker-Invariant Visual Features for Lipreading
di: Li, Yu, et al.
Pubblicazione: (2025)
di: Li, Yu, et al.
Pubblicazione: (2025)
Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
di: Bai, Tianyi, et al.
Pubblicazione: (2025)
di: Bai, Tianyi, et al.
Pubblicazione: (2025)
Ask Questions with Double Hints: Visual Question Generation with Answer-awareness and Region-reference
di: Shen, Kai, et al.
Pubblicazione: (2024)
di: Shen, Kai, et al.
Pubblicazione: (2024)
GEM: Context-Aware Gaze EstiMation with Visual Search Behavior Matching for Chest Radiograph
di: Liu, Shaonan, et al.
Pubblicazione: (2024)
di: Liu, Shaonan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
From Image Captioning to Visual Storytelling
di: Passadakis, Admitos, et al.
Pubblicazione: (2025) -
Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?
di: Song, Yingjin, et al.
Pubblicazione: (2025) -
Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication
di: Merlo, Filippo, et al.
Pubblicazione: (2025) -
Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding
di: Ignatev, Daniil, et al.
Pubblicazione: (2025) -
Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization
di: Deng, Haolin, et al.
Pubblicazione: (2026)