Is Your Image a Good Storyteller?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Xiujie, Pang, Xiaoyi, Tang, Haifeng, Wu, Mengyue, Zhu, Kenny Q. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generating Storytelling Images with Rich Chains-of-Reasoning
von: Song, Xiujie, et al.
Veröffentlicht: (2025)
von: Song, Xiujie, et al.
Veröffentlicht: (2025)
A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models
von: Song, Xiujie, et al.
Veröffentlicht: (2024)
von: Song, Xiujie, et al.
Veröffentlicht: (2024)
CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?
von: Zhang, Leixin, et al.
Veröffentlicht: (2024)
von: Zhang, Leixin, et al.
Veröffentlicht: (2024)
Are Large Vision Language Models Good Game Players?
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
von: Xing, Long, et al.
Veröffentlicht: (2025)
von: Xing, Long, et al.
Veröffentlicht: (2025)
Draw Your Mind: Personalized Generation via Condition-Level Modeling in Text-to-Image Diffusion Models
von: Kim, Hyungjin, et al.
Veröffentlicht: (2025)
von: Kim, Hyungjin, et al.
Veröffentlicht: (2025)
VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
von: Gado, Mohamed, et al.
Veröffentlicht: (2025)
von: Gado, Mohamed, et al.
Veröffentlicht: (2025)
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2024)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2024)
Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models
von: Xu, Shicheng, et al.
Veröffentlicht: (2024)
von: Xu, Shicheng, et al.
Veröffentlicht: (2024)
From Image Captioning to Visual Storytelling
von: Passadakis, Admitos, et al.
Veröffentlicht: (2025)
von: Passadakis, Admitos, et al.
Veröffentlicht: (2025)
How to Train Your Long-Context Visual Document Model
von: Veselka, Austin
Veröffentlicht: (2026)
von: Veselka, Austin
Veröffentlicht: (2026)
ESG Accountability Made Easy: DocQA at Your Service
von: Mishra, Lokesh, et al.
Veröffentlicht: (2023)
von: Mishra, Lokesh, et al.
Veröffentlicht: (2023)
When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification
von: Pang, Zirui, et al.
Veröffentlicht: (2025)
von: Pang, Zirui, et al.
Veröffentlicht: (2025)
Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data
von: Zhang, Chenhui, et al.
Veröffentlicht: (2024)
von: Zhang, Chenhui, et al.
Veröffentlicht: (2024)
Chatting with Images for Introspective Visual Thinking
von: Wu, Junfei, et al.
Veröffentlicht: (2026)
von: Wu, Junfei, et al.
Veröffentlicht: (2026)
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
von: Ryan, Yuriel, et al.
Veröffentlicht: (2025)
von: Ryan, Yuriel, et al.
Veröffentlicht: (2025)
StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant
von: Wang, Haibo, et al.
Veröffentlicht: (2025)
von: Wang, Haibo, et al.
Veröffentlicht: (2025)
Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
RestoreAgent: Autonomous Image Restoration Agent via Multimodal Large Language Models
von: Chen, Haoyu, et al.
Veröffentlicht: (2024)
von: Chen, Haoyu, et al.
Veröffentlicht: (2024)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
von: Zhang, Xiaoyi, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoyi, et al.
Veröffentlicht: (2025)
InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training
von: Zhang, Ziyun, et al.
Veröffentlicht: (2026)
von: Zhang, Ziyun, et al.
Veröffentlicht: (2026)
BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models
von: Arnould, Ludovic, et al.
Veröffentlicht: (2025)
von: Arnould, Ludovic, et al.
Veröffentlicht: (2025)
Can Large Vision-Language Models Detect Images Copyright Infringement from GenAI?
von: Xu, Qipan, et al.
Veröffentlicht: (2025)
von: Xu, Qipan, et al.
Veröffentlicht: (2025)
AlignGPT: Multi-modal Large Language Models with Adaptive Alignment Capability
von: Zhao, Fei, et al.
Veröffentlicht: (2024)
von: Zhao, Fei, et al.
Veröffentlicht: (2024)
Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2024)
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2024)
E$^2$AT: Multimodal Jailbreak Defense via Dynamic Joint Optimization for Multimodal Large Language Models
von: Lu, Liming, et al.
Veröffentlicht: (2025)
von: Lu, Liming, et al.
Veröffentlicht: (2025)
Enhancing Visual Dialog State Tracking through Iterative Object-Entity Alignment in Multi-Round Conversations
von: Pang, Wei, et al.
Veröffentlicht: (2024)
von: Pang, Wei, et al.
Veröffentlicht: (2024)
Mask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal Retrieval
von: Shen, Li-Cheng, et al.
Veröffentlicht: (2025)
von: Shen, Li-Cheng, et al.
Veröffentlicht: (2025)
Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
von: Ye, Junyan, et al.
Veröffentlicht: (2025)
von: Ye, Junyan, et al.
Veröffentlicht: (2025)
Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment
von: Xiao, Xin, et al.
Veröffentlicht: (2024)
von: Xiao, Xin, et al.
Veröffentlicht: (2024)
ORAL: Prompting Your Large-Scale LoRAs via Conditional Recurrent Diffusion
von: Khan, Rana Muhammad Shahroz, et al.
Veröffentlicht: (2025)
von: Khan, Rana Muhammad Shahroz, et al.
Veröffentlicht: (2025)
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
WsiCaption: Multiple Instance Generation of Pathology Reports for Gigapixel Whole-Slide Images
von: Chen, Pingyi, et al.
Veröffentlicht: (2023)
von: Chen, Pingyi, et al.
Veröffentlicht: (2023)
GEM: Empowering MLLM for Grounded ECG Understanding with Time Series and Images
von: Lan, Xiang, et al.
Veröffentlicht: (2025)
von: Lan, Xiang, et al.
Veröffentlicht: (2025)
KARL: Knowledge-Aware Reasoning and Reinforcement Learning for Knowledge-Intensive Visual Grounding
von: Ma, Xinyu, et al.
Veröffentlicht: (2025)
von: Ma, Xinyu, et al.
Veröffentlicht: (2025)
ReasonMap: Towards Fine-Grained Visual Reasoning from Transit Maps
von: Feng, Sicheng, et al.
Veröffentlicht: (2025)
von: Feng, Sicheng, et al.
Veröffentlicht: (2025)
MULTI: Multimodal Understanding Leaderboard with Text and Images
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
FOLIAGE: Towards Physical Intelligence World Models Via Unbounded Surface Evolution
von: Liu, Xiaoyi, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoyi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Generating Storytelling Images with Rich Chains-of-Reasoning
von: Song, Xiujie, et al.
Veröffentlicht: (2025) -
A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models
von: Song, Xiujie, et al.
Veröffentlicht: (2024) -
CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning
von: Li, Ming, et al.
Veröffentlicht: (2025) -
ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?
von: Zhang, Leixin, et al.
Veröffentlicht: (2024) -
Are Large Vision Language Models Good Game Players?
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)