Exploiting Pseudo Image Captions for Multimodal Summarization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Chaoya, Xie, Rui, Ye, Wei, Sun, Jinan, Zhang, Shikun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TiMix: Text-aware Image Mixing for Effective Vision-Language Pre-training
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
Enhancing In-Context Learning via Implicit Demonstration Augmentation
von: Zhou, Xiaoling, et al.
Veröffentlicht: (2024)
von: Zhou, Xiaoling, et al.
Veröffentlicht: (2024)
Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework
von: Jia, Hongrui, et al.
Veröffentlicht: (2025)
von: Jia, Hongrui, et al.
Veröffentlicht: (2025)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
From Blind Spots to Gains: Diagnostic-Driven Iterative Training for Large Multimodal Models
von: Jia, Hongrui, et al.
Veröffentlicht: (2026)
von: Jia, Hongrui, et al.
Veröffentlicht: (2026)
PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
von: Wang, Yidong, et al.
Veröffentlicht: (2023)
von: Wang, Yidong, et al.
Veröffentlicht: (2023)
SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types
von: Mou, Yutao, et al.
Veröffentlicht: (2024)
von: Mou, Yutao, et al.
Veröffentlicht: (2024)
EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations
von: Heng, Yongrui, et al.
Veröffentlicht: (2026)
von: Heng, Yongrui, et al.
Veröffentlicht: (2026)
Data Selection for Multi-turn Dialogue Instruction Tuning
von: Li, Bo, et al.
Veröffentlicht: (2026)
von: Li, Bo, et al.
Veröffentlicht: (2026)
Video Summarization: Towards Entity-Aware Captions
von: Ayyubi, Hammad A., et al.
Veröffentlicht: (2023)
von: Ayyubi, Hammad A., et al.
Veröffentlicht: (2023)
SumHiS: Extractive Summarization Exploiting Hidden Structure
von: Pavel, Tikhonov, et al.
Veröffentlicht: (2024)
von: Pavel, Tikhonov, et al.
Veröffentlicht: (2024)
SteerRM: Debiasing Reward Models via Sparse Autoencoders
von: Sun, Mengyuan, et al.
Veröffentlicht: (2026)
von: Sun, Mengyuan, et al.
Veröffentlicht: (2026)
BUS:Efficient and Effective Vision-language Pre-training with Bottom-Up Patch Summarization
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
TransportationGames: Benchmarking Transportation Knowledge of (Multimodal) Large Language Models
von: Zhang, Xue, et al.
Veröffentlicht: (2024)
von: Zhang, Xue, et al.
Veröffentlicht: (2024)
SaRO: Enhancing LLM Safety through Reasoning-based Alignment
von: Mou, Yutao, et al.
Veröffentlicht: (2025)
von: Mou, Yutao, et al.
Veröffentlicht: (2025)
Instruction Data Selection via Answer Divergence
von: Li, Bo, et al.
Veröffentlicht: (2026)
von: Li, Bo, et al.
Veröffentlicht: (2026)
EAMA : Entity-Aware Multimodal Alignment Based Approach for News Image Captioning
von: Zhang, Junzhe, et al.
Veröffentlicht: (2024)
von: Zhang, Junzhe, et al.
Veröffentlicht: (2024)
BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models
von: Zhang, Ziheng, et al.
Veröffentlicht: (2025)
von: Zhang, Ziheng, et al.
Veröffentlicht: (2025)
Wolf: Dense Video Captioning with a World Summarization Framework
von: Li, Boyi, et al.
Veröffentlicht: (2024)
von: Li, Boyi, et al.
Veröffentlicht: (2024)
Can You Really Trust Code Copilots? Evaluating Large Language Models from a Code Security Perspective
von: Mou, Yutao, et al.
Veröffentlicht: (2025)
von: Mou, Yutao, et al.
Veröffentlicht: (2025)
Decoupling Safety into Orthogonal Subspace: Cost-Efficient and Performance-Preserving Alignment for Large Language Models
von: Mou, Yutao, et al.
Veröffentlicht: (2025)
von: Mou, Yutao, et al.
Veröffentlicht: (2025)
Fine-grained and Explainable Factuality Evaluation for Multimodal Summarization
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
Leveraging Entity Information for Cross-Modality Correlation Learning: The Entity-Guided Multimodal Summarization
von: Zhang, Yanghai, et al.
Veröffentlicht: (2024)
von: Zhang, Yanghai, et al.
Veröffentlicht: (2024)
Semantic and Expressive Variation in Image Captions Across Languages
von: Ye, Andre, et al.
Veröffentlicht: (2023)
von: Ye, Andre, et al.
Veröffentlicht: (2023)
Supportiveness-based Knowledge Rewriting for Retrieval-augmented Language Modeling
von: Qiao, Zile, et al.
Veröffentlicht: (2024)
von: Qiao, Zile, et al.
Veröffentlicht: (2024)
SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization
von: Jia, Hongrui, et al.
Veröffentlicht: (2024)
von: Jia, Hongrui, et al.
Veröffentlicht: (2024)
Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
von: Sarto, Sara, et al.
Veröffentlicht: (2025)
von: Sarto, Sara, et al.
Veröffentlicht: (2025)
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
von: Bai, Longju, et al.
Veröffentlicht: (2024)
von: Bai, Longju, et al.
Veröffentlicht: (2024)
Polos: Multimodal Metric Learning from Human Feedback for Image Captioning
von: Wada, Yuiga, et al.
Veröffentlicht: (2024)
von: Wada, Yuiga, et al.
Veröffentlicht: (2024)
Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
von: Fonseca, Rui, et al.
Veröffentlicht: (2025)
von: Fonseca, Rui, et al.
Veröffentlicht: (2025)
TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought
von: Jiang, Chaoya, et al.
Veröffentlicht: (2025)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2025)
Improving Fairness of Large Language Models in Multi-document Summarization
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
Retrieval as Generation: A Unified Framework with Self-Triggered Information Planning
von: Li, Bo, et al.
Veröffentlicht: (2026)
von: Li, Bo, et al.
Veröffentlicht: (2026)
Cobra Effect in Reference-Free Image Captioning Metrics
von: Ma, Zheng, et al.
Veröffentlicht: (2024)
von: Ma, Zheng, et al.
Veröffentlicht: (2024)
Modeling Uncertainty Trends for Timely Retrieval in Dynamic RAG
von: Li, Bo, et al.
Veröffentlicht: (2025)
von: Li, Bo, et al.
Veröffentlicht: (2025)
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis
von: Bucciarelli, Davide, et al.
Veröffentlicht: (2024)
von: Bucciarelli, Davide, et al.
Veröffentlicht: (2024)
A Modular Approach for Multimodal Summarization of TV Shows
von: Mahon, Louis, et al.
Veröffentlicht: (2024)
von: Mahon, Louis, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TiMix: Text-aware Image Mixing for Effective Vision-Language Pre-training
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023) -
Enhancing In-Context Learning via Implicit Demonstration Augmentation
von: Zhou, Xiaoling, et al.
Veröffentlicht: (2024) -
Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework
von: Jia, Hongrui, et al.
Veröffentlicht: (2025) -
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024) -
Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)