CompCap: Improving Multimodal Large Language Models with Composite Captions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Xiaohui, Shukla, Satya Narayan, Azab, Mahmoud, Singh, Aashu, Wang, Qifan, Yang, David, Peng, ShengYun, Yu, Hanchao, Yan, Shen, Zhang, Xuewen, He, Baosheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Inference Compute-Optimal Video Vision Language Models
von: Wang, Peiqi, et al.
Veröffentlicht: (2025)
von: Wang, Peiqi, et al.
Veröffentlicht: (2025)
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
von: Yang, Yanlai, et al.
Veröffentlicht: (2025)
von: Yang, Yanlai, et al.
Veröffentlicht: (2025)
Think Then Embed: Generative Context Improves Multimodal Embedding
von: Cui, Xuanming, et al.
Veröffentlicht: (2025)
von: Cui, Xuanming, et al.
Veröffentlicht: (2025)
Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models
von: Peng, ShengYun, et al.
Veröffentlicht: (2024)
von: Peng, ShengYun, et al.
Veröffentlicht: (2024)
Heteroscedastic Temporal Variational Autoencoder For Irregular Time Series
von: Shukla, Satya Narayan, et al.
Veröffentlicht: (2021)
von: Shukla, Satya Narayan, et al.
Veröffentlicht: (2021)
Optimizing Recall or Relevance? A Multi-Task Multi-Head Approach for Item-to-Item Retrieval in Recommendation
von: Zhang, Jiang, et al.
Veröffentlicht: (2025)
von: Zhang, Jiang, et al.
Veröffentlicht: (2025)
Cultural and Historical Identity in Amitav Ghosh’s River of Smoke: A Postcolonial Perspective
von: Satya Narayan
Veröffentlicht: (2021)
von: Satya Narayan
Veröffentlicht: (2021)
Depicting Culture and Identity in Amitav Ghosh’s The Shadow Lines
von: Satya Narayan
Veröffentlicht: (2017)
von: Satya Narayan
Veröffentlicht: (2017)
Cultural and Historical Identity in Amitav Ghosh’s River of Smoke: A Postcolonial Perspective
von: Satya Narayan
Veröffentlicht: (2021)
von: Satya Narayan
Veröffentlicht: (2021)
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
von: Lee, Seongmin, et al.
Veröffentlicht: (2025)
von: Lee, Seongmin, et al.
Veröffentlicht: (2025)
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
Verifiable Reasoning for LLM-based Generative Recommendation
von: Lin, Xinyu, et al.
Veröffentlicht: (2026)
von: Lin, Xinyu, et al.
Veröffentlicht: (2026)
COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation
von: Deng, Xueqing, et al.
Veröffentlicht: (2025)
von: Deng, Xueqing, et al.
Veröffentlicht: (2025)
Self-Supervised Pre-Training for Table Structure Recognition Transformer
von: Peng, ShengYun, et al.
Veröffentlicht: (2024)
von: Peng, ShengYun, et al.
Veröffentlicht: (2024)
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles
von: Ng, Ho Yin 'Sam', et al.
Veröffentlicht: (2025)
von: Ng, Ho Yin 'Sam', et al.
Veröffentlicht: (2025)
Shape it Up! Restoring LLM Safety during Finetuning
von: Peng, ShengYun, et al.
Veröffentlicht: (2025)
von: Peng, ShengYun, et al.
Veröffentlicht: (2025)
Transfer between Modalities with MetaQueries
von: Pan, Xichen, et al.
Veröffentlicht: (2025)
von: Pan, Xichen, et al.
Veröffentlicht: (2025)
FingerCap: Fine-grained Finger-level Hand Motion Captioning
von: Shen, Xin, et al.
Veröffentlicht: (2025)
von: Shen, Xin, et al.
Veröffentlicht: (2025)
CapHDR2IR: Caption-Driven Transfer from Visible Light to Infrared Domain
von: Peng, Jingchao, et al.
Veröffentlicht: (2024)
von: Peng, Jingchao, et al.
Veröffentlicht: (2024)
Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation
von: Li, Chao, et al.
Veröffentlicht: (2026)
von: Li, Chao, et al.
Veröffentlicht: (2026)
ControlCap: Controllable Region-level Captioning
von: Zhao, Yuzhong, et al.
Veröffentlicht: (2024)
von: Zhao, Yuzhong, et al.
Veröffentlicht: (2024)
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
von: Fan, Tiehan, et al.
Veröffentlicht: (2024)
von: Fan, Tiehan, et al.
Veröffentlicht: (2024)
UniTable: Towards a Unified Framework for Table Recognition via Self-Supervised Pretraining
von: Peng, ShengYun, et al.
Veröffentlicht: (2024)
von: Peng, ShengYun, et al.
Veröffentlicht: (2024)
FigCaps-HF: A Figure-to-Caption Generative Framework and Benchmark with Human Feedback
von: Singh, Ashish, et al.
Veröffentlicht: (2023)
von: Singh, Ashish, et al.
Veröffentlicht: (2023)
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023
von: Hsu, Ting-Yao E., et al.
Veröffentlicht: (2025)
von: Hsu, Ting-Yao E., et al.
Veröffentlicht: (2025)
HyperCap: Hyperspectral Land Cover Captioning Dataset for Vision Language Models
von: Das, Aryan, et al.
Veröffentlicht: (2025)
von: Das, Aryan, et al.
Veröffentlicht: (2025)
A first-principles investigation of the diffusivities of oxygen and oxygen defects in ThO$_2$
von: Singh, Maniesha, et al.
Veröffentlicht: (2025)
von: Singh, Maniesha, et al.
Veröffentlicht: (2025)
SnapCap: Efficient Snapshot Compressive Video Captioning
von: Sun, Jianqiao, et al.
Veröffentlicht: (2024)
von: Sun, Jianqiao, et al.
Veröffentlicht: (2024)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning
von: Tosato, Lucrezia, et al.
Veröffentlicht: (2026)
von: Tosato, Lucrezia, et al.
Veröffentlicht: (2026)
PaveCap: The First Multimodal Framework for Comprehensive Pavement Condition Assessment with Dense Captioning and PCI Estimation
von: Kyem, Blessing Agyei, et al.
Veröffentlicht: (2024)
von: Kyem, Blessing Agyei, et al.
Veröffentlicht: (2024)
CycleCap: Improving VLMs Captioning Performance via Self-Supervised Cycle Consistency Fine-Tuning
von: Krestenitis, Marios, et al.
Veröffentlicht: (2026)
von: Krestenitis, Marios, et al.
Veröffentlicht: (2026)
LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
von: Phute, Mansi, et al.
Veröffentlicht: (2023)
von: Phute, Mansi, et al.
Veröffentlicht: (2023)
GroundCap: A Visually Grounded Image Captioning Dataset
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2025)
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2025)
IF-VidCap: Can Video Caption Models Follow Instructions?
von: Li, Shihao, et al.
Veröffentlicht: (2025)
von: Li, Shihao, et al.
Veröffentlicht: (2025)
Cap2Sum: Learning to Summarize Videos by Generating Captions
von: Zhao, Cairong, et al.
Veröffentlicht: (2024)
von: Zhao, Cairong, et al.
Veröffentlicht: (2024)
AlignCap: Aligning Speech Emotion Captioning to Human Preferences
von: Liang, Ziqi, et al.
Veröffentlicht: (2024)
von: Liang, Ziqi, et al.
Veröffentlicht: (2024)
XMeCap: Meme Caption Generation with Sub-Image Adaptability
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
von: Li, Yuying, et al.
Veröffentlicht: (2025)
von: Li, Yuying, et al.
Veröffentlicht: (2025)
ProCap: Projection-Aware Captioning for Spatial Augmented Reality
von: Cao, Zimo, et al.
Veröffentlicht: (2026)
von: Cao, Zimo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Inference Compute-Optimal Video Vision Language Models
von: Wang, Peiqi, et al.
Veröffentlicht: (2025) -
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
von: Yang, Yanlai, et al.
Veröffentlicht: (2025) -
Think Then Embed: Generative Context Improves Multimodal Embedding
von: Cui, Xuanming, et al.
Veröffentlicht: (2025) -
Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models
von: Peng, ShengYun, et al.
Veröffentlicht: (2024) -
Heteroscedastic Temporal Variational Autoencoder For Irregular Time Series
von: Shukla, Satya Narayan, et al.
Veröffentlicht: (2021)