MAMS: Model-Agnostic Module Selection Framework for Video Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Sangho, Chun, Il Yong, Park, Hogun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SToRM: Supervised Token Reduction for Multi-modal LLMs toward efficient end-to-end autonomous driving
by: Kim, Seo Hyun, et al.
Published: (2026)
by: Kim, Seo Hyun, et al.
Published: (2026)
VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis
by: Park, Jinho, et al.
Published: (2026)
by: Park, Jinho, et al.
Published: (2026)
JUDO: A Juxtaposed Domain-Oriented Multimodal Reasoner for Industrial Anomaly QA
by: Kang, Hyunju, et al.
Published: (2026)
by: Kang, Hyunju, et al.
Published: (2026)
QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models
by: Wang, Jiale, et al.
Published: (2026)
by: Wang, Jiale, et al.
Published: (2026)
Versatile Incremental Learning: Towards Class and Domain-Agnostic Incremental Learning
by: Park, Min-Yeong, et al.
Published: (2024)
by: Park, Min-Yeong, et al.
Published: (2024)
Leveraging Textual Compositional Reasoning for Robust Change Captioning
by: Park, Kyu Ri, et al.
Published: (2025)
by: Park, Kyu Ri, et al.
Published: (2025)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
Accurate and Fast Compressed Video Captioning
by: Shen, Yaojie, et al.
Published: (2023)
by: Shen, Yaojie, et al.
Published: (2023)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
Towards Multimodal Video Paragraph Captioning Models Robust to Missing Modality
by: Chen, Sishuo, et al.
Published: (2024)
by: Chen, Sishuo, et al.
Published: (2024)
FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model
by: Lee, Yebin, et al.
Published: (2024)
by: Lee, Yebin, et al.
Published: (2024)
CaptionFool: Universal Image Captioning Model Attacks
by: Parekh, Swapnil
Published: (2026)
by: Parekh, Swapnil
Published: (2026)
Task-Agnostic Noisy Label Detection via Standardized Loss Aggregation
by: Park, Inhyuk, et al.
Published: (2026)
by: Park, Inhyuk, et al.
Published: (2026)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
by: Jung, Mingi, et al.
Published: (2025)
by: Jung, Mingi, et al.
Published: (2025)
GEASS: Gated Evidence-Adaptive Selective Caption Trust for Vision-Language Models
by: Li, Zeshang, et al.
Published: (2026)
by: Li, Zeshang, et al.
Published: (2026)
Describe Anything: Detailed Localized Image and Video Captioning
by: Lian, Long, et al.
Published: (2025)
by: Lian, Long, et al.
Published: (2025)
Towards Fine-Grained Human Motion Video Captioning
by: Song, Guorui, et al.
Published: (2025)
by: Song, Guorui, et al.
Published: (2025)
Beyond Pedestrians: Caption-Guided CLIP Framework for High-Difficulty Video-based Person Re-Identification
by: Hamano, Shogo, et al.
Published: (2026)
by: Hamano, Shogo, et al.
Published: (2026)
SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning
by: Kim, Ye-Chan, et al.
Published: (2026)
by: Kim, Ye-Chan, et al.
Published: (2026)
Unveiling the Invisible: Captioning Videos with Metaphors
by: Kalarani, Abisek Rajakumar, et al.
Published: (2024)
by: Kalarani, Abisek Rajakumar, et al.
Published: (2024)
Dense Video Captioning using Graph-based Sentence Summarization
by: Zhang, Zhiwang, et al.
Published: (2025)
by: Zhang, Zhiwang, et al.
Published: (2025)
Controllable Hybrid Captioner for Improved Long-form Video Understanding
by: Sasse, Kuleen, et al.
Published: (2025)
by: Sasse, Kuleen, et al.
Published: (2025)
LaB-CL: Localized and Balanced Contrastive Learning for improving parking slot detection
by: Jeong, U Jin, et al.
Published: (2024)
by: Jeong, U Jin, et al.
Published: (2024)
Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization
by: Park, Jungin, et al.
Published: (2025)
by: Park, Jungin, et al.
Published: (2025)
Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
by: Lee, Wonjun, et al.
Published: (2025)
by: Lee, Wonjun, et al.
Published: (2025)
No Pose Estimation? No Problem: Pose-Agnostic and Instance-Aware Test-Time Adaptation for Monocular Depth Estimation
by: Sung, Mingyu, et al.
Published: (2025)
by: Sung, Mingyu, et al.
Published: (2025)
SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
by: Dang, Jisheng, et al.
Published: (2025)
by: Dang, Jisheng, et al.
Published: (2025)
A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images
by: Lee, Jaeseong, et al.
Published: (2025)
by: Lee, Jaeseong, et al.
Published: (2025)
Generating Accurate and Detailed Captions for High-Resolution Images
by: Lee, Hankyeol, et al.
Published: (2025)
by: Lee, Hankyeol, et al.
Published: (2025)
PASTA: Part-Aware Sketch-to-3D Shape Generation with Text-Aligned Prior
by: Lee, Seunggwan, et al.
Published: (2025)
by: Lee, Seunggwan, et al.
Published: (2025)
VC4VG: Optimizing Video Captions for Text-to-Video Generation
by: Du, Yang, et al.
Published: (2025)
by: Du, Yang, et al.
Published: (2025)
OW-VISCapTor: Abstractors for Open-World Video Instance Segmentation and Captioning
by: Choudhuri, Anwesa, et al.
Published: (2024)
by: Choudhuri, Anwesa, et al.
Published: (2024)
EVC-MF: End-to-end Video Captioning Network with Multi-scale Features
by: Niu, Tian-Zi, et al.
Published: (2024)
by: Niu, Tian-Zi, et al.
Published: (2024)
YTCommentQA: Video Question Answerability in Instructional Videos
by: Yang, Saelyne, et al.
Published: (2024)
by: Yang, Saelyne, et al.
Published: (2024)
Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations
by: Park, Jungin, et al.
Published: (2025)
by: Park, Jungin, et al.
Published: (2025)
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
by: Clark, Christopher, et al.
Published: (2026)
by: Clark, Christopher, et al.
Published: (2026)
AGIC: Attention-Guided Image Captioning to Improve Caption Relevance
by: Teja, L. D. M. S. Sai, et al.
Published: (2025)
by: Teja, L. D. M. S. Sai, et al.
Published: (2025)
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
by: Fan, Tiehan, et al.
Published: (2024)
by: Fan, Tiehan, et al.
Published: (2024)
VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks
by: Chen, Xinlong, et al.
Published: (2025)
by: Chen, Xinlong, et al.
Published: (2025)
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
by: Yang, Yanlai, et al.
Published: (2025)
by: Yang, Yanlai, et al.
Published: (2025)
Similar Items
-
SToRM: Supervised Token Reduction for Multi-modal LLMs toward efficient end-to-end autonomous driving
by: Kim, Seo Hyun, et al.
Published: (2026) -
VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis
by: Park, Jinho, et al.
Published: (2026) -
JUDO: A Juxtaposed Domain-Oriented Multimodal Reasoner for Industrial Anomaly QA
by: Kang, Hyunju, et al.
Published: (2026) -
QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models
by: Wang, Jiale, et al.
Published: (2026) -
Versatile Incremental Learning: Towards Class and Domain-Agnostic Incremental Learning
by: Park, Min-Yeong, et al.
Published: (2024)