OmniCaptioner: One Captioner to Rule Them All
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Yiting, Yuan, Jiakang, Li, Zhen, Zhao, Shitian, Qin, Qi, Li, Xinyue, Zhuo, Le, Wen, Licheng, Liu, Dongyang, Cao, Yuewen, Yan, Xiangchao, Li, Xin, Peng, Tianshuo, Zhang, Shufei, Shi, Botian, Chen, Tao, Chen, Zhibo, Bai, Lei, Gao, Peng, Zhang, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework
by: Bianchi, Lorenzo, et al.
Published: (2025)
by: Bianchi, Lorenzo, et al.
Published: (2025)
OmniQuality-R: Advancing Reward Models Through All-Encompassing Quality Assessment
by: Lu, Yiting, et al.
Published: (2025)
by: Lu, Yiting, et al.
Published: (2025)
One Ring to Rule Them All: Unifying Group-Based RL via Dynamic Power-Mean Geometry
by: Zhao, Weisong, et al.
Published: (2026)
by: Zhao, Weisong, et al.
Published: (2026)
Rule-driven News Captioning
by: Xu, Ning, et al.
Published: (2024)
by: Xu, Ning, et al.
Published: (2024)
Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
by: Ma, Ziyang, et al.
Published: (2025)
by: Ma, Ziyang, et al.
Published: (2025)
SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning
by: Zhang, Lin, et al.
Published: (2025)
by: Zhang, Lin, et al.
Published: (2025)
One SPACE to Rule Them All: Jointly Mitigating Factuality and Faithfulness Hallucinations in LLMs
by: Wang, Pengbo, et al.
Published: (2025)
by: Wang, Pengbo, et al.
Published: (2025)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
by: Wu, Peiran, et al.
Published: (2025)
by: Wu, Peiran, et al.
Published: (2025)
GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
by: Xia, Renqiu, et al.
Published: (2024)
by: Xia, Renqiu, et al.
Published: (2024)
SupeRANSAC: One RANSAC to Rule Them All
by: Barath, Daniel
Published: (2025)
by: Barath, Daniel
Published: (2025)
Set Prediction Guided by Semantic Concepts for Diverse Video Captioning
by: Lu, Yifan, et al.
Published: (2023)
by: Lu, Yifan, et al.
Published: (2023)
Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
by: Liu, Dongyang, et al.
Published: (2024)
by: Liu, Dongyang, et al.
Published: (2024)
IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring
by: Peng, Xinge, et al.
Published: (2026)
by: Peng, Xinge, et al.
Published: (2026)
Training-Free Adaptive Diffusion with Bounded Difference Approximation Strategy
by: Ye, Hancheng, et al.
Published: (2024)
by: Ye, Hancheng, et al.
Published: (2024)
All-in-One: Transferring Vision Foundation Models into Stereo Matching
by: Zhou, Jingyi, et al.
Published: (2024)
by: Zhou, Jingyi, et al.
Published: (2024)
The Role of Data Curation in Image Captioning
by: Li, Wenyan, et al.
Published: (2023)
by: Li, Wenyan, et al.
Published: (2023)
Region Guide Grid Cross Transformer for Image Caption
by: Jiayu Bai, et al.
Published: (2026)
by: Jiayu Bai, et al.
Published: (2026)
One Sample to Rule Them All: Extreme Data Efficiency in Multidiscipline Reasoning with Reinforcement Learning
by: Li, Yiyuan, et al.
Published: (2026)
by: Li, Yiyuan, et al.
Published: (2026)
One Prompt To Rule Them All: LLMs for Opinion Summary Evaluation
by: Siledar, Tejpalsingh, et al.
Published: (2024)
by: Siledar, Tejpalsingh, et al.
Published: (2024)
DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning
by: Wei, Yuancheng, et al.
Published: (2026)
by: Wei, Yuancheng, et al.
Published: (2026)
OmniLight: One Model to Rule All Lighting Conditions
by: Oh, Youngjin, et al.
Published: (2026)
by: Oh, Youngjin, et al.
Published: (2026)
CaptionQA: Is Your Caption as Useful as the Image Itself?
by: Yang, Shijia, et al.
Published: (2025)
by: Yang, Shijia, et al.
Published: (2025)
Scene Graph-guided SegCaptioning Transformer with Fine-grained Alignment for Controllable Video Segmentation and Captioning
by: Zhang, Xu, et al.
Published: (2026)
by: Zhang, Xu, et al.
Published: (2026)
One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning
by: Sun, Hao, et al.
Published: (2024)
by: Sun, Hao, et al.
Published: (2024)
Unveiling Effective In-Context Configurations for Image Captioning: An External & Internal Analysis
by: Li, Li, et al.
Published: (2025)
by: Li, Li, et al.
Published: (2025)
Dual-path Collaborative Generation Network for Emotional Video Captioning
by: Ye, Cheng, et al.
Published: (2024)
by: Ye, Cheng, et al.
Published: (2024)
BTCChat: Advancing Remote Sensing Bi-temporal Change Captioning with Multimodal Large Language Model
by: Li, Yujie, et al.
Published: (2025)
by: Li, Yujie, et al.
Published: (2025)
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
by: Lei, Zhenxin, et al.
Published: (2025)
by: Lei, Zhenxin, et al.
Published: (2025)
Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback
by: Yuan, Jiakang, et al.
Published: (2025)
by: Yuan, Jiakang, et al.
Published: (2025)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
by: Yao, Linli, et al.
Published: (2026)
by: Yao, Linli, et al.
Published: (2026)
Self-Explainable Affordance Learning with Embodied Caption
by: Zhang, Zhipeng, et al.
Published: (2024)
by: Zhang, Zhipeng, et al.
Published: (2024)
Benchmarking and Improving Detail Image Caption
by: Dong, Hongyuan, et al.
Published: (2024)
by: Dong, Hongyuan, et al.
Published: (2024)
AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
by: Ren, Yiming, et al.
Published: (2025)
by: Ren, Yiming, et al.
Published: (2025)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
by: Wei, Hongchen, et al.
Published: (2025)
by: Wei, Hongchen, et al.
Published: (2025)
One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers
by: Abagyan, Diana, et al.
Published: (2025)
by: Abagyan, Diana, et al.
Published: (2025)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
by: Li, Wenyan, et al.
Published: (2024)
by: Li, Wenyan, et al.
Published: (2024)
ViDiC: Video Difference Captioning
by: Wu, Jiangtao, et al.
Published: (2025)
by: Wu, Jiangtao, et al.
Published: (2025)
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
by: Lei, Jiayi, et al.
Published: (2025)
by: Lei, Jiayi, et al.
Published: (2025)
Technical Report for Soccernet 2023 -- Dense Video Captioning
by: Ruan, Zheng, et al.
Published: (2024)
by: Ruan, Zheng, et al.
Published: (2024)
L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training
by: Li, Li, et al.
Published: (2025)
by: Li, Li, et al.
Published: (2025)
Similar Items
-
One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework
by: Bianchi, Lorenzo, et al.
Published: (2025) -
OmniQuality-R: Advancing Reward Models Through All-Encompassing Quality Assessment
by: Lu, Yiting, et al.
Published: (2025) -
One Ring to Rule Them All: Unifying Group-Based RL via Dynamic Power-Mean Geometry
by: Zhao, Weisong, et al.
Published: (2026) -
Rule-driven News Captioning
by: Xu, Ning, et al.
Published: (2024) -
Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
by: Ma, Ziyang, et al.
Published: (2025)