AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Ren, Yiming, Lin, Zhiqiang, Li, Yu, Meng, Gao, Wang, Weiyun, Wang, Junjie, Lin, Zicheng, Dai, Jifeng, Yang, Yujiu, Wang, Wenhai, Chu, Ruihang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing
by: Chen, Yuqing, et al.
Published: (2025)
by: Chen, Yuqing, et al.
Published: (2025)
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
by: Ren, Yiming, et al.
Published: (2026)
by: Ren, Yiming, et al.
Published: (2026)
Mitigating the Reasoning Tax in Vision-Language Fine-Tuning with Input-Adaptive Depth Aggregation
by: Ren, Yiming, et al.
Published: (2026)
by: Ren, Yiming, et al.
Published: (2026)
OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis
by: Chen, Junting, et al.
Published: (2025)
by: Chen, Junting, et al.
Published: (2025)
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
by: Xu, Weiye, et al.
Published: (2025)
by: Xu, Weiye, et al.
Published: (2025)
Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
by: Ma, Ziyang, et al.
Published: (2025)
by: Ma, Ziyang, et al.
Published: (2025)
The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
by: Wang, Weiyun, et al.
Published: (2024)
by: Wang, Weiyun, et al.
Published: (2024)
Velocity-Space 3D Asset Editing
by: Liu, Hao, et al.
Published: (2026)
by: Liu, Hao, et al.
Published: (2026)
SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature
by: Ren, Yiming, et al.
Published: (2026)
by: Ren, Yiming, et al.
Published: (2026)
OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
by: Zhao, Xiangyu, et al.
Published: (2025)
by: Zhao, Xiangyu, et al.
Published: (2025)
UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
AR-Omni: A Unified Autoregressive Model for Any-to-Any Generation
by: Cheng, Dongjie, et al.
Published: (2026)
by: Cheng, Dongjie, et al.
Published: (2026)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
by: Tian, Changyao, et al.
Published: (2024)
by: Tian, Changyao, et al.
Published: (2024)
PTD-SQL: Partitioning and Targeted Drilling with LLMs in Text-to-SQL
by: Luo, Ruilin, et al.
Published: (2024)
by: Luo, Ruilin, et al.
Published: (2024)
VoCap: Video Object Captioning and Segmentation from Any Prompt
by: Uijlings, Jasper, et al.
Published: (2025)
by: Uijlings, Jasper, et al.
Published: (2025)
MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
by: Liu, Yangzhou, et al.
Published: (2024)
by: Liu, Yangzhou, et al.
Published: (2024)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
by: Lei, Zhenxin, et al.
Published: (2025)
by: Lei, Zhenxin, et al.
Published: (2025)
OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
by: Li, Qingyun, et al.
Published: (2024)
by: Li, Qingyun, et al.
Published: (2024)
OmniControl: Control Any Joint at Any Time for Human Motion Generation
by: Xie, Yiming, et al.
Published: (2023)
by: Xie, Yiming, et al.
Published: (2023)
Unlocking Multimodal Mathematical Reasoning via Process Reward Model
by: Luo, Ruilin, et al.
Published: (2025)
by: Luo, Ruilin, et al.
Published: (2025)
Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
by: Duan, Yuchen, et al.
Published: (2024)
by: Duan, Yuchen, et al.
Published: (2024)
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
by: Ding, Yang, et al.
Published: (2025)
by: Ding, Yang, et al.
Published: (2025)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
by: Lin, Zicheng, et al.
Published: (2024)
by: Lin, Zicheng, et al.
Published: (2024)
A Survey of Reasoning in Autonomous Driving Systems: Open Challenges and Emerging Paradigms
by: Yu, Kejin, et al.
Published: (2026)
by: Yu, Kejin, et al.
Published: (2026)
STUDY ON PREPARATION AND PROPERTIES OF PHENOL-FORMALDEHYDE-CHINESE FIR LIQUEFACTION COPOLYMER RESIN
by: Ruihang Lin
Published: (2014)
by: Ruihang Lin
Published: (2014)
OmniDiff: A Comprehensive Benchmark for Fine-grained Image Difference Captioning
by: Liu, Yuan, et al.
Published: (2025)
by: Liu, Yuan, et al.
Published: (2025)
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
by: Meng, Desen, et al.
Published: (2025)
by: Meng, Desen, et al.
Published: (2025)
CoMemo: LVLMs Need Image Context with Image Memory
by: Liu, Shi, et al.
Published: (2025)
by: Liu, Shi, et al.
Published: (2025)
Bounding Box Stability against Feature Dropout Reflects Detector Generalization across Environments
by: Yang, Yang, et al.
Published: (2024)
by: Yang, Yang, et al.
Published: (2024)
ProCap: Projection-Aware Captioning for Spatial Augmented Reality
by: Cao, Zimo, et al.
Published: (2026)
by: Cao, Zimo, et al.
Published: (2026)
OmniGen: Unified Image Generation
by: Xiao, Shitao, et al.
Published: (2024)
by: Xiao, Shitao, et al.
Published: (2024)
OmniEvent: Unified Event Representation Learning
by: Yan, Weiqi, et al.
Published: (2025)
by: Yan, Weiqi, et al.
Published: (2025)
InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
by: Tong, Wenwen, et al.
Published: (2025)
by: Tong, Wenwen, et al.
Published: (2025)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
by: Wu, Peiran, et al.
Published: (2025)
by: Wu, Peiran, et al.
Published: (2025)
AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
by: Chai, Wenhao, et al.
Published: (2024)
by: Chai, Wenhao, et al.
Published: (2024)
GroundCap: A Visually Grounded Image Captioning Dataset
by: Oliveira, Daniel A. P., et al.
Published: (2025)
by: Oliveira, Daniel A. P., et al.
Published: (2025)
OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts
by: Wang, Yuxuan, et al.
Published: (2025)
by: Wang, Yuxuan, et al.
Published: (2025)
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
by: Wang, Weiyun, et al.
Published: (2024)
by: Wang, Weiyun, et al.
Published: (2024)
Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning
by: Tosato, Lucrezia, et al.
Published: (2026)
by: Tosato, Lucrezia, et al.
Published: (2026)
Similar Items
-
O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing
by: Chen, Yuqing, et al.
Published: (2025) -
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
by: Ren, Yiming, et al.
Published: (2026) -
Mitigating the Reasoning Tax in Vision-Language Fine-Tuning with Input-Adaptive Depth Aggregation
by: Ren, Yiming, et al.
Published: (2026) -
OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis
by: Chen, Junting, et al.
Published: (2025) -
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
by: Xu, Weiye, et al.
Published: (2025)