Scaling Language-Centric Omnimodal Representation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Chenghao, Chan, Hou Pong, Zhang, Hao, Xu, Weiwen, Aljunied, Mahani, Rong, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
by: LASA Team, et al.
Published: (2025)
by: LASA Team, et al.
Published: (2025)
VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
by: Yuan, Ruifeng, et al.
Published: (2025)
by: Yuan, Ruifeng, et al.
Published: (2025)
SeaLLMs-Audio: Large Audio-Language Models for Southeast Asia
by: Liu, Chaoqun, et al.
Published: (2025)
by: Liu, Chaoqun, et al.
Published: (2025)
Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
by: Li, Yunxin, et al.
Published: (2025)
by: Li, Yunxin, et al.
Published: (2025)
Demographic and Linguistic Bias Evaluation in Omnimodal Language Models
by: Elobaid, Alaa
Published: (2026)
by: Elobaid, Alaa
Published: (2026)
Multimedia Generative Script Learning for Task Planning
by: Wang, Qingyun, et al.
Published: (2022)
by: Wang, Qingyun, et al.
Published: (2022)
Analyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal Representations
by: Xiao, Chenghao, et al.
Published: (2025)
by: Xiao, Chenghao, et al.
Published: (2025)
NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
by: Luo, Run, et al.
Published: (2025)
by: Luo, Run, et al.
Published: (2025)
MotionEdit: Benchmarking and Learning Motion-Centric Image Editing
by: Wan, Yixin, et al.
Published: (2025)
by: Wan, Yixin, et al.
Published: (2025)
From Pixels to Insights: A Survey on Automatic Chart Understanding in the Era of Large Foundation Models
by: Huang, Kung-Hsiang, et al.
Published: (2024)
by: Huang, Kung-Hsiang, et al.
Published: (2024)
Shifting AI Efficiency From Model-Centric to Data-Centric Compression
by: Liu, Xuyang, et al.
Published: (2025)
by: Liu, Xuyang, et al.
Published: (2025)
Allegory of the Cave: Measurement-Grounded Vision-Language Learning
by: Xu, Kepeng, et al.
Published: (2026)
by: Xu, Kepeng, et al.
Published: (2026)
VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning
by: Ma, Jingkun, et al.
Published: (2024)
by: Ma, Jingkun, et al.
Published: (2024)
Retrieval-based Disentangled Representation Learning with Natural Language Supervision
by: Zhou, Jiawei, et al.
Published: (2022)
by: Zhou, Jiawei, et al.
Published: (2022)
Babel: Open Multilingual Large Language Models Serving Over 90% of Global Speakers
by: Zhao, Yiran, et al.
Published: (2025)
by: Zhao, Yiran, et al.
Published: (2025)
Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
by: Lu, Meng, et al.
Published: (2025)
by: Lu, Meng, et al.
Published: (2025)
Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization
by: Wu, Jiulong, et al.
Published: (2025)
by: Wu, Jiulong, et al.
Published: (2025)
MindGYM: What Matters in Question Synthesis for Thinking-Centric Fine-Tuning?
by: Xu, Zhe, et al.
Published: (2025)
by: Xu, Zhe, et al.
Published: (2025)
CSLRConformer: A Data-Centric Conformer Approach for Continuous Arabic Sign Language Recognition on the Isharah Datase
by: Elden, Fatimah Mohamed Emad
Published: (2025)
by: Elden, Fatimah Mohamed Emad
Published: (2025)
Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
by: Zhang, Juntian, et al.
Published: (2025)
by: Zhang, Juntian, et al.
Published: (2025)
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
by: Kim, Youngmin, et al.
Published: (2025)
by: Kim, Youngmin, et al.
Published: (2025)
MLLMs-Augmented Visual-Language Representation Learning
by: Liu, Yanqing, et al.
Published: (2023)
by: Liu, Yanqing, et al.
Published: (2023)
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
by: Yu, Hao, et al.
Published: (2025)
by: Yu, Hao, et al.
Published: (2025)
Concept Lancet: Image Editing with Compositional Representation Transplant
by: Luo, Jinqi, et al.
Published: (2025)
by: Luo, Jinqi, et al.
Published: (2025)
Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning
by: Hu, Zhe, et al.
Published: (2025)
by: Hu, Zhe, et al.
Published: (2025)
VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models Reasoning
by: Xiao, Wenyi, et al.
Published: (2026)
by: Xiao, Wenyi, et al.
Published: (2026)
An Online Reference-Free Evaluation Framework for Flowchart Image-to-Code Generation
by: Nguyen, Giang Son, et al.
Published: (2026)
by: Nguyen, Giang Son, et al.
Published: (2026)
An Examination of the Compositionality of Large Generative Vision-Language Models
by: Ma, Teli, et al.
Published: (2023)
by: Ma, Teli, et al.
Published: (2023)
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
by: Ren, Shuhuai, et al.
Published: (2023)
by: Ren, Shuhuai, et al.
Published: (2023)
More Text, Less Point: Towards 3D Data-Efficient Point-Language Understanding
by: Tang, Yuan, et al.
Published: (2024)
by: Tang, Yuan, et al.
Published: (2024)
MemeMind: A Large-Scale Multimodal Dataset with Chain-of-Thought Reasoning for Harmful Meme Detection
by: Gu, Hexiang, et al.
Published: (2025)
by: Gu, Hexiang, et al.
Published: (2025)
Dyna-Mind: Learning to Simulate from Experience for Better AI Agents
by: Yu, Xiao, et al.
Published: (2025)
by: Yu, Xiao, et al.
Published: (2025)
VisualRWKV-HD and UHD: Advancing High-Resolution Processing for Visual Language Models
by: Li, Zihang, et al.
Published: (2024)
by: Li, Zihang, et al.
Published: (2024)
Superpixel Semantics Representation and Pre-training for Vision-Language Task
by: Zhang, Siyu, et al.
Published: (2023)
by: Zhang, Siyu, et al.
Published: (2023)
MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models
by: Hu, Wenbo, et al.
Published: (2024)
by: Hu, Wenbo, et al.
Published: (2024)
Brain-CLIPLM: Decoding Compressed Semantic Representations in EEG for Language Reconstruction
by: Yang, Xiaoli, et al.
Published: (2026)
by: Yang, Xiaoli, et al.
Published: (2026)
Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
by: Shi, Yang, et al.
Published: (2025)
by: Shi, Yang, et al.
Published: (2025)
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
by: Hong, Jack, et al.
Published: (2025)
by: Hong, Jack, et al.
Published: (2025)
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
by: Wang, Ziyang, et al.
Published: (2025)
by: Wang, Ziyang, et al.
Published: (2025)
VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
by: Li, Shicheng, et al.
Published: (2023)
by: Li, Shicheng, et al.
Published: (2023)
Similar Items
-
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
by: LASA Team, et al.
Published: (2025) -
VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
by: Yuan, Ruifeng, et al.
Published: (2025) -
SeaLLMs-Audio: Large Audio-Language Models for Southeast Asia
by: Liu, Chaoqun, et al.
Published: (2025) -
Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
by: Li, Yunxin, et al.
Published: (2025) -
Demographic and Linguistic Bias Evaluation in Omnimodal Language Models
by: Elobaid, Alaa
Published: (2026)