Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Guo, Lu, Lidong, Liu, Yicheng, Dong, Liangrui, Zou, Lidong, Lv, Jixin, Li, Zhenquan, Mao, Xinyi, Pei, Baoqi, Wang, Shihao, Li, Zhiqi, Sapra, Karan, Liu, Fuxiao, Zheng, Yin-Dong, Huang, Yifei, Wang, Limin, Yu, Zhiding, Tao, Andrew, Liu, Guilin, Lu, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AIDE: Agentically Improve Visual Language Model with Domain Experts
by: Chiu, Ming-Chang, et al.
Published: (2025)
by: Chiu, Ming-Chang, et al.
Published: (2025)
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
by: Lu, Lidong, et al.
Published: (2025)
by: Lu, Lidong, et al.
Published: (2025)
Learning Visual Affordance from Audio
by: Lu, Lidong, et al.
Published: (2025)
by: Lu, Lidong, et al.
Published: (2025)
Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision
by: He, Yuping, et al.
Published: (2025)
by: He, Yuping, et al.
Published: (2025)
Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
by: Chen, Guo, et al.
Published: (2025)
by: Chen, Guo, et al.
Published: (2025)
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
by: Chen, Guo, et al.
Published: (2024)
by: Chen, Guo, et al.
Published: (2024)
Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
by: Shi, Min, et al.
Published: (2024)
by: Shi, Min, et al.
Published: (2024)
Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
by: Chen, Guo, et al.
Published: (2024)
by: Chen, Guo, et al.
Published: (2024)
PhyCritic: Multimodal Critic Models for Physical AI
by: Xiong, Tianyi, et al.
Published: (2026)
by: Xiong, Tianyi, et al.
Published: (2026)
Toward a Green Revolution in soybean: The role of ultra‐high‐density planting
by: Chao Fang, et al.
Published: (2025)
by: Chao Fang, et al.
Published: (2025)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
by: Wang, Shihao, et al.
Published: (2025)
by: Wang, Shihao, et al.
Published: (2025)
General Information Theory: Time and Information
by: Liu, Yilun, et al.
Published: (2019)
by: Liu, Yilun, et al.
Published: (2019)
EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
by: Pei, Baoqi, et al.
Published: (2024)
by: Pei, Baoqi, et al.
Published: (2024)
Stateful Token Reduction for Long-Video Hybrid VLMs
by: Jiang, Jindong, et al.
Published: (2026)
by: Jiang, Jindong, et al.
Published: (2026)
Comment on “The association between COVID ‐19 and incident gestational diabetes: A population‐based case–control study of the National Health Insurance Research Database in Taiwan”
by: Lidong Gao
Published: (2026)
by: Lidong Gao
Published: (2026)
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
by: Wang, Shihao, et al.
Published: (2026)
by: Wang, Shihao, et al.
Published: (2026)
Insights Into Guanine Radical Cation Deprotonation Using the Quantum Mechanics and Quantum Mechanics/Molecular Mechanics (ABEEM) Methods
by: Yue Wang, et al.
Published: (2024)
by: Yue Wang, et al.
Published: (2024)
Adaptive Neural Fixed‐Time Command Filtered Control for Stochastic Nonlinear Systems With Input Quantization
by: Long Gu, et al.
Published: (2025)
by: Long Gu, et al.
Published: (2025)
EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World
by: Huang, Yifei, et al.
Published: (2024)
by: Huang, Yifei, et al.
Published: (2024)
LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering
by: Zhang, Hongjie, et al.
Published: (2023)
by: Zhang, Hongjie, et al.
Published: (2023)
Game-Theoretic Risk-Shaped Reinforcement Learning for Safe Autonomous Driving
by: Hu, Dong, et al.
Published: (2025)
by: Hu, Dong, et al.
Published: (2025)
Design, Control, and Clinical Applications of Magnetic Actuation Systems: Challenges and Opportunities
by: Yingxin Huo, et al.
Published: (2024)
by: Yingxin Huo, et al.
Published: (2024)
Lifelong Language-Conditioned Robotic Manipulation Learning
by: Wang, Xudong, et al.
Published: (2026)
by: Wang, Xudong, et al.
Published: (2026)
Lifelong Embodied Navigation Learning
by: Wang, Xudong, et al.
Published: (2026)
by: Wang, Xudong, et al.
Published: (2026)
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning
by: Yang, Zuhao, et al.
Published: (2026)
by: Yang, Zuhao, et al.
Published: (2026)
AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum
by: Li, Jiaqi, et al.
Published: (2026)
by: Li, Jiaqi, et al.
Published: (2026)
Yuan: Research on the Concept of Digital World Analogue Scientific Infrastructure and Science Popularization Communication Based on Suzhou Gardens Pattern
by: Lvyang, Zhang, et al.
Published: (2024)
by: Lvyang, Zhang, et al.
Published: (2024)
Maximum Power Reference Tracking Algorithm for Power Curtailment of Photovoltaic Systems
by: Paduani, Victor, et al.
Published: (2020)
by: Paduani, Victor, et al.
Published: (2020)
Stabilization of Nonlinear Systems with State-Dependent Representation: From Model-Based to Direct Data-Driven Control
by: Li, Lidong, et al.
Published: (2025)
by: Li, Lidong, et al.
Published: (2025)
DRL-Based Trajectory Tracking for Motion-Related Modules in Autonomous Driving
by: Xu, Yinda, et al.
Published: (2023)
by: Xu, Yinda, et al.
Published: (2023)
MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier
by: Yang, Zonglin, et al.
Published: (2026)
by: Yang, Zonglin, et al.
Published: (2026)
How Does CEO Cognitive Style and Board Characteristics Impact Strategic Risk‐Taking?
by: Xiaoyu Tian, et al.
Published: (2025)
by: Xiaoyu Tian, et al.
Published: (2025)
Cellular Cancer Immunotherapy in the Liver Transplant Population for HCC: An Attractive Therapeutic Option for the Next Decade
by: Dongdong Yu, et al.
Published: (2025)
by: Dongdong Yu, et al.
Published: (2025)
Recent Advances of Amorphous Nanomaterials: Synthesis and Applications
by: Lidong Li, et al.
Published: (2024)
by: Lidong Li, et al.
Published: (2024)
Evolving Prompts In-Context: An Open-ended, Self-replicating Perspective
by: Wang, Jianyu, et al.
Published: (2025)
by: Wang, Jianyu, et al.
Published: (2025)
Emerging Trends in Neoadjuvant Immunotherapy for Hepatocellular Carcinoma: A Focus on Liver Transplant Candidates
by: Dongdong Yu, et al.
Published: (2025)
by: Dongdong Yu, et al.
Published: (2025)
Slow-Fast Architecture for Video Multi-Modal Large Language Models
by: Shi, Min, et al.
Published: (2025)
by: Shi, Min, et al.
Published: (2025)
RubikSQL: Lifelong Learning Agentic Knowledge Base as an Industrial NL2SQL System
by: Chen, Zui, et al.
Published: (2025)
by: Chen, Zui, et al.
Published: (2025)
Asymmetric Coordination Modulating Co Spin State for Peroxymonosulfate Activation to Accelerate 1 O 2 Generation
by: Xiuze Li, et al.
Published: (2025)
by: Xiuze Li, et al.
Published: (2025)
Data-driven Control Against False Data Injection Attacks
by: Liu, Wenjie, et al.
Published: (2023)
by: Liu, Wenjie, et al.
Published: (2023)
Similar Items
-
AIDE: Agentically Improve Visual Language Model with Domain Experts
by: Chiu, Ming-Chang, et al.
Published: (2025) -
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
by: Lu, Lidong, et al.
Published: (2025) -
Learning Visual Affordance from Audio
by: Lu, Lidong, et al.
Published: (2025) -
Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision
by: He, Yuping, et al.
Published: (2025) -
Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
by: Chen, Guo, et al.
Published: (2025)