Exploring the Design Space of Visual Context Representation in Video MLLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Du, Yifan, Huo, Yuqi, Zhou, Kun, Zhao, Zijia, Lu, Haoyu, Huang, Han, Zhao, Wayne Xin, Wang, Bingning, Chen, Weipeng, Wen, Ji-Rong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Event-oriented Long Video Understanding
por: Du, Yifan, et al.
Publicado: (2024)
por: Du, Yifan, et al.
Publicado: (2024)
Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs
por: Zhao, Zijia, et al.
Publicado: (2024)
por: Zhao, Zijia, et al.
Publicado: (2024)
Efficient Motion-Aware Video MLLM
por: Zhao, Zijia, et al.
Publicado: (2025)
por: Zhao, Zijia, et al.
Publicado: (2025)
Unveiling the Flaws: Exploring Imperfections in Synthetic Data and Mitigation Strategies for Large Language Models
por: Chen, Jie, et al.
Publicado: (2024)
por: Chen, Jie, et al.
Publicado: (2024)
Beyond Filtering: Adaptive Image-Text Quality Enhancement for MLLM Pretraining
por: Huang, Han, et al.
Publicado: (2024)
por: Huang, Han, et al.
Publicado: (2024)
Virgo: A Preliminary Exploration on Reproducing o1-like MLLM
por: Du, Yifan, et al.
Publicado: (2025)
por: Du, Yifan, et al.
Publicado: (2025)
Extracting and Combining Abilities For Building Multi-lingual Ability-enhanced Large Language Models
por: Chen, Zhipeng, et al.
Publicado: (2024)
por: Chen, Zhipeng, et al.
Publicado: (2024)
Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models
por: Li, Yifan, et al.
Publicado: (2024)
por: Li, Yifan, et al.
Publicado: (2024)
Exploring Context Window of Large Language Models via Decomposed Positional Vectors
por: Dong, Zican, et al.
Publicado: (2024)
por: Dong, Zican, et al.
Publicado: (2024)
LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration Distillation
por: Dong, Zican, et al.
Publicado: (2025)
por: Dong, Zican, et al.
Publicado: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
por: Du, Yifan, et al.
Publicado: (2023)
por: Du, Yifan, et al.
Publicado: (2023)
Analyzing and Mitigating Object Hallucination: A Training Bias Perspective
por: Li, Yifan, et al.
Publicado: (2025)
por: Li, Yifan, et al.
Publicado: (2025)
Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models
por: Liu, Zikang, et al.
Publicado: (2025)
por: Liu, Zikang, et al.
Publicado: (2025)
Less is More: High-value Data Selection for Visual Instruction Tuning
por: Liu, Zikang, et al.
Publicado: (2024)
por: Liu, Zikang, et al.
Publicado: (2024)
DAWN-ICL: Strategic Planning of Problem-solving Trajectories for Zero-Shot In-Context Learning
por: Tang, Xinyu, et al.
Publicado: (2024)
por: Tang, Xinyu, et al.
Publicado: (2024)
Base of RoPE Bounds Context Length
por: Men, Xin, et al.
Publicado: (2024)
por: Men, Xin, et al.
Publicado: (2024)
Investigating the Pre-Training Dynamics of In-Context Learning: Task Recognition vs. Task Learning
por: Wang, Xiaolei, et al.
Publicado: (2024)
por: Wang, Xiaolei, et al.
Publicado: (2024)
Improving Vision-language Models with Perception-centric Process Reward Models
por: Min, Yingqian, et al.
Publicado: (2026)
por: Min, Yingqian, et al.
Publicado: (2026)
AVC-DPO: Aligned Video Captioning via Direct Preference Optimization
por: Tang, Jiyang, et al.
Publicado: (2025)
por: Tang, Jiyang, et al.
Publicado: (2025)
Not Everything is All You Need: Toward Low-Redundant Optimization for Large Language Model Alignment
por: Chen, Zhipeng, et al.
Publicado: (2024)
por: Chen, Zhipeng, et al.
Publicado: (2024)
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
por: Liu, Xiaolin, et al.
Publicado: (2026)
por: Liu, Xiaolin, et al.
Publicado: (2026)
ReasoningLM: Enabling Structural Subgraph Reasoning in Pre-trained Language Models for Question Answering over Knowledge Graph
por: Jiang, Jinhao, et al.
Publicado: (2023)
por: Jiang, Jinhao, et al.
Publicado: (2023)
Experience-Guided Reflective Co-Evolution of Prompts and Heuristics for Automatic Algorithm Design
por: Liu, Yihong, et al.
Publicado: (2025)
por: Liu, Yihong, et al.
Publicado: (2025)
MetaGPT: Merging Large Language Models Using Model Exclusive Task Arithmetic
por: Zhou, Yuyan, et al.
Publicado: (2024)
por: Zhou, Yuyan, et al.
Publicado: (2024)
KV Shifting Attention Enhances Language Modeling
por: Xu, Mingyu, et al.
Publicado: (2024)
por: Xu, Mingyu, et al.
Publicado: (2024)
ChainLM: Empowering Large Language Models with Improved Chain-of-Thought Prompting
por: Cheng, Xiaoxue, et al.
Publicado: (2024)
por: Cheng, Xiaoxue, et al.
Publicado: (2024)
Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinking
por: Cheng, Xiaoxue, et al.
Publicado: (2025)
por: Cheng, Xiaoxue, et al.
Publicado: (2025)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
por: Zheng, Naishan, et al.
Publicado: (2025)
por: Zheng, Naishan, et al.
Publicado: (2025)
Beyond Imitation: Leveraging Fine-grained Quality Signals for Alignment
por: Guo, Geyang, et al.
Publicado: (2023)
por: Guo, Geyang, et al.
Publicado: (2023)
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
por: Zhao, Jiahe, et al.
Publicado: (2025)
por: Zhao, Jiahe, et al.
Publicado: (2025)
Full-ECE: A Metric For Token-level Calibration on Large Language Models
por: Liu, Han, et al.
Publicado: (2024)
por: Liu, Han, et al.
Publicado: (2024)
Beyond the Last Frame: Process-aware Evaluation for Generative Video Reasoning
por: Li, Yifan, et al.
Publicado: (2025)
por: Li, Yifan, et al.
Publicado: (2025)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
por: Ouyang, Kun, et al.
Publicado: (2025)
por: Ouyang, Kun, et al.
Publicado: (2025)
Improving Conversational Recommendation Systems via Counterfactual Data Simulation
por: Wang, Xiaolei, et al.
Publicado: (2023)
por: Wang, Xiaolei, et al.
Publicado: (2023)
Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration
por: Chen, Zhipeng, et al.
Publicado: (2026)
por: Chen, Zhipeng, et al.
Publicado: (2026)
MMATH: A Multilingual Benchmark for Mathematical Reasoning
por: Luo, Wenyang, et al.
Publicado: (2025)
por: Luo, Wenyang, et al.
Publicado: (2025)
Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models
por: Sun, Haoxiang, et al.
Publicado: (2025)
por: Sun, Haoxiang, et al.
Publicado: (2025)
BAMBOO: A Comprehensive Benchmark for Evaluating Long Text Modeling Capacities of Large Language Models
por: Dong, Zican, et al.
Publicado: (2023)
por: Dong, Zican, et al.
Publicado: (2023)
Improving Large Language Models via Fine-grained Reinforcement Learning with Minimum Editing Constraint
por: Chen, Zhipeng, et al.
Publicado: (2024)
por: Chen, Zhipeng, et al.
Publicado: (2024)
KG-Agent: An Efficient Autonomous Agent Framework for Complex Reasoning over Knowledge Graph
por: Jiang, Jinhao, et al.
Publicado: (2024)
por: Jiang, Jinhao, et al.
Publicado: (2024)
Ejemplares similares
-
Towards Event-oriented Long Video Understanding
por: Du, Yifan, et al.
Publicado: (2024) -
Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs
por: Zhao, Zijia, et al.
Publicado: (2024) -
Efficient Motion-Aware Video MLLM
por: Zhao, Zijia, et al.
Publicado: (2025) -
Unveiling the Flaws: Exploring Imperfections in Synthetic Data and Mitigation Strategies for Large Language Models
por: Chen, Jie, et al.
Publicado: (2024) -
Beyond Filtering: Adaptive Image-Text Quality Enhancement for MLLM Pretraining
por: Huang, Han, et al.
Publicado: (2024)