AdaCodec: A Predictive Visual Code for Video MLLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hou, Haowen, Huang, Zhen, Liang, Zheming, Si, Qingyi, Li, Chenglin, Dong, Shuai, Shao, Kele, Li, Ruilin, Wang, Dianyi, Duan, Nan, Wang, Jiaqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VideoPro: Adaptive Program Reasoning for Long Video Understanding
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
Interleaved Latent Visual Reasoning with Selective Perceptual Modeling
von: Dong, Shuai, et al.
Veröffentlicht: (2025)
von: Dong, Shuai, et al.
Veröffentlicht: (2025)
AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
UniREditBench: A Unified Reasoning-based Image Editing Benchmark
von: Han, Feng, et al.
Veröffentlicht: (2025)
von: Han, Feng, et al.
Veröffentlicht: (2025)
CodePercept: Code-Grounded Visual STEM Perception for MLLMs
von: Guan, Tongkun, et al.
Veröffentlicht: (2026)
von: Guan, Tongkun, et al.
Veröffentlicht: (2026)
VisualRWKV-HD and UHD: Advancing High-Resolution Processing for Visual Language Models
von: Li, Zihang, et al.
Veröffentlicht: (2024)
von: Li, Zihang, et al.
Veröffentlicht: (2024)
EasyVideoR1: Easier RL for Video Understanding
von: Qin, Chuanyu, et al.
Veröffentlicht: (2026)
von: Qin, Chuanyu, et al.
Veröffentlicht: (2026)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
von: Wang, Dianyi, et al.
Veröffentlicht: (2025)
von: Wang, Dianyi, et al.
Veröffentlicht: (2025)
EtCon: Edit-then-Consolidate for Reliable Knowledge Editing
von: Li, Ruilin, et al.
Veröffentlicht: (2025)
von: Li, Ruilin, et al.
Veröffentlicht: (2025)
GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs
von: Duan, Yuxiang, et al.
Veröffentlicht: (2025)
von: Duan, Yuxiang, et al.
Veröffentlicht: (2025)
RLFR: Extending Reinforcement Learning for LLMs with Flow Environment
von: Zhang, Jinghao, et al.
Veröffentlicht: (2025)
von: Zhang, Jinghao, et al.
Veröffentlicht: (2025)
ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
Object Attribute Matters in Visual Question Answering
von: Li, Peize, et al.
Veröffentlicht: (2023)
von: Li, Peize, et al.
Veröffentlicht: (2023)
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
von: Wang, Siting, et al.
Veröffentlicht: (2025)
von: Wang, Siting, et al.
Veröffentlicht: (2025)
StreamingTOM: Streaming Token Compression for Efficient Video Understanding
von: Chen, Xueyi, et al.
Veröffentlicht: (2025)
von: Chen, Xueyi, et al.
Veröffentlicht: (2025)
Exploring the Design Space of Visual Context Representation in Video MLLMs
von: Du, Yifan, et al.
Veröffentlicht: (2024)
von: Du, Yifan, et al.
Veröffentlicht: (2024)
Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection
von: Miao, Ziqi, et al.
Veröffentlicht: (2025)
von: Miao, Ziqi, et al.
Veröffentlicht: (2025)
Octavius: Mitigating Task Interference in MLLMs via LoRA-MoE
von: Chen, Zeren, et al.
Veröffentlicht: (2023)
von: Chen, Zeren, et al.
Veröffentlicht: (2023)
Towards Flexible Evaluation for Generative Visual Question Answering
von: Ji, Huishan, et al.
Veröffentlicht: (2024)
von: Ji, Huishan, et al.
Veröffentlicht: (2024)
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
von: Liu, Xiaolin, et al.
Veröffentlicht: (2026)
von: Liu, Xiaolin, et al.
Veröffentlicht: (2026)
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
von: Zhao, Jiahe, et al.
Veröffentlicht: (2025)
von: Zhao, Jiahe, et al.
Veröffentlicht: (2025)
Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster Inference
von: Yin, Hao, et al.
Veröffentlicht: (2025)
von: Yin, Hao, et al.
Veröffentlicht: (2025)
MLLMs-Augmented Visual-Language Representation Learning
von: Liu, Yanqing, et al.
Veröffentlicht: (2023)
von: Liu, Yanqing, et al.
Veröffentlicht: (2023)
Substantial, Decomposable, and Invisible: Visual Context Misalignment in Instructional Videos for Physical Tasks
von: Li, Yayuan, et al.
Veröffentlicht: (2026)
von: Li, Yayuan, et al.
Veröffentlicht: (2026)
Self-Distilled RLVR
von: Yang, Chenxu, et al.
Veröffentlicht: (2026)
von: Yang, Chenxu, et al.
Veröffentlicht: (2026)
QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA
von: Li, Shuai, et al.
Veröffentlicht: (2025)
von: Li, Shuai, et al.
Veröffentlicht: (2025)
FDIM: A Feature-distance-based Generic Video Quality Metric for Versatile Codecs
von: Wang, Jiayi, et al.
Veröffentlicht: (2026)
von: Wang, Jiayi, et al.
Veröffentlicht: (2026)
Video-R1: Reinforcing Video Reasoning in MLLMs
von: Feng, Kaituo, et al.
Veröffentlicht: (2025)
von: Feng, Kaituo, et al.
Veröffentlicht: (2025)
RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition
von: Liu, Ziyu, et al.
Veröffentlicht: (2024)
von: Liu, Ziyu, et al.
Veröffentlicht: (2024)
AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs
von: Chang, Boyu, et al.
Veröffentlicht: (2026)
von: Chang, Boyu, et al.
Veröffentlicht: (2026)
Hierarchical Codec Diffusion for Video-to-Speech Generation
von: Ye, Jiaxin, et al.
Veröffentlicht: (2026)
von: Ye, Jiaxin, et al.
Veröffentlicht: (2026)
GeometryZero: Advancing Geometry Solving via Group Contrastive Policy Optimization
von: Wang, Yikun, et al.
Veröffentlicht: (2025)
von: Wang, Yikun, et al.
Veröffentlicht: (2025)
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
von: Huang, Jen-Tse, et al.
Veröffentlicht: (2025)
von: Huang, Jen-Tse, et al.
Veröffentlicht: (2025)
LoMo: Local Modality Substitution for Deeper Vision-Language Fusion
von: Han, Feng, et al.
Veröffentlicht: (2026)
von: Han, Feng, et al.
Veröffentlicht: (2026)
Revisiting MLLM Token Technology through the Lens of Classical Visual Coding
von: Liu, Jinming, et al.
Veröffentlicht: (2025)
von: Liu, Jinming, et al.
Veröffentlicht: (2025)
When MLLMs Meet Compression Distortion: A Coding Paradigm Tailored to MLLMs
von: Liu, Jinming, et al.
Veröffentlicht: (2025)
von: Liu, Jinming, et al.
Veröffentlicht: (2025)
Affordance Benchmark for MLLMs
von: Wang, Junying, et al.
Veröffentlicht: (2025)
von: Wang, Junying, et al.
Veröffentlicht: (2025)
Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
CurveStream: Boosting Streaming Video Understanding in MLLMs via Curvature-Aware Hierarchical Visual Memory Management
von: Wang, Chao, et al.
Veröffentlicht: (2026)
von: Wang, Chao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
VideoPro: Adaptive Program Reasoning for Long Video Understanding
von: Li, Chenglin, et al.
Veröffentlicht: (2025) -
Interleaved Latent Visual Reasoning with Selective Perceptual Modeling
von: Dong, Shuai, et al.
Veröffentlicht: (2025) -
AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
von: Wang, Xiao, et al.
Veröffentlicht: (2025) -
UniREditBench: A Unified Reasoning-based Image Editing Benchmark
von: Han, Feng, et al.
Veröffentlicht: (2025) -
CodePercept: Code-Grounded Visual STEM Perception for MLLMs
von: Guan, Tongkun, et al.
Veröffentlicht: (2026)