Dynamic Pyramid Network for Efficient Multimodal Large Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ai, Hao, Wang, Kunyi, Wang, Zezhou, Lu, Hao, Tian, Jin, Luo, Yaxin, Xing, Peng, Huang, Jen-Yuan, Li, Huaxia, luo, Gen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
InstantIR: Blind Image Restoration with Instant Generative Reference
von: Huang, Jen-Yuan, et al.
Veröffentlicht: (2024)
von: Huang, Jen-Yuan, et al.
Veröffentlicht: (2024)
Disambiguate Entity Matching using Large Language Models through Relation Discovery
von: Huang, Zezhou
Veröffentlicht: (2024)
von: Huang, Zezhou
Veröffentlicht: (2024)
$γ-$MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models
von: Luo, Yaxin, et al.
Veröffentlicht: (2024)
von: Luo, Yaxin, et al.
Veröffentlicht: (2024)
Understanding Spatio-Temporal Relations in Human-Object Interaction using Pyramid Graph Convolutional Network
von: Xing, Hao, et al.
Veröffentlicht: (2024)
von: Xing, Hao, et al.
Veröffentlicht: (2024)
PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
von: Xing, Long, et al.
Veröffentlicht: (2024)
von: Xing, Long, et al.
Veröffentlicht: (2024)
Optimizing Cross-Client Domain Coverage for Federated Instruction Tuning of Large Language Models
von: Wang, Zezhou, et al.
Veröffentlicht: (2024)
von: Wang, Zezhou, et al.
Veröffentlicht: (2024)
Cocoon: Semantic Table Profiling Using Large Language Models
von: Huang, Zezhou, et al.
Veröffentlicht: (2024)
von: Huang, Zezhou, et al.
Veröffentlicht: (2024)
HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices
von: HyperAI Team, et al.
Veröffentlicht: (2025)
von: HyperAI Team, et al.
Veröffentlicht: (2025)
Dual Traits in Probabilistic Reasoning of Large Language Models
von: Li, Shenxiong, et al.
Veröffentlicht: (2024)
von: Li, Shenxiong, et al.
Veröffentlicht: (2024)
Parameter-Inverted Image Pyramid Networks
von: Zhu, Xizhou, et al.
Veröffentlicht: (2024)
von: Zhu, Xizhou, et al.
Veröffentlicht: (2024)
Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
von: Li, Chenxu, et al.
Veröffentlicht: (2025)
von: Li, Chenxu, et al.
Veröffentlicht: (2025)
Cognitive Memory in Large Language Models
von: Shan, Lianlei, et al.
Veröffentlicht: (2025)
von: Shan, Lianlei, et al.
Veröffentlicht: (2025)
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
von: Luo, Gen, et al.
Veröffentlicht: (2025)
von: Luo, Gen, et al.
Veröffentlicht: (2025)
Efficient Pyramid Channel Attention Network for Pathological Myopia Recognition
von: Zhang, Xiaoqing, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaoqing, et al.
Veröffentlicht: (2023)
SDEval: Safety Dynamic Evaluation for Multimodal Large Language Models
von: Wang, Hanqing, et al.
Veröffentlicht: (2025)
von: Wang, Hanqing, et al.
Veröffentlicht: (2025)
Data Cleaning Using Large Language Models
von: Zhang, Shuo, et al.
Veröffentlicht: (2024)
von: Zhang, Shuo, et al.
Veröffentlicht: (2024)
Bridging Compressed Image Latents and Multimodal Large Language Models
von: Kao, Chia-Hao, et al.
Veröffentlicht: (2024)
von: Kao, Chia-Hao, et al.
Veröffentlicht: (2024)
MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
von: Chen, Feilong, et al.
Veröffentlicht: (2025)
von: Chen, Feilong, et al.
Veröffentlicht: (2025)
Delta -- Contrastive Decoding Mitigates Text Hallucinations in Large Language Models
von: Huang, Cheng Peng, et al.
Veröffentlicht: (2025)
von: Huang, Cheng Peng, et al.
Veröffentlicht: (2025)
Pyramidal Adaptive Cross-Gating for Multimodal Detection
von: Gu, Zidong, et al.
Veröffentlicht: (2025)
von: Gu, Zidong, et al.
Veröffentlicht: (2025)
Pyramidal Flow Matching for Efficient Video Generative Modeling
von: Jin, Yang, et al.
Veröffentlicht: (2024)
von: Jin, Yang, et al.
Veröffentlicht: (2024)
Pyramid-Driven Alignment: Pyramid Principle Guided Integration of Large Language Models and Knowledge Graphs
von: Sun, Lei, et al.
Veröffentlicht: (2024)
von: Sun, Lei, et al.
Veröffentlicht: (2024)
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model
von: Zhang, Yuting, et al.
Veröffentlicht: (2025)
von: Zhang, Yuting, et al.
Veröffentlicht: (2025)
NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
von: Tian, Changyao, et al.
Veröffentlicht: (2025)
von: Tian, Changyao, et al.
Veröffentlicht: (2025)
Elite360M: Efficient 360 Multi-task Learning via Bi-projection Fusion and Cross-task Collaboration
von: Ai, Hao, et al.
Veröffentlicht: (2024)
von: Ai, Hao, et al.
Veröffentlicht: (2024)
Elite360D: Towards Efficient 360 Depth Estimation via Semantic- and Distance-Aware Bi-Projection Fusion
von: Ai, Hao, et al.
Veröffentlicht: (2024)
von: Ai, Hao, et al.
Veröffentlicht: (2024)
Open Vocabulary Monocular 3D Object Detection
von: Yao, Jin, et al.
Veröffentlicht: (2024)
von: Yao, Jin, et al.
Veröffentlicht: (2024)
CompeteAI: Understanding the Competition Dynamics in Large Language Model-based Agents
von: Zhao, Qinlin, et al.
Veröffentlicht: (2023)
von: Zhao, Qinlin, et al.
Veröffentlicht: (2023)
Evaluating the Rehabilitation Needs of Stroke Patients in China: A Trend Analysis From 1990 to 2019
von: Peng Zhao, et al.
Veröffentlicht: (2025)
von: Peng Zhao, et al.
Veröffentlicht: (2025)
Equity in strategic exchange
von: Liu, Peng, et al.
Veröffentlicht: (2025)
von: Liu, Peng, et al.
Veröffentlicht: (2025)
A Survey on Benchmarks of Multimodal Large Language Models
von: Li, Jian, et al.
Veröffentlicht: (2024)
von: Li, Jian, et al.
Veröffentlicht: (2024)
MGKAN: Predicting Asymmetric Drug-Drug Interactions via a Multimodal Graph Kolmogorov-Arnold Network
von: Fan, Kunyi, et al.
Veröffentlicht: (2026)
von: Fan, Kunyi, et al.
Veröffentlicht: (2026)
InstantStyle-Plus: Style Transfer with Content-Preserving in Text-to-Image Generation
von: Wang, Haofan, et al.
Veröffentlicht: (2024)
von: Wang, Haofan, et al.
Veröffentlicht: (2024)
Towards Evaluating Proactive Risk Awareness of Multimodal Language Models
von: Yuan, Youliang, et al.
Veröffentlicht: (2025)
von: Yuan, Youliang, et al.
Veröffentlicht: (2025)
Diffusion Dynamics Models with Generative State Estimation for Cloth Manipulation
von: Tian, Tongxuan, et al.
Veröffentlicht: (2025)
von: Tian, Tongxuan, et al.
Veröffentlicht: (2025)
Pre-trained Visual Dynamics Representations for Efficient Policy Learning
von: Luo, Hao, et al.
Veröffentlicht: (2024)
von: Luo, Hao, et al.
Veröffentlicht: (2024)
SEA: Low-Resource Safety Alignment for Multimodal Large Language Models via Synthetic Embeddings
von: Lu, Weikai, et al.
Veröffentlicht: (2025)
von: Lu, Weikai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025) -
InstantIR: Blind Image Restoration with Instant Generative Reference
von: Huang, Jen-Yuan, et al.
Veröffentlicht: (2024) -
Disambiguate Entity Matching using Large Language Models through Relation Discovery
von: Huang, Zezhou
Veröffentlicht: (2024) -
$γ-$MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models
von: Luo, Yaxin, et al.
Veröffentlicht: (2024) -
Understanding Spatio-Temporal Relations in Human-Object Interaction using Pyramid Graph Convolutional Network
von: Xing, Hao, et al.
Veröffentlicht: (2024)