HiMix: Reducing Computational Complexity in Large Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Xuange, Li, Dengjie, Liu, Bo, Bao, Zenghao, Zhou, Yao, Yang, Baisong, Liu, Zhongying, Zhong, Yujie, Zhao, Zheng, Yuan, Tongtong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimizing Singular Spectrum for Large Language Model Compression
von: Li, Dengjie, et al.
Veröffentlicht: (2025)
von: Li, Dengjie, et al.
Veröffentlicht: (2025)
Manga Generation via Layout-controllable Diffusion
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
HiMix: Hierarchical Artifact-aware Mixup for Generalized Synthetic Image Detection
von: Zhou, Shuchang, et al.
Veröffentlicht: (2026)
von: Zhou, Shuchang, et al.
Veröffentlicht: (2026)
SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models
von: Liu, Bo, et al.
Veröffentlicht: (2025)
von: Liu, Bo, et al.
Veröffentlicht: (2025)
LinVT: Empower Your Image-level Large Language Model to Understand Videos
von: Gao, Lishuai, et al.
Veröffentlicht: (2024)
von: Gao, Lishuai, et al.
Veröffentlicht: (2024)
Advancing Visual Large Language Model for Multi-granular Versatile Perception
von: Xiang, Wentao, et al.
Veröffentlicht: (2025)
von: Xiang, Wentao, et al.
Veröffentlicht: (2025)
HiLa: Hierarchical Vision-Language Collaboration for Cancer Survival Prediction
von: Cui, Jiaqi, et al.
Veröffentlicht: (2025)
von: Cui, Jiaqi, et al.
Veröffentlicht: (2025)
TASR: Timestep-Aware Diffusion Model for Image Super-Resolution
von: Lin, Qinwei, et al.
Veröffentlicht: (2024)
von: Lin, Qinwei, et al.
Veröffentlicht: (2024)
Beyond Factual Correctness: Mitigating Preference-Inconsistent Explanations in Explainable Recommendation
von: Wang, Chengkai, et al.
Veröffentlicht: (2026)
von: Wang, Chengkai, et al.
Veröffentlicht: (2026)
Federated Cross-Domain Click-Through Rate Prediction With Large Language Model Augmentation
von: Qin, Jiangcheng, et al.
Veröffentlicht: (2025)
von: Qin, Jiangcheng, et al.
Veröffentlicht: (2025)
StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems
von: Ye, Jinhui, et al.
Veröffentlicht: (2026)
von: Ye, Jinhui, et al.
Veröffentlicht: (2026)
Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models
von: Xia, Xingyu, et al.
Veröffentlicht: (2026)
von: Xia, Xingyu, et al.
Veröffentlicht: (2026)
HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
von: Huang, Runhui, et al.
Veröffentlicht: (2024)
von: Huang, Runhui, et al.
Veröffentlicht: (2024)
CoFusion: Multispectral and Hyperspectral Image Fusion via Spectral Coordinate Attention
von: Li, Baisong
Veröffentlicht: (2026)
von: Li, Baisong
Veröffentlicht: (2026)
3DArticCyclists: Generating Synthetic Articulated 8D Pose-Controllable Cyclist Data for Computer Vision Applications
von: Corral-Soto, Eduardo R., et al.
Veröffentlicht: (2024)
von: Corral-Soto, Eduardo R., et al.
Veröffentlicht: (2024)
RFSR: Improving ISR Diffusion Models via Reward Feedback Learning
von: Sun, Xiaopeng, et al.
Veröffentlicht: (2024)
von: Sun, Xiaopeng, et al.
Veröffentlicht: (2024)
HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment
von: Wu, Ruijia, et al.
Veröffentlicht: (2025)
von: Wu, Ruijia, et al.
Veröffentlicht: (2025)
HyperSeg: Towards Universal Visual Segmentation with Large Language Model
von: Wei, Cong, et al.
Veröffentlicht: (2024)
von: Wei, Cong, et al.
Veröffentlicht: (2024)
Adversarial Alignment: Ensuring Value Consistency in Large Language Models for Sensitive Domains
von: Gao, Yuan, et al.
Veröffentlicht: (2026)
von: Gao, Yuan, et al.
Veröffentlicht: (2026)
InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models
von: Wei, Cong, et al.
Veröffentlicht: (2024)
von: Wei, Cong, et al.
Veröffentlicht: (2024)
Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation
von: Zhang, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaofeng, et al.
Veröffentlicht: (2024)
Activatable Janus Nanoparticles for Precise NIR‐II Bioimaging and Synergistic Cancer Therapy
von: Jiasheng Bao, et al.
Veröffentlicht: (2024)
von: Jiasheng Bao, et al.
Veröffentlicht: (2024)
WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models
von: Zhou, Runjie, et al.
Veröffentlicht: (2026)
von: Zhou, Runjie, et al.
Veröffentlicht: (2026)
HiCoGen: Hierarchical Compositional Text-to-Image Generation in Diffusion Models via Reinforcement Learning
von: Yang, Hongji, et al.
Veröffentlicht: (2025)
von: Yang, Hongji, et al.
Veröffentlicht: (2025)
HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies
von: Du, Zhiying, et al.
Veröffentlicht: (2025)
von: Du, Zhiying, et al.
Veröffentlicht: (2025)
Noisy Test-Time Adaptation in Vision-Language Models
von: Cao, Chentao, et al.
Veröffentlicht: (2025)
von: Cao, Chentao, et al.
Veröffentlicht: (2025)
OThink-SRR1: Search, Refine and Reasoning with Reinforced Learning for Large Language Models
von: Liang, Haijian, et al.
Veröffentlicht: (2026)
von: Liang, Haijian, et al.
Veröffentlicht: (2026)
Ada-K Routing: Boosting the Efficiency of MoE-based LLMs
von: Yue, Tongtian, et al.
Veröffentlicht: (2024)
von: Yue, Tongtian, et al.
Veröffentlicht: (2024)
HiKonv: Maximizing the Throughput of Quantized Convolution With Novel Bit-wise Management and Computation
von: Chen, Yao, et al.
Veröffentlicht: (2022)
von: Chen, Yao, et al.
Veröffentlicht: (2022)
HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models
von: Liu, Jizhihui, et al.
Veröffentlicht: (2025)
von: Liu, Jizhihui, et al.
Veröffentlicht: (2025)
Micro-Macro Retrieval: Reducing Long-Form Hallucination in Large Language Models
von: Feng, Yujie, et al.
Veröffentlicht: (2026)
von: Feng, Yujie, et al.
Veröffentlicht: (2026)
PLPP: Prompt Learning with Perplexity Is Self-Distillation for Vision-Language Models
von: Liu, Biao, et al.
Veröffentlicht: (2024)
von: Liu, Biao, et al.
Veröffentlicht: (2024)
HiPath: Hierarchical Vision-Language Alignment for Structured Pathology Report Prediction
von: Yuan, Ruicheng, et al.
Veröffentlicht: (2026)
von: Yuan, Ruicheng, et al.
Veröffentlicht: (2026)
Deciphering the Complexity of Step Profiles on Vicinal Si(001) Surfaces Through Multiscale Simulations
von: Pai Li, et al.
Veröffentlicht: (2025)
von: Pai Li, et al.
Veröffentlicht: (2025)
HiPER: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model Agents
von: Peng, Jiangweizhi, et al.
Veröffentlicht: (2026)
von: Peng, Jiangweizhi, et al.
Veröffentlicht: (2026)
The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering
von: Li, Zhuowei, et al.
Veröffentlicht: (2025)
von: Li, Zhuowei, et al.
Veröffentlicht: (2025)
Data Mixing Optimization for Supervised Fine-Tuning of Large Language Models
von: Li, Yuan, et al.
Veröffentlicht: (2025)
von: Li, Yuan, et al.
Veröffentlicht: (2025)
HiDe-LLaVA: Hierarchical Decoupling for Continual Instruction Tuning of Multimodal Large Language Model
von: Guo, Haiyang, et al.
Veröffentlicht: (2025)
von: Guo, Haiyang, et al.
Veröffentlicht: (2025)
A Survey of Vibe Coding with Large Language Models
von: Ge, Yuyao, et al.
Veröffentlicht: (2025)
von: Ge, Yuyao, et al.
Veröffentlicht: (2025)
Capture Global Feature Statistics for One-Shot Federated Learning
von: Guan, Zenghao, et al.
Veröffentlicht: (2025)
von: Guan, Zenghao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Optimizing Singular Spectrum for Large Language Model Compression
von: Li, Dengjie, et al.
Veröffentlicht: (2025) -
Manga Generation via Layout-controllable Diffusion
von: Chen, Siyu, et al.
Veröffentlicht: (2024) -
HiMix: Hierarchical Artifact-aware Mixup for Generalized Synthetic Image Detection
von: Zhou, Shuchang, et al.
Veröffentlicht: (2026) -
SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models
von: Liu, Bo, et al.
Veröffentlicht: (2025) -
LinVT: Empower Your Image-level Large Language Model to Understand Videos
von: Gao, Lishuai, et al.
Veröffentlicht: (2024)