Attention-aware Inference Optimizations for Large Vision-Language Models with Memory-efficient Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ilhan, Fatih, Liu, Gaowen, Kompella, Ramana Rao, Tekin, Selim Furkan, Huang, Tiansheng, Yahn, Zachary, Xu, Yichang, Liu, Ling |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Survey on Large Language Model-Based Game Agents
von: Hu, Sihao, et al.
Veröffentlicht: (2024)
von: Hu, Sihao, et al.
Veröffentlicht: (2024)
A Multi-Agent Perception-Action Alliance for Efficient Long Video Reasoning
von: Xu, Yichang, et al.
Veröffentlicht: (2026)
von: Xu, Yichang, et al.
Veröffentlicht: (2026)
A Neurosymbolic Agent System for Compositional Visual Reasoning
von: Xu, Yichang, et al.
Veröffentlicht: (2025)
von: Xu, Yichang, et al.
Veröffentlicht: (2025)
Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2025)
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2025)
Adversarial Attention Perturbations for Large Object Detection Transformers
von: Yahn, Zachary, et al.
Veröffentlicht: (2025)
von: Yahn, Zachary, et al.
Veröffentlicht: (2025)
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
von: Huang, Tiansheng, et al.
Veröffentlicht: (2025)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2025)
H3Fusion: Helpful, Harmless, Honest Fusion of Aligned LLMs
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2024)
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2024)
Vision Verification Enhanced Fusion of VLMs for Efficient Visual Reasoning
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2026)
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2026)
Personalized Face Privacy Protection From a Single Image
von: Yahn, Zachary, et al.
Veröffentlicht: (2026)
von: Yahn, Zachary, et al.
Veröffentlicht: (2026)
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
FedHFT: Efficient Federated Finetuning with Heterogeneous Edge Clients
von: Ilhan, Fatih, et al.
Veröffentlicht: (2025)
von: Ilhan, Fatih, et al.
Veröffentlicht: (2025)
Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
von: Huang, Tiansheng, et al.
Veröffentlicht: (2025)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2025)
Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
LLM-TOPLA: Efficient LLM Ensemble by Maximising Diversity
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2024)
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2024)
MELT: A Behavioral Trace Dataset for High-Risk Memecoin Launch Detection
von: Hu, Sihao, et al.
Veröffentlicht: (2026)
von: Hu, Sihao, et al.
Veröffentlicht: (2026)
Robust Few-Shot Ensemble Learning with Focal Diversity-Based Pruning
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2024)
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2024)
Large Language Model based Smart Contract Auditing with LLMBugScanner
von: Yuan, Yining, et al.
Veröffentlicht: (2025)
von: Yuan, Yining, et al.
Veröffentlicht: (2025)
Towards Vector Optimization on Low-Dimensional Vector Symbolic Architecture
von: Duan, Shijin, et al.
Veröffentlicht: (2025)
von: Duan, Shijin, et al.
Veröffentlicht: (2025)
Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR
von: Fan, Chongyu, et al.
Veröffentlicht: (2026)
von: Fan, Chongyu, et al.
Veröffentlicht: (2026)
VLA Knows Its Limits
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
Efficient Multitask Dense Predictor via Binarization
von: Shang, Yuzhang, et al.
Veröffentlicht: (2024)
von: Shang, Yuzhang, et al.
Veröffentlicht: (2024)
Targeted Forgetting of Image Subgroups in CLIP Models
von: Zhang, Zeliang, et al.
Veröffentlicht: (2025)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2025)
SwiftNDC: Fast Neural Depth Correction for High-Fidelity 3D Reconstruction
von: Han, Kang, et al.
Veröffentlicht: (2026)
von: Han, Kang, et al.
Veröffentlicht: (2026)
Orientation-anchored Hyper-Gaussian for 4D Reconstruction from Casual Videos
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
Urban Scene Diffusion through Semantic Occupancy Map
von: Zhang, Junge, et al.
Veröffentlicht: (2024)
von: Zhang, Junge, et al.
Veröffentlicht: (2024)
Motion Marionette: Rethinking Rigid Motion Transfer via Prior Guidance
von: Wang, Haoxuan, et al.
Veröffentlicht: (2025)
von: Wang, Haoxuan, et al.
Veröffentlicht: (2025)
Real-Time Robot Execution with Masked Action Chunking
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
Riemannian Multinomial Logistics Regression for SPD Neural Networks
von: Chen, Ziheng, et al.
Veröffentlicht: (2023)
von: Chen, Ziheng, et al.
Veröffentlicht: (2023)
Data Poisoning and Leakage Analysis in Federated Learning
von: Wei, Wenqi, et al.
Veröffentlicht: (2024)
von: Wei, Wenqi, et al.
Veröffentlicht: (2024)
Reversing the Forget-Retain Objectives: An Efficient LLM Unlearning Framework from Logit Difference
von: Ji, Jiabao, et al.
Veröffentlicht: (2024)
von: Ji, Jiabao, et al.
Veröffentlicht: (2024)
B+ANN: A Fast Billion-Scale Disk-based Nearest-Neighbor Index
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2025)
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2025)
Context Bootstrapped Reinforcement Learning
von: Agashe, Saaket, et al.
Veröffentlicht: (2026)
von: Agashe, Saaket, et al.
Veröffentlicht: (2026)
Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
GIFSplat: Generative Prior-Guided Iterative Feed-Forward 3D Gaussian Splatting from Sparse Views
von: Chen, Tianyu, et al.
Veröffentlicht: (2026)
von: Chen, Tianyu, et al.
Veröffentlicht: (2026)
ProDiF: Protecting Domain-Invariant Features to Secure Pre-Trained Models Against Extraction
von: Zhou, Tong, et al.
Veröffentlicht: (2025)
von: Zhou, Tong, et al.
Veröffentlicht: (2025)
Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems
von: Ji, Jiabao, et al.
Veröffentlicht: (2026)
von: Ji, Jiabao, et al.
Veröffentlicht: (2026)
Collision- and Reachability-Aware Multi-Robot Control with Grounded LLM Planners
von: Ji, Jiabao, et al.
Veröffentlicht: (2025)
von: Ji, Jiabao, et al.
Veröffentlicht: (2025)
ConQuER: Modular Architectures for Control and Bias Mitigation in IQP Quantum Generative Models
von: Zou, Xiaocheng, et al.
Veröffentlicht: (2025)
von: Zou, Xiaocheng, et al.
Veröffentlicht: (2025)
From Trojan Horses to Castle Walls: Unveiling Bilateral Data Poisoning Effects in Diffusion Models
von: Pan, Zhuoshi, et al.
Veröffentlicht: (2023)
von: Pan, Zhuoshi, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
A Survey on Large Language Model-Based Game Agents
von: Hu, Sihao, et al.
Veröffentlicht: (2024) -
A Multi-Agent Perception-Action Alliance for Efficient Long Video Reasoning
von: Xu, Yichang, et al.
Veröffentlicht: (2026) -
A Neurosymbolic Agent System for Compositional Visual Reasoning
von: Xu, Yichang, et al.
Veröffentlicht: (2025) -
Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2025) -
Adversarial Attention Perturbations for Large Object Detection Transformers
von: Yahn, Zachary, et al.
Veröffentlicht: (2025)