LatentLLM: Attention-Aware Joint Tensor Compression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Koike-Akino, Toshiaki, Chen, Xiangyu, Liu, Jing, Wang, Ye, Pu, Wang, Brand, Matthew |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SuperLoRA: Parameter-Efficient Unified Adaptation of Multi-Layer Attention Modules
von: Chen, Xiangyu, et al.
Veröffentlicht: (2024)
von: Chen, Xiangyu, et al.
Veröffentlicht: (2024)
Random Channel Ablation for Robust Hand Gesture Classification with Multimodal Biosignals
von: Bimbraw, Keshav, et al.
Veröffentlicht: (2024)
von: Bimbraw, Keshav, et al.
Veröffentlicht: (2024)
Quantum Implicit Neural Compression
von: Fujihashi, Takuya, et al.
Veröffentlicht: (2024)
von: Fujihashi, Takuya, et al.
Veröffentlicht: (2024)
Range Image-Based Implicit Neural Compression for LiDAR Point Clouds
von: Kuwabara, Akihiro, et al.
Veröffentlicht: (2025)
von: Kuwabara, Akihiro, et al.
Veröffentlicht: (2025)
GPT Sonograpy: Hand Gesture Decoding from Forearm Ultrasound Images via VLM
von: Bimbraw, Keshav, et al.
Veröffentlicht: (2024)
von: Bimbraw, Keshav, et al.
Veröffentlicht: (2024)
Adversarial Bi-Regressor Network for Domain Adaptive Regression
von: Xia, Haifeng, et al.
Veröffentlicht: (2022)
von: Xia, Haifeng, et al.
Veröffentlicht: (2022)
TTQ: Activation-Aware Test-Time Quantization to Accelerate LLM Inference On The Fly
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2026)
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2026)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
von: Wang, Ye, et al.
Veröffentlicht: (2026)
von: Wang, Ye, et al.
Veröffentlicht: (2026)
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025)
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025)
TuneComp: Joint Fine-tuning and Compression for Large Foundation Models
von: Chen, Xiangyu, et al.
Veröffentlicht: (2025)
von: Chen, Xiangyu, et al.
Veröffentlicht: (2025)
TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion Models
von: Ni, Haomiao, et al.
Veröffentlicht: (2024)
von: Ni, Haomiao, et al.
Veröffentlicht: (2024)
Exploring User-level Gradient Inversion with a Diffusion Prior
von: Li, Zhuohang, et al.
Veröffentlicht: (2024)
von: Li, Zhuohang, et al.
Veröffentlicht: (2024)
Directional Embedding Smoothing for Robust Vision Language Models
von: Wang, Ye, et al.
Veröffentlicht: (2026)
von: Wang, Ye, et al.
Veröffentlicht: (2026)
TriAttention: Efficient Long Reasoning with Trigonometric KV Compression
von: Mao, Weian, et al.
Veröffentlicht: (2026)
von: Mao, Weian, et al.
Veröffentlicht: (2026)
Latent Visual Reasoning
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
Prism: Spectral-Aware Block-Sparse Attention
von: Wang, Xinghao, et al.
Veröffentlicht: (2026)
von: Wang, Xinghao, et al.
Veröffentlicht: (2026)
LLM-PCGC: Large Language Model-based Point Cloud Geometry Compression
von: Ye, Yuqi, et al.
Veröffentlicht: (2024)
von: Ye, Yuqi, et al.
Veröffentlicht: (2024)
Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
von: Tian, Changyuan, et al.
Veröffentlicht: (2026)
von: Tian, Changyuan, et al.
Veröffentlicht: (2026)
Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
T3C: Test-Time Tensor Compression with Consistency Guarantees
von: Lamaakal, Ismail, et al.
Veröffentlicht: (2026)
von: Lamaakal, Ismail, et al.
Veröffentlicht: (2026)
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
von: Zheng, Duo, et al.
Veröffentlicht: (2024)
von: Zheng, Duo, et al.
Veröffentlicht: (2024)
Interleaved Latent Visual Reasoning with Selective Perceptual Modeling
von: Dong, Shuai, et al.
Veröffentlicht: (2025)
von: Dong, Shuai, et al.
Veröffentlicht: (2025)
Context Cascade Compression: Exploring the Upper Limits of Text Compression
von: Liu, Fanfan, et al.
Veröffentlicht: (2025)
von: Liu, Fanfan, et al.
Veröffentlicht: (2025)
Forest Before Trees: Latent Superposition for Efficient Visual Reasoning
von: Wang, Yubo, et al.
Veröffentlicht: (2026)
von: Wang, Yubo, et al.
Veröffentlicht: (2026)
VAEER: Visual Attention-Inspired Emotion Elicitation Reasoning
von: Man, Fanhang, et al.
Veröffentlicht: (2025)
von: Man, Fanhang, et al.
Veröffentlicht: (2025)
Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR
von: Hu, Ruina, et al.
Veröffentlicht: (2026)
von: Hu, Ruina, et al.
Veröffentlicht: (2026)
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
A Dual Semantic-Aware Recurrent Global-Adaptive Network For Vision-and-Language Navigation
von: Wang, Liuyi, et al.
Veröffentlicht: (2023)
von: Wang, Liuyi, et al.
Veröffentlicht: (2023)
Acceleration Multiple Heads Decoding for LLM via Dynamic Tree Attention
von: Zhang, Zhendong
Veröffentlicht: (2025)
von: Zhang, Zhendong
Veröffentlicht: (2025)
MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch Normalization
von: Guo, Jialong, et al.
Veröffentlicht: (2024)
von: Guo, Jialong, et al.
Veröffentlicht: (2024)
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
von: Jiang, Houcheng, et al.
Veröffentlicht: (2026)
von: Jiang, Houcheng, et al.
Veröffentlicht: (2026)
AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
von: Wang, Junyang, et al.
Veröffentlicht: (2023)
von: Wang, Junyang, et al.
Veröffentlicht: (2023)
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
von: Dai, Yifan, et al.
Veröffentlicht: (2026)
von: Dai, Yifan, et al.
Veröffentlicht: (2026)
MSR-Align: Policy-Grounded Multimodal Alignment for Safety-Aware Reasoning in Vision-Language Models
von: Xia, Yinan, et al.
Veröffentlicht: (2025)
von: Xia, Yinan, et al.
Veröffentlicht: (2025)
Vision-centric Token Compression in Large Language Model
von: Xing, Ling, et al.
Veröffentlicht: (2025)
von: Xing, Ling, et al.
Veröffentlicht: (2025)
LADR: Locality-Aware Dynamic Rescue for Efficient Text-to-Image Generation with Diffusion Large Language Models
von: Wang, Chenglin, et al.
Veröffentlicht: (2026)
von: Wang, Chenglin, et al.
Veröffentlicht: (2026)
Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
von: Ahmed, Mahmoud, et al.
Veröffentlicht: (2024)
von: Ahmed, Mahmoud, et al.
Veröffentlicht: (2024)
Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
von: Jian, Pu, et al.
Veröffentlicht: (2025)
von: Jian, Pu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SuperLoRA: Parameter-Efficient Unified Adaptation of Multi-Layer Attention Modules
von: Chen, Xiangyu, et al.
Veröffentlicht: (2024) -
Random Channel Ablation for Robust Hand Gesture Classification with Multimodal Biosignals
von: Bimbraw, Keshav, et al.
Veröffentlicht: (2024) -
Quantum Implicit Neural Compression
von: Fujihashi, Takuya, et al.
Veröffentlicht: (2024) -
Range Image-Based Implicit Neural Compression for LiDAR Point Clouds
von: Kuwabara, Akihiro, et al.
Veröffentlicht: (2025) -
GPT Sonograpy: Hand Gesture Decoding from Forearm Ultrasound Images via VLM
von: Bimbraw, Keshav, et al.
Veröffentlicht: (2024)