Is Less More? Exploring Token Condensation as Training-free Test-time Adaptation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zixin, Gong, Dong, Wang, Sen, Huang, Zi, Luo, Yadan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DPO: Dual-Perturbation Optimization for Test-time Adaptation in 3D Object Detection
von: Chen, Zhuoxiao, et al.
Veröffentlicht: (2024)
von: Chen, Zhuoxiao, et al.
Veröffentlicht: (2024)
Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
von: Li, Yulin, et al.
Veröffentlicht: (2025)
von: Li, Yulin, et al.
Veröffentlicht: (2025)
Open-CRB: Towards Open World Active Learning for 3D Object Detection
von: Chen, Zhuoxiao, et al.
Veröffentlicht: (2023)
von: Chen, Zhuoxiao, et al.
Veröffentlicht: (2023)
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
SCORE: Soft Label Compression-Centric Dataset Condensation via Coding Rate Optimization
von: Yuan, Bowen, et al.
Veröffentlicht: (2025)
von: Yuan, Bowen, et al.
Veröffentlicht: (2025)
Exploring the Potential of Encoder-free Architectures in 3D LMMs
von: Tang, Yiwen, et al.
Veröffentlicht: (2025)
von: Tang, Yiwen, et al.
Veröffentlicht: (2025)
MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
von: Wei, Lai, et al.
Veröffentlicht: (2023)
von: Wei, Lai, et al.
Veröffentlicht: (2023)
Why Only Text: Empowering Vision-and-Language Navigation with Multi-modal Prompts
von: Hong, Haodong, et al.
Veröffentlicht: (2024)
von: Hong, Haodong, et al.
Veröffentlicht: (2024)
General Scene Adaptation for Vision-and-Language Navigation
von: Hong, Haodong, et al.
Veröffentlicht: (2025)
von: Hong, Haodong, et al.
Veröffentlicht: (2025)
More Text, Less Point: Towards 3D Data-Efficient Point-Language Understanding
von: Tang, Yuan, et al.
Veröffentlicht: (2024)
von: Tang, Yuan, et al.
Veröffentlicht: (2024)
Find n' Propagate: Open-Vocabulary 3D Object Detection in Urban Environments
von: Etchegaray, Djamahl, et al.
Veröffentlicht: (2024)
von: Etchegaray, Djamahl, et al.
Veröffentlicht: (2024)
Box-QAymo: Box-Referring VQA Dataset for Autonomous Driving
von: Etchegaray, Djamahl, et al.
Veröffentlicht: (2025)
von: Etchegaray, Djamahl, et al.
Veröffentlicht: (2025)
VLMQ: Token Saliency-Driven Post-Training Quantization for Vision-language Models
von: Xue, Yufei, et al.
Veröffentlicht: (2025)
von: Xue, Yufei, et al.
Veröffentlicht: (2025)
OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation
von: Chen, Zhuoxiao, et al.
Veröffentlicht: (2025)
von: Chen, Zhuoxiao, et al.
Veröffentlicht: (2025)
CodeMerge: Codebook-Guided Model Merging for Robust Test-Time Adaptation in Autonomous Driving
von: Yang, Huitong, et al.
Veröffentlicht: (2025)
von: Yang, Huitong, et al.
Veröffentlicht: (2025)
Space Rotation with Basis Transformation for Training-free Test-Time Adaptation
von: Ding, Chenhao, et al.
Veröffentlicht: (2025)
von: Ding, Chenhao, et al.
Veröffentlicht: (2025)
Sparser Block-Sparse Attention via Token Permutation
von: Wang, Xinghao, et al.
Veröffentlicht: (2025)
von: Wang, Xinghao, et al.
Veröffentlicht: (2025)
On the Limits of Token Reduction for Efficient Unified Vision Language Training
von: Chen, Siyi, et al.
Veröffentlicht: (2026)
von: Chen, Siyi, et al.
Veröffentlicht: (2026)
Talk Less, Interact Better: Evaluating In-context Conversational Adaptation in Multimodal LLMs
von: Hua, Yilun, et al.
Veröffentlicht: (2024)
von: Hua, Yilun, et al.
Veröffentlicht: (2024)
Distributed Zero-Shot Learning for Visual Recognition
von: Chen, Zhi, et al.
Veröffentlicht: (2025)
von: Chen, Zhi, et al.
Veröffentlicht: (2025)
Less is More: High-value Data Selection for Visual Instruction Tuning
von: Liu, Zikang, et al.
Veröffentlicht: (2024)
von: Liu, Zikang, et al.
Veröffentlicht: (2024)
Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective
von: Yue, Zihao, et al.
Veröffentlicht: (2024)
von: Yue, Zihao, et al.
Veröffentlicht: (2024)
Historical Test-time Prompt Tuning for Vision Foundation Models
von: Zhang, Jingyi, et al.
Veröffentlicht: (2024)
von: Zhang, Jingyi, et al.
Veröffentlicht: (2024)
FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging
von: Fan, Ziyang, et al.
Veröffentlicht: (2026)
von: Fan, Ziyang, et al.
Veröffentlicht: (2026)
A General and Efficient Training for Transformer via Token Expansion
von: Huang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Huang, Wenxuan, et al.
Veröffentlicht: (2024)
SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs
von: Avogaro, Niccolo, et al.
Veröffentlicht: (2026)
von: Avogaro, Niccolo, et al.
Veröffentlicht: (2026)
Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark
von: Tu, Rong-Cheng, et al.
Veröffentlicht: (2024)
von: Tu, Rong-Cheng, et al.
Veröffentlicht: (2024)
Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
Less Is More? Selective Visual Attention to High-Importance Regions for Multimodal Radiology Summarization
von: Naznin, Mst. Fahmida Sultana, et al.
Veröffentlicht: (2026)
von: Naznin, Mst. Fahmida Sultana, et al.
Veröffentlicht: (2026)
Dynamic Context-oriented Decomposition for Task-aware Low-rank Adaptation with Less Forgetting and Faster Convergence
von: Yang, Yibo, et al.
Veröffentlicht: (2025)
von: Yang, Yibo, et al.
Veröffentlicht: (2025)
On the Interplay of Human-AI Alignment,Fairness, and Performance Trade-offs in Medical Imaging
von: Luo, Haozhe, et al.
Veröffentlicht: (2025)
von: Luo, Haozhe, et al.
Veröffentlicht: (2025)
LIME: Less Is More for MLLM Evaluation
von: Zhu, King, et al.
Veröffentlicht: (2024)
von: Zhu, King, et al.
Veröffentlicht: (2024)
Learning to Align Generative Appearance Priors for Fine-grained Image Retrieval
von: Wang, Shijie, et al.
Veröffentlicht: (2026)
von: Wang, Shijie, et al.
Veröffentlicht: (2026)
Test-Time 3D Occupancy Prediction
von: Zhang, Fengyi, et al.
Veröffentlicht: (2025)
von: Zhang, Fengyi, et al.
Veröffentlicht: (2025)
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
von: Zhong, Yiwu, et al.
Veröffentlicht: (2024)
von: Zhong, Yiwu, et al.
Veröffentlicht: (2024)
Latent Refinement via Flow Matching for Training-free Linear Inverse Problem Solving
von: Askari, Hossein, et al.
Veröffentlicht: (2025)
von: Askari, Hossein, et al.
Veröffentlicht: (2025)
On Epistemic Uncertainty of Visual Tokens for Object Hallucinations in Large Vision-Language Models
von: Seo, Hoigi, et al.
Veröffentlicht: (2025)
von: Seo, Hoigi, et al.
Veröffentlicht: (2025)
PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding
von: Wang, Nan, et al.
Veröffentlicht: (2026)
von: Wang, Nan, et al.
Veröffentlicht: (2026)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
von: Li, Jialu, et al.
Veröffentlicht: (2025)
von: Li, Jialu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DPO: Dual-Perturbation Optimization for Test-time Adaptation in 3D Object Detection
von: Chen, Zhuoxiao, et al.
Veröffentlicht: (2024) -
Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs
von: Song, Dingjie, et al.
Veröffentlicht: (2024) -
Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
von: Li, Yulin, et al.
Veröffentlicht: (2025) -
Open-CRB: Towards Open World Active Learning for 3D Object Detection
von: Chen, Zhuoxiao, et al.
Veröffentlicht: (2023) -
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)