Efficient Multi-modal Long Context Learning for Training-free Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Zehong, Zhang, Shiliang, Wei, Longhui, Tian, Qi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OVMR: Open-Vocabulary Recognition with Multi-Modal References
by: Ma, Zehong, et al.
Published: (2024)
by: Ma, Zehong, et al.
Published: (2024)
MagCache: Fast Video Generation with Magnitude-Aware Cache
by: Ma, Zehong, et al.
Published: (2025)
by: Ma, Zehong, et al.
Published: (2025)
DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation
by: Ma, Zehong, et al.
Published: (2025)
by: Ma, Zehong, et al.
Published: (2025)
Multi-modal Reference Learning for Fine-grained Text-to-Image Retrieval
by: Ma, Zehong, et al.
Published: (2025)
by: Ma, Zehong, et al.
Published: (2025)
PixelGen: Improving Pixel Diffusion with Perceptual Supervision
by: Ma, Zehong, et al.
Published: (2026)
by: Ma, Zehong, et al.
Published: (2026)
Decoupled Contrastive Learning for Long-Tailed Recognition
by: Xuan, Shiyu, et al.
Published: (2024)
by: Xuan, Shiyu, et al.
Published: (2024)
Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts
by: He, Xin, et al.
Published: (2025)
by: He, Xin, et al.
Published: (2025)
EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
by: He, Xin, et al.
Published: (2025)
by: He, Xin, et al.
Published: (2025)
Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs
by: Xuan, Shiyu, et al.
Published: (2023)
by: Xuan, Shiyu, et al.
Published: (2023)
AnyPhoto: Multi-Person Identity Preserving Image Generation with ID Adaptive Modulation on Location Canvas
by: Yuan, Longhui
Published: (2026)
by: Yuan, Longhui
Published: (2026)
Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models
by: He, Xin, et al.
Published: (2024)
by: He, Xin, et al.
Published: (2024)
ZeroSense:How Vision matters in Long Context Compression
by: Gao, Yonghan, et al.
Published: (2026)
by: Gao, Yonghan, et al.
Published: (2026)
Unlocking the Potential: Multi-task Deep Learning for Spaceborne Quantitative Monitoring of Fugitive Methane Plumes
by: Si, Guoxin, et al.
Published: (2024)
by: Si, Guoxin, et al.
Published: (2024)
HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
by: Peng, Zelin, et al.
Published: (2025)
by: Peng, Zelin, et al.
Published: (2025)
Efficient Large Multi-modal Models via Visual Context Compression
by: Chen, Jieneng, et al.
Published: (2024)
by: Chen, Jieneng, et al.
Published: (2024)
Feed-Forward 3D Gaussian Splatting Compression with Long-Context Modeling
by: Liu, Zhening, et al.
Published: (2025)
by: Liu, Zhening, et al.
Published: (2025)
PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards
by: Wang, Shulei, et al.
Published: (2025)
by: Wang, Shulei, et al.
Published: (2025)
Multi-modal Generation via Cross-Modal In-Context Learning
by: Kumar, Amandeep, et al.
Published: (2024)
by: Kumar, Amandeep, et al.
Published: (2024)
Multi-modal In-Context Learning Makes an Ego-evolving Scene Text Recognizer
by: Zhao, Zhen, et al.
Published: (2023)
by: Zhao, Zhen, et al.
Published: (2023)
Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuracy
by: Shen, Yunhang, et al.
Published: (2025)
by: Shen, Yunhang, et al.
Published: (2025)
Evolved Hierarchical Masking for Self-Supervised Learning
by: Feng, Zhanzhou, et al.
Published: (2025)
by: Feng, Zhanzhou, et al.
Published: (2025)
Training-free and Adaptive Sparse Attention for Efficient Long Video Generation
by: Xia, Yifei, et al.
Published: (2025)
by: Xia, Yifei, et al.
Published: (2025)
Learning What is Worth Learning: Active and Sequential Domain Adaptation for Multi-modal Gross Tumor Volume Segmentation
by: Yang, Jingyun, et al.
Published: (2025)
by: Yang, Jingyun, et al.
Published: (2025)
Training-Free Image Editing with Visual Context Integration and Concept Alignment
by: Song, Rui, et al.
Published: (2026)
by: Song, Rui, et al.
Published: (2026)
Image Anything: Towards Reasoning-coherent and Training-free Multi-modal Image Generation
by: Lyu, Yuanhuiyi, et al.
Published: (2024)
by: Lyu, Yuanhuiyi, et al.
Published: (2024)
Few-shot Adaptation of Multi-modal Foundation Models: A Survey
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
Visual Grounding with Multi-modal Conditional Adaptation
by: Yao, Ruilin, et al.
Published: (2024)
by: Yao, Ruilin, et al.
Published: (2024)
Training-Free Model Merging for Multi-target Domain Adaptation
by: Li, Wenyi, et al.
Published: (2024)
by: Li, Wenyi, et al.
Published: (2024)
LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models
by: Qiu, Han, et al.
Published: (2024)
by: Qiu, Han, et al.
Published: (2024)
Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning
by: Chen, Cheng, et al.
Published: (2025)
by: Chen, Cheng, et al.
Published: (2025)
VideoMerge: Towards Training-free Long Video Generation
by: Zhang, Siyang, et al.
Published: (2025)
by: Zhang, Siyang, et al.
Published: (2025)
EventDance: Unsupervised Source-free Cross-modal Adaptation for Event-based Object Recognition
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
Human-AI Collaborative Multi-modal Multi-rater Learning for Endometriosis Diagnosis
by: Wang, Hu, et al.
Published: (2024)
by: Wang, Hu, et al.
Published: (2024)
Boosting Segment Anything Model Towards Open-Vocabulary Learning
by: Han, Xumeng, et al.
Published: (2023)
by: Han, Xumeng, et al.
Published: (2023)
Efficient and Context-Aware Label Propagation for Zero-/Few-Shot Training-Free Adaptation of Vision-Language Model
by: Li, Yushu, et al.
Published: (2024)
by: Li, Yushu, et al.
Published: (2024)
Learnable Cross-modal Knowledge Distillation for Multi-modal Learning with Missing Modality
by: Wang, Hu, et al.
Published: (2023)
by: Wang, Hu, et al.
Published: (2023)
Anti-Forgetting Adaptation for Unsupervised Person Re-identification
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
by: Chow, Wei, et al.
Published: (2024)
by: Chow, Wei, et al.
Published: (2024)
Coherent and Multi-modality Image Inpainting via Latent Space Optimization
by: Pan, Lingzhi, et al.
Published: (2024)
by: Pan, Lingzhi, et al.
Published: (2024)
ZeroSmooth: Training-free Diffuser Adaptation for High Frame Rate Video Generation
by: Yang, Shaoshu, et al.
Published: (2024)
by: Yang, Shaoshu, et al.
Published: (2024)
Similar Items
-
OVMR: Open-Vocabulary Recognition with Multi-Modal References
by: Ma, Zehong, et al.
Published: (2024) -
MagCache: Fast Video Generation with Magnitude-Aware Cache
by: Ma, Zehong, et al.
Published: (2025) -
DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation
by: Ma, Zehong, et al.
Published: (2025) -
Multi-modal Reference Learning for Fine-grained Text-to-Image Retrieval
by: Ma, Zehong, et al.
Published: (2025) -
PixelGen: Improving Pixel Diffusion with Perceptual Supervision
by: Ma, Zehong, et al.
Published: (2026)