Adaptive Global and Fine-Grained Perceptual Fusion for MLLM Embeddings Compatible with Hard Negative Amplification
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Lexiang, Xue, Youze, Li, Dian, Liu, Gang, Lin, Zhouchen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying
by: Xue, Youze, et al.
Published: (2025)
by: Xue, Youze, et al.
Published: (2025)
DenseMLLM: Standard Multimodal LLMs for Dense Prediction
by: Li, Yi, et al.
Published: (2026)
by: Li, Yi, et al.
Published: (2026)
Globally Correlation-Aware Hard Negative Generation
by: Peng, Wenjie, et al.
Published: (2024)
by: Peng, Wenjie, et al.
Published: (2024)
Active Learning for Finely-Categorized Image-Text Retrieval by Selecting Hard Negative Unpaired Samples
by: Jo, Dae Ung, et al.
Published: (2024)
by: Jo, Dae Ung, et al.
Published: (2024)
FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data
by: Xu, Binqian, et al.
Published: (2024)
by: Xu, Binqian, et al.
Published: (2024)
Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models
by: Huang, Xin, et al.
Published: (2025)
by: Huang, Xin, et al.
Published: (2025)
Do MLLMs Exhibit Human-like Perceptual Behaviors? HVSBench: A Benchmark for MLLM Alignment with Human Perceptual Behavior
by: Lin, Jiaying, et al.
Published: (2024)
by: Lin, Jiaying, et al.
Published: (2024)
FineTec: Fine-Grained Action Recognition Under Temporal Corruption via Skeleton Decomposition and Sequence Completion
by: Shao, Dian, et al.
Published: (2025)
by: Shao, Dian, et al.
Published: (2025)
Are Face Embeddings Compatible Across Deep Neural Network Models?
by: Rubab, Fizza, et al.
Published: (2026)
by: Rubab, Fizza, et al.
Published: (2026)
Towards Compatible Fine-tuning for Vision-Language Model Updates
by: Wang, Zhengbo, et al.
Published: (2024)
by: Wang, Zhengbo, et al.
Published: (2024)
Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization
by: Wang, Zhicheng, et al.
Published: (2025)
by: Wang, Zhicheng, et al.
Published: (2025)
FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering
by: Zhong, Liangyu, et al.
Published: (2025)
by: Zhong, Liangyu, et al.
Published: (2025)
Look-Back: Implicit Visual Re-focusing in MLLM Reasoning
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
HIERAMP: Coarse-to-Fine Autoregressive Amplification for Generative Dataset Distillation
by: Zhao, Lin, et al.
Published: (2026)
by: Zhao, Lin, et al.
Published: (2026)
GEODE: Angle-Adaptive OOD Detection with Universal Scorer Compatibility
by: Abrahao, Bruno
Published: (2026)
by: Abrahao, Bruno
Published: (2026)
DiFiC: Your Diffusion Model Holds the Secret to Fine-Grained Clustering
by: Yang, Ruohong, et al.
Published: (2024)
by: Yang, Ruohong, et al.
Published: (2024)
FruitEnsemble: MLLM-Guided Arbitration for Heterogeneous ensemble in Fine-Grained Fruit Recognition
by: Yu, Enhui, et al.
Published: (2026)
by: Yu, Enhui, et al.
Published: (2026)
Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception
by: Shi, Yuheng, et al.
Published: (2025)
by: Shi, Yuheng, et al.
Published: (2025)
Unveiling Perceptual Artifacts: A Fine-Grained Benchmark for Interpretable AI-Generated Image Detection
by: Xiao, Yao, et al.
Published: (2026)
by: Xiao, Yao, et al.
Published: (2026)
Discovering Global False Negatives On the Fly for Self-supervised Contrastive Learning
by: Balmaseda, Vicente, et al.
Published: (2025)
by: Balmaseda, Vicente, et al.
Published: (2025)
New Dataset and Methods for Fine-Grained Compositional Referring Expression Comprehension via Specialist-MLLM Collaboration
by: Yang, Xuzheng, et al.
Published: (2025)
by: Yang, Xuzheng, et al.
Published: (2025)
BudgetFusion: Perceptually-Guided Adaptive Diffusion Models
by: Li, Qinchan, et al.
Published: (2024)
by: Li, Qinchan, et al.
Published: (2024)
Diffusion Model with Perceptual Loss
by: Lin, Shanchuan, et al.
Published: (2023)
by: Lin, Shanchuan, et al.
Published: (2023)
DAGLFNet: Deep Feature Attention Guided Global and Local Feature Fusion for Pseudo-Image Point Cloud Segmentation
by: Chen, Chuang, et al.
Published: (2025)
by: Chen, Chuang, et al.
Published: (2025)
Iterative Adversarial Attack on Image-guided Story Ending Generation
by: Wang, Youze, et al.
Published: (2023)
by: Wang, Youze, et al.
Published: (2023)
Neighborhood-Adaptive Generalized Linear Graph Embedding with Latent Pattern Mining
by: Peng, S., et al.
Published: (2025)
by: Peng, S., et al.
Published: (2025)
Perceptual Multi-Exposure Fusion
by: Liu, Xiaoning
Published: (2022)
by: Liu, Xiaoning
Published: (2022)
ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and Reasoning
by: Zhu, Wenjie, et al.
Published: (2025)
by: Zhu, Wenjie, et al.
Published: (2025)
DRIFT: Transferring Reasoning Priors for Efficient MLLM Fine-Tuning
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
Exploiting Fine-Grained Prototype Distribution for Boosting Unsupervised Class Incremental Learning
by: Liu, Jiaming, et al.
Published: (2024)
by: Liu, Jiaming, et al.
Published: (2024)
Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question Answering
by: Zou, Yuanhao, et al.
Published: (2025)
by: Zou, Yuanhao, et al.
Published: (2025)
Let's Split Up: Zero-Shot Classifier Edits for Fine-Grained Video Understanding
by: Liu, Kaiting, et al.
Published: (2026)
by: Liu, Kaiting, et al.
Published: (2026)
Adaptive Transformer Attention and Multi-Scale Fusion for Spine 3D Segmentation
by: Xiang, Yanlin, et al.
Published: (2025)
by: Xiang, Yanlin, et al.
Published: (2025)
MMGeoLM: Hard Negative Contrastive Learning for Fine-Grained Geometric Understanding in Large Multimodal Models
by: Sun, Kai, et al.
Published: (2025)
by: Sun, Kai, et al.
Published: (2025)
HNC: Leveraging Hard Negative Captions towards Models with Fine-Grained Visual-Linguistic Comprehension Capabilities
by: Dönmez, Esra, et al.
Published: (2026)
by: Dönmez, Esra, et al.
Published: (2026)
ADEM-VL: Adaptive and Embedded Fusion for Efficient Vision-Language Tuning
by: Hao, Zhiwei, et al.
Published: (2024)
by: Hao, Zhiwei, et al.
Published: (2024)
Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples
by: Wang, Yeyuan, et al.
Published: (2024)
by: Wang, Yeyuan, et al.
Published: (2024)
No Hard Negatives Required: Concept Centric Learning Leads to Compositionality without Degrading Zero-shot Capabilities of Contrastive Models
by: Pham, Hai X., et al.
Published: (2026)
by: Pham, Hai X., et al.
Published: (2026)
Negative Label Guided OOD Detection with Pretrained Vision-Language Models
by: Jiang, Xue, et al.
Published: (2024)
by: Jiang, Xue, et al.
Published: (2024)
Class-Incremental Learning with CLIP: Adaptive Representation Adjustment and Parameter Fusion
by: Huang, Linlan, et al.
Published: (2024)
by: Huang, Linlan, et al.
Published: (2024)
Similar Items
-
Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying
by: Xue, Youze, et al.
Published: (2025) -
DenseMLLM: Standard Multimodal LLMs for Dense Prediction
by: Li, Yi, et al.
Published: (2026) -
Globally Correlation-Aware Hard Negative Generation
by: Peng, Wenjie, et al.
Published: (2024) -
Active Learning for Finely-Categorized Image-Text Retrieval by Selecting Hard Negative Unpaired Samples
by: Jo, Dae Ung, et al.
Published: (2024) -
FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data
by: Xu, Binqian, et al.
Published: (2024)