Adaptive Global and Fine-Grained Perceptual Fusion for MLLM Embeddings Compatible with Hard Negative Amplification
Fuente:
arXiv
Guardado en:
| Autores principales: | Hu, Lexiang, Xue, Youze, Li, Dian, Liu, Gang, Lin, Zhouchen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying
por: Xue, Youze, et al.
Publicado: (2025)
por: Xue, Youze, et al.
Publicado: (2025)
DenseMLLM: Standard Multimodal LLMs for Dense Prediction
por: Li, Yi, et al.
Publicado: (2026)
por: Li, Yi, et al.
Publicado: (2026)
Globally Correlation-Aware Hard Negative Generation
por: Peng, Wenjie, et al.
Publicado: (2024)
por: Peng, Wenjie, et al.
Publicado: (2024)
Active Learning for Finely-Categorized Image-Text Retrieval by Selecting Hard Negative Unpaired Samples
por: Jo, Dae Ung, et al.
Publicado: (2024)
por: Jo, Dae Ung, et al.
Publicado: (2024)
FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data
por: Xu, Binqian, et al.
Publicado: (2024)
por: Xu, Binqian, et al.
Publicado: (2024)
Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models
por: Huang, Xin, et al.
Publicado: (2025)
por: Huang, Xin, et al.
Publicado: (2025)
Do MLLMs Exhibit Human-like Perceptual Behaviors? HVSBench: A Benchmark for MLLM Alignment with Human Perceptual Behavior
por: Lin, Jiaying, et al.
Publicado: (2024)
por: Lin, Jiaying, et al.
Publicado: (2024)
FineTec: Fine-Grained Action Recognition Under Temporal Corruption via Skeleton Decomposition and Sequence Completion
por: Shao, Dian, et al.
Publicado: (2025)
por: Shao, Dian, et al.
Publicado: (2025)
Are Face Embeddings Compatible Across Deep Neural Network Models?
por: Rubab, Fizza, et al.
Publicado: (2026)
por: Rubab, Fizza, et al.
Publicado: (2026)
Towards Compatible Fine-tuning for Vision-Language Model Updates
por: Wang, Zhengbo, et al.
Publicado: (2024)
por: Wang, Zhengbo, et al.
Publicado: (2024)
Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization
por: Wang, Zhicheng, et al.
Publicado: (2025)
por: Wang, Zhicheng, et al.
Publicado: (2025)
FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering
por: Zhong, Liangyu, et al.
Publicado: (2025)
por: Zhong, Liangyu, et al.
Publicado: (2025)
Look-Back: Implicit Visual Re-focusing in MLLM Reasoning
por: Yang, Shuo, et al.
Publicado: (2025)
por: Yang, Shuo, et al.
Publicado: (2025)
HIERAMP: Coarse-to-Fine Autoregressive Amplification for Generative Dataset Distillation
por: Zhao, Lin, et al.
Publicado: (2026)
por: Zhao, Lin, et al.
Publicado: (2026)
GEODE: Angle-Adaptive OOD Detection with Universal Scorer Compatibility
por: Abrahao, Bruno
Publicado: (2026)
por: Abrahao, Bruno
Publicado: (2026)
DiFiC: Your Diffusion Model Holds the Secret to Fine-Grained Clustering
por: Yang, Ruohong, et al.
Publicado: (2024)
por: Yang, Ruohong, et al.
Publicado: (2024)
FruitEnsemble: MLLM-Guided Arbitration for Heterogeneous ensemble in Fine-Grained Fruit Recognition
por: Yu, Enhui, et al.
Publicado: (2026)
por: Yu, Enhui, et al.
Publicado: (2026)
Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception
por: Shi, Yuheng, et al.
Publicado: (2025)
por: Shi, Yuheng, et al.
Publicado: (2025)
Unveiling Perceptual Artifacts: A Fine-Grained Benchmark for Interpretable AI-Generated Image Detection
por: Xiao, Yao, et al.
Publicado: (2026)
por: Xiao, Yao, et al.
Publicado: (2026)
Discovering Global False Negatives On the Fly for Self-supervised Contrastive Learning
por: Balmaseda, Vicente, et al.
Publicado: (2025)
por: Balmaseda, Vicente, et al.
Publicado: (2025)
New Dataset and Methods for Fine-Grained Compositional Referring Expression Comprehension via Specialist-MLLM Collaboration
por: Yang, Xuzheng, et al.
Publicado: (2025)
por: Yang, Xuzheng, et al.
Publicado: (2025)
BudgetFusion: Perceptually-Guided Adaptive Diffusion Models
por: Li, Qinchan, et al.
Publicado: (2024)
por: Li, Qinchan, et al.
Publicado: (2024)
Diffusion Model with Perceptual Loss
por: Lin, Shanchuan, et al.
Publicado: (2023)
por: Lin, Shanchuan, et al.
Publicado: (2023)
DAGLFNet: Deep Feature Attention Guided Global and Local Feature Fusion for Pseudo-Image Point Cloud Segmentation
por: Chen, Chuang, et al.
Publicado: (2025)
por: Chen, Chuang, et al.
Publicado: (2025)
Iterative Adversarial Attack on Image-guided Story Ending Generation
por: Wang, Youze, et al.
Publicado: (2023)
por: Wang, Youze, et al.
Publicado: (2023)
Neighborhood-Adaptive Generalized Linear Graph Embedding with Latent Pattern Mining
por: Peng, S., et al.
Publicado: (2025)
por: Peng, S., et al.
Publicado: (2025)
Perceptual Multi-Exposure Fusion
por: Liu, Xiaoning
Publicado: (2022)
por: Liu, Xiaoning
Publicado: (2022)
ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and Reasoning
por: Zhu, Wenjie, et al.
Publicado: (2025)
por: Zhu, Wenjie, et al.
Publicado: (2025)
DRIFT: Transferring Reasoning Priors for Efficient MLLM Fine-Tuning
por: Huang, Chao, et al.
Publicado: (2025)
por: Huang, Chao, et al.
Publicado: (2025)
Exploiting Fine-Grained Prototype Distribution for Boosting Unsupervised Class Incremental Learning
por: Liu, Jiaming, et al.
Publicado: (2024)
por: Liu, Jiaming, et al.
Publicado: (2024)
Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question Answering
por: Zou, Yuanhao, et al.
Publicado: (2025)
por: Zou, Yuanhao, et al.
Publicado: (2025)
Let's Split Up: Zero-Shot Classifier Edits for Fine-Grained Video Understanding
por: Liu, Kaiting, et al.
Publicado: (2026)
por: Liu, Kaiting, et al.
Publicado: (2026)
Adaptive Transformer Attention and Multi-Scale Fusion for Spine 3D Segmentation
por: Xiang, Yanlin, et al.
Publicado: (2025)
por: Xiang, Yanlin, et al.
Publicado: (2025)
MMGeoLM: Hard Negative Contrastive Learning for Fine-Grained Geometric Understanding in Large Multimodal Models
por: Sun, Kai, et al.
Publicado: (2025)
por: Sun, Kai, et al.
Publicado: (2025)
HNC: Leveraging Hard Negative Captions towards Models with Fine-Grained Visual-Linguistic Comprehension Capabilities
por: Dönmez, Esra, et al.
Publicado: (2026)
por: Dönmez, Esra, et al.
Publicado: (2026)
ADEM-VL: Adaptive and Embedded Fusion for Efficient Vision-Language Tuning
por: Hao, Zhiwei, et al.
Publicado: (2024)
por: Hao, Zhiwei, et al.
Publicado: (2024)
Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples
por: Wang, Yeyuan, et al.
Publicado: (2024)
por: Wang, Yeyuan, et al.
Publicado: (2024)
No Hard Negatives Required: Concept Centric Learning Leads to Compositionality without Degrading Zero-shot Capabilities of Contrastive Models
por: Pham, Hai X., et al.
Publicado: (2026)
por: Pham, Hai X., et al.
Publicado: (2026)
Negative Label Guided OOD Detection with Pretrained Vision-Language Models
por: Jiang, Xue, et al.
Publicado: (2024)
por: Jiang, Xue, et al.
Publicado: (2024)
Class-Incremental Learning with CLIP: Adaptive Representation Adjustment and Parameter Fusion
por: Huang, Linlan, et al.
Publicado: (2024)
por: Huang, Linlan, et al.
Publicado: (2024)
Ejemplares similares
-
Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying
por: Xue, Youze, et al.
Publicado: (2025) -
DenseMLLM: Standard Multimodal LLMs for Dense Prediction
por: Li, Yi, et al.
Publicado: (2026) -
Globally Correlation-Aware Hard Negative Generation
por: Peng, Wenjie, et al.
Publicado: (2024) -
Active Learning for Finely-Categorized Image-Text Retrieval by Selecting Hard Negative Unpaired Samples
por: Jo, Dae Ung, et al.
Publicado: (2024) -
FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data
por: Xu, Binqian, et al.
Publicado: (2024)