Salvato in:
| Autori principali: | Tong, Yijie, Hou, Yifan, Cui, Shaobo, Bosselut, Antoine, Sachan, Mrinmaya |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2605.30713 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Unveiling the Visual Counting Bottleneck in Vision-Language Models
di: Pang, Xingzhou, et al.
Pubblicazione: (2026)
di: Pang, Xingzhou, et al.
Pubblicazione: (2026)
Diversity-Guided MLP Reduction for Efficient Large Vision Transformers
di: Shen, Chengchao, et al.
Pubblicazione: (2025)
di: Shen, Chengchao, et al.
Pubblicazione: (2025)
Zero-shot image privacy classification with Vision-Language Models
di: Baia, Alina Elena, et al.
Pubblicazione: (2025)
di: Baia, Alina Elena, et al.
Pubblicazione: (2025)
Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages
di: Farina, Matteo, et al.
Pubblicazione: (2025)
di: Farina, Matteo, et al.
Pubblicazione: (2025)
Language Models as Black-Box Optimizers for Vision-Language Models
di: Liu, Shihong, et al.
Pubblicazione: (2023)
di: Liu, Shihong, et al.
Pubblicazione: (2023)
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
di: Sun, Zeyi, et al.
Pubblicazione: (2024)
di: Sun, Zeyi, et al.
Pubblicazione: (2024)
Detecting Content Rating Violations in Android Applications: A Vision-Language Approach
di: Denipitiyage, D., et al.
Pubblicazione: (2025)
di: Denipitiyage, D., et al.
Pubblicazione: (2025)
Test-Time Backdoor Attacks on Multimodal Large Language Models
di: Lu, Dong, et al.
Pubblicazione: (2024)
di: Lu, Dong, et al.
Pubblicazione: (2024)
Understanding the Fine-Grained Knowledge Capabilities of Vision-Language Models
di: Ghosh, Dhruba, et al.
Pubblicazione: (2026)
di: Ghosh, Dhruba, et al.
Pubblicazione: (2026)
RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models
di: Lin, Xiang, et al.
Pubblicazione: (2025)
di: Lin, Xiang, et al.
Pubblicazione: (2025)
Reducing Hallucinations in Vision-Language Models via Latent Space Steering
di: Liu, Sheng, et al.
Pubblicazione: (2024)
di: Liu, Sheng, et al.
Pubblicazione: (2024)
Cross-Modal Coordination Across a Diverse Set of Input Modalities
di: Sánchez, Jorge, et al.
Pubblicazione: (2024)
di: Sánchez, Jorge, et al.
Pubblicazione: (2024)
Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models
di: Zhang, Yabin, et al.
Pubblicazione: (2024)
di: Zhang, Yabin, et al.
Pubblicazione: (2024)
LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
di: Geng, Tiantian, et al.
Pubblicazione: (2024)
di: Geng, Tiantian, et al.
Pubblicazione: (2024)
Bridging Compressed Image Latents and Multimodal Large Language Models
di: Kao, Chia-Hao, et al.
Pubblicazione: (2024)
di: Kao, Chia-Hao, et al.
Pubblicazione: (2024)
Multimodal Transformer With a Low-Computational-Cost Guarantee
di: Park, Sungjin, et al.
Pubblicazione: (2024)
di: Park, Sungjin, et al.
Pubblicazione: (2024)
LinVT: Empower Your Image-level Large Language Model to Understand Videos
di: Gao, Lishuai, et al.
Pubblicazione: (2024)
di: Gao, Lishuai, et al.
Pubblicazione: (2024)
Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
di: Huang, Po-Hsuan, et al.
Pubblicazione: (2024)
di: Huang, Po-Hsuan, et al.
Pubblicazione: (2024)
Discover Your Neighbors: Advanced Stable Test-Time Adaptation in Dynamic World
di: Jiang, Qinting, et al.
Pubblicazione: (2024)
di: Jiang, Qinting, et al.
Pubblicazione: (2024)
From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
di: Chen, Yiming, et al.
Pubblicazione: (2025)
di: Chen, Yiming, et al.
Pubblicazione: (2025)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
di: Zhao, Shuai, et al.
Pubblicazione: (2023)
di: Zhao, Shuai, et al.
Pubblicazione: (2023)
COSMIC: Clique-Oriented Semantic Multi-space Integration for Robust CLIP Test-Time Adaptation
di: Huang, Fanding, et al.
Pubblicazione: (2025)
di: Huang, Fanding, et al.
Pubblicazione: (2025)
Words or Vision: Do Vision-Language Models Have Blind Faith in Text?
di: Deng, Ailin, et al.
Pubblicazione: (2025)
di: Deng, Ailin, et al.
Pubblicazione: (2025)
3DTV: A Feedforward Interpolation Network for Real-Time View Synthesis
di: Schulz, Stefan, et al.
Pubblicazione: (2026)
di: Schulz, Stefan, et al.
Pubblicazione: (2026)
Multimodal Learning on Low-Quality Data with Conformal Predictive Self-Calibration
di: Jiang, Xun, et al.
Pubblicazione: (2026)
di: Jiang, Xun, et al.
Pubblicazione: (2026)
Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video Retrieval
di: Li, Jun, et al.
Pubblicazione: (2026)
di: Li, Jun, et al.
Pubblicazione: (2026)
EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
di: Chen, Junyi, et al.
Pubblicazione: (2023)
di: Chen, Junyi, et al.
Pubblicazione: (2023)
Size Matters: Reconstructing Real-Scale 3D Models from Monocular Images for Food Portion Estimation
di: Vinod, Gautham, et al.
Pubblicazione: (2026)
di: Vinod, Gautham, et al.
Pubblicazione: (2026)
Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
di: Chen, Yang, et al.
Pubblicazione: (2024)
di: Chen, Yang, et al.
Pubblicazione: (2024)
Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
di: Han, Junlin, et al.
Pubblicazione: (2025)
di: Han, Junlin, et al.
Pubblicazione: (2025)
Image Complexity-Aware Adaptive Retrieval for Efficient Vision-Language Models
di: Williams-Lekuona, Mikel, et al.
Pubblicazione: (2025)
di: Williams-Lekuona, Mikel, et al.
Pubblicazione: (2025)
GroMo: Plant Growth Modeling with Multiview Images
di: Bhatt, Ruchi, et al.
Pubblicazione: (2025)
di: Bhatt, Ruchi, et al.
Pubblicazione: (2025)
Do Vision-Language Models Really Understand Visual Language?
di: Hou, Yifan, et al.
Pubblicazione: (2024)
di: Hou, Yifan, et al.
Pubblicazione: (2024)
Improving Long-Text Alignment for Text-to-Image Diffusion Models
di: Liu, Luping, et al.
Pubblicazione: (2024)
di: Liu, Luping, et al.
Pubblicazione: (2024)
Deep Video Codec Control for Vision Models
di: Reich, Christoph, et al.
Pubblicazione: (2023)
di: Reich, Christoph, et al.
Pubblicazione: (2023)
Unveiling Encoder-Free Vision-Language Models
di: Diao, Haiwen, et al.
Pubblicazione: (2024)
di: Diao, Haiwen, et al.
Pubblicazione: (2024)
Operationalizing Fairness in Text-to-Image Models: A Survey of Bias, Fairness Audits and Mitigation Strategies
di: Smith, Megan, et al.
Pubblicazione: (2026)
di: Smith, Megan, et al.
Pubblicazione: (2026)
TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
di: Qu, Leigang, et al.
Pubblicazione: (2025)
di: Qu, Leigang, et al.
Pubblicazione: (2025)
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
di: Deng, Ailin, et al.
Pubblicazione: (2024)
di: Deng, Ailin, et al.
Pubblicazione: (2024)
PlanLLM: Video Procedure Planning with Refinable Large Language Models
di: Yang, Dejie, et al.
Pubblicazione: (2024)
di: Yang, Dejie, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Unveiling the Visual Counting Bottleneck in Vision-Language Models
di: Pang, Xingzhou, et al.
Pubblicazione: (2026) -
Diversity-Guided MLP Reduction for Efficient Large Vision Transformers
di: Shen, Chengchao, et al.
Pubblicazione: (2025) -
Zero-shot image privacy classification with Vision-Language Models
di: Baia, Alina Elena, et al.
Pubblicazione: (2025) -
Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages
di: Farina, Matteo, et al.
Pubblicazione: (2025) -
Language Models as Black-Box Optimizers for Vision-Language Models
di: Liu, Shihong, et al.
Pubblicazione: (2023)