Guardado en:
| Autores principales: | Suissa, Omri, Ali, Muhiim, Azarbal, Ariana, Shen, Hui, Pradhan, Shekhar |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2503.13021 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Contrastive Learning with Enhanced Abstract Representations using Grouped Loss of Abstract Semantic Supervision
por: Suissa, Omri, et al.
Publicado: (2025)
por: Suissa, Omri, et al.
Publicado: (2025)
Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time
por: Berger, Uri, et al.
Publicado: (2025)
por: Berger, Uri, et al.
Publicado: (2025)
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
por: Beňová, Ivana, et al.
Publicado: (2024)
por: Beňová, Ivana, et al.
Publicado: (2024)
Generalised Medical Phrase Grounding
por: Zhang, Wenjun, et al.
Publicado: (2025)
por: Zhang, Wenjun, et al.
Publicado: (2025)
Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy, Trends and Metrics Analysis
por: Berger, Uri, et al.
Publicado: (2024)
por: Berger, Uri, et al.
Publicado: (2024)
Towards Robust Evaluation of Visual Activity Recognition: Resolving Verb Ambiguity with Sense Clustering
por: Yao, Louie Hong, et al.
Publicado: (2025)
por: Yao, Louie Hong, et al.
Publicado: (2025)
FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing
por: Tang, Zihan, et al.
Publicado: (2026)
por: Tang, Zihan, et al.
Publicado: (2026)
Token Sequence Compression for Efficient Multimodal Computing
por: Omri, Yasmine, et al.
Publicado: (2025)
por: Omri, Yasmine, et al.
Publicado: (2025)
Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning
por: Wei, Xinyu, et al.
Publicado: (2025)
por: Wei, Xinyu, et al.
Publicado: (2025)
FlairGPT: Repurposing LLMs for Interior Designs
por: Littlefair, Gabrielle, et al.
Publicado: (2025)
por: Littlefair, Gabrielle, et al.
Publicado: (2025)
GaussianVision: Vision-Language Alignment from Compressed Image Representations using 2D Gaussian Splatting
por: Omri, Yasmine, et al.
Publicado: (2025)
por: Omri, Yasmine, et al.
Publicado: (2025)
$\mathcal{V}isi\mathcal{P}runer$: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMs
por: Fan, Yingqi, et al.
Publicado: (2025)
por: Fan, Yingqi, et al.
Publicado: (2025)
Multi-Head Attention Driven Dynamic Visual-Semantic Embedding for Enhanced Image-Text Matching
por: Chen, Wenjing
Publicado: (2024)
por: Chen, Wenjing
Publicado: (2024)
Unveiling the Invisible: Captioning Videos with Metaphors
por: Kalarani, Abisek Rajakumar, et al.
Publicado: (2024)
por: Kalarani, Abisek Rajakumar, et al.
Publicado: (2024)
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey
por: Kuang, Jiayi, et al.
Publicado: (2024)
por: Kuang, Jiayi, et al.
Publicado: (2024)
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
por: Xing, Long, et al.
Publicado: (2025)
por: Xing, Long, et al.
Publicado: (2025)
CARINOX: Inference-time Scaling with Category-Aware Reward-based Initial Noise Optimization and Exploration
por: Kasaei, Seyed Amir, et al.
Publicado: (2025)
por: Kasaei, Seyed Amir, et al.
Publicado: (2025)
Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models
por: Wang, Zehao, et al.
Publicado: (2024)
por: Wang, Zehao, et al.
Publicado: (2024)
Few-shot Adaptation to Distribution Shifts By Mixing Source and Target Embeddings
por: Xue, Yihao, et al.
Publicado: (2023)
por: Xue, Yihao, et al.
Publicado: (2023)
RAVEL: Rare Concept Generation and Editing via Graph-driven Relational Guidance
por: Venkatesh, Kavana, et al.
Publicado: (2024)
por: Venkatesh, Kavana, et al.
Publicado: (2024)
MIEB: Massive Image Embedding Benchmark
por: Xiao, Chenghao, et al.
Publicado: (2025)
por: Xiao, Chenghao, et al.
Publicado: (2025)
Lost in Embeddings: Information Loss in Vision-Language Models
por: Li, Wenyan, et al.
Publicado: (2025)
por: Li, Wenyan, et al.
Publicado: (2025)
Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation
por: Zhang, Xiaofeng, et al.
Publicado: (2024)
por: Zhang, Xiaofeng, et al.
Publicado: (2024)
brat: Aligned Multi-View Embeddings for Brain MRI Analysis
por: Kayser, Maxime, et al.
Publicado: (2025)
por: Kayser, Maxime, et al.
Publicado: (2025)
Acceleration Multiple Heads Decoding for LLM via Dynamic Tree Attention
por: Zhang, Zhendong
Publicado: (2025)
por: Zhang, Zhendong
Publicado: (2025)
Inference Compute-Optimal Video Vision Language Models
por: Wang, Peiqi, et al.
Publicado: (2025)
por: Wang, Peiqi, et al.
Publicado: (2025)
KBE-DME: Dynamic Multimodal Evaluation via Knowledge Enhanced Benchmark Evolution
por: Zhang, Junzhe, et al.
Publicado: (2025)
por: Zhang, Junzhe, et al.
Publicado: (2025)
When Relations Break: Analyzing Relation Hallucination in Vision-Language Model Under Rotation and Noise
por: Shin, Philip Wootaek, et al.
Publicado: (2026)
por: Shin, Philip Wootaek, et al.
Publicado: (2026)
MatFormer: Nested Transformer for Elastic Inference
por: Devvrit, et al.
Publicado: (2023)
por: Devvrit, et al.
Publicado: (2023)
Mitigating Multimodal Hallucinations via Gradient-based Self-Reflection
por: Wang, Shan, et al.
Publicado: (2025)
por: Wang, Shan, et al.
Publicado: (2025)
FrEVL: Leveraging Frozen Pretrained Embeddings for Efficient Vision-Language Understanding
por: Bourigault, Emmanuelle, et al.
Publicado: (2025)
por: Bourigault, Emmanuelle, et al.
Publicado: (2025)
Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings
por: Agrawal, Aakriti, et al.
Publicado: (2025)
por: Agrawal, Aakriti, et al.
Publicado: (2025)
DoubleCCA: Improving Foundation Model Group Robustness with Random Sentence Embeddings
por: Liu, Hong, et al.
Publicado: (2024)
por: Liu, Hong, et al.
Publicado: (2024)
Natural Language Inference Improves Compositionality in Vision-Language Models
por: Cascante-Bonilla, Paola, et al.
Publicado: (2024)
por: Cascante-Bonilla, Paola, et al.
Publicado: (2024)
Inference-Time Structural Reasoning for Compositional Vision-Language Understanding
por: Bhattacharya, Amartya
Publicado: (2026)
por: Bhattacharya, Amartya
Publicado: (2026)
Expressive and Generalizable Low-rank Adaptation for Large Models via Slow Cascaded Learning
por: Li, Siwei, et al.
Publicado: (2024)
por: Li, Siwei, et al.
Publicado: (2024)
REBEL: Reinforcement Learning via Regressing Relative Rewards
por: Gao, Zhaolin, et al.
Publicado: (2024)
por: Gao, Zhaolin, et al.
Publicado: (2024)
VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
por: Meng, Rui, et al.
Publicado: (2025)
por: Meng, Rui, et al.
Publicado: (2025)
TI-JEPA: An Innovative Energy-based Joint Embedding Strategy for Text-Image Multimodal Systems
por: Vo, Khang H. N., et al.
Publicado: (2025)
por: Vo, Khang H. N., et al.
Publicado: (2025)
Improving Arabic Multi-Label Emotion Classification using Stacked Embeddings and Hybrid Loss Function
por: Aslam, Muhammad Azeem, et al.
Publicado: (2024)
por: Aslam, Muhammad Azeem, et al.
Publicado: (2024)
Ejemplares similares
-
Contrastive Learning with Enhanced Abstract Representations using Grouped Loss of Abstract Semantic Supervision
por: Suissa, Omri, et al.
Publicado: (2025) -
Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time
por: Berger, Uri, et al.
Publicado: (2025) -
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
por: Beňová, Ivana, et al.
Publicado: (2024) -
Generalised Medical Phrase Grounding
por: Zhang, Wenjun, et al.
Publicado: (2025) -
Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy, Trends and Metrics Analysis
por: Berger, Uri, et al.
Publicado: (2024)