FLAIR: VLM with Fine-grained Language-informed Image Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiao, Rui, Kim, Sanghwan, Georgescu, Mariana-Iuliana, Akata, Zeynep, Alaniz, Stephan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
von: Kim, Sanghwan, et al.
Veröffentlicht: (2024)
von: Kim, Sanghwan, et al.
Veröffentlicht: (2024)
FINER: MLLMs Hallucinate under Fine-grained Negative Queries
von: Xiao, Rui, et al.
Veröffentlicht: (2026)
von: Xiao, Rui, et al.
Veröffentlicht: (2026)
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
von: Kim, Sanghwan, et al.
Veröffentlicht: (2025)
von: Kim, Sanghwan, et al.
Veröffentlicht: (2025)
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
von: Wu, Boyong, et al.
Veröffentlicht: (2026)
von: Wu, Boyong, et al.
Veröffentlicht: (2026)
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval
von: Hummel, Thomas, et al.
Veröffentlicht: (2024)
von: Hummel, Thomas, et al.
Veröffentlicht: (2024)
SUB: Benchmarking CBM Generalization via Synthetic Attribute Substitutions
von: Bader, Jessica, et al.
Veröffentlicht: (2025)
von: Bader, Jessica, et al.
Veröffentlicht: (2025)
Concept-Guided Interpretability via Neural Chunking
von: Wu, Shuchen, et al.
Veröffentlicht: (2025)
von: Wu, Shuchen, et al.
Veröffentlicht: (2025)
Subspace-Boosted Model Merging
von: Skorobogat, Ronald, et al.
Veröffentlicht: (2025)
von: Skorobogat, Ronald, et al.
Veröffentlicht: (2025)
A Large Scale Analysis of Gender Biases in Text-to-Image Generative Models
von: Girrbach, Leander, et al.
Veröffentlicht: (2025)
von: Girrbach, Leander, et al.
Veröffentlicht: (2025)
LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
von: Kim, Jae Myung, et al.
Veröffentlicht: (2025)
von: Kim, Jae Myung, et al.
Veröffentlicht: (2025)
Road Obstacle Video Segmentation
von: Rai, Shyam Nandan, et al.
Veröffentlicht: (2025)
von: Rai, Shyam Nandan, et al.
Veröffentlicht: (2025)
FLAIR: Frequency- and Locality-Aware Implicit Neural Representations
von: Ko, Sukhun, et al.
Veröffentlicht: (2025)
von: Ko, Sukhun, et al.
Veröffentlicht: (2025)
Weight Copy and Low-Rank Adaptation for Few-Shot Distillation of Vision Transformers
von: Grigore, Diana-Nicoleta, et al.
Veröffentlicht: (2024)
von: Grigore, Diana-Nicoleta, et al.
Veröffentlicht: (2024)
X-Aligner: Composed Visual Retrieval without the Bells and Whistles
von: Zheng, Yuqian, et al.
Veröffentlicht: (2026)
von: Zheng, Yuqian, et al.
Veröffentlicht: (2026)
Identity-Aware U-Net: Fine-grained Cell Segmentation via Identity-Aware Representation Learning
von: Xiao, Rui
Veröffentlicht: (2026)
von: Xiao, Rui
Veröffentlicht: (2026)
Time Series Representations for Classification Lie Hidden in Pretrained Vision Transformers
von: Roschmann, Simon, et al.
Veröffentlicht: (2025)
von: Roschmann, Simon, et al.
Veröffentlicht: (2025)
Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation
von: Gu, Difei, et al.
Veröffentlicht: (2025)
von: Gu, Difei, et al.
Veröffentlicht: (2025)
Explaining CLIP Zero-shot Predictions Through Concepts
von: Ozdemir, Onat, et al.
Veröffentlicht: (2026)
von: Ozdemir, Onat, et al.
Veröffentlicht: (2026)
Improving Intervention Efficacy via Concept Realignment in Concept Bottleneck Models
von: Singhi, Nishad, et al.
Veröffentlicht: (2024)
von: Singhi, Nishad, et al.
Veröffentlicht: (2024)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
von: Pach, Mateusz, et al.
Veröffentlicht: (2025)
von: Pach, Mateusz, et al.
Veröffentlicht: (2025)
DataDream: Few-shot Guided Dataset Generation
von: Kim, Jae Myung, et al.
Veröffentlicht: (2024)
von: Kim, Jae Myung, et al.
Veröffentlicht: (2024)
MedFILIP: Medical Fine-grained Language-Image Pre-training
von: Liang, Xinjie, et al.
Veröffentlicht: (2025)
von: Liang, Xinjie, et al.
Veröffentlicht: (2025)
Pairwise Matching of Intermediate Representations for Fine-grained Explainability
von: Shrack, Lauren, et al.
Veröffentlicht: (2025)
von: Shrack, Lauren, et al.
Veröffentlicht: (2025)
VLM6D: VLM based 6Dof Pose Estimation based on RGB-D Images
von: Sarowar, Md Selim, et al.
Veröffentlicht: (2025)
von: Sarowar, Md Selim, et al.
Veröffentlicht: (2025)
Fine-grained Background Representation for Weakly Supervised Semantic Segmentation
von: Yin, Xu, et al.
Veröffentlicht: (2024)
von: Yin, Xu, et al.
Veröffentlicht: (2024)
M$^2$IV: Towards Efficient and Fine-grained Multimodal In-Context Learning via Representation Engineering
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
LensVLM: Selective Context Expansion for Compressed Visual Representation of Text
von: Xie, Roy, et al.
Veröffentlicht: (2026)
von: Xie, Roy, et al.
Veröffentlicht: (2026)
The Latent Color Subspace: Emergent Order in High-Dimensional Chaos
von: Pach, Mateusz, et al.
Veröffentlicht: (2026)
von: Pach, Mateusz, et al.
Veröffentlicht: (2026)
Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
von: Girrbach, Leander, et al.
Veröffentlicht: (2025)
von: Girrbach, Leander, et al.
Veröffentlicht: (2025)
Adaptive Self-training Framework for Fine-grained Scene Graph Generation
von: Kim, Kibum, et al.
Veröffentlicht: (2024)
von: Kim, Kibum, et al.
Veröffentlicht: (2024)
Multi-modal Representations for Fine-grained Multi-label Critical View of Safety Recognition
von: Baby, Britty, et al.
Veröffentlicht: (2025)
von: Baby, Britty, et al.
Veröffentlicht: (2025)
T2I-VeRW: Part-level Fine-grained Perception for Text-to-Image Vehicle Retrieval
von: Wang, Xiao, et al.
Veröffentlicht: (2026)
von: Wang, Xiao, et al.
Veröffentlicht: (2026)
Respect the model: Fine-grained and Robust Explanation with Sharing Ratio Decomposition
von: Han, Sangyu, et al.
Veröffentlicht: (2024)
von: Han, Sangyu, et al.
Veröffentlicht: (2024)
Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
von: Bader, Jessica, et al.
Veröffentlicht: (2025)
von: Bader, Jessica, et al.
Veröffentlicht: (2025)
A unified FLAIR hyperintensity segmentation model for various CNS tumor types and acquisition time points
von: Faanes, Mathilde Gajda, et al.
Veröffentlicht: (2025)
von: Faanes, Mathilde Gajda, et al.
Veröffentlicht: (2025)
FingER: Content Aware Fine-grained Evaluation with Reasoning for AI-Generated Videos
von: Chen, Rui, et al.
Veröffentlicht: (2025)
von: Chen, Rui, et al.
Veröffentlicht: (2025)
Lost in Space: Probing Fine-grained Spatial Understanding in Vision and Language Resamplers
von: Pantazopoulos, Georgios, et al.
Veröffentlicht: (2024)
von: Pantazopoulos, Georgios, et al.
Veröffentlicht: (2024)
A Closer Look at Multimodal Representation Collapse
von: Chaudhuri, Abhra, et al.
Veröffentlicht: (2025)
von: Chaudhuri, Abhra, et al.
Veröffentlicht: (2025)
Understanding and Defending VLM Jailbreaks via Jailbreak-Related Representation Shift
von: Wei, Zhihua, et al.
Veröffentlicht: (2026)
von: Wei, Zhihua, et al.
Veröffentlicht: (2026)
CARPE: Context-Aware Image Representation Prioritization via Ensemble for Large Vision-Language Models
von: Lee, Donghee, et al.
Veröffentlicht: (2026)
von: Lee, Donghee, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
von: Kim, Sanghwan, et al.
Veröffentlicht: (2024) -
FINER: MLLMs Hallucinate under Fine-grained Negative Queries
von: Xiao, Rui, et al.
Veröffentlicht: (2026) -
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
von: Kim, Sanghwan, et al.
Veröffentlicht: (2025) -
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
von: Wu, Boyong, et al.
Veröffentlicht: (2026) -
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval
von: Hummel, Thomas, et al.
Veröffentlicht: (2024)