Guardado en:
| Autores principales: | Liu, Li, Yang, Diji, Zhong, Sijia, Tholeti, Kalyana Suma Sree, Ding, Lei, Zhang, Yi, Gilpin, Leilani H. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2411.00394 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VFSI: Validity First Spatial Intelligence for Constraint-Guided Traffic Diffusion
por: Chauhan, Kargi, et al.
Publicado: (2025)
por: Chauhan, Kargi, et al.
Publicado: (2025)
Guaranteed Optimal Compositional Explanations for Neurons
por: La Rosa, Biagio, et al.
Publicado: (2025)
por: La Rosa, Biagio, et al.
Publicado: (2025)
Open Vocabulary Compositional Explanations for Neuron Alignment
por: La Rosa, Biagio, et al.
Publicado: (2025)
por: La Rosa, Biagio, et al.
Publicado: (2025)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
por: Zhang, Yue, et al.
Publicado: (2026)
por: Zhang, Yue, et al.
Publicado: (2026)
Explore the Loss space with Hill-ADAM
por: Manikandan, Meenakshi, et al.
Publicado: (2025)
por: Manikandan, Meenakshi, et al.
Publicado: (2025)
Autonomous Driving with Spiking Neural Networks
por: Zhu, Rui-Jie, et al.
Publicado: (2024)
por: Zhu, Rui-Jie, et al.
Publicado: (2024)
Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
por: Liu, Zhining, et al.
Publicado: (2025)
por: Liu, Zhining, et al.
Publicado: (2025)
LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer?
por: Cheng, Xueqi, et al.
Publicado: (2026)
por: Cheng, Xueqi, et al.
Publicado: (2026)
Bridging Legal Interpretation and Formal Logic: Faithfulness, Assumption, and the Future of AI Legal Reasoning
por: Wang, Olivia Peiyu, et al.
Publicado: (2026)
por: Wang, Olivia Peiyu, et al.
Publicado: (2026)
Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
por: Chen, Kaitao, et al.
Publicado: (2025)
por: Chen, Kaitao, et al.
Publicado: (2025)
VisualActBench: Can VLMs See and Act like a Human?
por: Zhang, Daoan, et al.
Publicado: (2025)
por: Zhang, Daoan, et al.
Publicado: (2025)
ProSLM : A Prolog Synergized Language Model for explainable Domain Specific Knowledge Based Question Answering
por: Vakharia, Priyesh, et al.
Publicado: (2024)
por: Vakharia, Priyesh, et al.
Publicado: (2024)
See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
por: Zhang, Shuoshuo, et al.
Publicado: (2025)
por: Zhang, Shuoshuo, et al.
Publicado: (2025)
Dual-Model Distillation for Efficient Action Classification with Hybrid Edge-Cloud Solution
por: Wei, Timothy, et al.
Publicado: (2024)
por: Wei, Timothy, et al.
Publicado: (2024)
Same Answer, Different Representations: Hidden instability in VLMs
por: Wani, Farooq Ahmad, et al.
Publicado: (2026)
por: Wani, Farooq Ahmad, et al.
Publicado: (2026)
Seeing the Evidence, Missing the Answer: Tool-Guided Vision-Language Models on Visual Illusions
por: Wang, Xuesong, et al.
Publicado: (2026)
por: Wang, Xuesong, et al.
Publicado: (2026)
Visual Question and Answer Using Medical Images
por: Dhanalakshmi, Kommu, et al.
Publicado: (2025)
por: Dhanalakshmi, Kommu, et al.
Publicado: (2025)
Learning to Draw ASCII Improves Spatial Reasoning in Language Models
por: Huang, Shiyuan, et al.
Publicado: (2026)
por: Huang, Shiyuan, et al.
Publicado: (2026)
Can Generalist Vision Language Models (VLMs) Rival Specialist Medical VLMs? Benchmarking and Strategic Insights
por: Zhong, Yuan, et al.
Publicado: (2025)
por: Zhong, Yuan, et al.
Publicado: (2025)
VLMs Guided Interpretable Decision Making for Autonomous Driving
por: Hu, Xin, et al.
Publicado: (2025)
por: Hu, Xin, et al.
Publicado: (2025)
Can Vision-Language Models Answer Face to Face Questions in the Real-World?
por: Pourreza, Reza, et al.
Publicado: (2025)
por: Pourreza, Reza, et al.
Publicado: (2025)
VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding
por: Yu, Haorui, et al.
Publicado: (2026)
por: Yu, Haorui, et al.
Publicado: (2026)
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
por: Shi, Chufan, et al.
Publicado: (2026)
por: Shi, Chufan, et al.
Publicado: (2026)
Quo Vadis, Anomaly Detection? LLMs and VLMs in the Spotlight
por: Ding, Xi, et al.
Publicado: (2024)
por: Ding, Xi, et al.
Publicado: (2024)
Can I Trust Your Answer? Visually Grounded Video Question Answering
por: Xiao, Junbin, et al.
Publicado: (2023)
por: Xiao, Junbin, et al.
Publicado: (2023)
GenIR: Generative Visual Feedback for Mental Image Retrieval
por: Yang, Diji, et al.
Publicado: (2025)
por: Yang, Diji, et al.
Publicado: (2025)
Decompose and Compare Consistency: Measuring VLMs' Answer Reliability via Task-Decomposition Consistency Comparison
por: Yang, Qian, et al.
Publicado: (2024)
por: Yang, Qian, et al.
Publicado: (2024)
See More, Change Less: Anatomy-Aware Diffusion for Contrast Enhancement
por: Liu, Junqi, et al.
Publicado: (2025)
por: Liu, Junqi, et al.
Publicado: (2025)
To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs
por: Hong, Rui, et al.
Publicado: (2026)
por: Hong, Rui, et al.
Publicado: (2026)
Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
por: Hong, Yuyang, et al.
Publicado: (2025)
por: Hong, Yuyang, et al.
Publicado: (2025)
PyFi: Toward Pyramid-like Financial Image Understanding for VLMs via Adversarial Agents
por: Zhang, Yuqun, et al.
Publicado: (2025)
por: Zhang, Yuqun, et al.
Publicado: (2025)
Less Precise Can Be More Reliable: A Systematic Evaluation of Quantization's Impact on VLMs Beyond Accuracy
por: Bouguerra, Aymen, et al.
Publicado: (2025)
por: Bouguerra, Aymen, et al.
Publicado: (2025)
VLMs Can Aggregate Scattered Training Patches
por: Zhou, Zhanhui, et al.
Publicado: (2025)
por: Zhou, Zhanhui, et al.
Publicado: (2025)
Many Minds from One Model: Bayesian-Inspired Transformers for Population Diversity
por: Yang, Diji, et al.
Publicado: (2025)
por: Yang, Diji, et al.
Publicado: (2025)
ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
por: Liang, Yijun, et al.
Publicado: (2025)
por: Liang, Yijun, et al.
Publicado: (2025)
Supervise Less, See More: Training-free Nuclear Instance Segmentation with Prototype-Guided Prompting
por: Zhang, Wen, et al.
Publicado: (2025)
por: Zhang, Wen, et al.
Publicado: (2025)
Vision-Language Models Can't See the Obvious
por: Dahou, Yasser, et al.
Publicado: (2025)
por: Dahou, Yasser, et al.
Publicado: (2025)
Learning to See More: UAS-Guided Super-Resolution of Satellite Imagery for Precision Agriculture
por: Masrur, Arif, et al.
Publicado: (2025)
por: Masrur, Arif, et al.
Publicado: (2025)
Can VLMs Detect and Localize Fine-Grained AI-Edited Images?
por: Sun, Zhen, et al.
Publicado: (2025)
por: Sun, Zhen, et al.
Publicado: (2025)
ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering?
por: Meshram, Pragati Shuddhodhan, et al.
Publicado: (2024)
por: Meshram, Pragati Shuddhodhan, et al.
Publicado: (2024)
Ejemplares similares
-
VFSI: Validity First Spatial Intelligence for Constraint-Guided Traffic Diffusion
por: Chauhan, Kargi, et al.
Publicado: (2025) -
Guaranteed Optimal Compositional Explanations for Neurons
por: La Rosa, Biagio, et al.
Publicado: (2025) -
Open Vocabulary Compositional Explanations for Neuron Alignment
por: La Rosa, Biagio, et al.
Publicado: (2025) -
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
por: Zhang, Yue, et al.
Publicado: (2026) -
Explore the Loss space with Hill-ADAM
por: Manikandan, Meenakshi, et al.
Publicado: (2025)