Saved in:
| Main Author: | Basu, Arkaprabha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2507.12195 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Grid2Matrix: Revealing Digital Agnosia in Vision-Language Models
by: Zhang, Yunkai, et al.
Published: (2026)
by: Zhang, Yunkai, et al.
Published: (2026)
Beyond Matching to Tiles: Bridging Unaligned Aerial and Satellite Views for Vision-Only UAV Navigation
by: Liu, Kejia, et al.
Published: (2026)
by: Liu, Kejia, et al.
Published: (2026)
IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
Adversarial Evasion Attacks on Computer Vision using SHAP Values
by: Mollard, Frank, et al.
Published: (2026)
by: Mollard, Frank, et al.
Published: (2026)
Unlocking the Essence of Beauty: Advanced Aesthetic Reasoning with Relative-Absolute Policy Optimization
by: Liu, Boyang, et al.
Published: (2025)
by: Liu, Boyang, et al.
Published: (2025)
HAVIR: HierArchical Vision to Image Reconstruction using CLIP-Guided Versatile Diffusion
by: Zhang, Shiyi, et al.
Published: (2025)
by: Zhang, Shiyi, et al.
Published: (2025)
HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
by: Nath, Sujoy, et al.
Published: (2025)
by: Nath, Sujoy, et al.
Published: (2025)
Freeze and Reveal: Exposing Modality Bias in Vision-Language Models
by: Kavuri, Vivek Hruday, et al.
Published: (2025)
by: Kavuri, Vivek Hruday, et al.
Published: (2025)
Representing Beauty: Towards a Participatory but Objective Latent Aesthetics
by: Rusnak, Alexander Michael
Published: (2025)
by: Rusnak, Alexander Michael
Published: (2025)
Revealing Multi-View Hallucination in Large Vision-Language Models
by: Park, Wooje, et al.
Published: (2026)
by: Park, Wooje, et al.
Published: (2026)
Image Based Character Recognition, Documentation System To Decode Inscription From Temple
by: G, Velmathi, et al.
Published: (2024)
by: G, Velmathi, et al.
Published: (2024)
Text Speaks Louder than Vision: ASCII Art Reveals Textual Biases in Vision-Language Models
by: Wang, Zhaochen, et al.
Published: (2025)
by: Wang, Zhaochen, et al.
Published: (2025)
Image Tiling for High-Resolution Reasoning: Balancing Local Detail with Global Context
by: de Margerie, Anatole Jacquin, et al.
Published: (2025)
by: de Margerie, Anatole Jacquin, et al.
Published: (2025)
KPCA-CAM: Visual Explainability of Deep Computer Vision Models using Kernel PCA
by: Karmani, Sachin, et al.
Published: (2024)
by: Karmani, Sachin, et al.
Published: (2024)
Feature-Optimized Vision for Adaptive 3D Scene Reconstruction
by: Liang, Eric
Published: (2026)
by: Liang, Eric
Published: (2026)
Visual Anagrams Reveal Hidden Differences in Holistic Shape Processing Across Vision Models
by: Doshi, Fenil R., et al.
Published: (2025)
by: Doshi, Fenil R., et al.
Published: (2025)
Vehicle detection from GSV imagery: Predicting travel behaviour for cycling and motorcycling using Computer Vision
by: Kyriaki, et al.
Published: (2025)
by: Kyriaki, et al.
Published: (2025)
Exploring Image Generation via Mutually Exclusive Probability Spaces and Local Correlation Hypothesis
by: Zhao, Chenqiu, et al.
Published: (2025)
by: Zhao, Chenqiu, et al.
Published: (2025)
Towards Ancient Plant Seed Classification: A Benchmark Dataset and Baseline Model
by: Xing, Rui, et al.
Published: (2025)
by: Xing, Rui, et al.
Published: (2025)
OBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?
by: Chen, Zijian, et al.
Published: (2024)
by: Chen, Zijian, et al.
Published: (2024)
HydroVision: Predicting Optically Active Parameters in Surface Water Using Computer Vision
by: Deshmukh, Shubham Laxmikant, et al.
Published: (2025)
by: Deshmukh, Shubham Laxmikant, et al.
Published: (2025)
Vision without Images: End-to-End Computer Vision from Single Compressive Measurements
by: Pan, Fengpu, et al.
Published: (2025)
by: Pan, Fengpu, et al.
Published: (2025)
A Hybrid Random Forest and CNN Framework for Tile-Wise Oil-Water Classification in Hyperspectral Images
by: Nickzamir, Mehdi, et al.
Published: (2025)
by: Nickzamir, Mehdi, et al.
Published: (2025)
ClaudesLens: Uncertainty Quantification in Computer Vision Models
by: Shaar, Mohamad Al, et al.
Published: (2024)
by: Shaar, Mohamad Al, et al.
Published: (2024)
AI Driven Soccer Analysis Using Computer Vision
by: Manchado, Adrian, et al.
Published: (2026)
by: Manchado, Adrian, et al.
Published: (2026)
What can Computer Vision learn from Ranganathan?
by: Bagchi, Mayukh, et al.
Published: (2026)
by: Bagchi, Mayukh, et al.
Published: (2026)
Leveraging Diffusion Perturbations for Measuring Fairness in Computer Vision
by: Lui, Nicholas, et al.
Published: (2023)
by: Lui, Nicholas, et al.
Published: (2023)
Suitability of KANs for Computer Vision: A preliminary investigation
by: Azam, Basim, et al.
Published: (2024)
by: Azam, Basim, et al.
Published: (2024)
Wake Vision: A Tailored Dataset and Benchmark Suite for TinyML Computer Vision Applications
by: Banbury, Colby, et al.
Published: (2024)
by: Banbury, Colby, et al.
Published: (2024)
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
by: Karamolegkou, Antonia, et al.
Published: (2026)
by: Karamolegkou, Antonia, et al.
Published: (2026)
On the Utility of Foundation Models for Fast MRI: Vision-Language-Guided Image Reconstruction
by: Feng, Ruimin, et al.
Published: (2025)
by: Feng, Ruimin, et al.
Published: (2025)
The Illusion of Clinical Reasoning: A Benchmark Reveals the Pervasive Gap in Vision-Language Models for Clinical Competency
by: Wang, Dingyu, et al.
Published: (2025)
by: Wang, Dingyu, et al.
Published: (2025)
How Texts Help? A Fine-grained Evaluation to Reveal the Role of Language in Vision-Language Tracking
by: Li, Xuchen, et al.
Published: (2024)
by: Li, Xuchen, et al.
Published: (2024)
Computer Vision based group activity detection and action spotting
by: Sivalingam, Narthana, et al.
Published: (2025)
by: Sivalingam, Narthana, et al.
Published: (2025)
Computer Vision and Deep Learning for 4D Augmented Reality
by: Shivashankar, Karthik
Published: (2025)
by: Shivashankar, Karthik
Published: (2025)
HAVIR: HierArchical Vision to Image Reconstruction using CLIP-Guided Versatile Diffusion
by: Zhang, Shiyi, et al.
Published: (2025)
by: Zhang, Shiyi, et al.
Published: (2025)
Application of Attention Mechanism with Bidirectional Long Short-Term Memory (BiLSTM) and CNN for Human Conflict Detection using Computer Vision
by: Farias, Erick da Silva, et al.
Published: (2025)
by: Farias, Erick da Silva, et al.
Published: (2025)
Learner Attentiveness and Engagement Analysis in Online Education Using Computer Vision
by: Gogawale, Sharva, et al.
Published: (2024)
by: Gogawale, Sharva, et al.
Published: (2024)
A Review on Discriminative Self-supervised Learning Methods in Computer Vision
by: Giakoumoglou, Nikolaos, et al.
Published: (2024)
by: Giakoumoglou, Nikolaos, et al.
Published: (2024)
Merlin: A Computed Tomography Vision-Language Foundation Model and Dataset
by: Blankemeier, Louis, et al.
Published: (2024)
by: Blankemeier, Louis, et al.
Published: (2024)
Similar Items
-
Grid2Matrix: Revealing Digital Agnosia in Vision-Language Models
by: Zhang, Yunkai, et al.
Published: (2026) -
Beyond Matching to Tiles: Bridging Unaligned Aerial and Satellite Views for Vision-Only UAV Navigation
by: Liu, Kejia, et al.
Published: (2026) -
IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning
by: Ghosal, Soumya Suvra, et al.
Published: (2024) -
Adversarial Evasion Attacks on Computer Vision using SHAP Values
by: Mollard, Frank, et al.
Published: (2026) -
Unlocking the Essence of Beauty: Advanced Aesthetic Reasoning with Relative-Absolute Policy Optimization
by: Liu, Boyang, et al.
Published: (2025)