On Evaluation of Vision Datasets and Models using Human Competency Frameworks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ramachandran, Rahul, Kulkarni, Tejal, Sharma, Charchit, Vijaykeerthy, Deepak, Balasubramanian, Vineeth N |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Foundation Model Priors Enhance Object Focus in Feature Space for Source-Free Object Detection
von: VCR, Sairam, et al.
Veröffentlicht: (2025)
von: VCR, Sairam, et al.
Veröffentlicht: (2025)
C2FDrone: Coarse-to-Fine Drone-to-Drone Detection using Vision Transformer Networks
von: Rebbapragada, Sairam VC, et al.
Veröffentlicht: (2024)
von: Rebbapragada, Sairam VC, et al.
Veröffentlicht: (2024)
Source-Free Domain Adaptation by Optimizing Batch-Wise Cosine Similarity
von: Pathak, Harsharaj, et al.
Veröffentlicht: (2026)
von: Pathak, Harsharaj, et al.
Veröffentlicht: (2026)
Understanding Task Transfer in Vision-Language Models
von: Sachdeva, Bhuvan, et al.
Veröffentlicht: (2025)
von: Sachdeva, Bhuvan, et al.
Veröffentlicht: (2025)
$\oslash$ Source Models Leak What They Shouldn't $\nrightarrow$: Unlearning Zero-Shot Transfer in Domain Adaptation Through Adversarial Optimization
von: Devalapally, Arnav, et al.
Veröffentlicht: (2026)
von: Devalapally, Arnav, et al.
Veröffentlicht: (2026)
Precise Event Spotting in Sports Videos: Solving Long-Range Dependency and Class Imbalance
von: Santra, Sanchayan, et al.
Veröffentlicht: (2025)
von: Santra, Sanchayan, et al.
Veröffentlicht: (2025)
Advancing Ante-Hoc Explainable Models through Generative Adversarial Networks
von: Garg, Tanmay, et al.
Veröffentlicht: (2024)
von: Garg, Tanmay, et al.
Veröffentlicht: (2024)
LogicCBMs: Logic-Enhanced Concept-Based Learning
von: Vemuri, Deepika SN, et al.
Veröffentlicht: (2025)
von: Vemuri, Deepika SN, et al.
Veröffentlicht: (2025)
Unifying Scientific Communication: Fine-Grained Correspondence Across Scientific Media
von: M, Megha Mariam K., et al.
Veröffentlicht: (2026)
von: M, Megha Mariam K., et al.
Veröffentlicht: (2026)
Open-Set Object Detection By Aligning Known Class Representations
von: Sarkar, Hiran, et al.
Veröffentlicht: (2024)
von: Sarkar, Hiran, et al.
Veröffentlicht: (2024)
Evaluation of Cultural Competence of Vision-Language Models
von: Yadav, Srishti, et al.
Veröffentlicht: (2025)
von: Yadav, Srishti, et al.
Veröffentlicht: (2025)
Can Better Text Semantics in Prompt Tuning Improve VLM Generalization?
von: Kuchibhotla, Hari Chandana, et al.
Veröffentlicht: (2024)
von: Kuchibhotla, Hari Chandana, et al.
Veröffentlicht: (2024)
GAViD: A Large-Scale Multimodal Dataset for Context-Aware Group Affect Recognition from Videos
von: Kumar, Deepak, et al.
Veröffentlicht: (2026)
von: Kumar, Deepak, et al.
Veröffentlicht: (2026)
iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
von: Mehrotra, Sarthak, et al.
Veröffentlicht: (2025)
von: Mehrotra, Sarthak, et al.
Veröffentlicht: (2025)
Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization
von: Kancheti, Sai Srinivas, et al.
Veröffentlicht: (2026)
von: Kancheti, Sai Srinivas, et al.
Veröffentlicht: (2026)
MicroVision: An Open Dataset and Benchmark Models for Detecting Vulnerable Road Users and Micromobility Vehicles
von: Rasch, Alexander, et al.
Veröffentlicht: (2026)
von: Rasch, Alexander, et al.
Veröffentlicht: (2026)
Can Reasons Help Improve Pedestrian Intent Estimation? A Cross-Modal Approach
von: Khindkar, Vaishnavi, et al.
Veröffentlicht: (2024)
von: Khindkar, Vaishnavi, et al.
Veröffentlicht: (2024)
Mitigate One, Skew Another? Tackling Intersectional Biases in Text-to-Image Models
von: Shukla, Pushkar, et al.
Veröffentlicht: (2025)
von: Shukla, Pushkar, et al.
Veröffentlicht: (2025)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
von: Kancheti, Sai Srinivas, et al.
Veröffentlicht: (2026)
von: Kancheti, Sai Srinivas, et al.
Veröffentlicht: (2026)
POET: Prompt Offset Tuning for Continual Human Action Adaptation
von: Garg, Prachi, et al.
Veröffentlicht: (2025)
von: Garg, Prachi, et al.
Veröffentlicht: (2025)
Walking the Web of Concept-Class Relationships in Incrementally Trained Interpretable Models
von: Agrawal, Susmit, et al.
Veröffentlicht: (2025)
von: Agrawal, Susmit, et al.
Veröffentlicht: (2025)
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
von: Ramachandran, Rahul, et al.
Veröffentlicht: (2025)
von: Ramachandran, Rahul, et al.
Veröffentlicht: (2025)
Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
von: Kuchibhotla, Hari Chandana, et al.
Veröffentlicht: (2025)
von: Kuchibhotla, Hari Chandana, et al.
Veröffentlicht: (2025)
Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs
von: Sinha, Rohit, et al.
Veröffentlicht: (2026)
von: Sinha, Rohit, et al.
Veröffentlicht: (2026)
Fiducial Focus Augmentation for Facial Landmark Detection
von: Kar, Purbayan, et al.
Veröffentlicht: (2024)
von: Kar, Purbayan, et al.
Veröffentlicht: (2024)
Interpreting Neurons in Deep Vision Networks with Language Models
von: Bai, Nicholas, et al.
Veröffentlicht: (2024)
von: Bai, Nicholas, et al.
Veröffentlicht: (2024)
VideoSAVi: Self-Aligned Video Language Models without Human Supervision
von: Kulkarni, Yogesh, et al.
Veröffentlicht: (2024)
von: Kulkarni, Yogesh, et al.
Veröffentlicht: (2024)
CRoPS: A Training-Free Hallucination Mitigation Framework for Vision-Language Models
von: Anand, Neeraj, et al.
Veröffentlicht: (2026)
von: Anand, Neeraj, et al.
Veröffentlicht: (2026)
Swift Sampling: Selecting Temporal Surprises via Taylor Series
von: Kim, Dahye, et al.
Veröffentlicht: (2026)
von: Kim, Dahye, et al.
Veröffentlicht: (2026)
Evaluation of Human Visual Privacy Protection: A Three-Dimensional Framework and Benchmark Dataset
von: Abdulaziz, Sara, et al.
Veröffentlicht: (2025)
von: Abdulaziz, Sara, et al.
Veröffentlicht: (2025)
Vision Transformers and Convolutional Neural Networks for Land Use Scene Classification
von: Kulkarni, Arun D.
Veröffentlicht: (2026)
von: Kulkarni, Arun D.
Veröffentlicht: (2026)
Human-Aligned Generative Perception: Bridging Psychophysics and Generative Models
von: Titikhsha, Antara, et al.
Veröffentlicht: (2025)
von: Titikhsha, Antara, et al.
Veröffentlicht: (2025)
The Urban Vision Hackathon Dataset and Models: Towards Image Annotations and Accurate Vision Models for Indian Traffic
von: Sharma, Akash, et al.
Veröffentlicht: (2025)
von: Sharma, Akash, et al.
Veröffentlicht: (2025)
DDFAV: Remote Sensing Large Vision Language Models Dataset and Evaluation Benchmark
von: Li, Haodong, et al.
Veröffentlicht: (2024)
von: Li, Haodong, et al.
Veröffentlicht: (2024)
On the Evaluation and Refinement of Vision-Language Instruction Tuning Datasets
von: Liao, Ning, et al.
Veröffentlicht: (2023)
von: Liao, Ning, et al.
Veröffentlicht: (2023)
Road Rage Reasoning with Vision-language Models (VLMs): Task Definition and Evaluation Dataset
von: Weng, Yibing, et al.
Veröffentlicht: (2025)
von: Weng, Yibing, et al.
Veröffentlicht: (2025)
Capsule Vision 2024 Challenge: Multi-Class Abnormality Classification for Video Capsule Endoscopy
von: Handa, Palak, et al.
Veröffentlicht: (2024)
von: Handa, Palak, et al.
Veröffentlicht: (2024)
Leveraging Human-Machine Interactions for Computer Vision Dataset Quality Enhancement
von: Anzaku, Esla Timothy, et al.
Veröffentlicht: (2024)
von: Anzaku, Esla Timothy, et al.
Veröffentlicht: (2024)
OuroMamba: A Data-Free Quantization Framework for Vision Mamba
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
Artifact Removal and Image Restoration in AFM:A Structured Mask-Guided Directional Inpainting Approach
von: Zhang, Juntao, et al.
Veröffentlicht: (2026)
von: Zhang, Juntao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Foundation Model Priors Enhance Object Focus in Feature Space for Source-Free Object Detection
von: VCR, Sairam, et al.
Veröffentlicht: (2025) -
C2FDrone: Coarse-to-Fine Drone-to-Drone Detection using Vision Transformer Networks
von: Rebbapragada, Sairam VC, et al.
Veröffentlicht: (2024) -
Source-Free Domain Adaptation by Optimizing Batch-Wise Cosine Similarity
von: Pathak, Harsharaj, et al.
Veröffentlicht: (2026) -
Understanding Task Transfer in Vision-Language Models
von: Sachdeva, Bhuvan, et al.
Veröffentlicht: (2025) -
$\oslash$ Source Models Leak What They Shouldn't $\nrightarrow$: Unlearning Zero-Shot Transfer in Domain Adaptation Through Adversarial Optimization
von: Devalapally, Arnav, et al.
Veröffentlicht: (2026)