VFA: Vision Frequency Analysis of Foundation Models and Human
Fuente:
arXiv
Saved in:
| Main Authors: | Darvishi-Bayazi, Mohammad-Javad, Arefin, Md Rifat, Faubert, Jocelyn, Rish, Irina |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Weak-Annotation of HAR Datasets using Vision Foundation Models
by: Bock, Marius, et al.
Published: (2024)
by: Bock, Marius, et al.
Published: (2024)
MedFoundationHub: A Lightweight and Secure Toolkit for Deploying Medical Vision Language Foundation Models
by: Li, Xiao, et al.
Published: (2025)
by: Li, Xiao, et al.
Published: (2025)
Amplifying Pathological Detection in EEG Signaling Pathways through Cross-Dataset Transfer Learning
by: Darvishi-Bayazi, Mohammad-Javad, et al.
Published: (2023)
by: Darvishi-Bayazi, Mohammad-Javad, et al.
Published: (2023)
A Dataset for Crucial Object Recognition in Blind and Low-Vision Individuals' Navigation
by: Islam, Md Touhidul, et al.
Published: (2024)
by: Islam, Md Touhidul, et al.
Published: (2024)
Visual Affect Analysis: Predicting Emotions of Image Viewers with Vision-Language Models
by: Nowicki, Filip, et al.
Published: (2026)
by: Nowicki, Filip, et al.
Published: (2026)
Scene-Aware Urban Design: A Human-AI Recommendation Framework Using Co-Occurrence Embeddings and Vision-Language Models
by: Gallardo, Rodrigo, et al.
Published: (2025)
by: Gallardo, Rodrigo, et al.
Published: (2025)
Vision Language Models as Values Detectors
by: Abbo, Giulio Antonio, et al.
Published: (2025)
by: Abbo, Giulio Antonio, et al.
Published: (2025)
PCNN: Probable-Class Nearest-Neighbor Explanations Improve Fine-Grained Image Classification Accuracy for AIs and Humans
by: Giang, et al.
Published: (2023)
by: Giang, et al.
Published: (2023)
Foundation Models for Zero-Shot Segmentation of Scientific Images without AI-Ready Data
by: Mukherjee, Shubhabrata, et al.
Published: (2025)
by: Mukherjee, Shubhabrata, et al.
Published: (2025)
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
by: Verma, Arnav, et al.
Published: (2025)
by: Verma, Arnav, et al.
Published: (2025)
LiDAR-based Human Activity Recognition through Laplacian Spectral Analysis
by: Sharifipour, Sasan, et al.
Published: (2025)
by: Sharifipour, Sasan, et al.
Published: (2025)
Do Vision Language Models Understand Human Engagement in Games?
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
Prompting with Sign Parameters for Low-resource Sign Language Instruction Generation
by: Tariquzzaman, Md, et al.
Published: (2025)
by: Tariquzzaman, Md, et al.
Published: (2025)
PoseDriver: A Unified Approach to Multi-Category Skeleton Detection for Autonomous Driving
by: Borhani, Yasamin, et al.
Published: (2026)
by: Borhani, Yasamin, et al.
Published: (2026)
An Egocentric Vision-Language Model based Portable Real-time Smart Assistant
by: Huang, Yifei, et al.
Published: (2025)
by: Huang, Yifei, et al.
Published: (2025)
Inclusive STEAM Education: A Framework for Teaching Cod-2 ing and Robotics to Students with Visually Impairment Using 3 Advanced Computer Vision
by: Hamash, Mahmoud, et al.
Published: (2025)
by: Hamash, Mahmoud, et al.
Published: (2025)
HarassGuard: Detecting Harassment Behaviors in Social Virtual Reality with Vision-Language Models
by: Lee, Junhee, et al.
Published: (2026)
by: Lee, Junhee, et al.
Published: (2026)
A General Model for Detecting Learner Engagement: Implementation and Evaluation
by: Malekshahi, Somayeh, et al.
Published: (2024)
by: Malekshahi, Somayeh, et al.
Published: (2024)
Towards Consumer-Grade Cybersickness Prediction: Multi-Model Alignment for Real-Time Vision-Only Inference
by: Zhu, Yitong, et al.
Published: (2025)
by: Zhu, Yitong, et al.
Published: (2025)
InterAnimate: Taming Region-aware Diffusion Model for Realistic Human Interaction Animation
by: Lin, Yukang, et al.
Published: (2025)
by: Lin, Yukang, et al.
Published: (2025)
"It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models
by: Garg, Kapil, et al.
Published: (2025)
by: Garg, Kapil, et al.
Published: (2025)
Leveraging Digital Perceptual Technologies for Remote Perception and Analysis of Human Biomechanical Processes: A Contactless Approach for Workload and Joint Force Assessment
by: Omidokun, Jesudara, et al.
Published: (2024)
by: Omidokun, Jesudara, et al.
Published: (2024)
L-WISE: Boosting Human Visual Category Learning Through Model-Based Image Selection and Enhancement
by: Talbot, Morgan B., et al.
Published: (2024)
by: Talbot, Morgan B., et al.
Published: (2024)
Constructive Apraxia: An Unexpected Limit of Instructible Vision-Language Models and Analog for Human Cognitive Disorders
by: Noever, David, et al.
Published: (2024)
by: Noever, David, et al.
Published: (2024)
MicroBi-ConvLSTM: An Ultra-Lightweight Efficient Model for Human Activity Recognition on Resource Constrained Devices
by: Mandal, Mridankan
Published: (2026)
by: Mandal, Mridankan
Published: (2026)
VISLIX: An XAI Framework for Validating Vision Models with Slice Discovery and Analysis
by: Yan, Xinyuan, et al.
Published: (2025)
by: Yan, Xinyuan, et al.
Published: (2025)
VisionCAD: An Integration-Free Radiology Copilot Framework
by: Li, Jiaming, et al.
Published: (2025)
by: Li, Jiaming, et al.
Published: (2025)
ViT-Explainer: An Interactive Walkthrough of the Vision Transformer Pipeline
by: Hernandez, Juan Manuel, et al.
Published: (2026)
by: Hernandez, Juan Manuel, et al.
Published: (2026)
BabyMamba-HAR: Lightweight Selective State Space Models for Efficient Human Activity Recognition on Resource Constrained Devices
by: Mandal, Mridankan
Published: (2026)
by: Mandal, Mridankan
Published: (2026)
Computer Vision for Objects used in Group Work: Challenges and Opportunities
by: Jung, Changsoo, et al.
Published: (2025)
by: Jung, Changsoo, et al.
Published: (2025)
Real-Time Cellist Postural Evaluation With On-Device Computer Vision
by: Wang, Paolo, et al.
Published: (2026)
by: Wang, Paolo, et al.
Published: (2026)
Machine Vision-Based Surgical Lighting System:Design and Implementation
by: Gharghabi, Amir, et al.
Published: (2025)
by: Gharghabi, Amir, et al.
Published: (2025)
Referring Human Pose and Mask Estimation in the Wild
by: Miao, Bo, et al.
Published: (2024)
by: Miao, Bo, et al.
Published: (2024)
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
by: Wu, Zhiyong, et al.
Published: (2024)
by: Wu, Zhiyong, et al.
Published: (2024)
SymbolSight: Minimizing Inter-Symbol Interference for Reading with Prosthetic Vision
by: Lesner, Jasmine, et al.
Published: (2026)
by: Lesner, Jasmine, et al.
Published: (2026)
iTrace: Click-Based Gaze Visualization on the Apple Vision Pro
by: Mehmedova, Esra, et al.
Published: (2025)
by: Mehmedova, Esra, et al.
Published: (2025)
Machine Learning-Based Jamun Leaf Disease Detection: A Comprehensive Review
by: Bhowmik, Auvick Chandra, et al.
Published: (2023)
by: Bhowmik, Auvick Chandra, et al.
Published: (2023)
Extracting Human Attention through Crowdsourced Patch Labeling
by: Chang, Minsuk, et al.
Published: (2024)
by: Chang, Minsuk, et al.
Published: (2024)
EgoPressure: A Dataset for Hand Pressure and Pose Estimation in Egocentric Vision
by: Zhao, Yiming, et al.
Published: (2024)
by: Zhao, Yiming, et al.
Published: (2024)
Detecting Clues for Skill Levels and Machine Operation Difficulty from Egocentric Vision
by: Long-fei, Chen, et al.
Published: (2019)
by: Long-fei, Chen, et al.
Published: (2019)
Similar Items
-
Weak-Annotation of HAR Datasets using Vision Foundation Models
by: Bock, Marius, et al.
Published: (2024) -
MedFoundationHub: A Lightweight and Secure Toolkit for Deploying Medical Vision Language Foundation Models
by: Li, Xiao, et al.
Published: (2025) -
Amplifying Pathological Detection in EEG Signaling Pathways through Cross-Dataset Transfer Learning
by: Darvishi-Bayazi, Mohammad-Javad, et al.
Published: (2023) -
A Dataset for Crucial Object Recognition in Blind and Low-Vision Individuals' Navigation
by: Islam, Md Touhidul, et al.
Published: (2024) -
Visual Affect Analysis: Predicting Emotions of Image Viewers with Vision-Language Models
by: Nowicki, Filip, et al.
Published: (2026)