SonoVision: A Computer Vision Approach for Helping Visually Challenged Individuals Locate Objects with the Help of Sound Cues
Fuente:
arXiv
Salvato in:
| Autori principali: | Zishan, Md Abu Obaida, Rasel, Annajiat Alim |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Supersampling Stable Diffusion and Beyond: A Seamless, Training-Free Approach for Scaling Neural Networks Using Common Interpolation Methods
di: Zishan, Md Abu Obaida, et al.
Pubblicazione: (2026)
di: Zishan, Md Abu Obaida, et al.
Pubblicazione: (2026)
Computer Vision for Clinical Gait Analysis: A Gait Abnormality Video Dataset
di: Ranjan, Rahm, et al.
Pubblicazione: (2024)
di: Ranjan, Rahm, et al.
Pubblicazione: (2024)
A Comprehensive Review of YOLO Architectures in Computer Vision: From YOLOv1 to YOLOv8 and YOLO-NAS
di: Terven, Juan, et al.
Pubblicazione: (2023)
di: Terven, Juan, et al.
Pubblicazione: (2023)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
BlindSight: Harnessing Sparsity for Efficient Vision-Language Models
di: Srikrishnan, Tharun Adithya, et al.
Pubblicazione: (2025)
di: Srikrishnan, Tharun Adithya, et al.
Pubblicazione: (2025)
MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
ViG-LRGC: Vision Graph Neural Networks with Learnable Reparameterized Graph Construction
di: Elsharkawi, Ismael, et al.
Pubblicazione: (2025)
di: Elsharkawi, Ismael, et al.
Pubblicazione: (2025)
Towards Hard and Soft Shadow Removal via Dual-Branch Separation Network and Vision Transformer
di: Liang, Jiajia
Pubblicazione: (2025)
di: Liang, Jiajia
Pubblicazione: (2025)
Video Event Reasoning and Prediction by Fusing World Knowledge from LLMs with Vision Foundation Models
di: Dubois, L'ea, et al.
Pubblicazione: (2025)
di: Dubois, L'ea, et al.
Pubblicazione: (2025)
Category-Agnostic Neural Object Rigging
di: He, Guangzhao, et al.
Pubblicazione: (2025)
di: He, Guangzhao, et al.
Pubblicazione: (2025)
Robust Visual Question Answering: Datasets, Methods, and Future Challenges
di: Ma, Jie, et al.
Pubblicazione: (2023)
di: Ma, Jie, et al.
Pubblicazione: (2023)
A Challenging Benchmark of Anime Style Recognition
di: Li, Haotang, et al.
Pubblicazione: (2022)
di: Li, Haotang, et al.
Pubblicazione: (2022)
CoMViT: An Efficient Vision Backbone for Supervised Classification in Medical Imaging
di: Safdar, Aon, et al.
Pubblicazione: (2025)
di: Safdar, Aon, et al.
Pubblicazione: (2025)
Context-Aware Indoor Point Cloud Object Generation through User Instructions
di: Luo, Yiyang, et al.
Pubblicazione: (2023)
di: Luo, Yiyang, et al.
Pubblicazione: (2023)
IMASHRIMP: Automatic White Shrimp (Penaeus vannamei) Biometrical Analysis from Laboratory Images Using Computer Vision and Deep Learning
di: González, Abiam Remache, et al.
Pubblicazione: (2025)
di: González, Abiam Remache, et al.
Pubblicazione: (2025)
Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models
di: Li, Jinhao, et al.
Pubblicazione: (2024)
di: Li, Jinhao, et al.
Pubblicazione: (2024)
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
di: Han, Yudong, et al.
Pubblicazione: (2026)
di: Han, Yudong, et al.
Pubblicazione: (2026)
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
di: Zeng, Zhitao, et al.
Pubblicazione: (2025)
di: Zeng, Zhitao, et al.
Pubblicazione: (2025)
Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering
di: Ma, Jie, et al.
Pubblicazione: (2024)
di: Ma, Jie, et al.
Pubblicazione: (2024)
VisRL: Intention-Driven Visual Perception via Reinforced Reasoning
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
A Vision-Language Model for Focal Liver Lesion Classification
di: Jian, Song, et al.
Pubblicazione: (2025)
di: Jian, Song, et al.
Pubblicazione: (2025)
Beyond Visual Understanding: Introducing PARROT-360V for Vision Language Model Benchmarking
di: Khurdula, Harsha Vardhan, et al.
Pubblicazione: (2024)
di: Khurdula, Harsha Vardhan, et al.
Pubblicazione: (2024)
Textual and Visual Guided Task Adaptation for Source-Free Cross-Domain Few-Shot Segmentation
di: Liu, Jianming, et al.
Pubblicazione: (2025)
di: Liu, Jianming, et al.
Pubblicazione: (2025)
An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images
di: Castrillón-Santana, Modesto, et al.
Pubblicazione: (2025)
di: Castrillón-Santana, Modesto, et al.
Pubblicazione: (2025)
One Leaf Reveals the Season: Occlusion-Based Contrastive Learning with Semantic-Aware Views for Efficient Visual Representation
di: Yang, Xiaoyu, et al.
Pubblicazione: (2024)
di: Yang, Xiaoyu, et al.
Pubblicazione: (2024)
A Survey on Dynamic Neural Networks: from Computer Vision to Multi-modal Sensor Fusion
di: Montello, Fabio, et al.
Pubblicazione: (2025)
di: Montello, Fabio, et al.
Pubblicazione: (2025)
Edge-Based Standing-Water Detection via FSM-Guided Tiering and Multi-Model Consensus
di: Larsen, Oliver Aleksander, et al.
Pubblicazione: (2026)
di: Larsen, Oliver Aleksander, et al.
Pubblicazione: (2026)
Prompt Sensitivity in Vision-Language Grounding: How Small Changes in Wording Affect Object Detection
di: Deka, Dawar Jyoti, et al.
Pubblicazione: (2026)
di: Deka, Dawar Jyoti, et al.
Pubblicazione: (2026)
Visual Graph Question Answering with ASP and LLMs for Language Parsing
di: Bauer, Jakob Johannes, et al.
Pubblicazione: (2025)
di: Bauer, Jakob Johannes, et al.
Pubblicazione: (2025)
MedVision: Dataset and Benchmark for Quantitative Medical Image Analysis
di: Yao, Yongcheng, et al.
Pubblicazione: (2025)
di: Yao, Yongcheng, et al.
Pubblicazione: (2025)
Zero-Shot Object Goal Visual Navigation With Class-Independent Relationship Network
di: Li, Xinting, et al.
Pubblicazione: (2023)
di: Li, Xinting, et al.
Pubblicazione: (2023)
From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation
di: Chen, Jingkun, et al.
Pubblicazione: (2025)
di: Chen, Jingkun, et al.
Pubblicazione: (2025)
SITUATE -- Synthetic Object Counting Dataset for VLM training
di: Peinl, René, et al.
Pubblicazione: (2026)
di: Peinl, René, et al.
Pubblicazione: (2026)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
di: Li, Huibin, et al.
Pubblicazione: (2025)
di: Li, Huibin, et al.
Pubblicazione: (2025)
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
di: Ge, Shiping, et al.
Pubblicazione: (2024)
di: Ge, Shiping, et al.
Pubblicazione: (2024)
Demo-Pose: Depth-Monocular Modality Fusion For Object Pose Estimation
di: Agarwal, Rachit, et al.
Pubblicazione: (2026)
di: Agarwal, Rachit, et al.
Pubblicazione: (2026)
HAAP: Vision-context Hierarchical Attention Autoregressive with Adaptive Permutation for Scene Text Recognition
di: Chen, Honghui, et al.
Pubblicazione: (2024)
di: Chen, Honghui, et al.
Pubblicazione: (2024)
Synthetic Industrial Object Detection: GenAI vs. Feature-Based Methods
di: Araya-Martinez, Jose Moises, et al.
Pubblicazione: (2025)
di: Araya-Martinez, Jose Moises, et al.
Pubblicazione: (2025)
DiffYOLO: Object Detection for Anti-Noise via YOLO and Diffusion Models
di: Liu, Yichen, et al.
Pubblicazione: (2024)
di: Liu, Yichen, et al.
Pubblicazione: (2024)
Learning Association via Track-Detection Matching for Multi-Object Tracking
di: Adžemović, Momir
Pubblicazione: (2025)
di: Adžemović, Momir
Pubblicazione: (2025)
Documenti analoghi
-
Supersampling Stable Diffusion and Beyond: A Seamless, Training-Free Approach for Scaling Neural Networks Using Common Interpolation Methods
di: Zishan, Md Abu Obaida, et al.
Pubblicazione: (2026) -
Computer Vision for Clinical Gait Analysis: A Gait Abnormality Video Dataset
di: Ranjan, Rahm, et al.
Pubblicazione: (2024) -
A Comprehensive Review of YOLO Architectures in Computer Vision: From YOLOv1 to YOLOv8 and YOLO-NAS
di: Terven, Juan, et al.
Pubblicazione: (2023) -
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025) -
BlindSight: Harnessing Sparsity for Efficient Vision-Language Models
di: Srikrishnan, Tharun Adithya, et al.
Pubblicazione: (2025)