OmniAcc: Personalized Accessibility Assistant Using Generative AI
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Karki, Siddhant, Han, Ethan, Mahmud, Nadim, Bhunia, Suman, Femiani, John, Raychoudhury, Vaskar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
Habitat Classification from Ground-Level Imagery Using Deep Neural Networks
von: Shi, Hongrui, et al.
Veröffentlicht: (2025)
von: Shi, Hongrui, et al.
Veröffentlicht: (2025)
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
von: Durrani, Hamza Ahmed, et al.
Veröffentlicht: (2026)
von: Durrani, Hamza Ahmed, et al.
Veröffentlicht: (2026)
From eye to AI: studying rodent social behavior in the era of machine Learning
von: Chindemi, Giuseppe, et al.
Veröffentlicht: (2025)
von: Chindemi, Giuseppe, et al.
Veröffentlicht: (2025)
Semi supervised GAN for smart microscopy, fast and data efficient cell cycle classification
von: Manick, Rajeev, et al.
Veröffentlicht: (2026)
von: Manick, Rajeev, et al.
Veröffentlicht: (2026)
MSPCaps: A Multi-Scale Patchify Capsule Network with Cross-Agreement Routing for Visual Recognition
von: Hu, Yudong, et al.
Veröffentlicht: (2025)
von: Hu, Yudong, et al.
Veröffentlicht: (2025)
Leum-VL Technical Report
von: He, Yuxuan, et al.
Veröffentlicht: (2026)
von: He, Yuxuan, et al.
Veröffentlicht: (2026)
Pixel-Level Pavement Distress Assessment Using Instance Segmentation
von: Dewick, Logan, et al.
Veröffentlicht: (2026)
von: Dewick, Logan, et al.
Veröffentlicht: (2026)
IMASHRIMP: Automatic White Shrimp (Penaeus vannamei) Biometrical Analysis from Laboratory Images Using Computer Vision and Deep Learning
von: González, Abiam Remache, et al.
Veröffentlicht: (2025)
von: González, Abiam Remache, et al.
Veröffentlicht: (2025)
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
von: Chen, Yuangong, et al.
Veröffentlicht: (2026)
von: Chen, Yuangong, et al.
Veröffentlicht: (2026)
Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation
von: Hou, Zhangcheng, et al.
Veröffentlicht: (2026)
von: Hou, Zhangcheng, et al.
Veröffentlicht: (2026)
Single-Shot Metric Depth from Focused Plenoptic Cameras
von: Lasheras-Hernandez, Blanca, et al.
Veröffentlicht: (2024)
von: Lasheras-Hernandez, Blanca, et al.
Veröffentlicht: (2024)
VLM-VPI: A Vision-Language Reasoning Framework for Improving Automated Vehicle-Pedestrian Interactions
von: Pu, Qingwen, et al.
Veröffentlicht: (2026)
von: Pu, Qingwen, et al.
Veröffentlicht: (2026)
Smooth regularization for efficient video recognition
von: Goldman, Gil, et al.
Veröffentlicht: (2025)
von: Goldman, Gil, et al.
Veröffentlicht: (2025)
Evaluating the Impact of Synthetic Data on Object Detection Tasks in Autonomous Driving
von: Özeren, Enes, et al.
Veröffentlicht: (2025)
von: Özeren, Enes, et al.
Veröffentlicht: (2025)
Neuromorphic Monocular Depth Estimation with Uncertainty Modeling
von: Bergkvist, Viktor, et al.
Veröffentlicht: (2026)
von: Bergkvist, Viktor, et al.
Veröffentlicht: (2026)
YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks
von: Bandyopadhyay, Saptarashmi, et al.
Veröffentlicht: (2025)
von: Bandyopadhyay, Saptarashmi, et al.
Veröffentlicht: (2025)
TD3Net: A temporal densely connected multi-dilated convolutional network for lipreading
von: Lee, Byung Hoon, et al.
Veröffentlicht: (2025)
von: Lee, Byung Hoon, et al.
Veröffentlicht: (2025)
Domain-Adaptive Pretraining Improves Primate Behavior Recognition
von: Mueller, Felix B., et al.
Veröffentlicht: (2025)
von: Mueller, Felix B., et al.
Veröffentlicht: (2025)
Visible Iris Area as a Quality Metric for Reliable Iris Recognition Under Pupil Dilation and Eyelid Occlusion
von: Pessaud, Jack, et al.
Veröffentlicht: (2025)
von: Pessaud, Jack, et al.
Veröffentlicht: (2025)
Context in object detection: a systematic literature review
von: Jamali, Mahtab, et al.
Veröffentlicht: (2025)
von: Jamali, Mahtab, et al.
Veröffentlicht: (2025)
Mask-Conditioned Voxel Diffusion for Joint Geometry and Color Inpainting
von: Sumuk, Aarya
Veröffentlicht: (2026)
von: Sumuk, Aarya
Veröffentlicht: (2026)
Pedestrian Detection in Low-Light Conditions: A Comprehensive Survey
von: Ghari, Bahareh, et al.
Veröffentlicht: (2024)
von: Ghari, Bahareh, et al.
Veröffentlicht: (2024)
PhysVideoGenerator: Towards Physically Aware Video Generation via Latent Physics Guidance
von: Satish, Siddarth Nilol Kundur, et al.
Veröffentlicht: (2026)
von: Satish, Siddarth Nilol Kundur, et al.
Veröffentlicht: (2026)
FlowIBR: Leveraging Pre-Training for Efficient Neural Image-Based Rendering of Dynamic Scenes
von: Büsching, Marcel, et al.
Veröffentlicht: (2023)
von: Büsching, Marcel, et al.
Veröffentlicht: (2023)
OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance
von: Wang, Chaoyi, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyi, et al.
Veröffentlicht: (2025)
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
von: Lee, Kyuho, et al.
Veröffentlicht: (2025)
von: Lee, Kyuho, et al.
Veröffentlicht: (2025)
Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens
von: Shen, Meng, et al.
Veröffentlicht: (2026)
von: Shen, Meng, et al.
Veröffentlicht: (2026)
Towards a Generalizable Fusion Architecture for Multimodal Object Detection
von: Berjawi, Jad, et al.
Veröffentlicht: (2025)
von: Berjawi, Jad, et al.
Veröffentlicht: (2025)
A Simple Baseline for Streaming Video Understanding
von: Shen, Yujiao, et al.
Veröffentlicht: (2026)
von: Shen, Yujiao, et al.
Veröffentlicht: (2026)
EmoVerse: A MLLMs-Driven Emotion Representation Dataset for Interpretable Visual Emotion Analysis
von: Guo, Yijie, et al.
Veröffentlicht: (2025)
von: Guo, Yijie, et al.
Veröffentlicht: (2025)
SimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World Model
von: Li, Xinqing, et al.
Veröffentlicht: (2025)
von: Li, Xinqing, et al.
Veröffentlicht: (2025)
SCA-Net: Spatial-Contextual Aggregation Network for Enhanced Small Building and Road Change Detection
von: Gholibeigi, Emad, et al.
Veröffentlicht: (2026)
von: Gholibeigi, Emad, et al.
Veröffentlicht: (2026)
Synthetic-Child: An AIGC-Based Synthetic Data Pipeline for Privacy-Preserving Child Posture Estimation
von: Zeng, Taowen
Veröffentlicht: (2026)
von: Zeng, Taowen
Veröffentlicht: (2026)
Exploring Surround-View Fisheye Camera 3D Object Detection
von: Li, Changcai, et al.
Veröffentlicht: (2025)
von: Li, Changcai, et al.
Veröffentlicht: (2025)
A Vision-Language Model for Focal Liver Lesion Classification
von: Jian, Song, et al.
Veröffentlicht: (2025)
von: Jian, Song, et al.
Veröffentlicht: (2025)
Car Object Counting and Position Estimation via Extension of the CLIP-EBC Framework
von: Jung, Seoik, et al.
Veröffentlicht: (2025)
von: Jung, Seoik, et al.
Veröffentlicht: (2025)
Vi-SAFE: A Spatial-Temporal Framework for Efficient Violence Detection in Public Surveillance
von: Chang, Ligang, et al.
Veröffentlicht: (2025)
von: Chang, Ligang, et al.
Veröffentlicht: (2025)
A Reverse Causal Framework to Mitigate Spurious Correlations for Debiasing Scene Graph Generation
von: Sun, Shuzhou, et al.
Veröffentlicht: (2025)
von: Sun, Shuzhou, et al.
Veröffentlicht: (2025)
Action Anticipation from SoccerNet Football Video Broadcasts
von: Dalal, Mohamad, et al.
Veröffentlicht: (2025)
von: Dalal, Mohamad, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025) -
Habitat Classification from Ground-Level Imagery Using Deep Neural Networks
von: Shi, Hongrui, et al.
Veröffentlicht: (2025) -
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
von: Durrani, Hamza Ahmed, et al.
Veröffentlicht: (2026) -
From eye to AI: studying rodent social behavior in the era of machine Learning
von: Chindemi, Giuseppe, et al.
Veröffentlicht: (2025) -
Semi supervised GAN for smart microscopy, fast and data efficient cell cycle classification
von: Manick, Rajeev, et al.
Veröffentlicht: (2026)