TEXT2TASTE: A Versatile Egocentric Vision System for Intelligent Reading Assistance Using Large Language Model
Fuente:
arXiv
Salvato in:
| Autori principali: | Mucha, Wiktor, Cuconasu, Florin, Etori, Naome A., Kalokyri, Valia, Trappolini, Giovanni |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
REST-HANDS: Rehabilitation with Egocentric Vision Using Smartglasses for Treatment of Hands after Surviving Stroke
di: Mucha, Wiktor, et al.
Pubblicazione: (2024)
di: Mucha, Wiktor, et al.
Pubblicazione: (2024)
In My Perspective, In My Hands: Accurate Egocentric 2D Hand Pose and Action Recognition
di: Mucha, Wiktor, et al.
Pubblicazione: (2024)
di: Mucha, Wiktor, et al.
Pubblicazione: (2024)
Towards Egocentric 3D Hand Pose Estimation in Unseen Domains
di: Mucha, Wiktor, et al.
Pubblicazione: (2026)
di: Mucha, Wiktor, et al.
Pubblicazione: (2026)
A Tale of Trust and Accuracy: Base vs. Instruct LLMs in RAG Systems
di: Cuconasu, Florin, et al.
Pubblicazione: (2024)
di: Cuconasu, Florin, et al.
Pubblicazione: (2024)
SHARP: Segmentation of Hands and Arms by Range using Pseudo-Depth for Enhanced Egocentric 3D Hand Pose Estimation and Action Recognition
di: Mucha, Wiktor, et al.
Pubblicazione: (2024)
di: Mucha, Wiktor, et al.
Pubblicazione: (2024)
Redefining Retrieval Evaluation in the Era of LLMs
di: Trappolini, Giovanni, et al.
Pubblicazione: (2025)
di: Trappolini, Giovanni, et al.
Pubblicazione: (2025)
Egocentric Vision Language Planning
di: Fang, Zhirui, et al.
Pubblicazione: (2024)
di: Fang, Zhirui, et al.
Pubblicazione: (2024)
Egocentric Bias in Vision-Language Models
di: Wang, Maijunxian, et al.
Pubblicazione: (2026)
di: Wang, Maijunxian, et al.
Pubblicazione: (2026)
An Outlook into the Future of Egocentric Vision
di: Plizzari, Chiara, et al.
Pubblicazione: (2023)
di: Plizzari, Chiara, et al.
Pubblicazione: (2023)
From Videos to Conversations: Egocentric Instructions for Task Assistance
di: Aggarwal, Lavisha, et al.
Pubblicazione: (2026)
di: Aggarwal, Lavisha, et al.
Pubblicazione: (2026)
AI-based Wearable Vision Assistance System for the Visually Impaired: Integrating Real-Time Object Recognition and Contextual Understanding Using Large Vision-Language Models
di: Baig, Mirza Samad Ahmed, et al.
Pubblicazione: (2024)
di: Baig, Mirza Samad Ahmed, et al.
Pubblicazione: (2024)
HEADS-UP: Head-Mounted Egocentric Dataset for Trajectory Prediction in Blind Assistance Systems
di: Haghighi, Yasaman, et al.
Pubblicazione: (2024)
di: Haghighi, Yasaman, et al.
Pubblicazione: (2024)
RRAML: Reinforced Retrieval Augmented Machine Learning
di: Bacciu, Andrea, et al.
Pubblicazione: (2023)
di: Bacciu, Andrea, et al.
Pubblicazione: (2023)
Egocentric Event-Based Vision for Ping Pong Ball Trajectory Prediction
di: Alberico, Ivan, et al.
Pubblicazione: (2025)
di: Alberico, Ivan, et al.
Pubblicazione: (2025)
EgoBlind: Towards Egocentric Visual Assistance for the Blind
di: Xiao, Junbin, et al.
Pubblicazione: (2025)
di: Xiao, Junbin, et al.
Pubblicazione: (2025)
RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance
di: Pellegrini, Chantal, et al.
Pubblicazione: (2023)
di: Pellegrini, Chantal, et al.
Pubblicazione: (2023)
Can Large Vision Language Models Read Maps Like a Human?
di: Xing, Shuo, et al.
Pubblicazione: (2025)
di: Xing, Shuo, et al.
Pubblicazione: (2025)
Fine-Tuning Vision-Language Models for Visual Navigation Assistance
di: Li, Xiao, et al.
Pubblicazione: (2025)
di: Li, Xiao, et al.
Pubblicazione: (2025)
Generating Dialogues from Egocentric Instructional Videos for Task Assistance: Dataset, Method and Benchmark
di: Aggarwal, Lavisha, et al.
Pubblicazione: (2025)
di: Aggarwal, Lavisha, et al.
Pubblicazione: (2025)
Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal Models
di: Jiao, Yang, et al.
Pubblicazione: (2024)
di: Jiao, Yang, et al.
Pubblicazione: (2024)
Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision
di: He, Yuping, et al.
Pubblicazione: (2025)
di: He, Yuping, et al.
Pubblicazione: (2025)
MoAI: Mixture of All Intelligence for Large Language and Vision Models
di: Lee, Byung-Kwan, et al.
Pubblicazione: (2024)
di: Lee, Byung-Kwan, et al.
Pubblicazione: (2024)
Beyond Language: Grounding Referring Expressions with Hand Pointing in Egocentric Vision
di: Li, Ling, et al.
Pubblicazione: (2026)
di: Li, Ling, et al.
Pubblicazione: (2026)
The Power of Noise: Redefining Retrieval for RAG Systems
di: Cuconasu, Florin, et al.
Pubblicazione: (2024)
di: Cuconasu, Florin, et al.
Pubblicazione: (2024)
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model
di: Huang, Yifei, et al.
Pubblicazione: (2024)
di: Huang, Yifei, et al.
Pubblicazione: (2024)
METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
di: Fu, Yankai, et al.
Pubblicazione: (2025)
di: Fu, Yankai, et al.
Pubblicazione: (2025)
MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models
di: Xia, Peng, et al.
Pubblicazione: (2024)
di: Xia, Peng, et al.
Pubblicazione: (2024)
EgoQR: Efficient QR Code Reading in Egocentric Settings
di: Moslehpour, Mohsen, et al.
Pubblicazione: (2024)
di: Moslehpour, Mohsen, et al.
Pubblicazione: (2024)
VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis
di: Pang, Chao, et al.
Pubblicazione: (2024)
di: Pang, Chao, et al.
Pubblicazione: (2024)
Generalizable Entity Grounding via Assistance of Large Language Model
di: Qi, Lu, et al.
Pubblicazione: (2024)
di: Qi, Lu, et al.
Pubblicazione: (2024)
Reading in the Dark with Foveated Event Vision
di: Brander, Carl, et al.
Pubblicazione: (2025)
di: Brander, Carl, et al.
Pubblicazione: (2025)
Unleashing the Capabilities of Large Vision-Language Models for Intelligent Perception of Roadside Infrastructure
di: Fu, Luxuan, et al.
Pubblicazione: (2026)
di: Fu, Luxuan, et al.
Pubblicazione: (2026)
EgoMotion: Hierarchical Reasoning and Diffusion for Egocentric Vision-Language Motion Generation
di: Hou, Ruibing, et al.
Pubblicazione: (2026)
di: Hou, Ruibing, et al.
Pubblicazione: (2026)
Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance
di: Zhang, Mingfang, et al.
Pubblicazione: (2025)
di: Zhang, Mingfang, et al.
Pubblicazione: (2025)
An Egocentric Vision-Language Model based Portable Real-time Smart Assistant
di: Huang, Yifei, et al.
Pubblicazione: (2025)
di: Huang, Yifei, et al.
Pubblicazione: (2025)
The State of Computer Vision Research in Africa
di: Omotayo, Abdul-Hakeem, et al.
Pubblicazione: (2024)
di: Omotayo, Abdul-Hakeem, et al.
Pubblicazione: (2024)
VIP: Versatile Image Outpainting Empowered by Multimodal Large Language Model
di: Yang, Jinze, et al.
Pubblicazione: (2024)
di: Yang, Jinze, et al.
Pubblicazione: (2024)
Ego-1K -- A Large-Scale Multiview Video Dataset for Egocentric Vision
di: Lee, Jae Yong, et al.
Pubblicazione: (2026)
di: Lee, Jae Yong, et al.
Pubblicazione: (2026)
On the Application of Egocentric Computer Vision to Industrial Scenarios
di: Chavan, Vivek, et al.
Pubblicazione: (2024)
di: Chavan, Vivek, et al.
Pubblicazione: (2024)
BioVL-QR: Egocentric Biochemical Vision-and-Language Dataset Using Micro QR Codes
di: Nishimoto, Tomohiro, et al.
Pubblicazione: (2024)
di: Nishimoto, Tomohiro, et al.
Pubblicazione: (2024)
Documenti analoghi
-
REST-HANDS: Rehabilitation with Egocentric Vision Using Smartglasses for Treatment of Hands after Surviving Stroke
di: Mucha, Wiktor, et al.
Pubblicazione: (2024) -
In My Perspective, In My Hands: Accurate Egocentric 2D Hand Pose and Action Recognition
di: Mucha, Wiktor, et al.
Pubblicazione: (2024) -
Towards Egocentric 3D Hand Pose Estimation in Unseen Domains
di: Mucha, Wiktor, et al.
Pubblicazione: (2026) -
A Tale of Trust and Accuracy: Base vs. Instruct LLMs in RAG Systems
di: Cuconasu, Florin, et al.
Pubblicazione: (2024) -
SHARP: Segmentation of Hands and Arms by Range using Pseudo-Depth for Enhanced Egocentric 3D Hand Pose Estimation and Action Recognition
di: Mucha, Wiktor, et al.
Pubblicazione: (2024)