Training a Vision Language Model as Smartphone Assistant
Fuente:
arXiv
Salvato in:
| Autori principali: | Dorka, Nicolai, Marecki, Janusz, Anwar, Ammar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
InterVLS: Interactive Model Understanding and Improvement with Vision-Language Surrogates
di: Huang, Jinbin, et al.
Pubblicazione: (2023)
di: Huang, Jinbin, et al.
Pubblicazione: (2023)
AI Guide Dog: Egocentric Path Prediction on Smartphone
di: Jadhav, Aishwarya, et al.
Pubblicazione: (2025)
di: Jadhav, Aishwarya, et al.
Pubblicazione: (2025)
Vision-Language Models for Ergonomic Assessment of Manual Lifting Tasks: Estimating Horizontal and Vertical Hand Distances from RGB Video
di: Rajabi, Mohammad Sadra, et al.
Pubblicazione: (2026)
di: Rajabi, Mohammad Sadra, et al.
Pubblicazione: (2026)
Predicting and Explaining Mobile UI Tappability with Vision Modeling and Saliency Analysis
di: Schoop, Eldon, et al.
Pubblicazione: (2022)
di: Schoop, Eldon, et al.
Pubblicazione: (2022)
DOTA: Distributional Test-Time Adaptation of Vision-Language Models
di: Han, Zongbo, et al.
Pubblicazione: (2024)
di: Han, Zongbo, et al.
Pubblicazione: (2024)
SpiderNets: Vision Models Predict Human Fear From Aversive Images
di: Pegler, Dominik, et al.
Pubblicazione: (2025)
di: Pegler, Dominik, et al.
Pubblicazione: (2025)
Bridging Human Concepts and Computer Vision for Explainable Face Verification
di: Doh, Miriam, et al.
Pubblicazione: (2024)
di: Doh, Miriam, et al.
Pubblicazione: (2024)
Screen2AX: Vision-Based Approach for Automatic macOS Accessibility Generation
di: Muryn, Viktor, et al.
Pubblicazione: (2025)
di: Muryn, Viktor, et al.
Pubblicazione: (2025)
BdSLW401: Transformer-Based Word-Level Bangla Sign Language Recognition Using Relative Quantization Encoding (RQE)
di: Rubaiyeat, Husne Ara, et al.
Pubblicazione: (2025)
di: Rubaiyeat, Husne Ara, et al.
Pubblicazione: (2025)
Semantic Approach to Quantifying the Consistency of Diffusion Model Image Generation
di: Bent, Brinnae
Pubblicazione: (2024)
di: Bent, Brinnae
Pubblicazione: (2024)
A Foundational Generative Model for Breast Ultrasound Image Analysis
di: Yu, Haojun, et al.
Pubblicazione: (2025)
di: Yu, Haojun, et al.
Pubblicazione: (2025)
LEDITS++: Limitless Image Editing using Text-to-Image Models
di: Brack, Manuel, et al.
Pubblicazione: (2023)
di: Brack, Manuel, et al.
Pubblicazione: (2023)
What's Producible May Not Be Reachable: Measuring the Steerability of Generative Models
di: Vafa, Keyon, et al.
Pubblicazione: (2025)
di: Vafa, Keyon, et al.
Pubblicazione: (2025)
ShelfHelp: Empowering Humans to Perform Vision-Independent Manipulation Tasks with a Socially Assistive Robotic Cane
di: Agrawal, Shivendra, et al.
Pubblicazione: (2024)
di: Agrawal, Shivendra, et al.
Pubblicazione: (2024)
I-CEE: Tailoring Explanations of Image Classification Models to User Expertise
di: Rong, Yao, et al.
Pubblicazione: (2023)
di: Rong, Yao, et al.
Pubblicazione: (2023)
Detoxifying Large Language Models via Knowledge Editing
di: Wang, Mengru, et al.
Pubblicazione: (2024)
di: Wang, Mengru, et al.
Pubblicazione: (2024)
Generating Synthetic Satellite Imagery for Rare Objects: An Empirical Comparison of Models and Metrics
di: Nguyen, Tuong Vy, et al.
Pubblicazione: (2024)
di: Nguyen, Tuong Vy, et al.
Pubblicazione: (2024)
HERO: Human-Feedback Efficient Reinforcement Learning for Online Diffusion Model Finetuning
di: Hiranaka, Ayano, et al.
Pubblicazione: (2024)
di: Hiranaka, Ayano, et al.
Pubblicazione: (2024)
Optimizing Small Language Models for In-Vehicle Function-Calling
di: Khiabani, Yahya Sowti, et al.
Pubblicazione: (2025)
di: Khiabani, Yahya Sowti, et al.
Pubblicazione: (2025)
ExpressEdit: Fast Editing of Stylized Facial Expressions with Diffusion Models in Photoshop
di: Tang, Kenan, et al.
Pubblicazione: (2026)
di: Tang, Kenan, et al.
Pubblicazione: (2026)
Interaction as Explanation: A User Interaction-based Method for Explaining Image Classification Models
di: Yun, Hyeonggeun
Pubblicazione: (2024)
di: Yun, Hyeonggeun
Pubblicazione: (2024)
Knowledge Mechanisms in Large Language Models: A Survey and Perspective
di: Wang, Mengru, et al.
Pubblicazione: (2024)
di: Wang, Mengru, et al.
Pubblicazione: (2024)
A Comprehensive Study of Knowledge Editing for Large Language Models
di: Zhang, Ningyu, et al.
Pubblicazione: (2024)
di: Zhang, Ningyu, et al.
Pubblicazione: (2024)
ReLearn: Unlearning via Learning for Large Language Models
di: Xu, Haoming, et al.
Pubblicazione: (2025)
di: Xu, Haoming, et al.
Pubblicazione: (2025)
Generating Synthetic Satellite Imagery With Deep-Learning Text-to-Image Models -- Technical Challenges and Implications for Monitoring and Verification
di: Nguyen, Tuong Vy, et al.
Pubblicazione: (2024)
di: Nguyen, Tuong Vy, et al.
Pubblicazione: (2024)
AutoTour: Automatic Photo Tour Guide with Smartphones and LLMs
di: Xu, Huatao, et al.
Pubblicazione: (2026)
di: Xu, Huatao, et al.
Pubblicazione: (2026)
Zero-shot Emotion Annotation in Facial Images Using Large Multimodal Models: Benchmarking and Prospects for Multi-Class, Multi-Frame Approaches
di: Zhang, He, et al.
Pubblicazione: (2025)
di: Zhang, He, et al.
Pubblicazione: (2025)
InstructEdit: Instruction-based Knowledge Editing for Large Language Models
di: Zhang, Ningyu, et al.
Pubblicazione: (2024)
di: Zhang, Ningyu, et al.
Pubblicazione: (2024)
Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
di: Natalie, Rosiana, et al.
Pubblicazione: (2025)
di: Natalie, Rosiana, et al.
Pubblicazione: (2025)
Quantitative Movement Testing: Measuring Patient Movements from a Single Smartphone Video
di: Mahajan, Pranav, et al.
Pubblicazione: (2026)
di: Mahajan, Pranav, et al.
Pubblicazione: (2026)
See-Control: A Multimodal Agent Framework for Smartphone Interaction with a Robotic Arm
di: Zhao, Haoyu, et al.
Pubblicazione: (2025)
di: Zhao, Haoyu, et al.
Pubblicazione: (2025)
EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models
di: Xu, Ziwen, et al.
Pubblicazione: (2025)
di: Xu, Ziwen, et al.
Pubblicazione: (2025)
Trust in Vision-Language Models: Insights from a Participatory User Workshop
di: Chiatti, Agnese, et al.
Pubblicazione: (2025)
di: Chiatti, Agnese, et al.
Pubblicazione: (2025)
Do Vision Language Models Understand Human Engagement in Games?
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
Intelligent Control of Robotic X-ray Devices using a Language-promptable Digital Twin
di: Killeen, Benjamin D., et al.
Pubblicazione: (2024)
di: Killeen, Benjamin D., et al.
Pubblicazione: (2024)
Looking for a better fit? An Incremental Learning Multimodal Object Referencing Framework adapting to Individual Drivers
di: Gomaa, Amr, et al.
Pubblicazione: (2024)
di: Gomaa, Amr, et al.
Pubblicazione: (2024)
Using Game Engines and Machine Learning to Create Synthetic Satellite Imagery for a Tabletop Verification Exercise
di: Hoster, Johannes, et al.
Pubblicazione: (2024)
di: Hoster, Johannes, et al.
Pubblicazione: (2024)
Gemini Goes to Med School: Exploring the Capabilities of Multimodal Large Language Models on Medical Challenge Problems & Hallucinations
di: Pal, Ankit, et al.
Pubblicazione: (2024)
di: Pal, Ankit, et al.
Pubblicazione: (2024)
Analysis of the 2024 BraTS Meningioma Radiotherapy Planning Automated Segmentation Challenge
di: LaBella, Dominic, et al.
Pubblicazione: (2024)
di: LaBella, Dominic, et al.
Pubblicazione: (2024)
How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training
di: Ou, Yixin, et al.
Pubblicazione: (2025)
di: Ou, Yixin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
InterVLS: Interactive Model Understanding and Improvement with Vision-Language Surrogates
di: Huang, Jinbin, et al.
Pubblicazione: (2023) -
AI Guide Dog: Egocentric Path Prediction on Smartphone
di: Jadhav, Aishwarya, et al.
Pubblicazione: (2025) -
Vision-Language Models for Ergonomic Assessment of Manual Lifting Tasks: Estimating Horizontal and Vertical Hand Distances from RGB Video
di: Rajabi, Mohammad Sadra, et al.
Pubblicazione: (2026) -
Predicting and Explaining Mobile UI Tappability with Vision Modeling and Saliency Analysis
di: Schoop, Eldon, et al.
Pubblicazione: (2022) -
DOTA: Distributional Test-Time Adaptation of Vision-Language Models
di: Han, Zongbo, et al.
Pubblicazione: (2024)