Training a Vision Language Model as Smartphone Assistant
Fuente:
arXiv
Saved in:
| Main Authors: | Dorka, Nicolai, Marecki, Janusz, Anwar, Ammar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InterVLS: Interactive Model Understanding and Improvement with Vision-Language Surrogates
by: Huang, Jinbin, et al.
Published: (2023)
by: Huang, Jinbin, et al.
Published: (2023)
AI Guide Dog: Egocentric Path Prediction on Smartphone
by: Jadhav, Aishwarya, et al.
Published: (2025)
by: Jadhav, Aishwarya, et al.
Published: (2025)
Vision-Language Models for Ergonomic Assessment of Manual Lifting Tasks: Estimating Horizontal and Vertical Hand Distances from RGB Video
by: Rajabi, Mohammad Sadra, et al.
Published: (2026)
by: Rajabi, Mohammad Sadra, et al.
Published: (2026)
Predicting and Explaining Mobile UI Tappability with Vision Modeling and Saliency Analysis
by: Schoop, Eldon, et al.
Published: (2022)
by: Schoop, Eldon, et al.
Published: (2022)
DOTA: Distributional Test-Time Adaptation of Vision-Language Models
by: Han, Zongbo, et al.
Published: (2024)
by: Han, Zongbo, et al.
Published: (2024)
SpiderNets: Vision Models Predict Human Fear From Aversive Images
by: Pegler, Dominik, et al.
Published: (2025)
by: Pegler, Dominik, et al.
Published: (2025)
Bridging Human Concepts and Computer Vision for Explainable Face Verification
by: Doh, Miriam, et al.
Published: (2024)
by: Doh, Miriam, et al.
Published: (2024)
Screen2AX: Vision-Based Approach for Automatic macOS Accessibility Generation
by: Muryn, Viktor, et al.
Published: (2025)
by: Muryn, Viktor, et al.
Published: (2025)
BdSLW401: Transformer-Based Word-Level Bangla Sign Language Recognition Using Relative Quantization Encoding (RQE)
by: Rubaiyeat, Husne Ara, et al.
Published: (2025)
by: Rubaiyeat, Husne Ara, et al.
Published: (2025)
Semantic Approach to Quantifying the Consistency of Diffusion Model Image Generation
by: Bent, Brinnae
Published: (2024)
by: Bent, Brinnae
Published: (2024)
A Foundational Generative Model for Breast Ultrasound Image Analysis
by: Yu, Haojun, et al.
Published: (2025)
by: Yu, Haojun, et al.
Published: (2025)
LEDITS++: Limitless Image Editing using Text-to-Image Models
by: Brack, Manuel, et al.
Published: (2023)
by: Brack, Manuel, et al.
Published: (2023)
What's Producible May Not Be Reachable: Measuring the Steerability of Generative Models
by: Vafa, Keyon, et al.
Published: (2025)
by: Vafa, Keyon, et al.
Published: (2025)
ShelfHelp: Empowering Humans to Perform Vision-Independent Manipulation Tasks with a Socially Assistive Robotic Cane
by: Agrawal, Shivendra, et al.
Published: (2024)
by: Agrawal, Shivendra, et al.
Published: (2024)
I-CEE: Tailoring Explanations of Image Classification Models to User Expertise
by: Rong, Yao, et al.
Published: (2023)
by: Rong, Yao, et al.
Published: (2023)
Detoxifying Large Language Models via Knowledge Editing
by: Wang, Mengru, et al.
Published: (2024)
by: Wang, Mengru, et al.
Published: (2024)
Generating Synthetic Satellite Imagery for Rare Objects: An Empirical Comparison of Models and Metrics
by: Nguyen, Tuong Vy, et al.
Published: (2024)
by: Nguyen, Tuong Vy, et al.
Published: (2024)
HERO: Human-Feedback Efficient Reinforcement Learning for Online Diffusion Model Finetuning
by: Hiranaka, Ayano, et al.
Published: (2024)
by: Hiranaka, Ayano, et al.
Published: (2024)
Optimizing Small Language Models for In-Vehicle Function-Calling
by: Khiabani, Yahya Sowti, et al.
Published: (2025)
by: Khiabani, Yahya Sowti, et al.
Published: (2025)
ExpressEdit: Fast Editing of Stylized Facial Expressions with Diffusion Models in Photoshop
by: Tang, Kenan, et al.
Published: (2026)
by: Tang, Kenan, et al.
Published: (2026)
Interaction as Explanation: A User Interaction-based Method for Explaining Image Classification Models
by: Yun, Hyeonggeun
Published: (2024)
by: Yun, Hyeonggeun
Published: (2024)
Knowledge Mechanisms in Large Language Models: A Survey and Perspective
by: Wang, Mengru, et al.
Published: (2024)
by: Wang, Mengru, et al.
Published: (2024)
A Comprehensive Study of Knowledge Editing for Large Language Models
by: Zhang, Ningyu, et al.
Published: (2024)
by: Zhang, Ningyu, et al.
Published: (2024)
ReLearn: Unlearning via Learning for Large Language Models
by: Xu, Haoming, et al.
Published: (2025)
by: Xu, Haoming, et al.
Published: (2025)
Generating Synthetic Satellite Imagery With Deep-Learning Text-to-Image Models -- Technical Challenges and Implications for Monitoring and Verification
by: Nguyen, Tuong Vy, et al.
Published: (2024)
by: Nguyen, Tuong Vy, et al.
Published: (2024)
AutoTour: Automatic Photo Tour Guide with Smartphones and LLMs
by: Xu, Huatao, et al.
Published: (2026)
by: Xu, Huatao, et al.
Published: (2026)
Zero-shot Emotion Annotation in Facial Images Using Large Multimodal Models: Benchmarking and Prospects for Multi-Class, Multi-Frame Approaches
by: Zhang, He, et al.
Published: (2025)
by: Zhang, He, et al.
Published: (2025)
InstructEdit: Instruction-based Knowledge Editing for Large Language Models
by: Zhang, Ningyu, et al.
Published: (2024)
by: Zhang, Ningyu, et al.
Published: (2024)
Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
by: Natalie, Rosiana, et al.
Published: (2025)
by: Natalie, Rosiana, et al.
Published: (2025)
Quantitative Movement Testing: Measuring Patient Movements from a Single Smartphone Video
by: Mahajan, Pranav, et al.
Published: (2026)
by: Mahajan, Pranav, et al.
Published: (2026)
See-Control: A Multimodal Agent Framework for Smartphone Interaction with a Robotic Arm
by: Zhao, Haoyu, et al.
Published: (2025)
by: Zhao, Haoyu, et al.
Published: (2025)
EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models
by: Xu, Ziwen, et al.
Published: (2025)
by: Xu, Ziwen, et al.
Published: (2025)
Trust in Vision-Language Models: Insights from a Participatory User Workshop
by: Chiatti, Agnese, et al.
Published: (2025)
by: Chiatti, Agnese, et al.
Published: (2025)
Do Vision Language Models Understand Human Engagement in Games?
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
Intelligent Control of Robotic X-ray Devices using a Language-promptable Digital Twin
by: Killeen, Benjamin D., et al.
Published: (2024)
by: Killeen, Benjamin D., et al.
Published: (2024)
Looking for a better fit? An Incremental Learning Multimodal Object Referencing Framework adapting to Individual Drivers
by: Gomaa, Amr, et al.
Published: (2024)
by: Gomaa, Amr, et al.
Published: (2024)
Using Game Engines and Machine Learning to Create Synthetic Satellite Imagery for a Tabletop Verification Exercise
by: Hoster, Johannes, et al.
Published: (2024)
by: Hoster, Johannes, et al.
Published: (2024)
Gemini Goes to Med School: Exploring the Capabilities of Multimodal Large Language Models on Medical Challenge Problems & Hallucinations
by: Pal, Ankit, et al.
Published: (2024)
by: Pal, Ankit, et al.
Published: (2024)
Analysis of the 2024 BraTS Meningioma Radiotherapy Planning Automated Segmentation Challenge
by: LaBella, Dominic, et al.
Published: (2024)
by: LaBella, Dominic, et al.
Published: (2024)
How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training
by: Ou, Yixin, et al.
Published: (2025)
by: Ou, Yixin, et al.
Published: (2025)
Similar Items
-
InterVLS: Interactive Model Understanding and Improvement with Vision-Language Surrogates
by: Huang, Jinbin, et al.
Published: (2023) -
AI Guide Dog: Egocentric Path Prediction on Smartphone
by: Jadhav, Aishwarya, et al.
Published: (2025) -
Vision-Language Models for Ergonomic Assessment of Manual Lifting Tasks: Estimating Horizontal and Vertical Hand Distances from RGB Video
by: Rajabi, Mohammad Sadra, et al.
Published: (2026) -
Predicting and Explaining Mobile UI Tappability with Vision Modeling and Saliency Analysis
by: Schoop, Eldon, et al.
Published: (2022) -
DOTA: Distributional Test-Time Adaptation of Vision-Language Models
by: Han, Zongbo, et al.
Published: (2024)