Evaluating Visual Prompts with Eye-Tracking Data for MLLM-Based Human Activity Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Choi, Jae Young, Kim, Seon Gyeom, Yoon, Hyungjun, Lee, Taeckyung, Lee, Donggun, Chung, Jaeryung, Kil, Jihyung, Rossi, Ryan, Lee, Sung-Ju, Lee, Tak Yeon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts
von: Kim, Seon Gyeom, et al.
Veröffentlicht: (2025)
von: Kim, Seon Gyeom, et al.
Veröffentlicht: (2025)
Understanding the Impact of Spatial Immersion in Web Data Stories
von: Kim, Seon Gyeom, et al.
Veröffentlicht: (2024)
von: Kim, Seon Gyeom, et al.
Veröffentlicht: (2024)
By My Eyes: Grounding Multimodal Large Language Models with Sensor Data via Visual Prompting
von: Yoon, Hyungjun, et al.
Veröffentlicht: (2024)
von: Yoon, Hyungjun, et al.
Veröffentlicht: (2024)
AETTA: Label-Free Accuracy Estimation for Test-Time Adaptation
von: Lee, Taeckyung, et al.
Veröffentlicht: (2024)
von: Lee, Taeckyung, et al.
Veröffentlicht: (2024)
Designing Prompt Analytics Dashboards to Analyze Student-ChatGPT Interactions in EFL Writing
von: Kim, Minsun, et al.
Veröffentlicht: (2024)
von: Kim, Minsun, et al.
Veröffentlicht: (2024)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
Bridging Bond Beyond Life: Designing VR Memorial Space with Stakeholder Collaboration via Research through Design
von: Bae, Heejae, et al.
Veröffentlicht: (2025)
von: Bae, Heejae, et al.
Veröffentlicht: (2025)
Extending Segment Anything Model into Auditory and Temporal Dimensions for Audio-Visual Segmentation
von: Seon, Juhyeong, et al.
Veröffentlicht: (2024)
von: Seon, Juhyeong, et al.
Veröffentlicht: (2024)
LLM-Driven Learning Analytics Dashboard for Teachers in EFL Writing Education
von: Kim, Minsun, et al.
Veröffentlicht: (2024)
von: Kim, Minsun, et al.
Veröffentlicht: (2024)
Beyond Hearing: Learning Task-Agnostic ExG Representations from Earphones via Physiology-Informed Tokenization
von: Yoon, Hyungjun, et al.
Veröffentlicht: (2025)
von: Yoon, Hyungjun, et al.
Veröffentlicht: (2025)
Building UI/UX Dataset for Dark Pattern Detection and YOLOv12x-based Real-Time Object Recognition Detection System
von: Jang, Se-Young, et al.
Veröffentlicht: (2025)
von: Jang, Se-Young, et al.
Veröffentlicht: (2025)
Mix from Failure: Confusion-Pairing Mixup for Long-Tailed Recognition
von: Yoon, Youngseok, et al.
Veröffentlicht: (2024)
von: Yoon, Youngseok, et al.
Veröffentlicht: (2024)
Designing and Evaluating In-Vehicle Temporal Decoupling Pointing System for Selecting External Object
von: Pyun, Jaehoon, et al.
Veröffentlicht: (2022)
von: Pyun, Jaehoon, et al.
Veröffentlicht: (2022)
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
von: Slyman, Eric, et al.
Veröffentlicht: (2025)
von: Slyman, Eric, et al.
Veröffentlicht: (2025)
Representation Shift: Unifying Token Compression with FlashAttention
von: Choi, Joonmyung, et al.
Veröffentlicht: (2025)
von: Choi, Joonmyung, et al.
Veröffentlicht: (2025)
Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval
von: lee, Jihyung, et al.
Veröffentlicht: (2026)
von: lee, Jihyung, et al.
Veröffentlicht: (2026)
ARES: Alternating Reinforcement Learning and Supervised Fine-Tuning for Enhanced Multi-Modal Chain-of-Thought Reasoning Through Diverse AI Feedback
von: Byun, Ju-Seung, et al.
Veröffentlicht: (2024)
von: Byun, Ju-Seung, et al.
Veröffentlicht: (2024)
SemCity: Semantic Scene Generation with Triplane Diffusion
von: Lee, Jumin, et al.
Veröffentlicht: (2024)
von: Lee, Jumin, et al.
Veröffentlicht: (2024)
DCG-SQL: Enhancing In-Context Learning for Text-to-SQL with Deep Contextual Schema Link Graph
von: Lee, Jihyung, et al.
Veröffentlicht: (2025)
von: Lee, Jihyung, et al.
Veröffentlicht: (2025)
BECoTTA: Input-dependent Online Blending of Experts for Continual Test-time Adaptation
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
Optimizing Data Delivery: Insights from User Preferences on Visuals, Tables, and Text
von: Luera, Reuben, et al.
Veröffentlicht: (2024)
von: Luera, Reuben, et al.
Veröffentlicht: (2024)
KoBLEX: Open Legal Question Answering with Multi-hop Reasoning
von: Lee, Jihyung, et al.
Veröffentlicht: (2025)
von: Lee, Jihyung, et al.
Veröffentlicht: (2025)
SCANNER: Knowledge-Enhanced Approach for Robust Multi-modal Named Entity Recognition of Unseen Entities
von: Ok, Hyunjong, et al.
Veröffentlicht: (2024)
von: Ok, Hyunjong, et al.
Veröffentlicht: (2024)
PySDTest: a Python/Stata Package for Stochastic Dominance Tests
von: Lee, Kyungho, et al.
Veröffentlicht: (2023)
von: Lee, Kyungho, et al.
Veröffentlicht: (2023)
FireSenseNet: A Dual-Branch CNN with Cross-Attentive Feature Interaction for Next-Day Wildfire Spread Prediction
von: Han, Jinzhen, et al.
Veröffentlicht: (2026)
von: Han, Jinzhen, et al.
Veröffentlicht: (2026)
DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning
von: Choi, Joonmyung, et al.
Veröffentlicht: (2026)
von: Choi, Joonmyung, et al.
Veröffentlicht: (2026)
A Distributed Consensus Algorithm for Prioritizing Autonomous Vehicle Passing at Unsignalized Intersections under Mixed Traffic
von: Lee, Younjeong, et al.
Veröffentlicht: (2025)
von: Lee, Younjeong, et al.
Veröffentlicht: (2025)
II-MMR: Identifying and Improving Multi-modal Multi-hop Reasoning in Visual Question Answering
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
Regularizing Dynamic Radiance Fields with Kinematic Fields
von: Im, Woobin, et al.
Veröffentlicht: (2024)
von: Im, Woobin, et al.
Veröffentlicht: (2024)
FedPOD: the deployable units of training for federated learning
von: Kim, Daewoon, et al.
Veröffentlicht: (2025)
von: Kim, Daewoon, et al.
Veröffentlicht: (2025)
Recover as It is Designed to Be: Recovering from Compatibility Mobile App Crashes by Reusing User Flows
von: Kim, Donghwi, et al.
Veröffentlicht: (2024)
von: Kim, Donghwi, et al.
Veröffentlicht: (2024)
GS-Scale: Unlocking Large-Scale 3D Gaussian Splatting Training via Host Offloading
von: Lee, Donghyun, et al.
Veröffentlicht: (2025)
von: Lee, Donghyun, et al.
Veröffentlicht: (2025)
FastPoint: Accelerating 3D Point Cloud Model Inference via Sample Point Distance Prediction
von: Lee, Donghyun, et al.
Veröffentlicht: (2025)
von: Lee, Donghyun, et al.
Veröffentlicht: (2025)
MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces
von: Luera, Reuben A., et al.
Veröffentlicht: (2025)
von: Luera, Reuben A., et al.
Veröffentlicht: (2025)
VERIA: Verification-Centric Multimodal Instance Augmentation for Long-Tailed 3D Object Detection
von: Lee, Jumin, et al.
Veröffentlicht: (2026)
von: Lee, Jumin, et al.
Veröffentlicht: (2026)
RBF-Solver: A Multistep Sampler for Diffusion Probabilistic Models via Radial Basis Functions
von: Park, Soochul, et al.
Veröffentlicht: (2026)
von: Park, Soochul, et al.
Veröffentlicht: (2026)
GFlowPO: Generative Flow Network as a Language Model Prompt Optimizer
von: Cho, Junmo, et al.
Veröffentlicht: (2026)
von: Cho, Junmo, et al.
Veröffentlicht: (2026)
The Effect of Empathic Expression Levels in Virtual Human Interaction: A Controlled Experiment
von: Park, Sung, et al.
Veröffentlicht: (2025)
von: Park, Sung, et al.
Veröffentlicht: (2025)
Harmful Suicide Content Detection
von: Park, Kyumin, et al.
Veröffentlicht: (2024)
von: Park, Kyumin, et al.
Veröffentlicht: (2024)
EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens
von: Hwang, Sunil, et al.
Veröffentlicht: (2022)
von: Hwang, Sunil, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts
von: Kim, Seon Gyeom, et al.
Veröffentlicht: (2025) -
Understanding the Impact of Spatial Immersion in Web Data Stories
von: Kim, Seon Gyeom, et al.
Veröffentlicht: (2024) -
By My Eyes: Grounding Multimodal Large Language Models with Sensor Data via Visual Prompting
von: Yoon, Hyungjun, et al.
Veröffentlicht: (2024) -
AETTA: Label-Free Accuracy Estimation for Test-Time Adaptation
von: Lee, Taeckyung, et al.
Veröffentlicht: (2024) -
Designing Prompt Analytics Dashboards to Analyze Student-ChatGPT Interactions in EFL Writing
von: Kim, Minsun, et al.
Veröffentlicht: (2024)