Guardado en:
| Autores principales: | Li, Chentao, Gao, Zirui, Gao, Mingze, Ren, Yinglian, Feng, Jianjiang, Zhou, Jie |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2604.21461 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MP-GUI: Modality Perception with MLLMs for GUI Understanding
por: Wang, Ziwei, et al.
Publicado: (2025)
por: Wang, Ziwei, et al.
Publicado: (2025)
SimVecVis: A Dataset for Enhancing MLLMs in Visualization Understanding
por: Liu, Can, et al.
Publicado: (2025)
por: Liu, Can, et al.
Publicado: (2025)
An Egocentric Vision-Language Model based Portable Real-time Smart Assistant
por: Huang, Yifei, et al.
Publicado: (2025)
por: Huang, Yifei, et al.
Publicado: (2025)
Detecting Clues for Skill Levels and Machine Operation Difficulty from Egocentric Vision
por: Long-fei, Chen, et al.
Publicado: (2019)
por: Long-fei, Chen, et al.
Publicado: (2019)
EgoPressure: A Dataset for Hand Pressure and Pose Estimation in Egocentric Vision
por: Zhao, Yiming, et al.
Publicado: (2024)
por: Zhao, Yiming, et al.
Publicado: (2024)
RWKV-UI: UI Understanding with Enhanced Perception and Reasoning
por: Yang, Jiaxi, et al.
Publicado: (2025)
por: Yang, Jiaxi, et al.
Publicado: (2025)
egoEMOTION: Egocentric Vision and Physiological Signals for Emotion and Personality Recognition in Real-World Tasks
por: Jammot, Matthias, et al.
Publicado: (2025)
por: Jammot, Matthias, et al.
Publicado: (2025)
Do Vision Language Models Understand Human Engagement in Games?
por: Wang, Ziyi, et al.
Publicado: (2026)
por: Wang, Ziyi, et al.
Publicado: (2026)
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
por: Wang, Siting, et al.
Publicado: (2025)
por: Wang, Siting, et al.
Publicado: (2025)
UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation
por: Han, Tianhao, et al.
Publicado: (2026)
por: Han, Tianhao, et al.
Publicado: (2026)
WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity Recognition
por: Bock, Marius, et al.
Publicado: (2023)
por: Bock, Marius, et al.
Publicado: (2023)
PEACE: Empowering Geologic Map Holistic Understanding with MLLMs
por: Huang, Yangyu, et al.
Publicado: (2025)
por: Huang, Yangyu, et al.
Publicado: (2025)
A Powered Prosthetic Hand with Vision System for Enhancing the Anthropopathic Grasp
por: Xu, Yansong, et al.
Publicado: (2024)
por: Xu, Yansong, et al.
Publicado: (2024)
DeltaDorsal: Enhancing Hand Pose Estimation with Dorsal Features in Egocentric Views
por: Huang, William, et al.
Publicado: (2026)
por: Huang, William, et al.
Publicado: (2026)
Interact with me: Joint Egocentric Forecasting of Intent to Interact, Attitude and Social Actions
por: Bian, Tongfei, et al.
Publicado: (2024)
por: Bian, Tongfei, et al.
Publicado: (2024)
Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning
por: Li, Xun, et al.
Publicado: (2025)
por: Li, Xun, et al.
Publicado: (2025)
Real Eyes Realize Faster: Gaze Stability and Pupil Novelty for Efficient Egocentric Learning
por: Subramanian, Ajan, et al.
Publicado: (2026)
por: Subramanian, Ajan, et al.
Publicado: (2026)
VLLMs Provide Better Context for Emotion Understanding Through Common Sense Reasoning
por: Xenos, Alexandros, et al.
Publicado: (2024)
por: Xenos, Alexandros, et al.
Publicado: (2024)
Detecting Activities of Daily Living in Egocentric Video to Contextualize Hand Use at Home in Outpatient Neurorehabilitation Settings
por: Kadambi, Adesh, et al.
Publicado: (2024)
por: Kadambi, Adesh, et al.
Publicado: (2024)
The Visual Experience Dataset: Over 200 Recorded Hours of Integrated Eye Movement, Odometry, and Egocentric Video
por: Greene, Michelle R., et al.
Publicado: (2024)
por: Greene, Michelle R., et al.
Publicado: (2024)
Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
por: Fei, Hao, et al.
Publicado: (2024)
por: Fei, Hao, et al.
Publicado: (2024)
Benchmarking Adaptive Intelligence and Computer Vision on Human-Robot Collaboration
por: Saraj, Salaar, et al.
Publicado: (2024)
por: Saraj, Salaar, et al.
Publicado: (2024)
Visually Grounded Narratives: Reducing Cognitive Burden in Researcher-Participant Interaction
por: Wu, Runtong, et al.
Publicado: (2025)
por: Wu, Runtong, et al.
Publicado: (2025)
Interactivity x Explainability: Toward Understanding How Interactivity Can Improve Computer Vision Explanations
por: Panigrahi, Indu, et al.
Publicado: (2025)
por: Panigrahi, Indu, et al.
Publicado: (2025)
Egocentric Co-Pilot: Web-Native Smart-Glasses Agents for Assistive Egocentric AI
por: Yang, Sicheng, et al.
Publicado: (2026)
por: Yang, Sicheng, et al.
Publicado: (2026)
Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation
por: Yang, Sicheng, et al.
Publicado: (2025)
por: Yang, Sicheng, et al.
Publicado: (2025)
Stratified Avatar Generation from Sparse Observations
por: Feng, Han, et al.
Publicado: (2024)
por: Feng, Han, et al.
Publicado: (2024)
Referring Human Pose and Mask Estimation in the Wild
por: Miao, Bo, et al.
Publicado: (2024)
por: Miao, Bo, et al.
Publicado: (2024)
Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding
por: Guo, Hao, et al.
Publicado: (2025)
por: Guo, Hao, et al.
Publicado: (2025)
FineSkiing: A Fine-grained Benchmark for Skiing Action Quality Assessment
por: Zhang, Yongji, et al.
Publicado: (2025)
por: Zhang, Yongji, et al.
Publicado: (2025)
EgoPoseFormer v2: Accurate Egocentric Human Motion Estimation for AR/VR
por: Li, Zhenyu, et al.
Publicado: (2026)
por: Li, Zhenyu, et al.
Publicado: (2026)
WristPP: A Wrist-Worn System for Hand Pose And Pressure Estimation
por: Xi, Ziheng, et al.
Publicado: (2026)
por: Xi, Ziheng, et al.
Publicado: (2026)
3DArticCyclists: Generating Synthetic Articulated 8D Pose-Controllable Cyclist Data for Computer Vision Applications
por: Corral-Soto, Eduardo R., et al.
Publicado: (2024)
por: Corral-Soto, Eduardo R., et al.
Publicado: (2024)
InterVLS: Interactive Model Understanding and Improvement with Vision-Language Surrogates
por: Huang, Jinbin, et al.
Publicado: (2023)
por: Huang, Jinbin, et al.
Publicado: (2023)
ThermoHands: A Benchmark for 3D Hand Pose Estimation from Egocentric Thermal Images
por: Ding, Fangqiang, et al.
Publicado: (2024)
por: Ding, Fangqiang, et al.
Publicado: (2024)
OMG-Bench: A New Challenging Benchmark for Skeleton-based Online Micro Hand Gesture Recognition
por: Chang, Haochen, et al.
Publicado: (2025)
por: Chang, Haochen, et al.
Publicado: (2025)
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
por: Verma, Arnav, et al.
Publicado: (2025)
por: Verma, Arnav, et al.
Publicado: (2025)
Towards Scalable Web Accessibility Audit with MLLMs as Copilots
por: Gu, Ming, et al.
Publicado: (2025)
por: Gu, Ming, et al.
Publicado: (2025)
VR-GS: A Physical Dynamics-Aware Interactive Gaussian Splatting System in Virtual Reality
por: Jiang, Ying, et al.
Publicado: (2024)
por: Jiang, Ying, et al.
Publicado: (2024)
Reasoning3D -- Grounding and Reasoning in 3D: Fine-Grained Zero-Shot Open-Vocabulary 3D Reasoning Part Segmentation via Large Vision-Language Models
por: Chen, Tianrun, et al.
Publicado: (2024)
por: Chen, Tianrun, et al.
Publicado: (2024)
Ejemplares similares
-
MP-GUI: Modality Perception with MLLMs for GUI Understanding
por: Wang, Ziwei, et al.
Publicado: (2025) -
SimVecVis: A Dataset for Enhancing MLLMs in Visualization Understanding
por: Liu, Can, et al.
Publicado: (2025) -
An Egocentric Vision-Language Model based Portable Real-time Smart Assistant
por: Huang, Yifei, et al.
Publicado: (2025) -
Detecting Clues for Skill Levels and Machine Operation Difficulty from Egocentric Vision
por: Long-fei, Chen, et al.
Publicado: (2019) -
EgoPressure: A Dataset for Hand Pressure and Pose Estimation in Egocentric Vision
por: Zhao, Yiming, et al.
Publicado: (2024)