Robust Egocentric Visual Attention Prediction Through Language-guided Scene Context-aware Learning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Park, Sungjune, Mao, Hongda, Chen, Qingshuang, Ro, Yong Man, Kim, Yelin |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Language-guided Learning for Object Detection Tackling Multiple Variations in Aerial Images
par: Park, Sungjune, et autres
Publié: (2025)
par: Park, Sungjune, et autres
Publié: (2025)
Integrating Language-Derived Appearance Elements with Visual Cues in Pedestrian Detection
par: Park, Sungjune, et autres
Publié: (2023)
par: Park, Sungjune, et autres
Publié: (2023)
Robust Pedestrian Detection via Constructing Versatile Pedestrian Knowledge Bank
par: Park, Sungjune, et autres
Publié: (2024)
par: Park, Sungjune, et autres
Publié: (2024)
DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes
par: Park, Sungjune, et autres
Publié: (2025)
par: Park, Sungjune, et autres
Publié: (2025)
studentSplat: Your Student Model Learns Single-view 3D Gaussian Splatting
par: Pan, Yimu, et autres
Publié: (2026)
par: Pan, Yimu, et autres
Publié: (2026)
Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling
par: Park, Sungjune, et autres
Publié: (2025)
par: Park, Sungjune, et autres
Publié: (2025)
Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues
par: Park, Beomchan, et autres
Publié: (2026)
par: Park, Beomchan, et autres
Publié: (2026)
Empathetic Response in Audio-Visual Conversations Using Emotion Preference Optimization and MambaCompressor
par: Kim, Yeonju, et autres
Publié: (2024)
par: Kim, Yeonju, et autres
Publié: (2024)
GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory
par: Yeo, Jeong Hun, et autres
Publié: (2025)
par: Yeo, Jeong Hun, et autres
Publié: (2025)
Towards Inclusive Communication: A Unified Framework for Generating Spoken Language from Sign, Lip, and Audio
par: Yeo, Jeong Hun, et autres
Publié: (2025)
par: Yeo, Jeong Hun, et autres
Publié: (2025)
EgoNav: Egocentric Scene-aware Human Trajectory Prediction
par: Wang, Weizhuo, et autres
Publié: (2024)
par: Wang, Weizhuo, et autres
Publié: (2024)
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding
par: Lee, Hosu, et autres
Publié: (2025)
par: Lee, Hosu, et autres
Publié: (2025)
Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models
par: Lee, Byung-Kwan, et autres
Publié: (2024)
par: Lee, Byung-Kwan, et autres
Publié: (2024)
CoLLaVO: Crayon Large Language and Vision mOdel
par: Lee, Byung-Kwan, et autres
Publié: (2024)
par: Lee, Byung-Kwan, et autres
Publié: (2024)
MoAI: Mixture of All Intelligence for Large Language and Vision Models
par: Lee, Byung-Kwan, et autres
Publié: (2024)
par: Lee, Byung-Kwan, et autres
Publié: (2024)
Where Visual Speech Meets Language: VSP-LLM Framework for Efficient and Context-Aware Visual Speech Processing
par: Yeo, Jeong Hun, et autres
Publié: (2024)
par: Yeo, Jeong Hun, et autres
Publié: (2024)
PRVQL: Progressive Knowledge-guided Refinement for Robust Egocentric Visual Query Localization
par: Fan, Bing, et autres
Publié: (2025)
par: Fan, Bing, et autres
Publié: (2025)
Phantom of Latent for Large Language and Vision Models
par: Lee, Byung-Kwan, et autres
Publié: (2024)
par: Lee, Byung-Kwan, et autres
Publié: (2024)
AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation
par: Choi, Jeongsoo, et autres
Publié: (2023)
par: Choi, Jeongsoo, et autres
Publié: (2023)
Revisiting Misalignment in Multispectral Pedestrian Detection: A Language-Driven Approach for Cross-modal Alignment Fusion
par: Kim, Taeheon, et autres
Publié: (2024)
par: Kim, Taeheon, et autres
Publié: (2024)
Efficient Training for Multilingual Visual Speech Recognition: Pre-training with Discretized Visual Speech Representation
par: Kim, Minsu, et autres
Publié: (2024)
par: Kim, Minsu, et autres
Publié: (2024)
Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
par: Yeo, Jeong Hun, et autres
Publié: (2025)
par: Yeo, Jeong Hun, et autres
Publié: (2025)
Visual Speech Recognition for Languages with Limited Labeled Data using Automatic Labels from Whisper
par: Yeo, Jeong Hun, et autres
Publié: (2023)
par: Yeo, Jeong Hun, et autres
Publié: (2023)
HENASY: Learning to Assemble Scene-Entities for Egocentric Video-Language Model
par: Vo, Khoa, et autres
Publié: (2024)
par: Vo, Khoa, et autres
Publié: (2024)
Scene-aware Human Motion Forecasting via Mutual Distance Prediction
par: Xing, Chaoyue, et autres
Publié: (2023)
par: Xing, Chaoyue, et autres
Publié: (2023)
HERO-VQL: Hierarchical, Egocentric and Robust Visual Query Localization
par: Chang, Joohyun, et autres
Publié: (2025)
par: Chang, Joohyun, et autres
Publié: (2025)
Exploring Phonetic Context-Aware Lip-Sync For Talking Face Generation
par: Park, Se Jin, et autres
Publié: (2023)
par: Park, Se Jin, et autres
Publié: (2023)
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking
par: Chung, Sangyun, et autres
Publié: (2024)
par: Chung, Sangyun, et autres
Publié: (2024)
AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues
par: Park, Se Jin, et autres
Publié: (2024)
par: Park, Se Jin, et autres
Publié: (2024)
CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal Models
par: Kim, Junho, et autres
Publié: (2024)
par: Kim, Junho, et autres
Publié: (2024)
MSCoTDet: Language-driven Multi-modal Fusion for Improved Multispectral Pedestrian Detection
par: Kim, Taeheon, et autres
Publié: (2024)
par: Kim, Taeheon, et autres
Publié: (2024)
LaserHuman: Language-guided Scene-aware Human Motion Generation in Free Environment
par: Cong, Peishan, et autres
Publié: (2024)
par: Cong, Peishan, et autres
Publié: (2024)
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
par: Jeong, Boseung, et autres
Publié: (2025)
par: Jeong, Boseung, et autres
Publié: (2025)
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
par: Kim, Junho, et autres
Publié: (2024)
par: Kim, Junho, et autres
Publié: (2024)
SPARK: Multi-Vision Sensor Perception and Reasoning Benchmark for Large-scale Vision-Language Models
par: Yu, Youngjoon, et autres
Publié: (2024)
par: Yu, Youngjoon, et autres
Publié: (2024)
Object-aware Sound Source Localization via Audio-Visual Scene Understanding
par: Um, Sung Jin, et autres
Publié: (2025)
par: Um, Sung Jin, et autres
Publié: (2025)
SCENIC: Scene-aware Semantic Navigation with Instruction-guided Control
par: Zhang, Xiaohan, et autres
Publié: (2024)
par: Zhang, Xiaohan, et autres
Publié: (2024)
Unified Reinforcement and Imitation Learning for Vision-Language Models
par: Lee, Byung-Kwan, et autres
Publié: (2025)
par: Lee, Byung-Kwan, et autres
Publié: (2025)
TroL: Traversal of Layers for Large Language and Vision Models
par: Lee, Byung-Kwan, et autres
Publié: (2024)
par: Lee, Byung-Kwan, et autres
Publié: (2024)
Spatially Prompted Visual Trajectory Prediction for Egocentric Manipulation
par: Li, Yifan, et autres
Publié: (2026)
par: Li, Yifan, et autres
Publié: (2026)
Documents similaires
-
Language-guided Learning for Object Detection Tackling Multiple Variations in Aerial Images
par: Park, Sungjune, et autres
Publié: (2025) -
Integrating Language-Derived Appearance Elements with Visual Cues in Pedestrian Detection
par: Park, Sungjune, et autres
Publié: (2023) -
Robust Pedestrian Detection via Constructing Versatile Pedestrian Knowledge Bank
par: Park, Sungjune, et autres
Publié: (2024) -
DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes
par: Park, Sungjune, et autres
Publié: (2025) -
studentSplat: Your Student Model Learns Single-view 3D Gaussian Splatting
par: Pan, Yimu, et autres
Publié: (2026)