Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span
Fuente:
arXiv
Guardado en:
| Autores principales: | Yun, Heeseung, Na, Joonil, Kim, Jaeyeon, Murdock, Calvin, Kim, Gunhee |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
por: Yun, Heeseung, et al.
Publicado: (2024)
por: Yun, Heeseung, et al.
Publicado: (2024)
Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates
por: Ahn, Jaewoo, et al.
Publicado: (2025)
por: Ahn, Jaewoo, et al.
Publicado: (2025)
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
por: Ahn, Jaewoo, et al.
Publicado: (2025)
por: Ahn, Jaewoo, et al.
Publicado: (2025)
Gaussian Blending: Rethinking Alpha Blending in 3D Gaussian Splatting
por: Koo, Junseo, et al.
Publicado: (2025)
por: Koo, Junseo, et al.
Publicado: (2025)
Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation
por: Lai, Bolin, et al.
Publicado: (2023)
por: Lai, Bolin, et al.
Publicado: (2023)
Improving Cone-Beam CT Image Quality with Knowledge Distillation-Enhanced Diffusion Model in Imbalanced Data Settings
por: Hwang, Joonil, et al.
Publicado: (2024)
por: Hwang, Joonil, et al.
Publicado: (2024)
MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering
por: Song, Seokwon, et al.
Publicado: (2025)
por: Song, Seokwon, et al.
Publicado: (2025)
Bi-directional Contextual Attention for 3D Dense Captioning
por: Kim, Minjung, et al.
Publicado: (2024)
por: Kim, Minjung, et al.
Publicado: (2024)
EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting
por: Choi, Jaeyoung, et al.
Publicado: (2026)
por: Choi, Jaeyoung, et al.
Publicado: (2026)
HalLoc: Token-level Localization of Hallucinations for Vision Language Models
por: Park, Eunkyu, et al.
Publicado: (2025)
por: Park, Eunkyu, et al.
Publicado: (2025)
See It All: Contextualized Late Aggregation for 3D Dense Captioning
por: Kim, Minjung, et al.
Publicado: (2024)
por: Kim, Minjung, et al.
Publicado: (2024)
Egocentric Gaze Estimation via Neck-Mounted Camera
por: Huang, Haoyu, et al.
Publicado: (2026)
por: Huang, Haoyu, et al.
Publicado: (2026)
ARGaze: Autoregressive Transformers for Online Egocentric Gaze Estimation
por: Li, Jia, et al.
Publicado: (2026)
por: Li, Jia, et al.
Publicado: (2026)
EgoCampus: Egocentric Pedestrian Eye Gaze Model and Dataset
por: John, Ronan, et al.
Publicado: (2025)
por: John, Ronan, et al.
Publicado: (2025)
In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation
por: Lai, Bolin, et al.
Publicado: (2022)
por: Lai, Bolin, et al.
Publicado: (2022)
Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents
por: Baik, Sangwon, et al.
Publicado: (2026)
por: Baik, Sangwon, et al.
Publicado: (2026)
Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities
por: Mazzamuto, Michele, et al.
Publicado: (2024)
por: Mazzamuto, Michele, et al.
Publicado: (2024)
GazeMotion: Gaze-guided Human Motion Forecasting
por: Hu, Zhiming, et al.
Publicado: (2024)
por: Hu, Zhiming, et al.
Publicado: (2024)
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
por: Pani, Anupam, et al.
Publicado: (2025)
por: Pani, Anupam, et al.
Publicado: (2025)
SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting
por: Liu, Ruicong, et al.
Publicado: (2025)
por: Liu, Ruicong, et al.
Publicado: (2025)
Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator
por: Shin, Chaehun, et al.
Publicado: (2024)
por: Shin, Chaehun, et al.
Publicado: (2024)
Exploring High-Order Self-Similarity for Video Understanding
por: Kim, Manjin, et al.
Publicado: (2026)
por: Kim, Manjin, et al.
Publicado: (2026)
ViSAGe: Video-to-Spatial Audio Generation
por: Kim, Jaeyeon, et al.
Publicado: (2025)
por: Kim, Jaeyeon, et al.
Publicado: (2025)
Can Language Models Laugh at YouTube Short-form Videos?
por: Ko, Dayoon, et al.
Publicado: (2023)
por: Ko, Dayoon, et al.
Publicado: (2023)
In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
por: Peng, Taiying, et al.
Publicado: (2025)
por: Peng, Taiying, et al.
Publicado: (2025)
Personalized Federated Learning for Egocentric Video Gaze Estimation with Comprehensive Parameter Frezzing
por: Feng, Yuhu, et al.
Publicado: (2025)
por: Feng, Yuhu, et al.
Publicado: (2025)
HERO-VQL: Hierarchical, Egocentric and Robust Visual Query Localization
por: Chang, Joohyun, et al.
Publicado: (2025)
por: Chang, Joohyun, et al.
Publicado: (2025)
Causal Representation-Based Domain Generalization on Gaze Estimation
por: Kim, Younghan, et al.
Publicado: (2024)
por: Kim, Younghan, et al.
Publicado: (2024)
GazeShift: Unsupervised Gaze Estimation and Dataset for VR
por: Shapira, Gil, et al.
Publicado: (2026)
por: Shapira, Gil, et al.
Publicado: (2026)
Gaze-Guided 3D Hand Motion Prediction for Detecting Intent in Egocentric Grasping Tasks
por: He, Yufei, et al.
Publicado: (2025)
por: He, Yufei, et al.
Publicado: (2025)
ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams
por: Kim, Chris Dongjoo, et al.
Publicado: (2025)
por: Kim, Chris Dongjoo, et al.
Publicado: (2025)
Neural Radiance and Gaze Fields for Visual Attention Modeling in 3D Environments
por: Chubarau, Andrei, et al.
Publicado: (2025)
por: Chubarau, Andrei, et al.
Publicado: (2025)
FedMeNF: Privacy-Preserving Federated Meta-Learning for Neural Fields
por: Yun, Junhyeog, et al.
Publicado: (2025)
por: Yun, Junhyeog, et al.
Publicado: (2025)
Eyes on Target: Gaze-Aware Object Detection in Egocentric Video
por: Lall, Vishakha, et al.
Publicado: (2025)
por: Lall, Vishakha, et al.
Publicado: (2025)
ESR-NeRF: Emissive Source Reconstruction Using LDR Multi-view Images
por: Jeong, Jinseo, et al.
Publicado: (2024)
por: Jeong, Jinseo, et al.
Publicado: (2024)
Pandora: Articulated 3D Scene Graphs from Egocentric Vision
por: Yu, Alan, et al.
Publicado: (2026)
por: Yu, Alan, et al.
Publicado: (2026)
Leveraging Gaze and Set-of-Mark in VLLMs for Human-Object Interaction Anticipation from Egocentric Videos
por: Materia, Daniele, et al.
Publicado: (2026)
por: Materia, Daniele, et al.
Publicado: (2026)
EyeCue: Driver Cognitive Distraction Detection via Gaze-Empowered Egocentric Video Understanding
por: Zhang, Lang, et al.
Publicado: (2026)
por: Zhang, Lang, et al.
Publicado: (2026)
GA3CE: Unconstrained 3D Gaze Estimation with Gaze-Aware 3D Context Encoding
por: Kawana, Yuki, et al.
Publicado: (2025)
por: Kawana, Yuki, et al.
Publicado: (2025)
Robust Egocentric Visual Attention Prediction Through Language-guided Scene Context-aware Learning
por: Park, Sungjune, et al.
Publicado: (2026)
por: Park, Sungjune, et al.
Publicado: (2026)
Ejemplares similares
-
Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
por: Yun, Heeseung, et al.
Publicado: (2024) -
Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates
por: Ahn, Jaewoo, et al.
Publicado: (2025) -
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
por: Ahn, Jaewoo, et al.
Publicado: (2025) -
Gaussian Blending: Rethinking Alpha Blending in 3D Gaussian Splatting
por: Koo, Junseo, et al.
Publicado: (2025) -
Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation
por: Lai, Bolin, et al.
Publicado: (2023)