Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
Fuente:
arXiv
Guardado en:
| Autores principales: | Yun, Heeseung, Gao, Ruohan, Ananthabhotla, Ishwarya, Kumar, Anurag, Donley, Jacob, Li, Chao, Kim, Gunhee, Ithapu, Vamsi Krishna, Murdock, Calvin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective
por: Jia, Wenqi, et al.
Publicado: (2023)
por: Jia, Wenqi, et al.
Publicado: (2023)
EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
Hearing Anywhere in Any Environment
por: Liu, Xiulong, et al.
Publicado: (2025)
por: Liu, Xiulong, et al.
Publicado: (2025)
Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span
por: Yun, Heeseung, et al.
Publicado: (2025)
por: Yun, Heeseung, et al.
Publicado: (2025)
Text-to-Stage: Spatial Layouts from Long-form Narratives
por: Hernandez, Jefferson, et al.
Publicado: (2026)
por: Hernandez, Jefferson, et al.
Publicado: (2026)
Hearing Loss Detection from Facial Expressions in One-on-one Conversations
por: Yin, Yufeng, et al.
Publicado: (2024)
por: Yin, Yufeng, et al.
Publicado: (2024)
ViSAGe: Video-to-Spatial Audio Generation
por: Kim, Jaeyeon, et al.
Publicado: (2025)
por: Kim, Jaeyeon, et al.
Publicado: (2025)
Towards Localizing Conversation Partners using Head Motion
por: Mohapatra, Payal, et al.
Publicado: (2026)
por: Mohapatra, Payal, et al.
Publicado: (2026)
Towards Perception-Informed Latent HRTF Representations
por: Zhang, You, et al.
Publicado: (2025)
por: Zhang, You, et al.
Publicado: (2025)
Scene-wide Acoustic Parameter Estimation
por: Falcon-Perez, Ricardo, et al.
Publicado: (2024)
por: Falcon-Perez, Ricardo, et al.
Publicado: (2024)
WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations
por: Kim, Jaeyeon, et al.
Publicado: (2025)
por: Kim, Jaeyeon, et al.
Publicado: (2025)
Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates
por: Ahn, Jaewoo, et al.
Publicado: (2025)
por: Ahn, Jaewoo, et al.
Publicado: (2025)
On HRTF Notch Frequency Prediction Using Anthropometric Features and Neural Networks
por: Arbel, Lior, et al.
Publicado: (2024)
por: Arbel, Lior, et al.
Publicado: (2024)
SonoWorld: From One Image to a 3D Audio-Visual Scene
por: Jin, Derong, et al.
Publicado: (2026)
por: Jin, Derong, et al.
Publicado: (2026)
Learning to Highlight Audio by Watching Movies
por: Huang, Chao, et al.
Publicado: (2025)
por: Huang, Chao, et al.
Publicado: (2025)
Modulating State Space Model with SlowFast Framework for Compute-Efficient Ultra Low-Latency Speech Enhancement
por: Cheng, Longbiao, et al.
Publicado: (2024)
por: Cheng, Longbiao, et al.
Publicado: (2024)
Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos
por: Majumder, Sagnik, et al.
Publicado: (2023)
por: Majumder, Sagnik, et al.
Publicado: (2023)
Language-Guided Joint Audio-Visual Editing via One-Shot Adaptation
por: Liang, Susan, et al.
Publicado: (2024)
por: Liang, Susan, et al.
Publicado: (2024)
WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation
por: Song, Quanjian, et al.
Publicado: (2025)
por: Song, Quanjian, et al.
Publicado: (2025)
Exploring Audio Hallucination in Egocentric Video Understanding
por: Seth, Ashish, et al.
Publicado: (2026)
por: Seth, Ashish, et al.
Publicado: (2026)
Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment
por: Hong, Joanna, et al.
Publicado: (2025)
por: Hong, Joanna, et al.
Publicado: (2025)
EgoAVU: Egocentric Audio-Visual Understanding
por: Seth, Ashish, et al.
Publicado: (2026)
por: Seth, Ashish, et al.
Publicado: (2026)
Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos
por: Cokelek, Mert, et al.
Publicado: (2025)
por: Cokelek, Mert, et al.
Publicado: (2025)
HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos
por: Wang, Zhi, et al.
Publicado: (2026)
por: Wang, Zhi, et al.
Publicado: (2026)
SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Binding
por: Chen, Mingfei, et al.
Publicado: (2025)
por: Chen, Mingfei, et al.
Publicado: (2025)
Network Engineering in the Era of AI: A Technical Review
por: Vamsi Krishna Gadireddy
Publicado: (2025)
por: Vamsi Krishna Gadireddy
Publicado: (2025)
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
por: Ahn, Jaewoo, et al.
Publicado: (2025)
por: Ahn, Jaewoo, et al.
Publicado: (2025)
Do Audio-Visual Large Language Models Really See and Hear?
por: Selvakumar, Ramaneswaran, et al.
Publicado: (2026)
por: Selvakumar, Ramaneswaran, et al.
Publicado: (2026)
EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos
por: Rai, Aashish, et al.
Publicado: (2024)
por: Rai, Aashish, et al.
Publicado: (2024)
Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation
por: Lai, Bolin, et al.
Publicado: (2023)
por: Lai, Bolin, et al.
Publicado: (2023)
ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams
por: Kim, Chris Dongjoo, et al.
Publicado: (2025)
por: Kim, Chris Dongjoo, et al.
Publicado: (2025)
HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos
por: Zhang, Jinglei, et al.
Publicado: (2025)
por: Zhang, Jinglei, et al.
Publicado: (2025)
Walk through Paintings: Egocentric World Models from Internet Priors
por: Bagchi, Anurag, et al.
Publicado: (2026)
por: Bagchi, Anurag, et al.
Publicado: (2026)
Gaze-Enhanced Multimodal Turn-Taking Prediction in Triadic Conversations
por: Heo, Seongsil, et al.
Publicado: (2025)
por: Heo, Seongsil, et al.
Publicado: (2025)
ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction
por: Wang, Qineng, et al.
Publicado: (2025)
por: Wang, Qineng, et al.
Publicado: (2025)
Closure to the PRISM equation derived from nonlinear response theory
por: Donley, James P.
Publicado: (2024)
por: Donley, James P.
Publicado: (2024)
Report of the Faculty Research Initiative Grant: Learning-Disabled Students and Academic Library Services.
por: Donley, Mary, et al.
Publicado: (1990)
por: Donley, Mary, et al.
Publicado: (1990)
Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time
por: Chowdhury, Sanjoy, et al.
Publicado: (2024)
por: Chowdhury, Sanjoy, et al.
Publicado: (2024)
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
SpreadLine: Visualizing Egocentric Dynamic Influence
por: Kuo, Yun-Hsin, et al.
Publicado: (2024)
por: Kuo, Yun-Hsin, et al.
Publicado: (2024)
Ejemplares similares
-
The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective
por: Jia, Wenqi, et al.
Publicado: (2023) -
EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception
por: Chowdhury, Sanjoy, et al.
Publicado: (2025) -
Hearing Anywhere in Any Environment
por: Liu, Xiulong, et al.
Publicado: (2025) -
Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span
por: Yun, Heeseung, et al.
Publicado: (2025) -
Text-to-Stage: Spatial Layouts from Long-form Narratives
por: Hernandez, Jefferson, et al.
Publicado: (2026)