EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception
Fuente:
arXiv
Guardado en:
| Autores principales: | Chowdhury, Sanjoy, Biswas, Subrata, Nag, Sayan, Nagarajan, Tushar, Murdock, Calvin, Ananthabhotla, Ishwarya, Qian, Yijun, Ithapu, Vamsi Krishna, Manocha, Dinesh, Gao, Ruohan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
por: Yun, Heeseung, et al.
Publicado: (2024)
por: Yun, Heeseung, et al.
Publicado: (2024)
The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective
por: Jia, Wenqi, et al.
Publicado: (2023)
por: Jia, Wenqi, et al.
Publicado: (2023)
Hearing Anywhere in Any Environment
por: Liu, Xiulong, et al.
Publicado: (2025)
por: Liu, Xiulong, et al.
Publicado: (2025)
Text-to-Stage: Spatial Layouts from Long-form Narratives
por: Hernandez, Jefferson, et al.
Publicado: (2026)
por: Hernandez, Jefferson, et al.
Publicado: (2026)
Aurelia: Test-time Reasoning Distillation in Audio-Visual LLMs
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
Hearing Loss Detection from Facial Expressions in One-on-one Conversations
por: Yin, Yufeng, et al.
Publicado: (2024)
por: Yin, Yufeng, et al.
Publicado: (2024)
Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time
por: Chowdhury, Sanjoy, et al.
Publicado: (2024)
por: Chowdhury, Sanjoy, et al.
Publicado: (2024)
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
EgoAdapt: Enhancing Robustness in Egocentric Interactive Speaker Detection Under Missing Modalities
por: Qian, Xinyuan, et al.
Publicado: (2026)
por: Qian, Xinyuan, et al.
Publicado: (2026)
Uncovering the Representation Geometry of Minimal Cores in Overcomplete Reasoning Traces
por: Chowdhury, Sanjoy, et al.
Publicado: (2026)
por: Chowdhury, Sanjoy, et al.
Publicado: (2026)
MeLFusion: Synthesizing Music from Image and Language Cues using Diffusion Models
por: Chowdhury, Sanjoy, et al.
Publicado: (2024)
por: Chowdhury, Sanjoy, et al.
Publicado: (2024)
EgoAdapt: A Multi-Scene Egocentric Adaptation Method for CVPR 2026 HD-EPIC VQA Challenge
por: Chen, Zhiwei, et al.
Publicado: (2026)
por: Chen, Zhiwei, et al.
Publicado: (2026)
Towards Localizing Conversation Partners using Head Motion
por: Mohapatra, Payal, et al.
Publicado: (2026)
por: Mohapatra, Payal, et al.
Publicado: (2026)
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
Can LLMs Generate Human-Like Wayfinding Instructions? Towards Platform-Agnostic Embodied Instruction Synthesis
por: Dorbala, Vishnu Sashank, et al.
Publicado: (2024)
por: Dorbala, Vishnu Sashank, et al.
Publicado: (2024)
Towards Perception-Informed Latent HRTF Representations
por: Zhang, You, et al.
Publicado: (2025)
por: Zhang, You, et al.
Publicado: (2025)
Multisensory machine intelligence
por: Ruohan Gao
Publicado: (2025)
por: Ruohan Gao
Publicado: (2025)
Scene-wide Acoustic Parameter Estimation
por: Falcon-Perez, Ricardo, et al.
Publicado: (2024)
por: Falcon-Perez, Ricardo, et al.
Publicado: (2024)
EgoSocial: Benchmarking Proactive Intervention Ability of Omnimodal LLMs via Egocentric Social Interaction Perception
por: Wang, Xijun, et al.
Publicado: (2025)
por: Wang, Xijun, et al.
Publicado: (2025)
EgoAVU: Egocentric Audio-Visual Understanding
por: Seth, Ashish, et al.
Publicado: (2026)
por: Seth, Ashish, et al.
Publicado: (2026)
AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
por: Galougah, Siminfar Samakoush, et al.
Publicado: (2025)
por: Galougah, Siminfar Samakoush, et al.
Publicado: (2025)
Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span
por: Yun, Heeseung, et al.
Publicado: (2025)
por: Yun, Heeseung, et al.
Publicado: (2025)
On HRTF Notch Frequency Prediction Using Anthropometric Features and Neural Networks
por: Arbel, Lior, et al.
Publicado: (2024)
por: Arbel, Lior, et al.
Publicado: (2024)
HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos
por: Wang, Zhi, et al.
Publicado: (2026)
por: Wang, Zhi, et al.
Publicado: (2026)
Implications of portal vector-like lepton on associated Higgs production at a multi-TeV muon collider
por: Tewary, Krishna, et al.
Publicado: (2025)
por: Tewary, Krishna, et al.
Publicado: (2025)
Modulating State Space Model with SlowFast Framework for Compute-Efficient Ultra Low-Latency Speech Enhancement
por: Cheng, Longbiao, et al.
Publicado: (2024)
por: Cheng, Longbiao, et al.
Publicado: (2024)
ASPIRE: Language-Guided Data Augmentation for Improving Robustness Against Spurious Correlations
por: Ghosh, Sreyan, et al.
Publicado: (2023)
por: Ghosh, Sreyan, et al.
Publicado: (2023)
Do Audio-Visual Large Language Models Really See and Hear?
por: Selvakumar, Ramaneswaran, et al.
Publicado: (2026)
por: Selvakumar, Ramaneswaran, et al.
Publicado: (2026)
VITED: Video Temporal Evidence Distillation
por: Lu, Yujie, et al.
Publicado: (2025)
por: Lu, Yujie, et al.
Publicado: (2025)
EgoGen: An Egocentric Synthetic Data Generator
por: Li, Gen, et al.
Publicado: (2024)
por: Li, Gen, et al.
Publicado: (2024)
EgoLife: Towards Egocentric Life Assistant
por: Yang, Jingkang, et al.
Publicado: (2025)
por: Yang, Jingkang, et al.
Publicado: (2025)
Step Differences in Instructional Video
por: Nagarajan, Tushar, et al.
Publicado: (2024)
por: Nagarajan, Tushar, et al.
Publicado: (2024)
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
por: Seth, Ashish, et al.
Publicado: (2025)
por: Seth, Ashish, et al.
Publicado: (2025)
CalibFree: Self-Supervised View Feature Separation for Calibration-Free Multi-Camera Multi-Object Tracking
por: Xian, Ruiqi, et al.
Publicado: (2026)
por: Xian, Ruiqi, et al.
Publicado: (2026)
PACE: Data-Driven Virtual Agent Interaction in Dense and Cluttered Environments
por: Mullen, James, et al.
Publicado: (2023)
por: Mullen, James, et al.
Publicado: (2023)
Listen2Scene: Interactive material-aware binaural sound propagation for reconstructed 3D scenes
por: Ratnarajah, Anton, et al.
Publicado: (2023)
por: Ratnarajah, Anton, et al.
Publicado: (2023)
Inst4DGS: Instance-Decomposed 4D Gaussian Splatting with Multi-Video Label Permutation Learning
por: Lee, Yonghan, et al.
Publicado: (2026)
por: Lee, Yonghan, et al.
Publicado: (2026)
EM-GANSim: Real-time and Accurate EM Simulation Using Conditional GANs for 3D Indoor Scenes
por: Wang, Ruichen, et al.
Publicado: (2024)
por: Wang, Ruichen, et al.
Publicado: (2024)
Network Engineering in the Era of AI: A Technical Review
por: Vamsi Krishna Gadireddy
Publicado: (2025)
por: Vamsi Krishna Gadireddy
Publicado: (2025)
QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding
por: Biswas, Subrata, et al.
Publicado: (2025)
por: Biswas, Subrata, et al.
Publicado: (2025)
Ejemplares similares
-
Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
por: Yun, Heeseung, et al.
Publicado: (2024) -
The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective
por: Jia, Wenqi, et al.
Publicado: (2023) -
Hearing Anywhere in Any Environment
por: Liu, Xiulong, et al.
Publicado: (2025) -
Text-to-Stage: Spatial Layouts from Long-form Narratives
por: Hernandez, Jefferson, et al.
Publicado: (2026) -
Aurelia: Test-time Reasoning Distillation in Audio-Visual LLMs
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)