EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Dongyan, Rust, Phillip, Corrales, Angel Villar, Tan, Alvin W. M., Luthra, Mahi, Saint-James, Charles-Éric, Moritz, Rashel, Krogh-Jespersen, Sheila, Stark, Vanessa, Parimi, Surya, Shen, Jiayi, Benchekroun, Youssef, Higuchi, Yosuke, Gleize, Martin, Fizycki, Tom, Hamilakis, Nicolas, Khentout, Manel, Tsuji, Sho, Kégl, Balázs, Pino, Juan, Frank, Michael C., Dupoux, Emmanuel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation
von: Luthra, Mahi, et al.
Veröffentlicht: (2025)
von: Luthra, Mahi, et al.
Veröffentlicht: (2025)
LongTail-Swap: benchmarking language models' abilities on rare words
von: Algayres, Robin, et al.
Veröffentlicht: (2025)
von: Algayres, Robin, et al.
Veröffentlicht: (2025)
MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery
von: Tandazo, Angelo Ortiz, et al.
Veröffentlicht: (2025)
von: Tandazo, Angelo Ortiz, et al.
Veröffentlicht: (2025)
SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision
von: Poli, Maxime, et al.
Veröffentlicht: (2025)
von: Poli, Maxime, et al.
Veröffentlicht: (2025)
DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units
von: Poli, Maxime, et al.
Veröffentlicht: (2026)
von: Poli, Maxime, et al.
Veröffentlicht: (2026)
Fifteen Years of Child-Centered Long-Form Recordings: Promises, Resources, and Remaining Challenges to Validity
von: Peurey, Loann, et al.
Veröffentlicht: (2025)
von: Peurey, Loann, et al.
Veröffentlicht: (2025)
Arqueología, neofascismo y supremacía blanca: descolonizando nuestras teorías, emprendiendo prácticas de resistencia
von: Yannis Hamilakis
Veröffentlicht: (2023)
von: Yannis Hamilakis
Veröffentlicht: (2023)
EgoAVU: Egocentric Audio-Visual Understanding
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
EgoGen: An Egocentric Synthetic Data Generator
von: Li, Gen, et al.
Veröffentlicht: (2024)
von: Li, Gen, et al.
Veröffentlicht: (2024)
EgoLife: Towards Egocentric Life Assistant
von: Yang, Jingkang, et al.
Veröffentlicht: (2025)
von: Yang, Jingkang, et al.
Veröffentlicht: (2025)
Les chiens s´approchent, et s´éloignent
von: Jean-Marie Gleize
Veröffentlicht: (2007)
von: Jean-Marie Gleize
Veröffentlicht: (2007)
EgoSound: Benchmarking Sound Understanding in Egocentric Videos
von: Zhu, Bingwen, et al.
Veröffentlicht: (2026)
von: Zhu, Bingwen, et al.
Veröffentlicht: (2026)
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding
von: Nagrani, Arsha, et al.
Veröffentlicht: (2026)
von: Nagrani, Arsha, et al.
Veröffentlicht: (2026)
EgoVLM: Policy Optimization for Egocentric Video Understanding
von: Vinod, Ashwin, et al.
Veröffentlicht: (2025)
von: Vinod, Ashwin, et al.
Veröffentlicht: (2025)
EgoPrompt: Prompt Learning for Egocentric Action Recognition
von: Lyu, Huaihai, et al.
Veröffentlicht: (2025)
von: Lyu, Huaihai, et al.
Veröffentlicht: (2025)
EgoCast: Forecasting Egocentric Human Pose in the Wild
von: Escobar, Maria, et al.
Veröffentlicht: (2024)
von: Escobar, Maria, et al.
Veröffentlicht: (2024)
EgoPoints: Advancing Point Tracking for Egocentric Videos
von: Darkhalil, Ahmad, et al.
Veröffentlicht: (2024)
von: Darkhalil, Ahmad, et al.
Veröffentlicht: (2024)
EgoBlind: Towards Egocentric Visual Assistance for the Blind
von: Xiao, Junbin, et al.
Veröffentlicht: (2025)
von: Xiao, Junbin, et al.
Veröffentlicht: (2025)
TeleEgo: Benchmarking Egocentric AI Assistants in the Wild
von: Yan, Jiaqi, et al.
Veröffentlicht: (2025)
von: Yan, Jiaqi, et al.
Veröffentlicht: (2025)
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
EgoForge: Goal-Directed Egocentric World Simulator
von: Shen, Yifan, et al.
Veröffentlicht: (2026)
von: Shen, Yifan, et al.
Veröffentlicht: (2026)
EgoSelf: From Memory to Personalized Egocentric Assistant
von: Wang, Yanshuo, et al.
Veröffentlicht: (2026)
von: Wang, Yanshuo, et al.
Veröffentlicht: (2026)
BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings
von: Charlot, Théo, et al.
Veröffentlicht: (2025)
von: Charlot, Théo, et al.
Veröffentlicht: (2025)
The Mounssif Transform and Its Applications in Number Theory
von: Benchekroun, Mounssif
Veröffentlicht: (2025)
von: Benchekroun, Mounssif
Veröffentlicht: (2025)
Precarious Motherhood
von: Benchekroun, Rachel
Veröffentlicht: (2025)
von: Benchekroun, Rachel
Veröffentlicht: (2025)
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
von: Kulkarni, Yogesh, et al.
Veröffentlicht: (2025)
von: Kulkarni, Yogesh, et al.
Veröffentlicht: (2025)
EgoQR: Efficient QR Code Reading in Egocentric Settings
von: Moslehpour, Mohsen, et al.
Veröffentlicht: (2024)
von: Moslehpour, Mohsen, et al.
Veröffentlicht: (2024)
EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
von: Pei, Baoqi, et al.
Veröffentlicht: (2025)
von: Pei, Baoqi, et al.
Veröffentlicht: (2025)
EgoSim: Egocentric World Simulator for Embodied Interaction Generation
von: Hao, Jinkun, et al.
Veröffentlicht: (2026)
von: Hao, Jinkun, et al.
Veröffentlicht: (2026)
EgoM2P: Egocentric Multimodal Multitask Pretraining
von: Li, Gen, et al.
Veröffentlicht: (2025)
von: Li, Gen, et al.
Veröffentlicht: (2025)
EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
von: Pei, Baoqi, et al.
Veröffentlicht: (2024)
von: Pei, Baoqi, et al.
Veröffentlicht: (2024)
EgoCogNav: Cognition-aware Human Egocentric Navigation
von: Qiu, Zhiwen, et al.
Veröffentlicht: (2025)
von: Qiu, Zhiwen, et al.
Veröffentlicht: (2025)
Ego-VPA: Egocentric Video Understanding with Parameter-efficient Adaptation
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2024)
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2024)
EgoMimic: Scaling Imitation Learning via Egocentric Video
von: Kareer, Simar, et al.
Veröffentlicht: (2024)
von: Kareer, Simar, et al.
Veröffentlicht: (2024)
EgoLCD: Egocentric Video Generation with Long Context Diffusion
von: Zhang, Liuzhou, et al.
Veröffentlicht: (2025)
von: Zhang, Liuzhou, et al.
Veröffentlicht: (2025)
EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos
von: Rai, Aashish, et al.
Veröffentlicht: (2024)
von: Rai, Aashish, et al.
Veröffentlicht: (2024)
EgoNav: Egocentric Scene-aware Human Trajectory Prediction
von: Wang, Weizhuo, et al.
Veröffentlicht: (2024)
von: Wang, Weizhuo, et al.
Veröffentlicht: (2024)
EgoLM: Multi-Modal Language Model of Egocentric Motions
von: Hong, Fangzhou, et al.
Veröffentlicht: (2024)
von: Hong, Fangzhou, et al.
Veröffentlicht: (2024)
EgoGraph: Temporal Knowledge Graph for Egocentric Video Understanding
von: Sun, Shitong, et al.
Veröffentlicht: (2026)
von: Sun, Shitong, et al.
Veröffentlicht: (2026)
EgoCampus: Egocentric Pedestrian Eye Gaze Model and Dataset
von: John, Ronan, et al.
Veröffentlicht: (2025)
von: John, Ronan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation
von: Luthra, Mahi, et al.
Veröffentlicht: (2025) -
LongTail-Swap: benchmarking language models' abilities on rare words
von: Algayres, Robin, et al.
Veröffentlicht: (2025) -
MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery
von: Tandazo, Angelo Ortiz, et al.
Veröffentlicht: (2025) -
SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision
von: Poli, Maxime, et al.
Veröffentlicht: (2025) -
DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units
von: Poli, Maxime, et al.
Veröffentlicht: (2026)