VL-KnG: Persistent Spatiotemporal Knowledge Graphs from Egocentric Video for Embodied Scene Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Mdfaa, Mohamad Al, Lukina, Svetlana, Akhtyamov, Timur, Nigmatzyanov, Arthur, Nalberskii, Dmitrii, Zagoruyko, Sergey, Ferrer, Gonzalo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mapping the Unseen: Unified Promptable Panoptic Mapping with Dynamic Labeling using Foundation Models
by: Mdfaa, Mohamad Al, et al.
Published: (2024)
by: Mdfaa, Mohamad Al, et al.
Published: (2024)
EgoWalk: A Multimodal Dataset for Robot Navigation in the Wild
by: Akhtyamov, Timur, et al.
Published: (2025)
by: Akhtyamov, Timur, et al.
Published: (2025)
PixelNav: Towards Model-based Vision-Only Navigation with Topological Graphs
by: Bakulin, Sergey, et al.
Published: (2025)
by: Bakulin, Sergey, et al.
Published: (2025)
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
Quranic Audio Dataset: Crowdsourced and Labeled Recitation from Non-Arabic Speakers
by: Salameh, Raghad, et al.
Published: (2024)
by: Salameh, Raghad, et al.
Published: (2024)
KnFu: Effective Knowledge Fusion
by: Seyedmohammadi, S. Jamal, et al.
Published: (2024)
by: Seyedmohammadi, S. Jamal, et al.
Published: (2024)
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding
by: Nagrani, Arsha, et al.
Published: (2026)
by: Nagrani, Arsha, et al.
Published: (2026)
On the word-representability of $K_m$-$K_n$ graphs
by: Chen, Herman Z. Q., et al.
Published: (2025)
by: Chen, Herman Z. Q., et al.
Published: (2025)
Non-Hausdorff germinal groupoids for actions of countable groups
by: Lukina, Olga
Published: (2023)
by: Lukina, Olga
Published: (2023)
ValentinJeutner, The Reasonable Person: A Legal Biography, Cambridge University Press, 2024, 252 pp, hb £95.00
by: Anna Lukina
Published: (2025)
by: Anna Lukina
Published: (2025)
Eigen-Factors a Bilevel Optimization for Plane SLAM of 3D Point Clouds
by: Ferrer, Gonzalo, et al.
Published: (2023)
by: Ferrer, Gonzalo, et al.
Published: (2023)
AGORA: Adversarial Generation Of Real-time Animatable 3D Gaussian Head Avatars
by: Fazylov, Ramazan, et al.
Published: (2025)
by: Fazylov, Ramazan, et al.
Published: (2025)
Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions
by: Dong, Ze, et al.
Published: (2026)
by: Dong, Ze, et al.
Published: (2026)
VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI
by: Cheng, Sijie, et al.
Published: (2024)
by: Cheng, Sijie, et al.
Published: (2024)
SNOW: Spatio-Temporal Scene Understanding with World Knowledge for Open-World Embodied Reasoning
by: Sohn, Tin Stribor, et al.
Published: (2025)
by: Sohn, Tin Stribor, et al.
Published: (2025)
AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding
by: Suglia, Alessandro, et al.
Published: (2024)
by: Suglia, Alessandro, et al.
Published: (2024)
Skew-product systems over infinite interval exchange transformations
by: Bruin, Henk, et al.
Published: (2024)
by: Bruin, Henk, et al.
Published: (2024)
Type invariants for non-abelian odometers
by: Hurder, Steven, et al.
Published: (2023)
by: Hurder, Steven, et al.
Published: (2023)
Prime spectrum and dynamics for nilpotent Cantor actions
by: Hurder, Steven, et al.
Published: (2023)
by: Hurder, Steven, et al.
Published: (2023)
Fine-grained Spatiotemporal Grounding on Egocentric Videos
by: Liang, Shuo, et al.
Published: (2025)
by: Liang, Shuo, et al.
Published: (2025)
EgoGraph: Temporal Knowledge Graph for Egocentric Video Understanding
by: Sun, Shitong, et al.
Published: (2026)
by: Sun, Shitong, et al.
Published: (2026)
EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
by: Wu, Yuqi, et al.
Published: (2024)
by: Wu, Yuqi, et al.
Published: (2024)
On the crossing profile of rectilinear drawings of $K_n$
by: Chen, Isaac, et al.
Published: (2025)
by: Chen, Isaac, et al.
Published: (2025)
The rectilinear local crossing number of $K_n$
by: Ábrego, Bernardo M., et al.
Published: (2015)
by: Ábrego, Bernardo M., et al.
Published: (2015)
BioVL-QR: Egocentric Biochemical Vision-and-Language Dataset Using Micro QR Codes
by: Nishimoto, Tomohiro, et al.
Published: (2024)
by: Nishimoto, Tomohiro, et al.
Published: (2024)
Scene-Driven Multimodal Knowledge Graph Construction for Embodied AI
by: Yaoxian, Song, et al.
Published: (2023)
by: Yaoxian, Song, et al.
Published: (2023)
Pelican-VL 1.0: A Foundation Brain Model for Embodied Intelligence
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
Spatiotemporal Persistence Landscapes
by: Flammer, Martina, et al.
Published: (2024)
by: Flammer, Martina, et al.
Published: (2024)
Embodied Scene Understanding for Vision Language Models via MetaVQA
by: Wang, Weizhen, et al.
Published: (2025)
by: Wang, Weizhen, et al.
Published: (2025)
From Scan to Action: Leveraging Realistic Scans for Embodied Scene Understanding
by: Halacheva, Anna-Maria, et al.
Published: (2025)
by: Halacheva, Anna-Maria, et al.
Published: (2025)
Investigating Simple Drawings of $K_n$ using SAT
by: Bergold, Helena, et al.
Published: (2025)
by: Bergold, Helena, et al.
Published: (2025)
Leaky forcing and resilience of Cartesian products of $K_n$
by: Herrman, Rebekah, et al.
Published: (2024)
by: Herrman, Rebekah, et al.
Published: (2024)
On the geometric $k$-colored crossing number of $K_n$
by: Hahn, Benedikt, et al.
Published: (2025)
by: Hahn, Benedikt, et al.
Published: (2025)
Star arboricity relaxed book thickness of $K_n$
by: Kainen, Paul C.
Published: (2024)
by: Kainen, Paul C.
Published: (2024)
ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark
by: Dang, Ronghao, et al.
Published: (2025)
by: Dang, Ronghao, et al.
Published: (2025)
EgoSim: Egocentric World Simulator for Embodied Interaction Generation
by: Hao, Jinkun, et al.
Published: (2026)
by: Hao, Jinkun, et al.
Published: (2026)
ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction
by: Wang, Qineng, et al.
Published: (2025)
by: Wang, Qineng, et al.
Published: (2025)
EgoGaussian: Dynamic Scene Understanding from Egocentric Video with 3D Gaussian Splatting
by: Zhang, Daiwei, et al.
Published: (2024)
by: Zhang, Daiwei, et al.
Published: (2024)
VL-Reader: Vision and Language Reconstructor is an Effective Scene Text Recognizer
by: Zhong, Humen, et al.
Published: (2024)
by: Zhong, Humen, et al.
Published: (2024)
Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding
by: Zhang, Haoyu, et al.
Published: (2025)
by: Zhang, Haoyu, et al.
Published: (2025)
Similar Items
-
Mapping the Unseen: Unified Promptable Panoptic Mapping with Dynamic Labeling using Foundation Models
by: Mdfaa, Mohamad Al, et al.
Published: (2024) -
EgoWalk: A Multimodal Dataset for Robot Navigation in the Wild
by: Akhtyamov, Timur, et al.
Published: (2025) -
PixelNav: Towards Model-based Vision-Only Navigation with Topological Graphs
by: Bakulin, Sergey, et al.
Published: (2025) -
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding
by: Fan, Yue, et al.
Published: (2024) -
Quranic Audio Dataset: Crowdsourced and Labeled Recitation from Non-Arabic Speakers
by: Salameh, Raghad, et al.
Published: (2024)