ALGO: Object-Grounded Visual Commonsense Reasoning for Open-World Egocentric Action Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Kundu, Sanjoy, Trehan, Shubham, Aakur, Sathyanarayanan N. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Discovering Novel Actions from Open World Egocentric Videos with Object-Grounded Visual Commonsense Reasoning
por: Kundu, Sanjoy, et al.
Publicado: (2023)
por: Kundu, Sanjoy, et al.
Publicado: (2023)
A Probabilistic Jump-Diffusion Framework for Open-World Egocentric Activity Recognition
por: Kundu, Sanjoy, et al.
Publicado: (2025)
por: Kundu, Sanjoy, et al.
Publicado: (2025)
ProbRes: Probabilistic Jump Diffusion for Open-World Egocentric Activity Recognition
por: Kundu, Sanjoy, et al.
Publicado: (2025)
por: Kundu, Sanjoy, et al.
Publicado: (2025)
Self-supervised Multi-actor Social Activity Understanding in Streaming Videos
por: Trehan, Shubham, et al.
Publicado: (2024)
por: Trehan, Shubham, et al.
Publicado: (2024)
Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection
por: Vellamcheti, Shanmukha, et al.
Publicado: (2025)
por: Vellamcheti, Shanmukha, et al.
Publicado: (2025)
FSP-DETR: Few-Shot Prototypical Parasitic Ova Detection
por: Trehan, Shubham, et al.
Publicado: (2025)
por: Trehan, Shubham, et al.
Publicado: (2025)
EASE: Embodied Active Event Perception via Self-Supervised Energy Minimization
por: Chen, Zhou, et al.
Publicado: (2025)
por: Chen, Zhou, et al.
Publicado: (2025)
CRAFT: A Neuro-Symbolic Framework for Visual Functional Affordance Grounding
por: Chen, Zhou, et al.
Publicado: (2025)
por: Chen, Zhou, et al.
Publicado: (2025)
Generalized Event Partonomy Inference with Structured Hierarchical Predictive Learning
por: Chen, Zhou, et al.
Publicado: (2025)
por: Chen, Zhou, et al.
Publicado: (2025)
STaTS: Structure-Aware Temporal Sequence Summarization via Statistical Window Merging
por: Bhowmick, Disharee, et al.
Publicado: (2025)
por: Bhowmick, Disharee, et al.
Publicado: (2025)
Capturing Temporal Components for Time Series Classification
por: Vavilthota, Venkata Ragavendra, et al.
Publicado: (2024)
por: Vavilthota, Venkata Ragavendra, et al.
Publicado: (2024)
CVT-Bench: Counterfactual Viewpoint Transformations Reveal Unstable Spatial Representations in Multimodal LLMs
por: Vellamcheti, Shanmukha, et al.
Publicado: (2026)
por: Vellamcheti, Shanmukha, et al.
Publicado: (2026)
WHOLE: World-Grounded Hand-Object Lifted from Egocentric Videos
por: Ye, Yufei, et al.
Publicado: (2026)
por: Ye, Yufei, et al.
Publicado: (2026)
A Study of Commonsense Reasoning over Visual Object Properties
por: Kolari, Abhishek, et al.
Publicado: (2025)
por: Kolari, Abhishek, et al.
Publicado: (2025)
LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model
por: Gao, Quankai, et al.
Publicado: (2026)
por: Gao, Quankai, et al.
Publicado: (2026)
Visual Intention Grounding for Egocentric Assistants
por: Sun, Pengzhan, et al.
Publicado: (2025)
por: Sun, Pengzhan, et al.
Publicado: (2025)
Augmented Commonsense Knowledge for Remote Object Grounding
por: Mohammadi, Bahram, et al.
Publicado: (2024)
por: Mohammadi, Bahram, et al.
Publicado: (2024)
Domain Generalization using Action Sequences for Egocentric Action Recognition
por: Nasirimajd, Amirshayan, et al.
Publicado: (2025)
por: Nasirimajd, Amirshayan, et al.
Publicado: (2025)
Causal Debiasing for Visual Commonsense Reasoning
por: Zou, Jiayi, et al.
Publicado: (2025)
por: Zou, Jiayi, et al.
Publicado: (2025)
EgoPrompt: Prompt Learning for Egocentric Action Recognition
por: Lyu, Huaihai, et al.
Publicado: (2025)
por: Lyu, Huaihai, et al.
Publicado: (2025)
OpenGround: Active Cognition-based Reasoning for Open-World 3D Visual Grounding
por: Huang, Wenyuan, et al.
Publicado: (2025)
por: Huang, Wenyuan, et al.
Publicado: (2025)
Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition
por: Liang, Zichen, et al.
Publicado: (2025)
por: Liang, Zichen, et al.
Publicado: (2025)
Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
por: Yuan, Haobo, et al.
Publicado: (2025)
por: Yuan, Haobo, et al.
Publicado: (2025)
Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts
por: Wei, Jiude, et al.
Publicado: (2025)
por: Wei, Jiude, et al.
Publicado: (2025)
Interaction Region Visual Transformer for Egocentric Action Anticipation
por: Roy, Debaditya, et al.
Publicado: (2022)
por: Roy, Debaditya, et al.
Publicado: (2022)
Simultaneous Detection and Interaction Reasoning for Object-Centric Action Recognition
por: Li, Xunsong, et al.
Publicado: (2024)
por: Li, Xunsong, et al.
Publicado: (2024)
Efficient Egocentric Action Recognition with Multimodal Data
por: Calzavara, Marco, et al.
Publicado: (2025)
por: Calzavara, Marco, et al.
Publicado: (2025)
Object Aware Egocentric Online Action Detection
por: An, Joungbin, et al.
Publicado: (2024)
por: An, Joungbin, et al.
Publicado: (2024)
Multimodal Cross-Domain Few-Shot Learning for Egocentric Action Recognition
por: Hatano, Masashi, et al.
Publicado: (2024)
por: Hatano, Masashi, et al.
Publicado: (2024)
Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition
por: Zhang, Mingfang, et al.
Publicado: (2024)
por: Zhang, Mingfang, et al.
Publicado: (2024)
Multimodal Knowledge Distillation for Egocentric Action Recognition Robust to Missing Modalities
por: Santos-Villafranca, Maria, et al.
Publicado: (2025)
por: Santos-Villafranca, Maria, et al.
Publicado: (2025)
Cross-view Action Recognition Understanding From Exocentric to Egocentric Perspective
por: Truong, Thanh-Dat, et al.
Publicado: (2023)
por: Truong, Thanh-Dat, et al.
Publicado: (2023)
Object-Shot Enhanced Grounding Network for Egocentric Video
por: Feng, Yisen, et al.
Publicado: (2025)
por: Feng, Yisen, et al.
Publicado: (2025)
Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
por: Yun, Heeseung, et al.
Publicado: (2024)
por: Yun, Heeseung, et al.
Publicado: (2024)
Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
por: Ren, Tianhe, et al.
Publicado: (2024)
por: Ren, Tianhe, et al.
Publicado: (2024)
EgoActor: Grounding Task Planning into Spatial-aware Egocentric Actions for Humanoid Robots via Visual-Language Models
por: Bai, Yu, et al.
Publicado: (2026)
por: Bai, Yu, et al.
Publicado: (2026)
Egocentric World Model for Photorealistic Hand-Object Interaction Synthesis
por: Li, Dayou, et al.
Publicado: (2026)
por: Li, Dayou, et al.
Publicado: (2026)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
por: Liu, Huabin, et al.
Publicado: (2025)
por: Liu, Huabin, et al.
Publicado: (2025)
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
por: Yuan, Yuqian, et al.
Publicado: (2025)
por: Yuan, Yuqian, et al.
Publicado: (2025)
CalibFree: Self-Supervised View Feature Separation for Calibration-Free Multi-Camera Multi-Object Tracking
por: Xian, Ruiqi, et al.
Publicado: (2026)
por: Xian, Ruiqi, et al.
Publicado: (2026)
Ejemplares similares
-
Discovering Novel Actions from Open World Egocentric Videos with Object-Grounded Visual Commonsense Reasoning
por: Kundu, Sanjoy, et al.
Publicado: (2023) -
A Probabilistic Jump-Diffusion Framework for Open-World Egocentric Activity Recognition
por: Kundu, Sanjoy, et al.
Publicado: (2025) -
ProbRes: Probabilistic Jump Diffusion for Open-World Egocentric Activity Recognition
por: Kundu, Sanjoy, et al.
Publicado: (2025) -
Self-supervised Multi-actor Social Activity Understanding in Streaming Videos
por: Trehan, Shubham, et al.
Publicado: (2024) -
Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection
por: Vellamcheti, Shanmukha, et al.
Publicado: (2025)