Text-driven Affordance Learning from Egocentric Vision
Fuente:
arXiv
Guardado en:
| Autores principales: | Yoshida, Tomoya, Kurita, Shuhei, Nishimura, Taichi, Mori, Shinsuke |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision
por: Yoshida, Tomoya, et al.
Publicado: (2025)
por: Yoshida, Tomoya, et al.
Publicado: (2025)
Developing Vision-Language-Action Model from Egocentric Videos
por: Yoshida, Tomoya, et al.
Publicado: (2025)
por: Yoshida, Tomoya, et al.
Publicado: (2025)
EgoOops: A Dataset for Mistake Action Detection from Egocentric Videos referring to Procedural Texts
por: Haneji, Yuto, et al.
Publicado: (2024)
por: Haneji, Yuto, et al.
Publicado: (2024)
STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models
por: Ukai, Mahiro, et al.
Publicado: (2025)
por: Ukai, Mahiro, et al.
Publicado: (2025)
Vision Language Model-based Caption Evaluation Method Leveraging Visual Context Extraction
por: Maeda, Koki, et al.
Publicado: (2024)
por: Maeda, Koki, et al.
Publicado: (2024)
Grounding 3D Scene Affordance From Egocentric Interactions
por: Liu, Cuiyu, et al.
Publicado: (2024)
por: Liu, Cuiyu, et al.
Publicado: (2024)
PhysQuantAgent: An Inference Pipeline of Mass Estimation for Vision-Language Models
por: Yokomizo, Hisayuki, et al.
Publicado: (2026)
por: Yokomizo, Hisayuki, et al.
Publicado: (2026)
BioVL-QR: Egocentric Biochemical Vision-and-Language Dataset Using Micro QR Codes
por: Nishimoto, Tomohiro, et al.
Publicado: (2024)
por: Nishimoto, Tomohiro, et al.
Publicado: (2024)
Informative Text-Image Alignment for Visual Affordance Learning with Foundation Models
por: Zhang, Qian, et al.
Publicado: (2025)
por: Zhang, Qian, et al.
Publicado: (2025)
Referring Expression Comprehension for Small Objects
por: Goto, Kanoko, et al.
Publicado: (2025)
por: Goto, Kanoko, et al.
Publicado: (2025)
Egocentric Bias in Vision-Language Models
por: Wang, Maijunxian, et al.
Publicado: (2026)
por: Wang, Maijunxian, et al.
Publicado: (2026)
Challenges and Trends in Egocentric Vision: A Survey
por: Li, Xiang, et al.
Publicado: (2025)
por: Li, Xiang, et al.
Publicado: (2025)
AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question Answering
por: Ukai, Mahiro, et al.
Publicado: (2024)
por: Ukai, Mahiro, et al.
Publicado: (2024)
Self-Explainable Affordance Learning with Embodied Caption
por: Zhang, Zhipeng, et al.
Publicado: (2024)
por: Zhang, Zhipeng, et al.
Publicado: (2024)
CityNav: A Large-Scale Dataset for Real-World Aerial Navigation
por: Lee, Jungdae, et al.
Publicado: (2024)
por: Lee, Jungdae, et al.
Publicado: (2024)
Recipe Generation from Unsegmented Cooking Videos
por: Nishimura, Taichi, et al.
Publicado: (2022)
por: Nishimura, Taichi, et al.
Publicado: (2022)
Selective Contrastive Learning for Weakly Supervised Affordance Grounding
por: Moon, WonJun, et al.
Publicado: (2025)
por: Moon, WonJun, et al.
Publicado: (2025)
SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video
por: Valdez, Hector A., et al.
Publicado: (2024)
por: Valdez, Hector A., et al.
Publicado: (2024)
Comparing Learning Paradigms for Egocentric Video Summarization
por: Wen, Daniel
Publicado: (2025)
por: Wen, Daniel
Publicado: (2025)
EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
por: Zhang, Deheng, et al.
Publicado: (2025)
por: Zhang, Deheng, et al.
Publicado: (2025)
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
por: Yuan, Wentao, et al.
Publicado: (2024)
por: Yuan, Wentao, et al.
Publicado: (2024)
Information-driven Affordance Discovery for Efficient Robotic Manipulation
por: Mazzaglia, Pietro, et al.
Publicado: (2024)
por: Mazzaglia, Pietro, et al.
Publicado: (2024)
General Flow as Foundation Affordance for Scalable Robot Learning
por: Yuan, Chengbo, et al.
Publicado: (2024)
por: Yuan, Chengbo, et al.
Publicado: (2024)
SYNTHIA: Novel Concept Design with Affordance Composition
por: Ha, Hyeonjeong, et al.
Publicado: (2025)
por: Ha, Hyeonjeong, et al.
Publicado: (2025)
Learning Egocentric In-Hand Object Segmentation through Weak Supervision from Human Narrations
por: Messina, Nicola, et al.
Publicado: (2025)
por: Messina, Nicola, et al.
Publicado: (2025)
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
por: Kim, Junhyeok, et al.
Publicado: (2025)
por: Kim, Junhyeok, et al.
Publicado: (2025)
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
por: Yang, Ruihan, et al.
Publicado: (2025)
por: Yang, Ruihan, et al.
Publicado: (2025)
WorldAfford: Affordance Grounding based on Natural Language Instructions
por: Chen, Changmao, et al.
Publicado: (2024)
por: Chen, Changmao, et al.
Publicado: (2024)
TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning
por: Spigler, Giacomo
Publicado: (2026)
por: Spigler, Giacomo
Publicado: (2026)
Object Aware Egocentric Online Action Detection
por: An, Joungbin, et al.
Publicado: (2024)
por: An, Joungbin, et al.
Publicado: (2024)
EgoGen: An Egocentric Synthetic Data Generator
por: Li, Gen, et al.
Publicado: (2024)
por: Li, Gen, et al.
Publicado: (2024)
EAGLE: Egocentric AGgregated Language-video Engine
por: Bi, Jing, et al.
Publicado: (2024)
por: Bi, Jing, et al.
Publicado: (2024)
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
por: Seth, Ashish, et al.
Publicado: (2025)
por: Seth, Ashish, et al.
Publicado: (2025)
Efficient Egocentric Action Recognition with Multimodal Data
por: Calzavara, Marco, et al.
Publicado: (2025)
por: Calzavara, Marco, et al.
Publicado: (2025)
Exploring Audio Hallucination in Egocentric Video Understanding
por: Seth, Ashish, et al.
Publicado: (2026)
por: Seth, Ashish, et al.
Publicado: (2026)
Evaluating Multimodal Large Language Models on Vertically Written Japanese Text
por: Sasagawa, Keito, et al.
Publicado: (2025)
por: Sasagawa, Keito, et al.
Publicado: (2025)
When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models
por: Yakun, Cui, et al.
Publicado: (2026)
por: Yakun, Cui, et al.
Publicado: (2026)
Like Humans to Few-Shot Learning through Knowledge Permeation of Vision and Text
por: Jia, Yuyu, et al.
Publicado: (2024)
por: Jia, Yuyu, et al.
Publicado: (2024)
EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric Videos
por: Li, Yuxuan, et al.
Publicado: (2025)
por: Li, Yuxuan, et al.
Publicado: (2025)
Capturing Fine-Grained Alignments Improves 3D Affordance Detection
por: Tokumitsu, Junsei, et al.
Publicado: (2025)
por: Tokumitsu, Junsei, et al.
Publicado: (2025)
Ejemplares similares
-
Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision
por: Yoshida, Tomoya, et al.
Publicado: (2025) -
Developing Vision-Language-Action Model from Egocentric Videos
por: Yoshida, Tomoya, et al.
Publicado: (2025) -
EgoOops: A Dataset for Mistake Action Detection from Egocentric Videos referring to Procedural Texts
por: Haneji, Yuto, et al.
Publicado: (2024) -
STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models
por: Ukai, Mahiro, et al.
Publicado: (2025) -
Vision Language Model-based Caption Evaluation Method Leveraging Visual Context Extraction
por: Maeda, Koki, et al.
Publicado: (2024)