Self-Explainable Affordance Learning with Embodied Caption
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Zhipeng, Wei, Zhimin, Sun, Guolei, Wang, Peng, Van Gool, Luc |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Learning Local and Global Temporal Contexts for Video Semantic Segmentation
por: Sun, Guolei, et al.
Publicado: (2022)
por: Sun, Guolei, et al.
Publicado: (2022)
ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives
por: Fu, Yuqian, et al.
Publicado: (2024)
por: Fu, Yuqian, et al.
Publicado: (2024)
Learning Generative Interactive Environments By Trained Agent Exploration
por: Kazemi, Naser, et al.
Publicado: (2024)
por: Kazemi, Naser, et al.
Publicado: (2024)
TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
por: Motamed, Saman, et al.
Publicado: (2025)
por: Motamed, Saman, et al.
Publicado: (2025)
EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
por: Zhang, Deheng, et al.
Publicado: (2025)
por: Zhang, Deheng, et al.
Publicado: (2025)
Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025
por: Fu, Yuqian, et al.
Publicado: (2025)
por: Fu, Yuqian, et al.
Publicado: (2025)
Informative Text-Image Alignment for Visual Affordance Learning with Foundation Models
por: Zhang, Qian, et al.
Publicado: (2025)
por: Zhang, Qian, et al.
Publicado: (2025)
Language-Guided Instance-Aware Domain-Adaptive Panoptic Segmentation
por: Mansour, Elham Amin, et al.
Publicado: (2024)
por: Mansour, Elham Amin, et al.
Publicado: (2024)
Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models
por: Zhong, Yi, et al.
Publicado: (2026)
por: Zhong, Yi, et al.
Publicado: (2026)
SYNTHIA: Novel Concept Design with Affordance Composition
por: Ha, Hyeonjeong, et al.
Publicado: (2025)
por: Ha, Hyeonjeong, et al.
Publicado: (2025)
Samba: Synchronized Set-of-Sequences Modeling for Multiple Object Tracking
por: Segu, Mattia, et al.
Publicado: (2024)
por: Segu, Mattia, et al.
Publicado: (2024)
VOID: Video Object and Interaction Deletion
por: Motamed, Saman, et al.
Publicado: (2026)
por: Motamed, Saman, et al.
Publicado: (2026)
MedVSR: Medical Video Super-Resolution with Cross State-Space Propagation
por: Liu, Xinyu, et al.
Publicado: (2025)
por: Liu, Xinyu, et al.
Publicado: (2025)
Probabilistic Sampling of Balanced K-Means using Adiabatic Quantum Computing
por: Zaech, Jan-Nico, et al.
Publicado: (2023)
por: Zaech, Jan-Nico, et al.
Publicado: (2023)
CCCaption: Dual-Reward Reinforcement Learning for Complete and Correct Image Captioning
por: Tang, Zhijiang, et al.
Publicado: (2026)
por: Tang, Zhijiang, et al.
Publicado: (2026)
Explainable, Multi-modal Wound Infection Classification from Images Augmented with Generated Captions
por: Busaranuvong, Palawat, et al.
Publicado: (2025)
por: Busaranuvong, Palawat, et al.
Publicado: (2025)
EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations
por: Kim, Hyunjong, et al.
Publicado: (2025)
por: Kim, Hyunjong, et al.
Publicado: (2025)
Text-driven Affordance Learning from Egocentric Vision
por: Yoshida, Tomoya, et al.
Publicado: (2024)
por: Yoshida, Tomoya, et al.
Publicado: (2024)
Selective Contrastive Learning for Weakly Supervised Affordance Grounding
por: Moon, WonJun, et al.
Publicado: (2025)
por: Moon, WonJun, et al.
Publicado: (2025)
Vision Transformers with Hierarchical Attention
por: Liu, Yun, et al.
Publicado: (2021)
por: Liu, Yun, et al.
Publicado: (2021)
General Flow as Foundation Affordance for Scalable Robot Learning
por: Yuan, Chengbo, et al.
Publicado: (2024)
por: Yuan, Chengbo, et al.
Publicado: (2024)
HalCECE: A Framework for Explainable Hallucination Detection through Conceptual Counterfactuals in Image Captioning
por: Lymperaiou, Maria, et al.
Publicado: (2025)
por: Lymperaiou, Maria, et al.
Publicado: (2025)
Learning Additively Compositional Latent Actions for Embodied AI
por: Wei, Hangxing, et al.
Publicado: (2026)
por: Wei, Hangxing, et al.
Publicado: (2026)
EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering
por: Li, Yanjun, et al.
Publicado: (2025)
por: Li, Yanjun, et al.
Publicado: (2025)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
por: Wu, Shengqiong, et al.
Publicado: (2025)
por: Wu, Shengqiong, et al.
Publicado: (2025)
Hierarchical Dual-Change Collaborative Learning for UAV Scene Change Captioning
por: Chen, Fuhai, et al.
Publicado: (2026)
por: Chen, Fuhai, et al.
Publicado: (2026)
Rethinking Global Context in Crowd Counting
por: Sun, Guolei, et al.
Publicado: (2021)
por: Sun, Guolei, et al.
Publicado: (2021)
CamSAM2: Segment Anything Accurately in Camouflaged Videos
por: Zhou, Yuli, et al.
Publicado: (2025)
por: Zhou, Yuli, et al.
Publicado: (2025)
Grounding 3D Scene Affordance From Egocentric Interactions
por: Liu, Cuiyu, et al.
Publicado: (2024)
por: Liu, Cuiyu, et al.
Publicado: (2024)
Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning
por: Zhao, Baining, et al.
Publicado: (2025)
por: Zhao, Baining, et al.
Publicado: (2025)
Top-Down Semantic Refinement for Image Captioning
por: Zhang, Jusheng, et al.
Publicado: (2025)
por: Zhang, Jusheng, et al.
Publicado: (2025)
Sam-Guided Enhanced Fine-Grained Encoding with Mixed Semantic Learning for Medical Image Captioning
por: Zhang, Zhenyu, et al.
Publicado: (2023)
por: Zhang, Zhenyu, et al.
Publicado: (2023)
KTVIC: A Vietnamese Image Captioning Dataset on the Life Domain
por: Pham, Anh-Cuong, et al.
Publicado: (2024)
por: Pham, Anh-Cuong, et al.
Publicado: (2024)
Vision-Language Navigation with Embodied Intelligence: A Survey
por: Gao, Peng, et al.
Publicado: (2024)
por: Gao, Peng, et al.
Publicado: (2024)
CaptionFool: Universal Image Captioning Model Attacks
por: Parekh, Swapnil
Publicado: (2026)
por: Parekh, Swapnil
Publicado: (2026)
Self-Corrected Image Generation with Explainable Latent Rewards
por: Luo, Yinyi, et al.
Publicado: (2026)
por: Luo, Yinyi, et al.
Publicado: (2026)
When SAM2 Meets Video Camouflaged Object Segmentation: A Comprehensive Evaluation and Adaptation
por: Zhou, Yuli, et al.
Publicado: (2024)
por: Zhou, Yuli, et al.
Publicado: (2024)
World to Code: Multi-modal Data Generation via Self-Instructed Compositional Captioning and Filtering
por: Wang, Jiacong, et al.
Publicado: (2024)
por: Wang, Jiacong, et al.
Publicado: (2024)
SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning
por: Kim, Ye-Chan, et al.
Publicado: (2026)
por: Kim, Ye-Chan, et al.
Publicado: (2026)
Towards Learning a Generalist Model for Embodied Navigation
por: Zheng, Duo, et al.
Publicado: (2023)
por: Zheng, Duo, et al.
Publicado: (2023)
Ejemplares similares
-
Learning Local and Global Temporal Contexts for Video Semantic Segmentation
por: Sun, Guolei, et al.
Publicado: (2022) -
ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives
por: Fu, Yuqian, et al.
Publicado: (2024) -
Learning Generative Interactive Environments By Trained Agent Exploration
por: Kazemi, Naser, et al.
Publicado: (2024) -
TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
por: Motamed, Saman, et al.
Publicado: (2025) -
EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
por: Zhang, Deheng, et al.
Publicado: (2025)