Capturing Fine-Grained Alignments Improves 3D Affordance Detection
Fuente:
arXiv
Guardado en:
| Autores principales: | Tokumitsu, Junsei, Wada, Yuiga |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ZINA: Multimodal Fine-grained Hallucination Detection and Editing
por: Wada, Yuiga, et al.
Publicado: (2025)
por: Wada, Yuiga, et al.
Publicado: (2025)
DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning
por: Matsuda, Kazuki, et al.
Publicado: (2024)
por: Matsuda, Kazuki, et al.
Publicado: (2024)
Polos: Multimodal Metric Learning from Human Feedback for Image Captioning
por: Wada, Yuiga, et al.
Publicado: (2024)
por: Wada, Yuiga, et al.
Publicado: (2024)
VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
por: Matsuda, Kazuki, et al.
Publicado: (2025)
por: Matsuda, Kazuki, et al.
Publicado: (2025)
AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection
por: Wang, Hao, et al.
Publicado: (2026)
por: Wang, Hao, et al.
Publicado: (2026)
FG-CLIP: Fine-Grained Visual and Textual Alignment
por: Xie, Chunyu, et al.
Publicado: (2025)
por: Xie, Chunyu, et al.
Publicado: (2025)
Informative Text-Image Alignment for Visual Affordance Learning with Foundation Models
por: Zhang, Qian, et al.
Publicado: (2025)
por: Zhang, Qian, et al.
Publicado: (2025)
Grounding 3D Scene Affordance From Egocentric Interactions
por: Liu, Cuiyu, et al.
Publicado: (2024)
por: Liu, Cuiyu, et al.
Publicado: (2024)
Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding
por: Kim, Namho, et al.
Publicado: (2025)
por: Kim, Namho, et al.
Publicado: (2025)
Decomposed Vision-Language Alignment for Fine-Grained Open-Vocabulary Segmentation
por: Wang, Chenhao, et al.
Publicado: (2026)
por: Wang, Chenhao, et al.
Publicado: (2026)
Fine-Grained Pillar Feature Encoding Via Spatio-Temporal Virtual Grid for 3D Object Detection
por: Park, Konyul, et al.
Publicado: (2024)
por: Park, Konyul, et al.
Publicado: (2024)
Improving Fine-Grained Rice Leaf Disease Detection via Angular-Compactness Dual Loss Learning
por: Mia, Md. Rokon, et al.
Publicado: (2026)
por: Mia, Md. Rokon, et al.
Publicado: (2026)
COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts
por: Wang, Bingli, et al.
Publicado: (2026)
por: Wang, Bingli, et al.
Publicado: (2026)
RAAP: Retrieval-Augmented Affordance Prediction with Cross-Image Action Alignment
por: Zhuang, Qiyuan, et al.
Publicado: (2026)
por: Zhuang, Qiyuan, et al.
Publicado: (2026)
FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs
por: Asokan, Mothilal, et al.
Publicado: (2025)
por: Asokan, Mothilal, et al.
Publicado: (2025)
Tooth-Diffusion: Guided 3D CBCT Synthesis with Fine-Grained Tooth Conditioning
por: Said, Said Djafar, et al.
Publicado: (2025)
por: Said, Said Djafar, et al.
Publicado: (2025)
Novel Extraction of Discriminative Fine-Grained Feature to Improve Retinal Vessel Segmentation
por: Zeng, Shuang, et al.
Publicado: (2025)
por: Zeng, Shuang, et al.
Publicado: (2025)
GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding
por: Shao, Yawen, et al.
Publicado: (2024)
por: Shao, Yawen, et al.
Publicado: (2024)
SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model
por: Yu, Chunlin, et al.
Publicado: (2024)
por: Yu, Chunlin, et al.
Publicado: (2024)
Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training
por: Qiu, Longtian, et al.
Publicado: (2024)
por: Qiu, Longtian, et al.
Publicado: (2024)
Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
por: Zhu, He, et al.
Publicado: (2025)
por: Zhu, He, et al.
Publicado: (2025)
SSR3D-LLM: Structured Spatial Reasoning via Latent Steps for Fine-Grained Grounding in Unified 3D-LLMs
por: Li, Jiawei, et al.
Publicado: (2026)
por: Li, Jiawei, et al.
Publicado: (2026)
ITSELF: Attention Guided Fine-Grained Alignment for Vision-Language Retrieval
por: Nguyen, Tien-Huy, et al.
Publicado: (2026)
por: Nguyen, Tien-Huy, et al.
Publicado: (2026)
Learning Relative Representations for Fine-Grained Multimodal Alignment with Limited Data
por: Kim, Shiwon, et al.
Publicado: (2026)
por: Kim, Shiwon, et al.
Publicado: (2026)
Unlocking Vision-Language Models for Video Anomaly Detection via Fine-Grained Prompting
por: Zou, Shu, et al.
Publicado: (2025)
por: Zou, Shu, et al.
Publicado: (2025)
FB-CLIP: Fine-Grained Zero-Shot Anomaly Detection with Foreground-Background Disentanglement
por: Hu, Ming, et al.
Publicado: (2026)
por: Hu, Ming, et al.
Publicado: (2026)
Learning Environment-Aware Affordance for 3D Articulated Object Manipulation under Occlusions
por: Wu, Ruihai, et al.
Publicado: (2023)
por: Wu, Ruihai, et al.
Publicado: (2023)
Zooming into Comics: Region-Aware RL Improves Fine-Grained Comic Understanding in Vision-Language Models
por: Chen, Yule, et al.
Publicado: (2025)
por: Chen, Yule, et al.
Publicado: (2025)
Medical Image Synthesis via Fine-Grained Image-Text Alignment and Anatomy-Pathology Prompting
por: Chen, Wenting, et al.
Publicado: (2024)
por: Chen, Wenting, et al.
Publicado: (2024)
MiSCHiEF: A Benchmark in Minimal-Pairs of Safety and Culture for Holistic Evaluation of Fine-Grained Image-Caption Alignment
por: Banerjee, Sagarika, et al.
Publicado: (2026)
por: Banerjee, Sagarika, et al.
Publicado: (2026)
Region-Grounded Report Generation for 3D Medical Imaging: A Fine-Grained Dataset and Graph-Enhanced Framework
por: Nguyen, Cong Huy, et al.
Publicado: (2026)
por: Nguyen, Cong Huy, et al.
Publicado: (2026)
ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling
por: Yan, Siming, et al.
Publicado: (2024)
por: Yan, Siming, et al.
Publicado: (2024)
SYNTHIA: Novel Concept Design with Affordance Composition
por: Ha, Hyeonjeong, et al.
Publicado: (2025)
por: Ha, Hyeonjeong, et al.
Publicado: (2025)
Self-Explainable Affordance Learning with Embodied Caption
por: Zhang, Zhipeng, et al.
Publicado: (2024)
por: Zhang, Zhipeng, et al.
Publicado: (2024)
Look Before You Fuse: 2D-Guided Cross-Modal Alignment for Robust 3D Detection
por: Li, Xiang, et al.
Publicado: (2025)
por: Li, Xiang, et al.
Publicado: (2025)
Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report Generation
por: Chen, Wenting, et al.
Publicado: (2023)
por: Chen, Wenting, et al.
Publicado: (2023)
Fine-Grained Representation for Lane Topology Reasoning
por: Xu, Guoqing, et al.
Publicado: (2025)
por: Xu, Guoqing, et al.
Publicado: (2025)
Saccadic Vision for Fine-Grained Visual Classification
por: Schmidt, Johann, et al.
Publicado: (2025)
por: Schmidt, Johann, et al.
Publicado: (2025)
Fine-Grained ImageNet Classification in the Wild
por: Lymperaiou, Maria, et al.
Publicado: (2023)
por: Lymperaiou, Maria, et al.
Publicado: (2023)
FILA: Fine-Grained Vision Language Models
por: Zhu, Shiding, et al.
Publicado: (2024)
por: Zhu, Shiding, et al.
Publicado: (2024)
Ejemplares similares
-
ZINA: Multimodal Fine-grained Hallucination Detection and Editing
por: Wada, Yuiga, et al.
Publicado: (2025) -
DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning
por: Matsuda, Kazuki, et al.
Publicado: (2024) -
Polos: Multimodal Metric Learning from Human Feedback for Image Captioning
por: Wada, Yuiga, et al.
Publicado: (2024) -
VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
por: Matsuda, Kazuki, et al.
Publicado: (2025) -
AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection
por: Wang, Hao, et al.
Publicado: (2026)