HAMMER: Harnessing MLLM via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Lei, Chen, Yong, Su, Yuejiao, Wang, Yi, Liu, Moyun, Chau, Lap-Pui |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
by: Yao, Lei, et al.
Published: (2025)
by: Yao, Lei, et al.
Published: (2025)
SGIFormer: Semantic-guided and Geometric-enhanced Interleaving Transformer for 3D Instance Segmentation
by: Yao, Lei, et al.
Published: (2024)
by: Yao, Lei, et al.
Published: (2024)
LaSSM: Efficient Semantic-Spatial Query Decoding via Local Aggregation and State Space Models for 3D Instance Segmentation
by: Yao, Lei, et al.
Published: (2026)
by: Yao, Lei, et al.
Published: (2026)
CaRe-Ego: Contact-aware Relationship Modeling for Egocentric Interactive Hand-object Segmentation
by: Su, Yuejiao, et al.
Published: (2024)
by: Su, Yuejiao, et al.
Published: (2024)
ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric Interaction
by: Su, Yuejiao, et al.
Published: (2025)
by: Su, Yuejiao, et al.
Published: (2025)
EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding
by: Su, Yuejiao, et al.
Published: (2026)
by: Su, Yuejiao, et al.
Published: (2026)
Interaction-aware Representation Modeling with Co-occurrence Consistency for Egocentric Hand-Object Parsing
by: Su, Yuejiao, et al.
Published: (2026)
by: Su, Yuejiao, et al.
Published: (2026)
GVSynergy-Det: Synergistic Gaussian-Voxel Representations for Multi-View 3D Object Detection
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning
by: Li, Junlong, et al.
Published: (2026)
by: Li, Junlong, et al.
Published: (2026)
HSNet: Heterogeneous Subgraph Network for Single Image Super-resolution
by: Hu, Qiongyang, et al.
Published: (2025)
by: Hu, Qiongyang, et al.
Published: (2025)
3DGeoDet: General-purpose Geometry-aware Image-based 3D Object Detection
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
A Survey of Embodied Learning for Object-Centric Robotic Manipulation
by: Zheng, Ying, et al.
Published: (2024)
by: Zheng, Ying, et al.
Published: (2024)
Symmetric Multi-Similarity Loss for EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2024
by: Wang, Xiaoqi, et al.
Published: (2024)
by: Wang, Xiaoqi, et al.
Published: (2024)
Weakly-supervised Part-Attention and Mentored Networks for Vehicle Re-Identification
by: Tang, Lisha, et al.
Published: (2021)
by: Tang, Lisha, et al.
Published: (2021)
Egocentric Human-Object Interaction Detection: A New Benchmark and Method
by: Deng, Kunyuan, et al.
Published: (2025)
by: Deng, Kunyuan, et al.
Published: (2025)
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization
by: Wang, Xiaoqi, et al.
Published: (2025)
by: Wang, Xiaoqi, et al.
Published: (2025)
OccProphet: Pushing Efficiency Frontier of Camera-Only 4D Occupancy Forecasting with Observer-Forecaster-Refiner Framework
by: Chen, Junliang, et al.
Published: (2025)
by: Chen, Junliang, et al.
Published: (2025)
ProCal: Probability Calibration for Neighborhood-Guided Source-Free Domain Adaptation
by: Zheng, Ying, et al.
Published: (2026)
by: Zheng, Ying, et al.
Published: (2026)
Evolution-Inspired Sample Competition for Deep Neural Network Optimization
by: Zheng, Ying, et al.
Published: (2026)
by: Zheng, Ying, et al.
Published: (2026)
SignEye: Traffic Sign Interpretation from Vehicle First-Person View
by: Yang, Chuang, et al.
Published: (2024)
by: Yang, Chuang, et al.
Published: (2024)
Fuzzy-aware Loss for Source-free Domain Adaptation in Visual Emotion Recognition
by: Zheng, Ying, et al.
Published: (2025)
by: Zheng, Ying, et al.
Published: (2025)
GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency
by: Lu, Dongyue, et al.
Published: (2024)
by: Lu, Dongyue, et al.
Published: (2024)
HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis
by: Chen, Mingjin, et al.
Published: (2026)
by: Chen, Mingjin, et al.
Published: (2026)
A Survey on Occupancy Perception for Autonomous Driving: The Information Fusion Perspective
by: Xu, Huaiyuan, et al.
Published: (2024)
by: Xu, Huaiyuan, et al.
Published: (2024)
GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding
by: Shao, Yawen, et al.
Published: (2024)
by: Shao, Yawen, et al.
Published: (2024)
Learning 2D Invariant Affordance Knowledge for 3D Affordance Grounding
by: Gao, Xianqiang, et al.
Published: (2024)
by: Gao, Xianqiang, et al.
Published: (2024)
EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
MASS: Mesh-inellipse Aligned Deformable Surfel Splatting for Hand Reconstruction and Rendering from Egocentric Monocular Video
by: Zhu, Haoyu, et al.
Published: (2026)
by: Zhu, Haoyu, et al.
Published: (2026)
Grounding by Remembering: Cross-Scene and In-Scene Memory for 3D Functional Affordances
by: Wang, Qirui, et al.
Published: (2026)
by: Wang, Qirui, et al.
Published: (2026)
VAGNet: Grounding 3D Affordance from Human-Object Interactions in Videos
by: Mao, Aihua, et al.
Published: (2026)
by: Mao, Aihua, et al.
Published: (2026)
3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds
by: Chu, Hengshuo, et al.
Published: (2025)
by: Chu, Hengshuo, et al.
Published: (2025)
Affostruction: 3D Affordance Grounding with Generative Reconstruction
by: Park, Chunghyun, et al.
Published: (2026)
by: Park, Chunghyun, et al.
Published: (2026)
CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects
by: Li, Jingliang, et al.
Published: (2026)
by: Li, Jingliang, et al.
Published: (2026)
Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment
by: Zhao, Pengfei, et al.
Published: (2025)
by: Zhao, Pengfei, et al.
Published: (2025)
AffordanceSAM: Segment Anything Once More in Affordance Grounding
by: Jiang, Dengyang, et al.
Published: (2025)
by: Jiang, Dengyang, et al.
Published: (2025)
Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
by: Liu, Tianyi, et al.
Published: (2025)
by: Liu, Tianyi, et al.
Published: (2025)
PADetBench: Towards Benchmarking Physical Attacks against Object Detection
by: Lian, Jiawei, et al.
Published: (2024)
by: Lian, Jiawei, et al.
Published: (2024)
AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis
by: Wu, Xiaofei, et al.
Published: (2026)
by: Wu, Xiaofei, et al.
Published: (2026)
Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges
by: Li, Junlong, et al.
Published: (2025)
by: Li, Junlong, et al.
Published: (2025)
PromptSR: Cascade Prompting for Lightweight Image Super-Resolution
by: Liu, Wenyang, et al.
Published: (2025)
by: Liu, Wenyang, et al.
Published: (2025)
Similar Items
-
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
by: Yao, Lei, et al.
Published: (2025) -
SGIFormer: Semantic-guided and Geometric-enhanced Interleaving Transformer for 3D Instance Segmentation
by: Yao, Lei, et al.
Published: (2024) -
LaSSM: Efficient Semantic-Spatial Query Decoding via Local Aggregation and State Space Models for 3D Instance Segmentation
by: Yao, Lei, et al.
Published: (2026) -
CaRe-Ego: Contact-aware Relationship Modeling for Egocentric Interactive Hand-object Segmentation
by: Su, Yuejiao, et al.
Published: (2024) -
ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric Interaction
by: Su, Yuejiao, et al.
Published: (2025)