Gespeichert in:
| Hauptverfasser: | Wang, Zehao, Liu, Xinpeng, Zhang, Yudonglin, Wu, Xiaoqian, Fang, Zhou, Fang, Yifan, Pu, Junfu, Lu, Cewu, Li, Yong-Lu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2412.04939 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Isolated Islands to Pangea: Unifying Semantic Space for Human Action Understanding
von: Li, Yong-Lu, et al.
Veröffentlicht: (2023)
von: Li, Yong-Lu, et al.
Veröffentlicht: (2023)
Dynamic Relation Inference via Verb Embeddings
von: Suissa, Omri, et al.
Veröffentlicht: (2025)
von: Suissa, Omri, et al.
Veröffentlicht: (2025)
Bridging the Gap between Human Motion and Action Semantics via Kinematic Phrases
von: Liu, Xinpeng, et al.
Veröffentlicht: (2023)
von: Liu, Xinpeng, et al.
Veröffentlicht: (2023)
Revisit Human-Scene Interaction via Space Occupancy
von: Liu, Xinpeng, et al.
Veröffentlicht: (2023)
von: Liu, Xinpeng, et al.
Veröffentlicht: (2023)
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
von: Beňová, Ivana, et al.
Veröffentlicht: (2024)
von: Beňová, Ivana, et al.
Veröffentlicht: (2024)
VerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awareness
von: Cha, SeungJu, et al.
Veröffentlicht: (2025)
von: Cha, SeungJu, et al.
Veröffentlicht: (2025)
Digital Gene: Learning about the Physical World through Analytic Concepts
von: Sun, Jianhua, et al.
Veröffentlicht: (2025)
von: Sun, Jianhua, et al.
Veröffentlicht: (2025)
Mirage: Unveiling Hidden Artifacts in Synthetic Images with Large Vision-Language Models
von: Sharma, Pranav, et al.
Veröffentlicht: (2025)
von: Sharma, Pranav, et al.
Veröffentlicht: (2025)
Homogeneous Dynamics Space for Heterogeneous Humans
von: Liu, Xinpeng, et al.
Veröffentlicht: (2024)
von: Liu, Xinpeng, et al.
Veröffentlicht: (2024)
ImDy: Human Inverse Dynamics from Imitated Observations
von: Liu, Xinpeng, et al.
Veröffentlicht: (2024)
von: Liu, Xinpeng, et al.
Veröffentlicht: (2024)
Low-Rank Similarity Mining for Multimodal Dataset Distillation
von: Xu, Yue, et al.
Veröffentlicht: (2024)
von: Xu, Yue, et al.
Veröffentlicht: (2024)
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
von: He, Zhentao, et al.
Veröffentlicht: (2025)
von: He, Zhentao, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
von: Lu, Yifan, et al.
Veröffentlicht: (2025)
von: Lu, Yifan, et al.
Veröffentlicht: (2025)
Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink
von: Wang, Yining, et al.
Veröffentlicht: (2025)
von: Wang, Yining, et al.
Veröffentlicht: (2025)
Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts
von: Wei, Jiude, et al.
Veröffentlicht: (2025)
von: Wei, Jiude, et al.
Veröffentlicht: (2025)
DeformPAM: Data-Efficient Learning for Long-horizon Deformable Object Manipulation via Preference-based Action Alignment
von: Chen, Wendi, et al.
Veröffentlicht: (2024)
von: Chen, Wendi, et al.
Veröffentlicht: (2024)
LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment
von: Xu, Yifu, et al.
Veröffentlicht: (2026)
von: Xu, Yifu, et al.
Veröffentlicht: (2026)
Towards Robust Evaluation of Visual Activity Recognition: Resolving Verb Ambiguity with Sense Clustering
von: Yao, Louie Hong, et al.
Veröffentlicht: (2025)
von: Yao, Louie Hong, et al.
Veröffentlicht: (2025)
The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
von: Leng, Sicong, et al.
Veröffentlicht: (2024)
von: Leng, Sicong, et al.
Veröffentlicht: (2024)
Graspness Discovery in Clutters for Fast and Accurate Grasp Detection
von: Wang, Chenxi, et al.
Veröffentlicht: (2024)
von: Wang, Chenxi, et al.
Veröffentlicht: (2024)
AnyDexGrasp: General Dexterous Grasping for Different Hands with Human-level Learning Efficiency
von: Fang, Hao-Shu, et al.
Veröffentlicht: (2025)
von: Fang, Hao-Shu, et al.
Veröffentlicht: (2025)
Dense Policy: Bidirectional Autoregressive Learning of Actions
von: Su, Yue, et al.
Veröffentlicht: (2025)
von: Su, Yue, et al.
Veröffentlicht: (2025)
Interacted Object Grounding in Spatio-Temporal Human-Object Interactions
von: Liu, Xiaoyang, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoyang, et al.
Veröffentlicht: (2024)
Grounded Chain-of-Thought for Multimodal Large Language Models
von: Wu, Qiong, et al.
Veröffentlicht: (2025)
von: Wu, Qiong, et al.
Veröffentlicht: (2025)
Dancing with Still Images: Video Distillation via Static-Dynamic Disentanglement
von: Wang, Ziyu, et al.
Veröffentlicht: (2023)
von: Wang, Ziyu, et al.
Veröffentlicht: (2023)
Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
von: Yin, Jianghao, et al.
Veröffentlicht: (2026)
von: Yin, Jianghao, et al.
Veröffentlicht: (2026)
Hallucination of Multimodal Large Language Models: A Survey
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models
von: Dong, Xinpeng, et al.
Veröffentlicht: (2026)
von: Dong, Xinpeng, et al.
Veröffentlicht: (2026)
DISCO: Embodied Navigation and Interaction via Differentiable Scene Semantics and Dual-level Control
von: Xu, Xinyu, et al.
Veröffentlicht: (2024)
von: Xu, Xinyu, et al.
Veröffentlicht: (2024)
RSHallu: Dual-Mode Hallucination Evaluation for Remote-Sensing Multimodal Large Language Models with Domain-Tailored Mitigation
von: Zhou, Zihui, et al.
Veröffentlicht: (2026)
von: Zhou, Zihui, et al.
Veröffentlicht: (2026)
Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
PACE: A Large-Scale Dataset with Pose Annotations in Cluttered Environments
von: You, Yang, et al.
Veröffentlicht: (2023)
von: You, Yang, et al.
Veröffentlicht: (2023)
Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
von: Pu, Bowei, et al.
Veröffentlicht: (2025)
von: Pu, Bowei, et al.
Veröffentlicht: (2025)
Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
von: Li, Jiale, et al.
Veröffentlicht: (2025)
von: Li, Jiale, et al.
Veröffentlicht: (2025)
Graph Integrated Multimodal Concept Bottleneck Model
von: Lin, Jiakai, et al.
Veröffentlicht: (2025)
von: Lin, Jiakai, et al.
Veröffentlicht: (2025)
Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models
von: Zhang, Jinrui, et al.
Veröffentlicht: (2024)
von: Zhang, Jinrui, et al.
Veröffentlicht: (2024)
SSL-OTA: Unveiling Backdoor Threats in Self-Supervised Learning for Object Detection
von: Wang, Qiannan, et al.
Veröffentlicht: (2023)
von: Wang, Qiannan, et al.
Veröffentlicht: (2023)
Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models
von: Zhou, Haoran, et al.
Veröffentlicht: (2025)
von: Zhou, Haoran, et al.
Veröffentlicht: (2025)
ChatGarment: Garment Estimation, Generation and Editing via Large Language Models
von: Bian, Siyuan, et al.
Veröffentlicht: (2024)
von: Bian, Siyuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
From Isolated Islands to Pangea: Unifying Semantic Space for Human Action Understanding
von: Li, Yong-Lu, et al.
Veröffentlicht: (2023) -
Dynamic Relation Inference via Verb Embeddings
von: Suissa, Omri, et al.
Veröffentlicht: (2025) -
Bridging the Gap between Human Motion and Action Semantics via Kinematic Phrases
von: Liu, Xinpeng, et al.
Veröffentlicht: (2023) -
Revisit Human-Scene Interaction via Space Occupancy
von: Liu, Xinpeng, et al.
Veröffentlicht: (2023) -
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
von: Beňová, Ivana, et al.
Veröffentlicht: (2024)