RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Jie, Hou, Ruibing, Zhao, Jiahe, Chang, Hong, Shan, Shiguang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding
von: Zhao, Jiahe, et al.
Veröffentlicht: (2025)
von: Zhao, Jiahe, et al.
Veröffentlicht: (2025)
HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
von: Li, Keliang, et al.
Veröffentlicht: (2024)
von: Li, Keliang, et al.
Veröffentlicht: (2024)
UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing
von: Li, Yiheng, et al.
Veröffentlicht: (2024)
von: Li, Yiheng, et al.
Veröffentlicht: (2024)
un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP
von: Li, Yinqi, et al.
Veröffentlicht: (2025)
von: Li, Yinqi, et al.
Veröffentlicht: (2025)
MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
von: Hou, Ruibing, et al.
Veröffentlicht: (2025)
von: Hou, Ruibing, et al.
Veröffentlicht: (2025)
CLIP-Guided Adaptable Self-Supervised Learning for Human-Centric Visual Tasks
von: Luo, Mingshuang, et al.
Veröffentlicht: (2026)
von: Luo, Mingshuang, et al.
Veröffentlicht: (2026)
Clothes-Changing Person Re-Identification with Feasibility-Aware Intermediary Matching
von: Zhao, Jiahe, et al.
Veröffentlicht: (2024)
von: Zhao, Jiahe, et al.
Veröffentlicht: (2024)
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
von: Li, Keliang, et al.
Veröffentlicht: (2025)
von: Li, Keliang, et al.
Veröffentlicht: (2025)
DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks
von: Li, Yinqi, et al.
Veröffentlicht: (2025)
von: Li, Yinqi, et al.
Veröffentlicht: (2025)
Component-Based Out-of-Distribution Detection
von: Liu, Wenrui, et al.
Veröffentlicht: (2026)
von: Liu, Wenrui, et al.
Veröffentlicht: (2026)
UMFC: Unsupervised Multi-Domain Feature Calibration for Vision-Language Models
von: Liang, Jiachen, et al.
Veröffentlicht: (2024)
von: Liang, Jiachen, et al.
Veröffentlicht: (2024)
Revisiting Logit Distributions for Reliable Out-of-Distribution Detection
von: Liang, Jiachen, et al.
Veröffentlicht: (2025)
von: Liang, Jiachen, et al.
Veröffentlicht: (2025)
Generalized Semi-Supervised Learning via Self-Supervised Feature Adaptation
von: Liang, Jiachen, et al.
Veröffentlicht: (2024)
von: Liang, Jiachen, et al.
Veröffentlicht: (2024)
M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation
von: Luo, Mingshuang, et al.
Veröffentlicht: (2024)
von: Luo, Mingshuang, et al.
Veröffentlicht: (2024)
AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling
von: Li, Yiheng, et al.
Veröffentlicht: (2026)
von: Li, Yiheng, et al.
Veröffentlicht: (2026)
RefAV: Towards Planning-Centric Scenario Mining
von: Davidson, Cainan, et al.
Veröffentlicht: (2025)
von: Davidson, Cainan, et al.
Veröffentlicht: (2025)
EgoMotion: Hierarchical Reasoning and Diffusion for Egocentric Vision-Language Motion Generation
von: Hou, Ruibing, et al.
Veröffentlicht: (2026)
von: Hou, Ruibing, et al.
Veröffentlicht: (2026)
RefComp: A Reference-guided Unified Framework for Unpaired Point Cloud Completion
von: Yang, Yixuan, et al.
Veröffentlicht: (2025)
von: Yang, Yixuan, et al.
Veröffentlicht: (2025)
OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling
von: Xiao, Linhui, et al.
Veröffentlicht: (2024)
von: Xiao, Linhui, et al.
Veröffentlicht: (2024)
Revisiting Multimodal Positional Encoding in Vision-Language Models
von: Huang, Jie, et al.
Veröffentlicht: (2025)
von: Huang, Jie, et al.
Veröffentlicht: (2025)
Face-MLLM: A Large Face Perception Model
von: Sun, Haomiao, et al.
Veröffentlicht: (2024)
von: Sun, Haomiao, et al.
Veröffentlicht: (2024)
Morph: A Motion-free Physics Optimization Framework for Human Motion Generation
von: Li, Zhuo, et al.
Veröffentlicht: (2024)
von: Li, Zhuo, et al.
Veröffentlicht: (2024)
RefCut: Interactive Segmentation with Reference Guidance
von: Lin, Zheng, et al.
Veröffentlicht: (2025)
von: Lin, Zheng, et al.
Veröffentlicht: (2025)
Assimilation Matters: Model-level Backdoor Detection in Vision-Language Pretrained Models
von: Wang, Zhongqi, et al.
Veröffentlicht: (2025)
von: Wang, Zhongqi, et al.
Veröffentlicht: (2025)
Anonymization Prompt Learning for Facial Privacy-Preserving Text-to-Image Generation
von: Shi, Liang, et al.
Veröffentlicht: (2024)
von: Shi, Liang, et al.
Veröffentlicht: (2024)
Generalized Face Liveness Detection via De-fake Face Generator
von: Long, Xingming, et al.
Veröffentlicht: (2024)
von: Long, Xingming, et al.
Veröffentlicht: (2024)
Confidence Aware Learning for Reliable Face Anti-spoofing
von: Long, Xingming, et al.
Veröffentlicht: (2024)
von: Long, Xingming, et al.
Veröffentlicht: (2024)
Steering Vision-Language Pre-trained Models for Incremental Face Presentation Attack Detection
von: Li, Haoze, et al.
Veröffentlicht: (2025)
von: Li, Haoze, et al.
Veröffentlicht: (2025)
BIMM: Brain Inspired Masked Modeling for Video Representation Learning
von: Wan, Zhifan, et al.
Veröffentlicht: (2024)
von: Wan, Zhifan, et al.
Veröffentlicht: (2024)
Dynamic Attention Analysis for Backdoor Detection in Text-to-Image Diffusion Models
von: Wang, Zhongqi, et al.
Veröffentlicht: (2025)
von: Wang, Zhongqi, et al.
Veröffentlicht: (2025)
VOPE: Revisiting Hallucination of Vision-Language Models in Voluntary Imagination Task
von: Long, Xingming, et al.
Veröffentlicht: (2025)
von: Long, Xingming, et al.
Veröffentlicht: (2025)
T2IShield: Defending Against Backdoors on Text-to-Image Diffusion Models
von: Wang, Zhongqi, et al.
Veröffentlicht: (2024)
von: Wang, Zhongqi, et al.
Veröffentlicht: (2024)
RefTok: Reference-Based Tokenization for Video Generation
von: Fan, Xiang, et al.
Veröffentlicht: (2025)
von: Fan, Xiang, et al.
Veröffentlicht: (2025)
RefAlign: Representation Alignment for Reference-to-Video Generation
von: Wang, Lei, et al.
Veröffentlicht: (2026)
von: Wang, Lei, et al.
Veröffentlicht: (2026)
Pre-trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness
von: Wang, Sibo, et al.
Veröffentlicht: (2024)
von: Wang, Sibo, et al.
Veröffentlicht: (2024)
UniVision: A Unified Framework for Vision-Centric 3D Perception
von: Hong, Yu, et al.
Veröffentlicht: (2024)
von: Hong, Yu, et al.
Veröffentlicht: (2024)
Dysca: A Dynamic and Scalable Benchmark for Evaluating Perception Ability of LVLMs
von: Zhang, Jie, et al.
Veröffentlicht: (2024)
von: Zhang, Jie, et al.
Veröffentlicht: (2024)
RefVSR++: Exploiting Reference Inputs for Reference-based Video Super-resolution
von: Zou, Han, et al.
Veröffentlicht: (2023)
von: Zou, Han, et al.
Veröffentlicht: (2023)
LatRef-Diff: Latent and Reference-Guided Diffusion for Facial Attribute Editing and Style Manipulation
von: Huang, Wenmin, et al.
Veröffentlicht: (2026)
von: Huang, Wenmin, et al.
Veröffentlicht: (2026)
RefFusion: Reference Adapted Diffusion Models for 3D Scene Inpainting
von: Mirzaei, Ashkan, et al.
Veröffentlicht: (2024)
von: Mirzaei, Ashkan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding
von: Zhao, Jiahe, et al.
Veröffentlicht: (2025) -
HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
von: Li, Keliang, et al.
Veröffentlicht: (2024) -
UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing
von: Li, Yiheng, et al.
Veröffentlicht: (2024) -
un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP
von: Li, Yinqi, et al.
Veröffentlicht: (2025) -
MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
von: Hou, Ruibing, et al.
Veröffentlicht: (2025)