You Only Learn One Query: Learning Unified Human Query for Single-Stage Multi-Person Multi-Task Human-Centric Perception
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Sheng, Li, Shuhuai, Li, Tong, Liu, Wentao, Qian, Chen, Luo, Ping |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries
by: Xue, Wangyu, et al.
Published: (2024)
by: Xue, Wangyu, et al.
Published: (2024)
When Pedestrian Detection Meets Multi-Modal Learning: Generalist Model and Benchmark Dataset
by: Zhang, Yi, et al.
Published: (2024)
by: Zhang, Yi, et al.
Published: (2024)
OneReward: Unified Mask-Guided Image Generation via Multi-Task Human Preference Learning
by: Gong, Yuan, et al.
Published: (2025)
by: Gong, Yuan, et al.
Published: (2025)
Category Query Learning for Human-Object Interaction Classification
by: Xie, Chi, et al.
Published: (2023)
by: Xie, Chi, et al.
Published: (2023)
You Only Need One Stage: Novel-View Synthesis From A Single Blind Face Image
by: Wang, Taoyue, et al.
Published: (2026)
by: Wang, Taoyue, et al.
Published: (2026)
EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation
by: Meng, Rang, et al.
Published: (2025)
by: Meng, Rang, et al.
Published: (2025)
NADER: Neural Architecture Design via Multi-Agent Collaboration
by: Yang, Zekang, et al.
Published: (2024)
by: Yang, Zekang, et al.
Published: (2024)
UniParser: Multi-Human Parsing with Unified Correlation Representation Learning
by: Chu, Jiaming, et al.
Published: (2023)
by: Chu, Jiaming, et al.
Published: (2023)
CLIP-Guided Adaptable Self-Supervised Learning for Human-Centric Visual Tasks
by: Luo, Mingshuang, et al.
Published: (2026)
by: Luo, Mingshuang, et al.
Published: (2026)
You Only Look at Once for Real-time and Generic Multi-Task
by: Wang, Jiayuan, et al.
Published: (2023)
by: Wang, Jiayuan, et al.
Published: (2023)
UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image Personalization
by: He, Junjie, et al.
Published: (2024)
by: He, Junjie, et al.
Published: (2024)
RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios
by: Huang, Jie, et al.
Published: (2024)
by: Huang, Jie, et al.
Published: (2024)
Self-supervised One-Stage Learning for RF-based Multi-Person Pose Estimation
by: Shin, Seunghwan, et al.
Published: (2025)
by: Shin, Seunghwan, et al.
Published: (2025)
MMTL-UniAD: A Unified Framework for Multimodal and Multi-Task Learning in Assistive Driving Perception
by: Liu, Wenzhuo, et al.
Published: (2025)
by: Liu, Wenzhuo, et al.
Published: (2025)
Hierarchical Matching and Reasoning for Multi-Query Image Retrieval
by: Ji, Zhong, et al.
Published: (2023)
by: Ji, Zhong, et al.
Published: (2023)
All-in-One Transferring Image Compression from Human Perception to Multi-Machine Perception
by: Zhao, Jiancheng, et al.
Published: (2025)
by: Zhao, Jiancheng, et al.
Published: (2025)
INSTINCT: Instance-Level Interaction Architecture for Query-Based Collaborative Perception
by: Xu, Yunjiang, et al.
Published: (2025)
by: Xu, Yunjiang, et al.
Published: (2025)
RayFormer: Improving Query-Based Multi-Camera 3D Object Detection via Ray-Centric Strategies
by: Chu, Xiaomeng, et al.
Published: (2024)
by: Chu, Xiaomeng, et al.
Published: (2024)
Data Augmentation in Human-Centric Vision
by: Jiang, Wentao, et al.
Published: (2024)
by: Jiang, Wentao, et al.
Published: (2024)
UniFS: Universal Few-shot Instance Perception with Point Representations
by: Jin, Sheng, et al.
Published: (2024)
by: Jin, Sheng, et al.
Published: (2024)
StageInteractor: Query-based Object Detector with Cross-stage Interaction
by: Teng, Yao, et al.
Published: (2023)
by: Teng, Yao, et al.
Published: (2023)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
by: Chen, Liyang, et al.
Published: (2025)
by: Chen, Liyang, et al.
Published: (2025)
Evolving Without Ending: Unifying Multimodal Incremental Learning for Continual Panoptic Perception
by: Yuan, Bo, et al.
Published: (2026)
by: Yuan, Bo, et al.
Published: (2026)
PanopticQuery: Unified Query-Time Reasoning for 4D Scenes
by: Tang, Ruilin, et al.
Published: (2026)
by: Tang, Ruilin, et al.
Published: (2026)
RTMO: Towards High-Performance One-Stage Real-Time Multi-Person Pose Estimation
by: Lu, Peng, et al.
Published: (2023)
by: Lu, Peng, et al.
Published: (2023)
AutoMMLab: Automatically Generating Deployable Models from Language Instructions for Computer Vision Tasks
by: Yang, Zekang, et al.
Published: (2024)
by: Yang, Zekang, et al.
Published: (2024)
Learning Human Visual Attention on 3D Surfaces through Geometry-Queried Semantic Priors
by: Pahari, Soham, et al.
Published: (2026)
by: Pahari, Soham, et al.
Published: (2026)
Adaptive Query Prompting for Multi-Domain Landmark Detection
by: Li, Yuhui, et al.
Published: (2024)
by: Li, Yuhui, et al.
Published: (2024)
Efficient Action Counting with Dynamic Queries
by: Ma, Xiaoxuan, et al.
Published: (2024)
by: Ma, Xiaoxuan, et al.
Published: (2024)
Multi-HMR: Multi-Person Whole-Body Human Mesh Recovery in a Single Shot
by: Baradel, Fabien, et al.
Published: (2024)
by: Baradel, Fabien, et al.
Published: (2024)
DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation
by: Guo, Xu, et al.
Published: (2026)
by: Guo, Xu, et al.
Published: (2026)
Group-On: Boosting One-Shot Segmentation with Supportive Query
by: Zhou, Hanjing, et al.
Published: (2024)
by: Zhou, Hanjing, et al.
Published: (2024)
QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection
by: Wang, Yuxiao, et al.
Published: (2025)
by: Wang, Yuxiao, et al.
Published: (2025)
ActFormer: Scalable Collaborative Perception via Active Queries
by: Huang, Suozhi, et al.
Published: (2024)
by: Huang, Suozhi, et al.
Published: (2024)
Llama Learns to Direct: DirectorLLM for Human-Centric Video Generation
by: Song, Kunpeng, et al.
Published: (2024)
by: Song, Kunpeng, et al.
Published: (2024)
GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human Data
by: Wang, Wentao, et al.
Published: (2024)
by: Wang, Wentao, et al.
Published: (2024)
UV-M3TL: A Unified and Versatile Multimodal Multi-Task Learning Framework for Assistive Driving Perception
by: Liu, Wenzhuo, et al.
Published: (2026)
by: Liu, Wenzhuo, et al.
Published: (2026)
GKGNet: Group K-Nearest Neighbor based Graph Convolutional Network for Multi-Label Image Recognition
by: Yao, Ruijie, et al.
Published: (2023)
by: Yao, Ruijie, et al.
Published: (2023)
Coherent Human-Scene Reconstruction from Multi-Person Multi-View Video in a Single Pass
by: Kim, Sangmin, et al.
Published: (2026)
by: Kim, Sangmin, et al.
Published: (2026)
Cross-Domain Multi-Person Human Activity Recognition via Near-Field Wi-Fi Sensing
by: Li, Xin, et al.
Published: (2025)
by: Li, Xin, et al.
Published: (2025)
Similar Items
-
ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries
by: Xue, Wangyu, et al.
Published: (2024) -
When Pedestrian Detection Meets Multi-Modal Learning: Generalist Model and Benchmark Dataset
by: Zhang, Yi, et al.
Published: (2024) -
OneReward: Unified Mask-Guided Image Generation via Multi-Task Human Preference Learning
by: Gong, Yuan, et al.
Published: (2025) -
Category Query Learning for Human-Object Interaction Classification
by: Xie, Chi, et al.
Published: (2023) -
You Only Need One Stage: Novel-View Synthesis From A Single Blind Face Image
by: Wang, Taoyue, et al.
Published: (2026)