IDSelect: A RL-Based Cost-Aware Selection Agent for Video-based Multi-Modal Person Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Ji, Yuyang, Shen, Yixuan, Nguyen, Kien, Zhou, Lifeng, Liu, Feng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Person Recognition in Aerial Surveillance: A Decade Survey
by: Nguyen, Kien, et al.
Published: (2025)
by: Nguyen, Kien, et al.
Published: (2025)
AG-VPReID: A Challenging Large-Scale Benchmark for Aerial-Ground Video-based Person Re-Identification
by: Nguyen, Huy, et al.
Published: (2025)
by: Nguyen, Huy, et al.
Published: (2025)
AG-VPReID.VIR: Bridging Aerial and Ground Platforms for Video-based Visible-Infrared Person Re-ID
by: Nguyen, Huy, et al.
Published: (2025)
by: Nguyen, Huy, et al.
Published: (2025)
From 3D Pose to Prose: Biomechanics-Grounded Vision--Language Coaching
by: Ji, Yuyang, et al.
Published: (2026)
by: Ji, Yuyang, et al.
Published: (2026)
Non-Colliding Biometric Identities for Digital Entities: Geometry, Capacity, and Million-Scale Virtual Identity Provisioning
by: Ji, Yuyang, et al.
Published: (2026)
by: Ji, Yuyang, et al.
Published: (2026)
SGTA: Scene-Graph Based Multi-Modal Traffic Agent for Video Understanding
by: Zhou, Xingcheng, et al.
Published: (2026)
by: Zhou, Xingcheng, et al.
Published: (2026)
AG-ReID.v2: Bridging Aerial and Ground Views for Person Re-identification
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
Revisiting the Role of Texture in 3D Person Re-identification
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
by: Huang, Jiehui, et al.
Published: (2025)
by: Huang, Jiehui, et al.
Published: (2025)
OmniGCD: Abstracting Generalized Category Discovery for Modality Agnosticism
by: Shipard, Jordan, et al.
Published: (2026)
by: Shipard, Jordan, et al.
Published: (2026)
h-Edit: Effective and Flexible Diffusion-Based Editing via Doob's h-Transform
by: Nguyen, Toan, et al.
Published: (2025)
by: Nguyen, Toan, et al.
Published: (2025)
Socratic Chart: Cooperating Multiple Agents for Robust SVG Chart Understanding
by: Ji, Yuyang, et al.
Published: (2025)
by: Ji, Yuyang, et al.
Published: (2025)
FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification
by: Sun, Zhen, et al.
Published: (2025)
by: Sun, Zhen, et al.
Published: (2025)
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
by: Li, Liyang, et al.
Published: (2026)
by: Li, Liyang, et al.
Published: (2026)
CoSTA$\ast$: Cost-Sensitive Toolpath Agent for Multi-turn Image Editing
by: Gupta, Advait, et al.
Published: (2025)
by: Gupta, Advait, et al.
Published: (2025)
Filling the Gaps: A Multitask Hybrid Multiscale Generative Framework for Missing Modality in Remote Sensing Semantic Segmentation
by: Kieu, Nhi, et al.
Published: (2025)
by: Kieu, Nhi, et al.
Published: (2025)
DIS2: Disentanglement Meets Distillation with Classwise Attention for Robust Remote Sensing Segmentation under Missing Modalities
by: Kieu, Nhi, et al.
Published: (2026)
by: Kieu, Nhi, et al.
Published: (2026)
Modality-Aware Bias Mitigation and Invariance Learning for Unsupervised Visible-Infrared Person Re-Identification
by: Wang, Menglin, et al.
Published: (2025)
by: Wang, Menglin, et al.
Published: (2025)
Action Selection Learning for Multi-label Multi-view Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2024)
by: Nguyen, Trung Thanh, et al.
Published: (2024)
Recognition of Daily Activities through Multi-Modal Deep Learning: A Video, Pose, and Object-Aware Approach for Ambient Assisted Living
by: Hashemifard, Kooshan, et al.
Published: (2026)
by: Hashemifard, Kooshan, et al.
Published: (2026)
FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation
by: Le, Minh Khoa, et al.
Published: (2026)
by: Le, Minh Khoa, et al.
Published: (2026)
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition
by: Cong, Kaixuan, et al.
Published: (2025)
by: Cong, Kaixuan, et al.
Published: (2025)
MultiCOIN: Multi-Modal COntrollable Video INbetweening
by: Tanveer, Maham, et al.
Published: (2025)
by: Tanveer, Maham, et al.
Published: (2025)
BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration
by: Gao, Bo, et al.
Published: (2026)
by: Gao, Bo, et al.
Published: (2026)
VideoAgent2: Enhancing the LLM-Based Agent System for Long-Form Video Understanding by Uncertainty-Aware CoT
by: Zhi, Zhuo, et al.
Published: (2025)
by: Zhi, Zhuo, et al.
Published: (2025)
UniSOT: A Unified Framework for Multi-Modality Single Object Tracking
by: Ma, Yinchao, et al.
Published: (2025)
by: Ma, Yinchao, et al.
Published: (2025)
Event-based Video Person Re-identification via Cross-Modality and Temporal Collaboration
by: Li, Renkai, et al.
Published: (2025)
by: Li, Renkai, et al.
Published: (2025)
AG-VPReID 2025: Aerial-Ground Video-based Person Re-identification Challenge Results
by: Nguyen, Kien, et al.
Published: (2025)
by: Nguyen, Kien, et al.
Published: (2025)
Multi-Modality Co-Learning for Efficient Skeleton-based Action Recognition
by: Liu, Jinfu, et al.
Published: (2024)
by: Liu, Jinfu, et al.
Published: (2024)
LLandMark: A Multi-Agent Framework for Landmark-Aware Multimodal Interactive Video Retrieval
by: Phung, Minh-Chi, et al.
Published: (2026)
by: Phung, Minh-Chi, et al.
Published: (2026)
Learning Language-Driven Sequence-Level Modal-Invariant Representations for Video-Based Visible-Infrared Person Re-Identification
by: Yang, Xiaomei, et al.
Published: (2026)
by: Yang, Xiaomei, et al.
Published: (2026)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
by: Wang, Junyang, et al.
Published: (2024)
by: Wang, Junyang, et al.
Published: (2024)
End-to-End Multi-Person Pose Estimation with Pose-Aware Video Transformer
by: Yu, Yonghui, et al.
Published: (2025)
by: Yu, Yonghui, et al.
Published: (2025)
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
by: Wu, Peiran, et al.
Published: (2025)
by: Wu, Peiran, et al.
Published: (2025)
Multi-modal Speech Emotion Recognition via Feature Distribution Adaptation Network
by: Li, Shaokai, et al.
Published: (2024)
by: Li, Shaokai, et al.
Published: (2024)
ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
by: Pan, Panwang, et al.
Published: (2025)
by: Pan, Panwang, et al.
Published: (2025)
Mono3DV: Monocular 3D Object Detection with 3D-Aware Bipartite Matching and Variational Query DeNoising
by: Vu, Kiet Dang, et al.
Published: (2026)
by: Vu, Kiet Dang, et al.
Published: (2026)
MMGait: Towards Multi-Modal Gait Recognition
by: Wang, Chenye, et al.
Published: (2026)
by: Wang, Chenye, et al.
Published: (2026)
SuMa: A Subspace Mapping Approach for Robust and Effective Concept Erasure in Text-to-Image Diffusion Models
by: Nguyen, Kien, et al.
Published: (2025)
by: Nguyen, Kien, et al.
Published: (2025)
Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification
by: Wang, Xiao, et al.
Published: (2026)
by: Wang, Xiao, et al.
Published: (2026)
Similar Items
-
Person Recognition in Aerial Surveillance: A Decade Survey
by: Nguyen, Kien, et al.
Published: (2025) -
AG-VPReID: A Challenging Large-Scale Benchmark for Aerial-Ground Video-based Person Re-Identification
by: Nguyen, Huy, et al.
Published: (2025) -
AG-VPReID.VIR: Bridging Aerial and Ground Platforms for Video-based Visible-Infrared Person Re-ID
by: Nguyen, Huy, et al.
Published: (2025) -
From 3D Pose to Prose: Biomechanics-Grounded Vision--Language Coaching
by: Ji, Yuyang, et al.
Published: (2026) -
Non-Colliding Biometric Identities for Digital Entities: Geometry, Capacity, and Million-Scale Virtual Identity Provisioning
by: Ji, Yuyang, et al.
Published: (2026)