FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Jie, Guo, Xiao, Su, Yiyang, Jain, Anil, Liu, Xiaoming |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Quality-Guided Mixture of Score-Fusion Experts Framework for Human Recognition
by: Zhu, Jie, et al.
Published: (2025)
by: Zhu, Jie, et al.
Published: (2025)
On the Holistic Approach for Detecting Human Image Forgery
by: Guo, Xiao, et al.
Published: (2026)
by: Guo, Xiao, et al.
Published: (2026)
KeyPoint Relative Position Encoding for Face Recognition
by: Kim, Minchul, et al.
Published: (2024)
by: Kim, Minchul, et al.
Published: (2024)
Open-Set Biometrics: Beyond Good Closed-Set Models
by: Su, Yiyang, et al.
Published: (2024)
by: Su, Yiyang, et al.
Published: (2024)
LocalScore: Local Density-Aware Similarity Scoring for Biometrics
by: Su, Yiyang, et al.
Published: (2026)
by: Su, Yiyang, et al.
Published: (2026)
SapiensID: Foundation for Human Recognition
by: Kim, Minchul, et al.
Published: (2025)
by: Kim, Minchul, et al.
Published: (2025)
50 Years of Automated Face Recognition
by: Kim, Minchul, et al.
Published: (2025)
by: Kim, Minchul, et al.
Published: (2025)
Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?
by: Zhu, Jie, et al.
Published: (2026)
by: Zhu, Jie, et al.
Published: (2026)
Person Recognition at Altitude and Range: Fusion of Face, Body Shape and Gait
by: Liu, Feng, et al.
Published: (2025)
by: Liu, Feng, et al.
Published: (2025)
DepthAgent: Towards Better Universal Depth Estimation via Sample-wise Expert Selection
by: Zhu, Jie, et al.
Published: (2026)
by: Zhu, Jie, et al.
Published: (2026)
Interpretable Perception and Reasoning for Audiovisual Geolocation
by: Su, Yiyang, et al.
Published: (2026)
by: Su, Yiyang, et al.
Published: (2026)
AffectAgent: Collaborative Multi-Agent Reasoning for Retrieval-Augmented Multimodal Emotion Recognition
by: Wang, Zeheng, et al.
Published: (2026)
by: Wang, Zeheng, et al.
Published: (2026)
Foundation versus Domain-specific Models: Performance Comparison, Fusion, and Explainability in Face Recognition
by: Sony, Redwan, et al.
Published: (2025)
by: Sony, Redwan, et al.
Published: (2025)
HAMoBE: Hierarchical and Adaptive Mixture of Biometric Experts for Video-based Person ReID
by: Su, Yiyang, et al.
Published: (2025)
by: Su, Yiyang, et al.
Published: (2025)
Adversarial Watermarking for Face Recognition
by: Yao, Yuguang, et al.
Published: (2024)
by: Yao, Yuguang, et al.
Published: (2024)
Towards Adaptive Fusion of Multimodal Deep Networks for Human Action Recognition
by: Yudistira, Novanto
Published: (2025)
by: Yudistira, Novanto
Published: (2025)
Hide and Seek: How Does Watermarking Impact Face Recognition?
by: Yao, Yuguang, et al.
Published: (2024)
by: Yao, Yuguang, et al.
Published: (2024)
Unsupervised Gait Recognition with Selective Fusion
by: Ren, Xuqian, et al.
Published: (2023)
by: Ren, Xuqian, et al.
Published: (2023)
Non-Colliding Biometric Identities for Digital Entities: Geometry, Capacity, and Million-Scale Virtual Identity Provisioning
by: Ji, Yuyang, et al.
Published: (2026)
by: Ji, Yuyang, et al.
Published: (2026)
ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding
by: Zhou, Yiyang, et al.
Published: (2025)
by: Zhou, Yiyang, et al.
Published: (2025)
AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios
by: Su, Zhaochen, et al.
Published: (2026)
by: Su, Zhaochen, et al.
Published: (2026)
Confidence-driven Gradient Modulation for Multimodal Human Activity Recognition: A Dynamic Contrastive Dual-Path Learning Approach
by: Ji, Panpan, et al.
Published: (2025)
by: Ji, Panpan, et al.
Published: (2025)
Universal Fingerprint Generation: Controllable Diffusion Model with Multimodal Conditions
by: Grosz, Steven A., et al.
Published: (2024)
by: Grosz, Steven A., et al.
Published: (2024)
AgentPose: Progressive Distribution Alignment via Feature Agent for Human Pose Distillation
by: Zhang, Feng, et al.
Published: (2025)
by: Zhang, Feng, et al.
Published: (2025)
HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics
by: Li, Weiqi, et al.
Published: (2025)
by: Li, Weiqi, et al.
Published: (2025)
Mobile Contactless Palmprint Recognition: Use of Multiscale, Multimodel Embeddings
by: Grosz, Steven A., et al.
Published: (2024)
by: Grosz, Steven A., et al.
Published: (2024)
Data Agent: Learning to Select Data via End-to-End Dynamic Optimization
by: Yang, Suorong, et al.
Published: (2026)
by: Yang, Suorong, et al.
Published: (2026)
SNN-Driven Multimodal Human Action Recognition via Sparse Spatial-Temporal Data Fusion
by: Zheng, Naichuan, et al.
Published: (2025)
by: Zheng, Naichuan, et al.
Published: (2025)
EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion
by: Wei, Jiangchuan, et al.
Published: (2025)
by: Wei, Jiangchuan, et al.
Published: (2025)
MM-Gesture: Towards Precise Micro-Gesture Recognition through Multimodal Fusion
by: Gu, Jihao, et al.
Published: (2025)
by: Gu, Jihao, et al.
Published: (2025)
Forge-and-Quench: Enhancing Image Generation for Higher Fidelity in Unified Multimodal Models
by: Zeng, Yanbing, et al.
Published: (2026)
by: Zeng, Yanbing, et al.
Published: (2026)
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
Application of Multimodal Fusion Deep Learning Model in Disease Recognition
by: Liu, Xiaoyi, et al.
Published: (2024)
by: Liu, Xiaoyi, et al.
Published: (2024)
IDSelect: A RL-Based Cost-Aware Selection Agent for Video-based Multi-Modal Person Recognition
by: Ji, Yuyang, et al.
Published: (2026)
by: Ji, Yuyang, et al.
Published: (2026)
Unbiased Dynamic Multimodal Fusion
by: Wei, Shicai, et al.
Published: (2026)
by: Wei, Shicai, et al.
Published: (2026)
Conflict-Aware Multimodal Fusion for Ambivalence and Hesitancy Recognition
by: Bekhouche, Salah Eddine, et al.
Published: (2026)
by: Bekhouche, Salah Eddine, et al.
Published: (2026)
A Trustworthy Method for Multimodal Emotion Recognition
by: Xue, Junxiao, et al.
Published: (2025)
by: Xue, Junxiao, et al.
Published: (2025)
Bringing Multimodal Large Language Models to Infrared-Visible Image Fusion Quality Assessment
by: Guo, Yuchen, et al.
Published: (2026)
by: Guo, Yuchen, et al.
Published: (2026)
OracleAgent: A Multimodal Reasoning Agent for Oracle Bone Script Research
by: Li, Caoshuo, et al.
Published: (2025)
by: Li, Caoshuo, et al.
Published: (2025)
Tracing Hyperparameter Dependencies for Model Parsing via Learnable Graph Pooling Network
by: Guo, Xiao, et al.
Published: (2023)
by: Guo, Xiao, et al.
Published: (2023)
Similar Items
-
A Quality-Guided Mixture of Score-Fusion Experts Framework for Human Recognition
by: Zhu, Jie, et al.
Published: (2025) -
On the Holistic Approach for Detecting Human Image Forgery
by: Guo, Xiao, et al.
Published: (2026) -
KeyPoint Relative Position Encoding for Face Recognition
by: Kim, Minchul, et al.
Published: (2024) -
Open-Set Biometrics: Beyond Good Closed-Set Models
by: Su, Yiyang, et al.
Published: (2024) -
LocalScore: Local Density-Aware Similarity Scoring for Biometrics
by: Su, Yiyang, et al.
Published: (2026)