ViMoNet: A Multimodal Vision-Language Framework for Human Behavior Understanding from Motion and Video
Fuente:
arXiv
Saved in:
| Main Authors: | Gupta, Rajan Das, Wei, Lei, Rahat, Md Yeasin, Fahad, Nafiz, Ahmed, Abir, Hui, Liew Tze |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GANDiff FR: Hybrid GAN Diffusion Synthesis for Causal Bias Attribution in Face Recognition
by: Reaj, Md Asgor Hossain, et al.
Published: (2025)
by: Reaj, Md Asgor Hossain, et al.
Published: (2025)
From Explanations to Architecture: Explainability-Driven CNN Refinement for Brain Tumor Classification in MRI
by: Gupta, Rajan Das, et al.
Published: (2025)
by: Gupta, Rajan Das, et al.
Published: (2025)
Exploring the Convergence of HCI and Evolving Technologies in Information Systems
by: Gupta, Rajan Das, et al.
Published: (2025)
by: Gupta, Rajan Das, et al.
Published: (2025)
Cross-Lingual Probing and Community-Grounded Analysis of Gender Bias in Low-Resource Bengali
by: Reaj, Md Asgor Hossain, et al.
Published: (2026)
by: Reaj, Md Asgor Hossain, et al.
Published: (2026)
Real-Time Confidence Detection through Facial Expressions and Hand Gestures
by: Sakib, Tanjil Hasan, et al.
Published: (2025)
by: Sakib, Tanjil Hasan, et al.
Published: (2025)
BRAINS: A Retrieval-Augmented System for Alzheimer's Detection and Monitoring
by: Gupta, Rajan Das, et al.
Published: (2025)
by: Gupta, Rajan Das, et al.
Published: (2025)
Advancing Exchange Rate Forecasting: Leveraging Machine Learning and AI for Enhanced Accuracy in Global Financial Markets
by: Rahat, Md. Yeasin, et al.
Published: (2025)
by: Rahat, Md. Yeasin, et al.
Published: (2025)
MotionLLM: Understanding Human Behaviors from Human Motions and Videos
by: Chen, Ling-Hao, et al.
Published: (2024)
by: Chen, Ling-Hao, et al.
Published: (2024)
DyCAF-Net: Dynamic Class-Aware Fusion Network
by: Jahin, Md Abrar, et al.
Published: (2025)
by: Jahin, Md Abrar, et al.
Published: (2025)
ViMo: Generating Motions from Casual Videos
by: Qiu, Liangdong, et al.
Published: (2024)
by: Qiu, Liangdong, et al.
Published: (2024)
CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos
by: Zhao, Chengfeng, et al.
Published: (2026)
by: Zhao, Chengfeng, et al.
Published: (2026)
MIC: Medical Image Classification Using Chest X-ray (COVID-19 and Pneumonia) Dataset with the Help of CNN and Customized CNN
by: Fahad, Nafiz, et al.
Published: (2024)
by: Fahad, Nafiz, et al.
Published: (2024)
Eco‐Friendly Supply Chains: Unveiling the Keys to Sustainable Success in the Textile Industry of an Emerging Economy
by: Md. Yeasin, et al.
Published: (2025)
by: Md. Yeasin, et al.
Published: (2025)
Multimodal Programming in Computer Science with Interactive Assistance Powered by Large Language Model
by: Gupta, Rajan Das, et al.
Published: (2025)
by: Gupta, Rajan Das, et al.
Published: (2025)
VLAgeBench: Benchmarking Large Vision-Language Models for Zero-Shot Human Age Estimation
by: Sajib, Rakib Hossain, et al.
Published: (2026)
by: Sajib, Rakib Hossain, et al.
Published: (2026)
Vision Transformers for End-to-End Quark-Gluon Jet Classification from Calorimeter Images
by: Jahin, Md Abrar, et al.
Published: (2025)
by: Jahin, Md Abrar, et al.
Published: (2025)
ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos
by: Chen, Zerui, et al.
Published: (2024)
by: Chen, Zerui, et al.
Published: (2024)
Predictive Analytics for Dementia: Machine Learning on Healthcare Data
by: Opee, Shafiul Ajam, et al.
Published: (2026)
by: Opee, Shafiul Ajam, et al.
Published: (2026)
HMVLM: Human Motion-Vision-Lanuage Model via MoE LoRA
by: Hu, Lei, et al.
Published: (2025)
by: Hu, Lei, et al.
Published: (2025)
A Minimalistic Approach to Predict and Understand the Relation of App Usage with Students' Academic Performances
by: Ahmed, Md Sabbir, et al.
Published: (2025)
by: Ahmed, Md Sabbir, et al.
Published: (2025)
AdeptHEQ-FL: Adaptive Homomorphic Encryption for Federated Learning of Hybrid Classical-Quantum Models with Dynamic Layer Sparing
by: Jahin, Md Abrar, et al.
Published: (2025)
by: Jahin, Md Abrar, et al.
Published: (2025)
Multilingual Question Answering in Low-Resource Settings: A Dzongkha-English Benchmark for Foundation Models
by: Hosain, Md. Tanzib, et al.
Published: (2025)
by: Hosain, Md. Tanzib, et al.
Published: (2025)
Exploring Chaotic Motion of a Particle in the Centre of a Galaxy with a Prolate Halo
by: Nag, Uditi, et al.
Published: (2026)
by: Nag, Uditi, et al.
Published: (2026)
KinMo: Kinematic-aware Human Motion Understanding and Generation
by: Zhang, Pengfei, et al.
Published: (2024)
by: Zhang, Pengfei, et al.
Published: (2024)
HuMoCon: Concept Discovery for Human Motion Understanding
by: Fang, Qihang, et al.
Published: (2025)
by: Fang, Qihang, et al.
Published: (2025)
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
by: Maaz, Muhammad, et al.
Published: (2023)
by: Maaz, Muhammad, et al.
Published: (2023)
HyCARD-Net: A Synergistic Hybrid Intelligence Framework for Cardiovascular Disease Diagnosis
by: Gupta, Rajan Das, et al.
Published: (2026)
by: Gupta, Rajan Das, et al.
Published: (2026)
ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video Understanding
by: Yashima, Daichi, et al.
Published: (2026)
by: Yashima, Daichi, et al.
Published: (2026)
Improving Human Motion Plausibility with Body Momentum
by: Nguyen, Ha Linh, et al.
Published: (2025)
by: Nguyen, Ha Linh, et al.
Published: (2025)
MoViE: Mobile Diffusion for Video Editing
by: Karjauv, Adil, et al.
Published: (2024)
by: Karjauv, Adil, et al.
Published: (2024)
Vision-Language Models for Automated Chest X-ray Interpretation: Leveraging ViT and GPT-2
by: Islam, Md. Rakibul, et al.
Published: (2025)
by: Islam, Md. Rakibul, et al.
Published: (2025)
MoViD: View-Invariant 3D Human Pose Estimation via Motion-View Disentanglement
by: Liu, Yejia, et al.
Published: (2026)
by: Liu, Yejia, et al.
Published: (2026)
The Longest Common Bitonic Subsequence: A Match-Sensitive Dynamic Programming Approach
by: Rahat, Md. Tanzeem, et al.
Published: (2025)
by: Rahat, Md. Tanzeem, et al.
Published: (2025)
Joint Optimization of RU Allocation and C-SR in Multi-AP Coordinated Wi-Fi Systems
by: Hasan, Md Rahat, et al.
Published: (2025)
by: Hasan, Md Rahat, et al.
Published: (2025)
Vision-Guided Chunking Is All You Need: Enhancing RAG with Multimodal Document Understanding
by: Tripathi, Vishesh, et al.
Published: (2025)
by: Tripathi, Vishesh, et al.
Published: (2025)
ViMU: Benchmarking Video Metaphorical Understanding
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
by: Rasheed, Hanoona, et al.
Published: (2025)
by: Rasheed, Hanoona, et al.
Published: (2025)
Towards Speaker Identification with Minimal Dataset and Constrained Resources using 1D-Convolution Neural Network
by: Shahan, Irfan Nafiz, et al.
Published: (2024)
by: Shahan, Irfan Nafiz, et al.
Published: (2024)
FreeMotion: MoCap-Free Human Motion Synthesis with Multimodal Large Language Models
by: Zhang, Zhikai, et al.
Published: (2024)
by: Zhang, Zhikai, et al.
Published: (2024)
Privacy-Preserving Chest X-ray Report Generation via Multimodal Federated Learning with ViT and GPT-2
by: Hossain, Md. Zahid, et al.
Published: (2025)
by: Hossain, Md. Zahid, et al.
Published: (2025)
Similar Items
-
GANDiff FR: Hybrid GAN Diffusion Synthesis for Causal Bias Attribution in Face Recognition
by: Reaj, Md Asgor Hossain, et al.
Published: (2025) -
From Explanations to Architecture: Explainability-Driven CNN Refinement for Brain Tumor Classification in MRI
by: Gupta, Rajan Das, et al.
Published: (2025) -
Exploring the Convergence of HCI and Evolving Technologies in Information Systems
by: Gupta, Rajan Das, et al.
Published: (2025) -
Cross-Lingual Probing and Community-Grounded Analysis of Gender Bias in Low-Resource Bengali
by: Reaj, Md Asgor Hossain, et al.
Published: (2026) -
Real-Time Confidence Detection through Facial Expressions and Hand Gestures
by: Sakib, Tanjil Hasan, et al.
Published: (2025)