Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Matsuishi, Koki, Ukita, Kosuke, Okita, Tsuyoshi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-instance Learning as Downstream Task of Self-Supervised Learning-based Pre-trained Model
by: Matsuishi, Koki, et al.
Published: (2025)
by: Matsuishi, Koki, et al.
Published: (2025)
Image Classification Using a Diffusion Model as a Pre-Training Model
by: Ukita, Kosuke, et al.
Published: (2025)
by: Ukita, Kosuke, et al.
Published: (2025)
High-Performance Self-Supervised Learning by Joint Training of Flow Matching
by: Ukita, Kosuke, et al.
Published: (2025)
by: Ukita, Kosuke, et al.
Published: (2025)
Brain Hematoma Marker Recognition Using Multitask Learning: SwinTransformer and Swin-Unet
by: Hirata, Kodai, et al.
Published: (2025)
by: Hirata, Kodai, et al.
Published: (2025)
COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition
by: Chen, Baiyu, et al.
Published: (2025)
by: Chen, Baiyu, et al.
Published: (2025)
Robust Multimodal Learning via Cross-Modal Proxy Tokens
by: Reza, Md Kaykobad, et al.
Published: (2025)
by: Reza, Md Kaykobad, et al.
Published: (2025)
Diffusion Model-based Activity Completion for AI Motion Capture from Videos
by: Huayu, Gao, et al.
Published: (2025)
by: Huayu, Gao, et al.
Published: (2025)
(Almost) Free Modality Stitching of Foundation Models
by: Singh, Jaisidh, et al.
Published: (2025)
by: Singh, Jaisidh, et al.
Published: (2025)
MIND: Modality-Informed Knowledge Distillation Framework for Multimodal Clinical Prediction Tasks
by: Guerra-Manzanares, Alejandro, et al.
Published: (2025)
by: Guerra-Manzanares, Alejandro, et al.
Published: (2025)
Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning
by: Hu, Tao, et al.
Published: (2026)
by: Hu, Tao, et al.
Published: (2026)
Optimization-Free Test-Time Adaptation for Cross-Person Activity Recognition
by: Wang, Shuoyuan, et al.
Published: (2023)
by: Wang, Shuoyuan, et al.
Published: (2023)
Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
by: Zhang, Yuhui, et al.
Published: (2024)
by: Zhang, Yuhui, et al.
Published: (2024)
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
by: Ramachandran, Rahul, et al.
Published: (2025)
by: Ramachandran, Rahul, et al.
Published: (2025)
Post-pre-training for Modality Alignment in Vision-Language Foundation Models
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
Data Augmentation Techniques for Cross-Domain WiFi CSI-based Human Activity Recognition
by: Strohmayer, Julian, et al.
Published: (2024)
by: Strohmayer, Julian, et al.
Published: (2024)
Diagnosing and Mitigating Modality Interference in Multimodal Large Language Models
by: Cai, Rui, et al.
Published: (2025)
by: Cai, Rui, et al.
Published: (2025)
CroMe: Multimodal Fake News Detection using Cross-Modal Tri-Transformer and Metric Learning
by: Choi, Eunjee, et al.
Published: (2025)
by: Choi, Eunjee, et al.
Published: (2025)
Understanding Retrieval-Augmented Task Adaptation for Vision-Language Models
by: Ming, Yifei, et al.
Published: (2024)
by: Ming, Yifei, et al.
Published: (2024)
4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
by: Bachmann, Roman, et al.
Published: (2024)
by: Bachmann, Roman, et al.
Published: (2024)
MORPH: PDE Foundation Models with Arbitrary Data Modality
by: Rautela, Mahindra Singh, et al.
Published: (2025)
by: Rautela, Mahindra Singh, et al.
Published: (2025)
Transformer-based Models to Deal with Heterogeneous Environments in Human Activity Recognition
by: EK, Sannara, et al.
Published: (2022)
by: EK, Sannara, et al.
Published: (2022)
Mars-Bench: A Benchmark for Evaluating Foundation Models for Mars Science Tasks
by: Purohit, Mirali, et al.
Published: (2025)
by: Purohit, Mirali, et al.
Published: (2025)
Explanation Bottleneck Models
by: Yamaguchi, Shin'ya, et al.
Published: (2024)
by: Yamaguchi, Shin'ya, et al.
Published: (2024)
Collaborative Temporal Feature Generation via Critic-Free Reinforcement Learning for Cross-User Sensor-Based Activity Recognition
by: Ye, Xiaozhou, et al.
Published: (2026)
by: Ye, Xiaozhou, et al.
Published: (2026)
Thoughts on Objectives of Sparse and Hierarchical Masked Image Model
by: Miyazaki, Asahi, et al.
Published: (2025)
by: Miyazaki, Asahi, et al.
Published: (2025)
Are Deep Learning Models Robust to Partial Object Occlusion in Visual Recognition Tasks?
by: Kassaw, Kaleb, et al.
Published: (2024)
by: Kassaw, Kaleb, et al.
Published: (2024)
Hierarchical Network Fusion for Multi-Modal Electron Micrograph Representation Learning with Foundational Large Language Models
by: Srinivas, Sakhinana Sagar, et al.
Published: (2024)
by: Srinivas, Sakhinana Sagar, et al.
Published: (2024)
Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning
by: Yaras, Can, et al.
Published: (2024)
by: Yaras, Can, et al.
Published: (2024)
Deep Multimodal Learning with Missing Modality: A Survey
by: Wu, Renjie, et al.
Published: (2024)
by: Wu, Renjie, et al.
Published: (2024)
HEMM: Holistic Evaluation of Multimodal Foundation Models
by: Liang, Paul Pu, et al.
Published: (2024)
by: Liang, Paul Pu, et al.
Published: (2024)
RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition
by: Liu, Ziyu, et al.
Published: (2024)
by: Liu, Ziyu, et al.
Published: (2024)
Cross-Domain Generalization Limits of Vision Foundation Models in Facial Deepfake Detection
by: Delibasoglu, Ibrahim
Published: (2026)
by: Delibasoglu, Ibrahim
Published: (2026)
PFM-VEPAR: Prompting Foundation Models for RGB-Event Camera based Pedestrian Attribute Recognition
by: Xu, Minghe, et al.
Published: (2026)
by: Xu, Minghe, et al.
Published: (2026)
Hyperdimensional Cross-Modal Alignment of Frozen Language and Image Models for Efficient Image Captioning
by: Dalvi, Abhishek, et al.
Published: (2026)
by: Dalvi, Abhishek, et al.
Published: (2026)
Hierarchical Hypercomplex Network for Multimodal Emotion Recognition
by: Lopez, Eleonora, et al.
Published: (2024)
by: Lopez, Eleonora, et al.
Published: (2024)
Transferability-Guided Cross-Domain Cross-Task Transfer Learning
by: Tan, Yang, et al.
Published: (2022)
by: Tan, Yang, et al.
Published: (2022)
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
by: Lai, Zhengfeng, et al.
Published: (2024)
by: Lai, Zhengfeng, et al.
Published: (2024)
Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis
by: Moon, Jihyun, et al.
Published: (2025)
by: Moon, Jihyun, et al.
Published: (2025)
CLASH: A Benchmark for Cross-Modal Contradiction Detection
by: Popordanoska, Teodora, et al.
Published: (2025)
by: Popordanoska, Teodora, et al.
Published: (2025)
Detecting Informative Channels: ActionFormer
by: Zhao, Kunpeng, et al.
Published: (2025)
by: Zhao, Kunpeng, et al.
Published: (2025)
Similar Items
-
Multi-instance Learning as Downstream Task of Self-Supervised Learning-based Pre-trained Model
by: Matsuishi, Koki, et al.
Published: (2025) -
Image Classification Using a Diffusion Model as a Pre-Training Model
by: Ukita, Kosuke, et al.
Published: (2025) -
High-Performance Self-Supervised Learning by Joint Training of Flow Matching
by: Ukita, Kosuke, et al.
Published: (2025) -
Brain Hematoma Marker Recognition Using Multitask Learning: SwinTransformer and Swin-Unet
by: Hirata, Kodai, et al.
Published: (2025) -
COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition
by: Chen, Baiyu, et al.
Published: (2025)