A Multimodal-Multitask Framework with Cross-modal Relation and Hierarchical Interactive Attention for Semantic Comprehension
Fuente:
arXiv
Saved in:
| Main Authors: | Rehman, Mohammad Zia Ur, Raghuvanshi, Devraj, Jain, Umang, Bansal, Shubhi, Kumar, Nagendra |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ImpliHateVid: A Benchmark Dataset and Two-stage Contrastive Learning Framework for Implicit Hate Speech Detection in Videos
by: Rehman, Mohammad Zia Ur, et al.
Published: (2025)
by: Rehman, Mohammad Zia Ur, et al.
Published: (2025)
KisanQRS: A Deep Learning-based Automated Query-Response System for Agricultural Decision-Making
by: Rehman, Mohammad Zia Ur, et al.
Published: (2024)
by: Rehman, Mohammad Zia Ur, et al.
Published: (2024)
CoTCoNet: An Optimized Coupled Transformer-Convolutional Network with an Adaptive Graph Reconstruction for Leukemia Detection
by: Raghaw, Chandravardhan Singh, et al.
Published: (2024)
by: Raghaw, Chandravardhan Singh, et al.
Published: (2024)
A Context-aware Attention and Graph Neural Network-based Multimodal Framework for Misogyny Detection
by: Rehman, Mohammad Zia Ur, et al.
Published: (2025)
by: Rehman, Mohammad Zia Ur, et al.
Published: (2025)
T-MPEDNet: Unveiling the Synergy of Transformer-aware Multiscale Progressive Encoder-Decoder Network with Feature Recalibration for Tumor and Liver Segmentation
by: Raghaw, Chandravardhan Singh, et al.
Published: (2025)
by: Raghaw, Chandravardhan Singh, et al.
Published: (2025)
A Comprehensive Survey of Mamba Architectures for Medical Image Analysis: Classification, Segmentation, Restoration and Beyond
by: Bansal, Shubhi, et al.
Published: (2024)
by: Bansal, Shubhi, et al.
Published: (2024)
GCM-Net: Graph-enhanced Cross-Modal Infusion with a Metaheuristic-Driven Network for Video Sentiment and Emotion Analysis
by: Chaudhari, Prasad, et al.
Published: (2024)
by: Chaudhari, Prasad, et al.
Published: (2024)
Emotion-aware Dual Cross-Attentive Neural Network with Label Fusion for Stance Detection in Misinformative Social Media Content
by: Pangtey, Lata, et al.
Published: (2025)
by: Pangtey, Lata, et al.
Published: (2025)
An Explainable Deep Neural Network with Frequency-Aware Channel and Spatial Refinement for Flood Prediction in Sustainable Cities
by: Dar, Shahid Shafi, et al.
Published: (2025)
by: Dar, Shahid Shafi, et al.
Published: (2025)
An Explainable Contrastive-based Dilated Convolutional Network with Transformer for Pediatric Pneumonia Detection
by: Raghaw, Chandravardhan Singh, et al.
Published: (2024)
by: Raghaw, Chandravardhan Singh, et al.
Published: (2024)
D-HUMOR: Dark Humor Understanding via Multimodal Open-ended Reasoning -- A Benchmark Dataset and Method
by: Kasu, Sai Kartheek Reddy, et al.
Published: (2025)
by: Kasu, Sai Kartheek Reddy, et al.
Published: (2025)
A Hybrid Filtering for Micro-video Hashtag Recommendation using Graph-based Deep Neural Network
by: Bansal, Shubhi, et al.
Published: (2024)
by: Bansal, Shubhi, et al.
Published: (2024)
A multi-temporal multi-spectral attention-augmented deep convolution neural network with contrastive learning for crop yield prediction
by: Dangi, Shalini, et al.
Published: (2025)
by: Dangi, Shalini, et al.
Published: (2025)
An Adaptive Supervised Contrastive Learning Framework for Implicit Sexism Detection in Digital Social Networks
by: Rehman, Mohammad Zia Ur, et al.
Published: (2025)
by: Rehman, Mohammad Zia Ur, et al.
Published: (2025)
A Multimodal Framework for Depression Detection during Covid-19 via Harvesting Social Media: A Novel Dataset and Method
by: Anshul, Ashutosh, et al.
Published: (2025)
by: Anshul, Ashutosh, et al.
Published: (2025)
A Multilateral Attention-enhanced Deep Neural Network for Disease Outbreak Forecasting: A Case Study on COVID-19
by: Anshul, Ashutosh, et al.
Published: (2024)
by: Anshul, Ashutosh, et al.
Published: (2024)
M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation
by: Luo, Mingshuang, et al.
Published: (2024)
by: Luo, Mingshuang, et al.
Published: (2024)
SmartScan: An AI-based Interactive Framework for Automated Region Extraction from Satellite Images
by: Nagendra, Savinay, et al.
Published: (2025)
by: Nagendra, Savinay, et al.
Published: (2025)
MeD-3D: A Multimodal Deep Learning Framework for Precise Recurrence Prediction in Clear Cell Renal Cell Carcinoma (ccRCC)
by: Maqsood, Hasaan, et al.
Published: (2025)
by: Maqsood, Hasaan, et al.
Published: (2025)
Dual-Encoder Transformer-Based Multimodal Learning for Ischemic Stroke Lesion Segmentation Using Diffusion MRI
by: Usman, Muhammad, et al.
Published: (2025)
by: Usman, Muhammad, et al.
Published: (2025)
FissionFusion: Fast Geometric Generation and Hierarchical Souping for Medical Image Analysis
by: Sanjeev, Santosh, et al.
Published: (2024)
by: Sanjeev, Santosh, et al.
Published: (2024)
Envisioning MedCLIP: A Deep Dive into Explainability for Medical Vision-Language Models
by: Hashmi, Anees Ur Rehman, et al.
Published: (2024)
by: Hashmi, Anees Ur Rehman, et al.
Published: (2024)
Skin Cancer Classification: Hybrid CNN-Transformer Models with KAN-Based Fusion
by: Agarwal, Shubhi, et al.
Published: (2025)
by: Agarwal, Shubhi, et al.
Published: (2025)
Self-paced Multi-grained Cross-modal Interaction Modeling for Referring Expression Comprehension
by: Miao, Peihan, et al.
Published: (2022)
by: Miao, Peihan, et al.
Published: (2022)
A Human-Centered Approach for Improving Supervised Learning
by: Bansal, Shubhi, et al.
Published: (2024)
by: Bansal, Shubhi, et al.
Published: (2024)
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
by: Fu, Jinlan, et al.
Published: (2025)
by: Fu, Jinlan, et al.
Published: (2025)
Uncertainty-aware Semi-supervised Ensemble Teacher Framework for Multilingual Depression Detection
by: Rehman, Mohammad Zia Ur, et al.
Published: (2025)
by: Rehman, Mohammad Zia Ur, et al.
Published: (2025)
MMSF: Multitask and Multimodal Supervised Framework for WSI Classification and Survival Analysis
by: She, Chengying, et al.
Published: (2026)
by: She, Chengying, et al.
Published: (2026)
Sentiment and Hashtag-aware Attentive Deep Neural Network for Multimodal Post Popularity Prediction
by: Bansal, Shubhi, et al.
Published: (2024)
by: Bansal, Shubhi, et al.
Published: (2024)
SeCG: Semantic-Enhanced 3D Visual Grounding via Cross-modal Graph Attention
by: Xiao, Feng, et al.
Published: (2024)
by: Xiao, Feng, et al.
Published: (2024)
Integrating Feedback Loss from Bi-modal Sarcasm Detector for Sarcastic Speech Synthesis
by: Li, Zhu, et al.
Published: (2025)
by: Li, Zhu, et al.
Published: (2025)
MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation
by: Yao, Jihan, et al.
Published: (2025)
by: Yao, Jihan, et al.
Published: (2025)
MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
by: Xu, Shuolin, et al.
Published: (2025)
by: Xu, Shuolin, et al.
Published: (2025)
Hierarchical Multi-modal Transformer for Cross-modal Long Document Classification
by: Liu, Tengfei, et al.
Published: (2024)
by: Liu, Tengfei, et al.
Published: (2024)
CrossWeaver: Cross-modal Weaving for Arbitrary-Modality Semantic Segmentation
by: Zhang, Zelin, et al.
Published: (2026)
by: Zhang, Zelin, et al.
Published: (2026)
Hierarchical Cross-modal Prompt Learning for Vision-Language Models
by: Zheng, Hao, et al.
Published: (2025)
by: Zheng, Hao, et al.
Published: (2025)
TAR: Text Semantic Assisted Cross-modal Image Registration Framework for Optical and SAR Images
by: Cai, Zhuoyu, et al.
Published: (2026)
by: Cai, Zhuoyu, et al.
Published: (2026)
SWARD: Stochastic Window-Attention-Based Relational Distillation for Cross-Architectural Semantic Segmentation
by: Makineni, Aditya, et al.
Published: (2026)
by: Makineni, Aditya, et al.
Published: (2026)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
by: Zhang, Ming, et al.
Published: (2024)
by: Zhang, Ming, et al.
Published: (2024)
Efficient Inter-Task Attention for Multitask Transformer Models
by: Bohn, Christian, et al.
Published: (2025)
by: Bohn, Christian, et al.
Published: (2025)
Similar Items
-
ImpliHateVid: A Benchmark Dataset and Two-stage Contrastive Learning Framework for Implicit Hate Speech Detection in Videos
by: Rehman, Mohammad Zia Ur, et al.
Published: (2025) -
KisanQRS: A Deep Learning-based Automated Query-Response System for Agricultural Decision-Making
by: Rehman, Mohammad Zia Ur, et al.
Published: (2024) -
CoTCoNet: An Optimized Coupled Transformer-Convolutional Network with an Adaptive Graph Reconstruction for Leukemia Detection
by: Raghaw, Chandravardhan Singh, et al.
Published: (2024) -
A Context-aware Attention and Graph Neural Network-based Multimodal Framework for Misogyny Detection
by: Rehman, Mohammad Zia Ur, et al.
Published: (2025) -
T-MPEDNet: Unveiling the Synergy of Transformer-aware Multiscale Progressive Encoder-Decoder Network with Feature Recalibration for Tumor and Liver Segmentation
by: Raghaw, Chandravardhan Singh, et al.
Published: (2025)