Hierarchical Multi-modal Transformer for Cross-modal Long Document Classification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Tengfei, Hu, Yongli, Gao, Junbin, Sun, Yanfeng, Yin, Baocai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Awesome Multi-modal Object Tracking
von: Zhang, Chunhui, et al.
Veröffentlicht: (2024)
von: Zhang, Chunhui, et al.
Veröffentlicht: (2024)
HAMLET-FFD: Hierarchical Adaptive Multi-modal Learning Embeddings Transformation for Face Forgery Detection
von: Cui, Jialei, et al.
Veröffentlicht: (2025)
von: Cui, Jialei, et al.
Veröffentlicht: (2025)
Tighnari: Multi-modal Plant Species Prediction Based on Hierarchical Cross-Attention Using Graph-Based and Vision Backbone-Extracted Features
von: Liu, Haixu, et al.
Veröffentlicht: (2025)
von: Liu, Haixu, et al.
Veröffentlicht: (2025)
EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical Alignment
von: Xing, Yifei, et al.
Veröffentlicht: (2024)
von: Xing, Yifei, et al.
Veröffentlicht: (2024)
Cross-modality Guidance-aided Multi-modal Learning with Dual Attention for MRI Brain Tumor Grading
von: Xu, Dunyuan, et al.
Veröffentlicht: (2024)
von: Xu, Dunyuan, et al.
Veröffentlicht: (2024)
e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings
von: Chen, Haonan, et al.
Veröffentlicht: (2026)
von: Chen, Haonan, et al.
Veröffentlicht: (2026)
HC-LLM: Historical-Constrained Large Language Models for Radiology Report Generation
von: Liu, Tengfei, et al.
Veröffentlicht: (2024)
von: Liu, Tengfei, et al.
Veröffentlicht: (2024)
Multi-modal Deepfake Detection and Localization with FPN-Transformer
von: Zheng, Chende, et al.
Veröffentlicht: (2025)
von: Zheng, Chende, et al.
Veröffentlicht: (2025)
Multi-modality Anomaly Segmentation on the Road
von: Gao, Heng, et al.
Veröffentlicht: (2025)
von: Gao, Heng, et al.
Veröffentlicht: (2025)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
von: Wang, Xin, et al.
Veröffentlicht: (2024)
von: Wang, Xin, et al.
Veröffentlicht: (2024)
TACFN: Transformer-based Adaptive Cross-modal Fusion Network for Multimodal Emotion Recognition
von: Liu, Feng, et al.
Veröffentlicht: (2025)
von: Liu, Feng, et al.
Veröffentlicht: (2025)
Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation
von: Yue, Junrong, et al.
Veröffentlicht: (2025)
von: Yue, Junrong, et al.
Veröffentlicht: (2025)
Fusion-Mamba for Cross-modality Object Detection
von: Dong, Wenhao, et al.
Veröffentlicht: (2024)
von: Dong, Wenhao, et al.
Veröffentlicht: (2024)
Simultaneous Long-tailed Recognition and Multi-modal Fusion for Highly Imbalanced Multi-modal Data
von: Yoon, Heegeon, et al.
Veröffentlicht: (2026)
von: Yoon, Heegeon, et al.
Veröffentlicht: (2026)
Bridging Modality Gap for Visual Grounding with Effecitve Cross-modal Distillation
von: Wang, Jiaxi, et al.
Veröffentlicht: (2023)
von: Wang, Jiaxi, et al.
Veröffentlicht: (2023)
Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
von: Yu, Ting, et al.
Veröffentlicht: (2024)
von: Yu, Ting, et al.
Veröffentlicht: (2024)
A Multimodal-Multitask Framework with Cross-modal Relation and Hierarchical Interactive Attention for Semantic Comprehension
von: Rehman, Mohammad Zia Ur, et al.
Veröffentlicht: (2025)
von: Rehman, Mohammad Zia Ur, et al.
Veröffentlicht: (2025)
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance
von: Li, Zhang, et al.
Veröffentlicht: (2025)
von: Li, Zhang, et al.
Veröffentlicht: (2025)
Variational Adapter for Cross-modal Similarity Representation
von: Wei, WenZhang, et al.
Veröffentlicht: (2026)
von: Wei, WenZhang, et al.
Veröffentlicht: (2026)
Cross-Fundus Transformer for Multi-modal Diabetic Retinopathy Grading with Cataract
von: Xiao, Fan, et al.
Veröffentlicht: (2024)
von: Xiao, Fan, et al.
Veröffentlicht: (2024)
Cross-domain Few-shot Object Detection with Multi-modal Textual Enrichment
von: Shangguan, Zeyu, et al.
Veröffentlicht: (2025)
von: Shangguan, Zeyu, et al.
Veröffentlicht: (2025)
Multi-modal Spatio-Temporal Transformer for High-resolution Land Subsidence Prediction
von: Yao, Wendong, et al.
Veröffentlicht: (2025)
von: Yao, Wendong, et al.
Veröffentlicht: (2025)
Spectral Discrepancy and Cross-modal Semantic Consistency Learning for Object Detection in Hyperspectral Image
von: He, Xiao, et al.
Veröffentlicht: (2025)
von: He, Xiao, et al.
Veröffentlicht: (2025)
Explainable, Multi-modal Wound Infection Classification from Images Augmented with Generated Captions
von: Busaranuvong, Palawat, et al.
Veröffentlicht: (2025)
von: Busaranuvong, Palawat, et al.
Veröffentlicht: (2025)
UniEmoX: Cross-modal Semantic-Guided Large-Scale Pretraining for Universal Scene Emotion Perception
von: Chen, Chuang, et al.
Veröffentlicht: (2024)
von: Chen, Chuang, et al.
Veröffentlicht: (2024)
MCAD: Multi-teacher Cross-modal Alignment Distillation for efficient image-text retrieval
von: Lei, Youbo, et al.
Veröffentlicht: (2023)
von: Lei, Youbo, et al.
Veröffentlicht: (2023)
ViewSAM: Learning View-aware Cross-modal Semantics for Weakly Supervised Cross-view Referring Multi-Object Tracking
von: Ge, Jiawei, et al.
Veröffentlicht: (2026)
von: Ge, Jiawei, et al.
Veröffentlicht: (2026)
Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Benefit from Reference: Retrieval-Augmented Cross-modal Point Cloud Completion
von: Hou, Hongye, et al.
Veröffentlicht: (2025)
von: Hou, Hongye, et al.
Veröffentlicht: (2025)
ViT-Lens: Towards Omni-modal Representations
von: Lei, Weixian, et al.
Veröffentlicht: (2023)
von: Lei, Weixian, et al.
Veröffentlicht: (2023)
Diffexplainer: Towards Cross-modal Global Explanations with Diffusion Models
von: Pennisi, Matteo, et al.
Veröffentlicht: (2024)
von: Pennisi, Matteo, et al.
Veröffentlicht: (2024)
Energy-Latency Manipulation of Multi-modal Large Language Models via Verbose Samples
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
Integrating Text and Image Pre-training for Multi-modal Algorithmic Reasoning
von: Zhang, Zijian, et al.
Veröffentlicht: (2024)
von: Zhang, Zijian, et al.
Veröffentlicht: (2024)
AnomalyControl: Learning Cross-modal Semantic Features for Controllable Anomaly Synthesis
von: He, Shidan, et al.
Veröffentlicht: (2024)
von: He, Shidan, et al.
Veröffentlicht: (2024)
Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration
von: Mena, Francisco, et al.
Veröffentlicht: (2025)
von: Mena, Francisco, et al.
Veröffentlicht: (2025)
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
von: Yu, Shi, et al.
Veröffentlicht: (2024)
von: Yu, Shi, et al.
Veröffentlicht: (2024)
MGHFT: Multi-Granularity Hierarchical Fusion Transformer for Cross-Modal Sticker Emotion Recognition
von: Chen, Jian, et al.
Veröffentlicht: (2025)
von: Chen, Jian, et al.
Veröffentlicht: (2025)
DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
von: Zhang, Yudong, et al.
Veröffentlicht: (2024)
von: Zhang, Yudong, et al.
Veröffentlicht: (2024)
FreeEdit: Mask-free Reference-based Image Editing with Multi-modal Instruction
von: He, Runze, et al.
Veröffentlicht: (2024)
von: He, Runze, et al.
Veröffentlicht: (2024)
Multi-Modality Distillation via Learning the teacher's modality-level Gram Matrix
von: Liu, Peng
Veröffentlicht: (2021)
von: Liu, Peng
Veröffentlicht: (2021)
Ähnliche Einträge
-
Awesome Multi-modal Object Tracking
von: Zhang, Chunhui, et al.
Veröffentlicht: (2024) -
HAMLET-FFD: Hierarchical Adaptive Multi-modal Learning Embeddings Transformation for Face Forgery Detection
von: Cui, Jialei, et al.
Veröffentlicht: (2025) -
Tighnari: Multi-modal Plant Species Prediction Based on Hierarchical Cross-Attention Using Graph-Based and Vision Backbone-Extracted Features
von: Liu, Haixu, et al.
Veröffentlicht: (2025) -
EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical Alignment
von: Xing, Yifei, et al.
Veröffentlicht: (2024) -
Cross-modality Guidance-aided Multi-modal Learning with Dual Attention for MRI Brain Tumor Grading
von: Xu, Dunyuan, et al.
Veröffentlicht: (2024)