Long-tailed Medical Diagnosis with Relation-aware Representation Learning and Iterative Classifier Calibration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pan, Li, Zhang, Yupei, Yang, Qiushi, Li, Tan, Chen, Zhen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Attributes-aware Visual Emotion Representation Learning
von: Maharjan, Rahul Singh, et al.
Veröffentlicht: (2025)
von: Maharjan, Rahul Singh, et al.
Veröffentlicht: (2025)
ASR-enhanced Multimodal Representation Learning for Cross-Domain Product Retrieval
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2024)
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2024)
Learning Segment Similarity and Alignment in Large-Scale Content Based Video Retrieval
von: Jiang, Chen, et al.
Veröffentlicht: (2023)
von: Jiang, Chen, et al.
Veröffentlicht: (2023)
URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection
von: Wang, Zhenyu, et al.
Veröffentlicht: (2026)
von: Wang, Zhenyu, et al.
Veröffentlicht: (2026)
Focus on Focus: Focus-oriented Representation Learning and Multi-view Cross-modal Alignment for Glioma Grading
von: Pan, Li, et al.
Veröffentlicht: (2024)
von: Pan, Li, et al.
Veröffentlicht: (2024)
Knowledge-enhanced Multi-perspective Video Representation Learning for Scene Recognition
von: Yu, Xuzheng, et al.
Veröffentlicht: (2024)
von: Yu, Xuzheng, et al.
Veröffentlicht: (2024)
Enhancing Visible-Infrared Person Re-identification with Modality- and Instance-aware Visual Prompt Learning
von: Wu, Ruiqi, et al.
Veröffentlicht: (2024)
von: Wu, Ruiqi, et al.
Veröffentlicht: (2024)
Learning to Generate Conditional Tri-plane for 3D-aware Expression Controllable Portrait Animation
von: Ki, Taekyung, et al.
Veröffentlicht: (2024)
von: Ki, Taekyung, et al.
Veröffentlicht: (2024)
MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
von: Zhang, Yuang, et al.
Veröffentlicht: (2024)
von: Zhang, Yuang, et al.
Veröffentlicht: (2024)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
von: Wang, Yun, et al.
Veröffentlicht: (2025)
von: Wang, Yun, et al.
Veröffentlicht: (2025)
EvRepSL: Event-Stream Representation via Self-Supervised Learning for Event-Based Vision
von: Qu, Qiang, et al.
Veröffentlicht: (2024)
von: Qu, Qiang, et al.
Veröffentlicht: (2024)
Towards Efficient Low-rate Image Compression with Frequency-aware Diffusion Prior Refinement
von: Xia, Yichong, et al.
Veröffentlicht: (2026)
von: Xia, Yichong, et al.
Veröffentlicht: (2026)
LLM-EvRep: Learning an LLM-Compatible Event Representation Using a Self-Supervised Framework
von: Yu, Zongyou, et al.
Veröffentlicht: (2025)
von: Yu, Zongyou, et al.
Veröffentlicht: (2025)
Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
von: Li, Chengzhi, et al.
Veröffentlicht: (2025)
von: Li, Chengzhi, et al.
Veröffentlicht: (2025)
CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learning
von: Meng, Chunlei, et al.
Veröffentlicht: (2026)
von: Meng, Chunlei, et al.
Veröffentlicht: (2026)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2024)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
von: Yu, Jiashuo, et al.
Veröffentlicht: (2025)
von: Yu, Jiashuo, et al.
Veröffentlicht: (2025)
Distilling Generative-Discriminative Representations for Very Low-Resolution Face Recognition
von: Zhang, Junzheng, et al.
Veröffentlicht: (2024)
von: Zhang, Junzheng, et al.
Veröffentlicht: (2024)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
von: Park, Kyu Ri, et al.
Veröffentlicht: (2024)
von: Park, Kyu Ri, et al.
Veröffentlicht: (2024)
Harmonizing Attention: Training-free Texture-aware Geometry Transfer
von: Ikuta, Eito, et al.
Veröffentlicht: (2024)
von: Ikuta, Eito, et al.
Veröffentlicht: (2024)
Scaling Up Single Image Dehazing Algorithm by Cross-Data Vision Alignment for Richer Representation Learning and Beyond
von: Shi, Yukai, et al.
Veröffentlicht: (2024)
von: Shi, Yukai, et al.
Veröffentlicht: (2024)
Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
von: Li, Jie, et al.
Veröffentlicht: (2025)
von: Li, Jie, et al.
Veröffentlicht: (2025)
VidCtx: Context-aware Video Question Answering with Image Models
von: Goulas, Andreas, et al.
Veröffentlicht: (2024)
von: Goulas, Andreas, et al.
Veröffentlicht: (2024)
MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention
von: Wang, Tianyi, et al.
Veröffentlicht: (2025)
von: Wang, Tianyi, et al.
Veröffentlicht: (2025)
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
von: Yao, Lei, et al.
Veröffentlicht: (2025)
von: Yao, Lei, et al.
Veröffentlicht: (2025)
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
von: Yan, Xin, et al.
Veröffentlicht: (2024)
von: Yan, Xin, et al.
Veröffentlicht: (2024)
Diagnosing and Re-learning for Balanced Multimodal Learning
von: Wei, Yake, et al.
Veröffentlicht: (2024)
von: Wei, Yake, et al.
Veröffentlicht: (2024)
DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing
von: Li, Ke, et al.
Veröffentlicht: (2026)
von: Li, Ke, et al.
Veröffentlicht: (2026)
Robust Latent Representation Tuning for Image-text Classification
von: Sun, Hao, et al.
Veröffentlicht: (2024)
von: Sun, Hao, et al.
Veröffentlicht: (2024)
Multimodal Chaptering for Long-Form TV Newscast Video
von: Guetari, Khalil, et al.
Veröffentlicht: (2024)
von: Guetari, Khalil, et al.
Veröffentlicht: (2024)
Mitigating Easy Option Bias in Multiple-Choice Question Answering
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment
von: Chen, Zhen, et al.
Veröffentlicht: (2025)
von: Chen, Zhen, et al.
Veröffentlicht: (2025)
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers
von: Yang, Chenglin, et al.
Veröffentlicht: (2023)
von: Yang, Chenglin, et al.
Veröffentlicht: (2023)
Cross-domain Few-shot In-context Learning for Enhancing Traffic Sign Recognition
von: Gan, Yaozong, et al.
Veröffentlicht: (2024)
von: Gan, Yaozong, et al.
Veröffentlicht: (2024)
Relational Retrieval: Leveraging Known-Novel Interactions for Generalized Category Discovery
von: Xu, Yulin, et al.
Veröffentlicht: (2026)
von: Xu, Yulin, et al.
Veröffentlicht: (2026)
DualComp: End-to-End Learning of a Unified Dual-Modality Lossless Compressor
von: Zhao, Yan, et al.
Veröffentlicht: (2025)
von: Zhao, Yan, et al.
Veröffentlicht: (2025)
Exploring the latent space of diffusion models directly through singular value decomposition
von: Wang, Li, et al.
Veröffentlicht: (2025)
von: Wang, Li, et al.
Veröffentlicht: (2025)
EvAnimate: Event-conditioned Image-to-Video Generation for Human Animation
von: Qu, Qiang, et al.
Veröffentlicht: (2025)
von: Qu, Qiang, et al.
Veröffentlicht: (2025)
Official-NV: An LLM-Generated News Video Dataset for Multimodal Fake News Detection
von: Wang, Yihao, et al.
Veröffentlicht: (2024)
von: Wang, Yihao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Attributes-aware Visual Emotion Representation Learning
von: Maharjan, Rahul Singh, et al.
Veröffentlicht: (2025) -
ASR-enhanced Multimodal Representation Learning for Cross-Domain Product Retrieval
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2024) -
Learning Segment Similarity and Alignment in Large-Scale Content Based Video Retrieval
von: Jiang, Chen, et al.
Veröffentlicht: (2023) -
URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection
von: Wang, Zhenyu, et al.
Veröffentlicht: (2026) -
Focus on Focus: Focus-oriented Representation Learning and Multi-view Cross-modal Alignment for Glioma Grading
von: Pan, Li, et al.
Veröffentlicht: (2024)