Representation Learning with Semantic-aware Instance and Sparse Token Alignments
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bui, Phuoc-Nguyen, Nguyen, Toan Duc, Bum, Junghyun, Le, Duc-Tai, Choo, Hyunseung |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-scale Feature Enhancement in Multi-task Learning for Medical Image Analysis
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2024)
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2024)
Frequency Adapter with SAM for Generalized Medical Image Segmentation
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2026)
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2026)
Clinical Graph-Mediated Distillation for Unpaired MRI-to-CFI Hypertension Prediction
von: Imans, Dillan, et al.
Veröffentlicht: (2026)
von: Imans, Dillan, et al.
Veröffentlicht: (2026)
Unsupervised Domain Adaptation with SAM-RefiSeR for Enhanced Brain Tumor Segmentation
von: Imans, Dillan, et al.
Veröffentlicht: (2026)
von: Imans, Dillan, et al.
Veröffentlicht: (2026)
Cross Feature Fusion of Fundus Image and Generated Lesion Map for Referable Diabetic Retinopathy Classification
von: Mok, Dahyun, et al.
Veröffentlicht: (2024)
von: Mok, Dahyun, et al.
Veröffentlicht: (2024)
Adaptive Cache Enhancement for Test-Time Adaptation of Vision-Language Models
von: Nguyen, Khanh-Binh, et al.
Veröffentlicht: (2025)
von: Nguyen, Khanh-Binh, et al.
Veröffentlicht: (2025)
Accelerating Conditional Prompt Learning via Masked Image Modeling for Vision-Language Models
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2025)
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2025)
Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2025)
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2025)
Response-Aware Multimodal Learning for Post-Treatment Visual Acuity Forecasting
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2026)
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2026)
h-Edit: Effective and Flexible Diffusion-Based Editing via Doob's h-Transform
von: Nguyen, Toan, et al.
Veröffentlicht: (2025)
von: Nguyen, Toan, et al.
Veröffentlicht: (2025)
Bidirectional Diffusion Bridge Models
von: Kieu, Duc, et al.
Veröffentlicht: (2025)
von: Kieu, Duc, et al.
Veröffentlicht: (2025)
Semi-supervised 3D Semantic Scene Completion with 2D Vision Foundation Model Guidance
von: Pham, Duc-Hai, et al.
Veröffentlicht: (2024)
von: Pham, Duc-Hai, et al.
Veröffentlicht: (2024)
Collaborative Perceiver: Elevating Vision-based 3D Object Detection via Local Density-Aware Spatial Occupancy
von: Yuan, Jicheng, et al.
Veröffentlicht: (2025)
von: Yuan, Jicheng, et al.
Veröffentlicht: (2025)
Domain Generalization through Spatial Relation Induction over Visual Primitives
von: Nguyen, Dat, et al.
Veröffentlicht: (2026)
von: Nguyen, Dat, et al.
Veröffentlicht: (2026)
Enhanced Generative Data Augmentation for Semantic Segmentation via Stronger Guidance
von: Che, Quang-Huy, et al.
Veröffentlicht: (2024)
von: Che, Quang-Huy, et al.
Veröffentlicht: (2024)
YOWOv3: An Efficient and Generalized Framework for Human Action Detection and Recognition
von: Dang, Duc Manh Nguyen, et al.
Veröffentlicht: (2024)
von: Dang, Duc Manh Nguyen, et al.
Veröffentlicht: (2024)
MoBind: Motion Binding for Fine-Grained IMU-Video Pose Alignment
von: Nguyen, Duc Duy, et al.
Veröffentlicht: (2026)
von: Nguyen, Duc Duy, et al.
Veröffentlicht: (2026)
Improved Training Technique for Shortcut Models
von: Nguyen, Anh, et al.
Veröffentlicht: (2025)
von: Nguyen, Anh, et al.
Veröffentlicht: (2025)
VisionKG: Unleashing the Power of Visual Datasets via Knowledge Graph
von: Yuan, Jicheng, et al.
Veröffentlicht: (2023)
von: Yuan, Jicheng, et al.
Veröffentlicht: (2023)
Exploring the Practicality of Federated Learning: A Survey Towards the Communication Perspective
von: Le, Khiem, et al.
Veröffentlicht: (2024)
von: Le, Khiem, et al.
Veröffentlicht: (2024)
A comparison of extended object tracking with multi-modal sensors in indoor environment
von: Shuai, Jiangtao, et al.
Veröffentlicht: (2024)
von: Shuai, Jiangtao, et al.
Veröffentlicht: (2024)
A model-agnostic active learning approach for animal detection from camera traps
von: Nguyen, Thi Thu Thuy, et al.
Veröffentlicht: (2025)
von: Nguyen, Thi Thu Thuy, et al.
Veröffentlicht: (2025)
MatSSL: Robust Self-Supervised Representation Learning for Metallographic Image Segmentation
von: Nguyen, Hoang Hai Nam, et al.
Veröffentlicht: (2025)
von: Nguyen, Hoang Hai Nam, et al.
Veröffentlicht: (2025)
PADM: A Physics-aware Diffusion Model for Attenuation Correction
von: Pham, Trung Kien, et al.
Veröffentlicht: (2025)
von: Pham, Trung Kien, et al.
Veröffentlicht: (2025)
The Art of Camouflage: Few-Shot Learning for Animal Detection and Segmentation
von: Nguyen, Thanh-Danh, et al.
Veröffentlicht: (2023)
von: Nguyen, Thanh-Danh, et al.
Veröffentlicht: (2023)
Multi-Surrogate-Teacher Assistance for Representation Alignment in Fingerprint-based Indoor Localization
von: Nguyen, Son Minh, et al.
Veröffentlicht: (2024)
von: Nguyen, Son Minh, et al.
Veröffentlicht: (2024)
BALM: A Model-Agnostic Framework for Balanced Multimodal Learning under Imbalanced Missing Rates
von: Nguyen, Phuong-Anh, et al.
Veröffentlicht: (2026)
von: Nguyen, Phuong-Anh, et al.
Veröffentlicht: (2026)
Color Alignment in Diffusion
von: Shum, Ka Chun, et al.
Veröffentlicht: (2025)
von: Shum, Ka Chun, et al.
Veröffentlicht: (2025)
Instance-dependent Noisy-label Learning with Graphical Model Based Noise-rate Estimation
von: Garg, Arpit, et al.
Veröffentlicht: (2023)
von: Garg, Arpit, et al.
Veröffentlicht: (2023)
Dual Strategies for Test-Time Adaptation
von: Phuong, Nam Nguyen, et al.
Veröffentlicht: (2026)
von: Phuong, Nam Nguyen, et al.
Veröffentlicht: (2026)
Multi-Perspective Data Augmentation for Few-shot Object Detection
von: Vu, Anh-Khoa Nguyen, et al.
Veröffentlicht: (2025)
von: Vu, Anh-Khoa Nguyen, et al.
Veröffentlicht: (2025)
Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance
von: Nguyen, Phuc D. A., et al.
Veröffentlicht: (2023)
von: Nguyen, Phuc D. A., et al.
Veröffentlicht: (2023)
LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models
von: Debnath, Soumyaratna, et al.
Veröffentlicht: (2026)
von: Debnath, Soumyaratna, et al.
Veröffentlicht: (2026)
Detection Fire in Camera RGB-NIR
von: Khai, Nguyen Truong, et al.
Veröffentlicht: (2025)
von: Khai, Nguyen Truong, et al.
Veröffentlicht: (2025)
Anatomical Attention Alignment representation for Radiology Report Generation
von: Nguyen, Quang Vinh, et al.
Veröffentlicht: (2025)
von: Nguyen, Quang Vinh, et al.
Veröffentlicht: (2025)
PDIWS: Thermal Imaging Dataset for Person Detection in Intrusion Warning Systems
von: Thuan, Nguyen Duc, et al.
Veröffentlicht: (2023)
von: Thuan, Nguyen Duc, et al.
Veröffentlicht: (2023)
Enhancing person re-identification via Uncertainty Feature Fusion Method and Auto-weighted Measure Combination
von: Che, Quang-Huy, et al.
Veröffentlicht: (2024)
von: Che, Quang-Huy, et al.
Veröffentlicht: (2024)
FurniMAS: Language-Guided Furniture Decoration using Multi-Agent System
von: Nguyen, Toan, et al.
Veröffentlicht: (2025)
von: Nguyen, Toan, et al.
Veröffentlicht: (2025)
Deep-Wide Learning Assistance for Insect Pest Classification
von: Nguyen, Toan, et al.
Veröffentlicht: (2024)
von: Nguyen, Toan, et al.
Veröffentlicht: (2024)
TwinLiteNet+: An Enhanced Multi-Task Segmentation Model for Autonomous Driving
von: Che, Quang-Huy, et al.
Veröffentlicht: (2024)
von: Che, Quang-Huy, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Multi-scale Feature Enhancement in Multi-task Learning for Medical Image Analysis
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2024) -
Frequency Adapter with SAM for Generalized Medical Image Segmentation
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2026) -
Clinical Graph-Mediated Distillation for Unpaired MRI-to-CFI Hypertension Prediction
von: Imans, Dillan, et al.
Veröffentlicht: (2026) -
Unsupervised Domain Adaptation with SAM-RefiSeR for Enhanced Brain Tumor Segmentation
von: Imans, Dillan, et al.
Veröffentlicht: (2026) -
Cross Feature Fusion of Fundus Image and Generated Lesion Map for Referable Diabetic Retinopathy Classification
von: Mok, Dahyun, et al.
Veröffentlicht: (2024)