Improving Generalization in Visual Reasoning via Self-Ensemble
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nguyen, Tien-Huy, Tran, Quang-Khai, Quang-Hoang, Anh-Tuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TriLiteNet: Lightweight Model for Multi-Task Visual Perception
von: Che, Quang-Huy, et al.
Veröffentlicht: (2025)
von: Che, Quang-Huy, et al.
Veröffentlicht: (2025)
Cycle Training with Semi-Supervised Domain Adaptation: Bridging Accuracy and Efficiency for Real-Time Mobile Scene Detection
von: Phan-Nguyen, Huu-Phong, et al.
Veröffentlicht: (2025)
von: Phan-Nguyen, Huu-Phong, et al.
Veröffentlicht: (2025)
RelWitness: Open-Vocabulary 3D Scene Graph Generation with Visual-Geometric Relation Witnesses
von: Nguyen, Minh Anh, et al.
Veröffentlicht: (2026)
von: Nguyen, Minh Anh, et al.
Veröffentlicht: (2026)
CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
von: Nguyen, Quang-Binh, et al.
Veröffentlicht: (2025)
von: Nguyen, Quang-Binh, et al.
Veröffentlicht: (2025)
Label-Efficient Cross-Modality Generalization for Liver Segmentation in Multi-Phase MRI
von: Bui-Tran, Quang-Khai, et al.
Veröffentlicht: (2025)
von: Bui-Tran, Quang-Khai, et al.
Veröffentlicht: (2025)
Hybrid, Unified and Iterative: A Novel Framework for Text-based Person Anomaly Retrieval
von: Nguyen, Tien-Huy, et al.
Veröffentlicht: (2025)
von: Nguyen, Tien-Huy, et al.
Veröffentlicht: (2025)
Enhanced Generative Data Augmentation for Semantic Segmentation via Stronger Guidance
von: Che, Quang-Huy, et al.
Veröffentlicht: (2024)
von: Che, Quang-Huy, et al.
Veröffentlicht: (2024)
Stable Messenger: Steganography for Message-Concealed Image Generation
von: Nguyen, Quang, et al.
Veröffentlicht: (2023)
von: Nguyen, Quang, et al.
Veröffentlicht: (2023)
Enhancing person re-identification via Uncertainty Feature Fusion Method and Auto-weighted Measure Combination
von: Che, Quang-Huy, et al.
Veröffentlicht: (2024)
von: Che, Quang-Huy, et al.
Veröffentlicht: (2024)
SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step Diffusion
von: Nguyen, Trong-Tung, et al.
Veröffentlicht: (2024)
von: Nguyen, Trong-Tung, et al.
Veröffentlicht: (2024)
Towards Efficient and Robust Moment Retrieval System: A Unified Framework for Multi-Granularity Models and Temporal Reranking
von: Tran, Huu-Loc, et al.
Veröffentlicht: (2025)
von: Tran, Huu-Loc, et al.
Veröffentlicht: (2025)
When Less Is More: A Sparse Facial Motion Structure For Listening Motion Learning
von: Nguyen, Tri Tung Nguyen, et al.
Veröffentlicht: (2025)
von: Nguyen, Tri Tung Nguyen, et al.
Veröffentlicht: (2025)
ViOCRVQA: Novel Benchmark Dataset and Vision Reader for Visual Question Answering by Understanding Vietnamese Text in Images
von: Pham, Huy Quang, et al.
Veröffentlicht: (2024)
von: Pham, Huy Quang, et al.
Veröffentlicht: (2024)
PainDiffusion: Learning to Express Pain
von: Dam, Quang Tien, et al.
Veröffentlicht: (2024)
von: Dam, Quang Tien, et al.
Veröffentlicht: (2024)
ScriptHOI: Learning Scripted State Transitions for Open-Vocabulary Human-Object Interaction Detection
von: Nguyen, Minh Anh, et al.
Veröffentlicht: (2026)
von: Nguyen, Minh Anh, et al.
Veröffentlicht: (2026)
A Lightweight Moment Retrieval System with Global Re-Ranking and Robust Adaptive Bidirectional Temporal Search
von: Nguyen-Nhu, Tinh-Anh, et al.
Veröffentlicht: (2025)
von: Nguyen-Nhu, Tinh-Anh, et al.
Veröffentlicht: (2025)
Aligning What You Separate: Denoised Patch Mixing for Source-Free Domain Adaptation in Medical Image Segmentation
von: Bui-Tran, Quang-Khai, et al.
Veröffentlicht: (2025)
von: Bui-Tran, Quang-Khai, et al.
Veröffentlicht: (2025)
Domain-invariant Mixed-domain Semi-supervised Medical Image Segmentation with Clustered Maximum Mean Discrepancy Alignment
von: Lam, Ba-Thinh, et al.
Veröffentlicht: (2026)
von: Lam, Ba-Thinh, et al.
Veröffentlicht: (2026)
SDPA++: A General Framework for Self-Supervised Denoising with Patch Aggregation
von: Nguyen, Huy Minh Nhat, et al.
Veröffentlicht: (2025)
von: Nguyen, Huy Minh Nhat, et al.
Veröffentlicht: (2025)
Stratified Domain Adaptation: A Progressive Self-Training Approach for Scene Text Recognition
von: Le, Kha Nhat, et al.
Veröffentlicht: (2024)
von: Le, Kha Nhat, et al.
Veröffentlicht: (2024)
Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments
von: Nguyen, Van Quang
Veröffentlicht: (2026)
von: Nguyen, Van Quang
Veröffentlicht: (2026)
TF-SASM: Training-free Spatial-aware Sparse Memory for Multi-object Tracking
von: Nguyen-Quang, Thuc, et al.
Veröffentlicht: (2024)
von: Nguyen-Quang, Thuc, et al.
Veröffentlicht: (2024)
MambaU-Lite: A Lightweight Model based on Mamba and Integrated Channel-Spatial Attention for Skin Lesion Segmentation
von: Nguyen, Thi-Nhu-Quynh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thi-Nhu-Quynh, et al.
Veröffentlicht: (2024)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
von: Anh, Duy Le Dinh, et al.
Veröffentlicht: (2024)
von: Anh, Duy Le Dinh, et al.
Veröffentlicht: (2024)
UIT-DarkCow team at ImageCLEFmedical Caption 2024: Diagnostic Captioning for Radiology Images Efficiency with Transformer Models
von: Van Nguyen, Quan, et al.
Veröffentlicht: (2024)
von: Van Nguyen, Quan, et al.
Veröffentlicht: (2024)
TwinLiteNet+: An Enhanced Multi-Task Segmentation Model for Autonomous Driving
von: Che, Quang-Huy, et al.
Veröffentlicht: (2024)
von: Che, Quang-Huy, et al.
Veröffentlicht: (2024)
TaleForge: Interactive Multimodal System for Personalized Story Creation
von: Nguyen, Minh-Loi, et al.
Veröffentlicht: (2025)
von: Nguyen, Minh-Loi, et al.
Veröffentlicht: (2025)
Novel sparse PCA method via Runge Kutta numerical method(s) for face recognition
von: Tran, Loc Hoang, et al.
Veröffentlicht: (2025)
von: Tran, Loc Hoang, et al.
Veröffentlicht: (2025)
Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts
von: Le, Minh, et al.
Veröffentlicht: (2025)
von: Le, Minh, et al.
Veröffentlicht: (2025)
A-SCoRe: Attention-based Scene Coordinate Regression for wide-ranging scenarios
von: Bui, Huy-Hoang, et al.
Veröffentlicht: (2025)
von: Bui, Huy-Hoang, et al.
Veröffentlicht: (2025)
Semi-Supervised Semantic Segmentation using Redesigned Self-Training for White Blood Cells
von: Luu, Vinh Quoc, et al.
Veröffentlicht: (2024)
von: Luu, Vinh Quoc, et al.
Veröffentlicht: (2024)
SwiftBrush: One-Step Text-to-Image Diffusion Model with Variational Score Distillation
von: Nguyen, Thuan Hoang, et al.
Veröffentlicht: (2023)
von: Nguyen, Thuan Hoang, et al.
Veröffentlicht: (2023)
Digital FAST: An AI-Driven Multimodal Framework for Rapid and Early Stroke Screening
von: Hoang, Ngoc-Khai, et al.
Veröffentlicht: (2026)
von: Hoang, Ngoc-Khai, et al.
Veröffentlicht: (2026)
FurniMAS: Language-Guided Furniture Decoration using Multi-Agent System
von: Nguyen, Toan, et al.
Veröffentlicht: (2025)
von: Nguyen, Toan, et al.
Veröffentlicht: (2025)
Scalable Group Choreography via Variational Phase Manifold Learning
von: Le, Nhat, et al.
Veröffentlicht: (2024)
von: Le, Nhat, et al.
Veröffentlicht: (2024)
Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models
von: Nguyen, Minh Khoi, et al.
Veröffentlicht: (2026)
von: Nguyen, Minh Khoi, et al.
Veröffentlicht: (2026)
KTVIC: A Vietnamese Image Captioning Dataset on the Life Domain
von: Pham, Anh-Cuong, et al.
Veröffentlicht: (2024)
von: Pham, Anh-Cuong, et al.
Veröffentlicht: (2024)
KAN-Based Fusion of Dual-Domain for Audio-Driven Facial Landmarks Generation
von: Vo-Thanh, Hoang-Son, et al.
Veröffentlicht: (2024)
von: Vo-Thanh, Hoang-Son, et al.
Veröffentlicht: (2024)
Lightweight Temporal Transformer Decomposition for Federated Autonomous Driving
von: Do, Tuong, et al.
Veröffentlicht: (2025)
von: Do, Tuong, et al.
Veröffentlicht: (2025)
ShowFlow: From Robust Single Concept to Condition-Free Multi-Concept Generation
von: Hoang, Trong-Vu, et al.
Veröffentlicht: (2025)
von: Hoang, Trong-Vu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TriLiteNet: Lightweight Model for Multi-Task Visual Perception
von: Che, Quang-Huy, et al.
Veröffentlicht: (2025) -
Cycle Training with Semi-Supervised Domain Adaptation: Bridging Accuracy and Efficiency for Real-Time Mobile Scene Detection
von: Phan-Nguyen, Huu-Phong, et al.
Veröffentlicht: (2025) -
RelWitness: Open-Vocabulary 3D Scene Graph Generation with Visual-Geometric Relation Witnesses
von: Nguyen, Minh Anh, et al.
Veröffentlicht: (2026) -
CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
von: Nguyen, Quang-Binh, et al.
Veröffentlicht: (2025) -
Label-Efficient Cross-Modality Generalization for Liver Segmentation in Multi-Phase MRI
von: Bui-Tran, Quang-Khai, et al.
Veröffentlicht: (2025)