TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Seongah, Tran, Dinh Phu, Hwang, Hyeontaek, Wazir, Saad, Minh, Duc Do, Kim, Daeyoung |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SAT: Selective Aggregation Transformer for Image Super-Resolution
by: Tran, Dinh Phu, et al.
Published: (2026)
by: Tran, Dinh Phu, et al.
Published: (2026)
Knowing When to Answer: Adaptive Confidence Refinement for Reliable Audio-Visual Question Answering
by: Tran, Dinh Phu, et al.
Published: (2026)
by: Tran, Dinh Phu, et al.
Published: (2026)
Reference-Based Post-OCR Processing with LLM for Precise Diacritic Text in Historical Document Recognition
by: Do, Thao, et al.
Published: (2024)
by: Do, Thao, et al.
Published: (2024)
Rethinking the Nested U-Net Approach: Enhancing Biomarker Segmentation with Attention Mechanisms and Multiscale Feature Fusion
by: Wazir, Saad, et al.
Published: (2025)
by: Wazir, Saad, et al.
Published: (2025)
Rethinking Decoder Design: Improving Biomarker Segmentation Using Depth-to-Space Restoration and Residual Linear Attention
by: Wazir, Saad, et al.
Published: (2025)
by: Wazir, Saad, et al.
Published: (2025)
VSRM: A Robust Mamba-Based Framework for Video Super-Resolution
by: Tran, Dinh Phu, et al.
Published: (2025)
by: Tran, Dinh Phu, et al.
Published: (2025)
Channel-Partitioned Windowed Attention And Frequency Learning for Single Image Super-Resolution
by: Tran, Dinh Phu, et al.
Published: (2024)
by: Tran, Dinh Phu, et al.
Published: (2024)
Model-Dowser: Data-Free Importance Probing to Mitigate Catastrophic Forgetting in Multimodal Large Language Models
by: Hwang, Hyeontaek, et al.
Published: (2026)
by: Hwang, Hyeontaek, et al.
Published: (2026)
Time-Efficient and Identity-Consistent Virtual Try-On Using A Variant of Altered Diffusion Models
by: Dam, Phuong, et al.
Published: (2024)
by: Dam, Phuong, et al.
Published: (2024)
LungCRCT: Causal Representation based Lung CT Processing for Lung Cancer Treatment
by: Kim, Daeyoung
Published: (2026)
by: Kim, Daeyoung
Published: (2026)
LightHCG: a Lightweight yet powerful HSIC Disentanglement based Causal Glaucoma Detection Model framework
by: Kim, Daeyoung
Published: (2025)
by: Kim, Daeyoung
Published: (2025)
LooComp: Leverage Leave-One-Out Strategy to Encoder-only Transformer for Efficient Query-aware Context Compression
by: Do, Thao, et al.
Published: (2026)
by: Do, Thao, et al.
Published: (2026)
GCVAMD: A Modified CausalVAE Model for Causal Age-related Macular Degeneration Risk Factor Detection and Prediction
by: Kim, Daeyoung
Published: (2025)
by: Kim, Daeyoung
Published: (2025)
FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation
by: Le, Minh Khoa, et al.
Published: (2026)
by: Le, Minh Khoa, et al.
Published: (2026)
Trans2Unet: Neural fusion for Nuclei Semantic Segmentation
by: Tran, Dinh-Phu, et al.
Published: (2024)
by: Tran, Dinh-Phu, et al.
Published: (2024)
Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-shot Skeleton-based Action Recognition
by: Do, Jeonghyeok, et al.
Published: (2024)
by: Do, Jeonghyeok, et al.
Published: (2024)
Interactive Interface For Semantic Segmentation Dataset Synthesis
by: Tran, Ngoc-Do, et al.
Published: (2025)
by: Tran, Ngoc-Do, et al.
Published: (2025)
UGGNet: Bridging U-Net and VGG for Advanced Breast Cancer Diagnosis
by: Minh, Tran Cao, et al.
Published: (2024)
by: Minh, Tran Cao, et al.
Published: (2024)
Online 3D Multi-Camera Perception through Robust 2D Tracking and Depth-based Late Aggregation
by: Le, Vu-Minh, et al.
Published: (2025)
by: Le, Vu-Minh, et al.
Published: (2025)
Multispectral Detection Transformer with Infrared-Centric Feature Fusion
by: Hwang, Seongmin, et al.
Published: (2025)
by: Hwang, Seongmin, et al.
Published: (2025)
DG-DETR: Toward Domain Generalized Detection Transformer
by: Hwang, Seongmin, et al.
Published: (2025)
by: Hwang, Seongmin, et al.
Published: (2025)
TwinMixing: A Shuffle-Aware Feature Interaction Model for Multi-Task Segmentation
by: Do, Minh-Khoi, et al.
Published: (2026)
by: Do, Minh-Khoi, et al.
Published: (2026)
Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery
by: Tran, Chi-Nguyen, et al.
Published: (2026)
by: Tran, Chi-Nguyen, et al.
Published: (2026)
OurDB: Ouroboric Domain Bridging for Multi-Target Domain Adaptive Semantic Segmentation
by: Woo, Seungbeom, et al.
Published: (2024)
by: Woo, Seungbeom, et al.
Published: (2024)
Persistent Test-time Adaptation in Recurring Testing Scenarios
by: Hoang, Trung-Hieu, et al.
Published: (2023)
by: Hoang, Trung-Hieu, et al.
Published: (2023)
WaveDH: Wavelet Sub-bands Guided ConvNet for Efficient Image Dehazing
by: Hwang, Seongmin, et al.
Published: (2024)
by: Hwang, Seongmin, et al.
Published: (2024)
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization
by: Nguyen, Son, et al.
Published: (2025)
by: Nguyen, Son, et al.
Published: (2025)
Visual Instruction-Finetuned Language Model for Versatile Brain MR Image Tasks
by: Kim, Jonghun, et al.
Published: (2026)
by: Kim, Jonghun, et al.
Published: (2026)
ECLIPSE: Efficient Continual Learning in Panoptic Segmentation with Visual Prompt Tuning
by: Kim, Beomyoung, et al.
Published: (2024)
by: Kim, Beomyoung, et al.
Published: (2024)
Enhancing Dataset Distillation via Non-Critical Region Refinement
by: Tran, Minh-Tuan, et al.
Published: (2025)
by: Tran, Minh-Tuan, et al.
Published: (2025)
Rethinking Saliency-Guided Weakly-Supervised Semantic Segmentation
by: Kim, Beomyoung, et al.
Published: (2024)
by: Kim, Beomyoung, et al.
Published: (2024)
Semantic-Fast-SAM: Efficient Semantic Segmenter
by: Kim, Byunghyun
Published: (2026)
by: Kim, Byunghyun
Published: (2026)
Test-Time Instance-Specific Parameter Composition: A New Paradigm for Adaptive Generative Modeling
by: Tran, Minh-Tuan, et al.
Published: (2026)
by: Tran, Minh-Tuan, et al.
Published: (2026)
Parameter Efficient Fine Tuning for Multi-scanner PET to PET Reconstruction
by: Kim, Yumin, et al.
Published: (2024)
by: Kim, Yumin, et al.
Published: (2024)
D3T: Distinctive Dual-Domain Teacher Zigzagging Across RGB-Thermal Gap for Domain-Adaptive Object Detection
by: Do, Dinh Phat, et al.
Published: (2024)
by: Do, Dinh Phat, et al.
Published: (2024)
KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts
by: Hwang, Taebaek, et al.
Published: (2025)
by: Hwang, Taebaek, et al.
Published: (2025)
VARCO-VISION: Expanding Frontiers in Korean Vision-Language Models
by: Ju, Jeongho, et al.
Published: (2024)
by: Ju, Jeongho, et al.
Published: (2024)
Audio-3DVG: Unified Audio -- Point Cloud Fusion for 3D Visual Grounding
by: Cao-Dinh, Duc, et al.
Published: (2025)
by: Cao-Dinh, Duc, et al.
Published: (2025)
NAYER: Noisy Layer Data Generation for Efficient and Effective Data-free Knowledge Distillation
by: Tran, Minh-Tuan, et al.
Published: (2023)
by: Tran, Minh-Tuan, et al.
Published: (2023)
Semantic Anchoring for Robust Personalization in Text-to-Image Diffusion Models
by: Yang, Seoyun, et al.
Published: (2025)
by: Yang, Seoyun, et al.
Published: (2025)
Similar Items
-
SAT: Selective Aggregation Transformer for Image Super-Resolution
by: Tran, Dinh Phu, et al.
Published: (2026) -
Knowing When to Answer: Adaptive Confidence Refinement for Reliable Audio-Visual Question Answering
by: Tran, Dinh Phu, et al.
Published: (2026) -
Reference-Based Post-OCR Processing with LLM for Precise Diacritic Text in Historical Document Recognition
by: Do, Thao, et al.
Published: (2024) -
Rethinking the Nested U-Net Approach: Enhancing Biomarker Segmentation with Attention Mechanisms and Multiscale Feature Fusion
by: Wazir, Saad, et al.
Published: (2025) -
Rethinking Decoder Design: Improving Biomarker Segmentation Using Depth-to-Space Restoration and Residual Linear Attention
by: Wazir, Saad, et al.
Published: (2025)