Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Tran, Tuyen, Le, Thao Minh, Le, Quang-Hung, Tran, Truyen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Agentic AI for Multimodal-Guided Video Object Segmentation
by: Tran, Tuyen, et al.
Published: (2025)
by: Tran, Tuyen, et al.
Published: (2025)
Unified Framework with Consistency across Modalities for Human Activity Recognition
by: Tran, Tuyen, et al.
Published: (2024)
by: Tran, Tuyen, et al.
Published: (2024)
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
by: Le, Quang-Hung, et al.
Published: (2024)
by: Le, Quang-Hung, et al.
Published: (2024)
The 2nd Solution for LSVOS Challenge RVOS Track: Spatial-temporal Refinement for Consistent Semantic Segmentation
by: Tran, Tuyen
Published: (2024)
by: Tran, Tuyen
Published: (2024)
SADL: An Effective In-Context Learning Method for Compositional Visual QA
by: Dang, Long Hoang, et al.
Published: (2024)
by: Dang, Long Hoang, et al.
Published: (2024)
Finding the Trigger: Causal Abductive Reasoning on Video Events
by: Le, Thao Minh, et al.
Published: (2025)
by: Le, Thao Minh, et al.
Published: (2025)
EDGER: EDge-Guided with HEatmap Refinement for Generalizable Image Forgery Localization
by: Le-Phan, Minh-Khoa, et al.
Published: (2026)
by: Le-Phan, Minh-Khoa, et al.
Published: (2026)
FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation
by: Le, Minh Khoa, et al.
Published: (2026)
by: Le, Minh Khoa, et al.
Published: (2026)
Enhancing Dataset Distillation via Non-Critical Region Refinement
by: Tran, Minh-Tuan, et al.
Published: (2025)
by: Tran, Minh-Tuan, et al.
Published: (2025)
FaR: Enhancing Multi-Concept Text-to-Image Diffusion via Concept Fusion and Localized Refinement
by: Tran, Gia-Nghia, et al.
Published: (2025)
by: Tran, Gia-Nghia, et al.
Published: (2025)
HENASY: Learning to Assemble Scene-Entities for Egocentric Video-Language Model
by: Vo, Khoa, et al.
Published: (2024)
by: Vo, Khoa, et al.
Published: (2024)
Test-Time Instance-Specific Parameter Composition: A New Paradigm for Adaptive Generative Modeling
by: Tran, Minh-Tuan, et al.
Published: (2026)
by: Tran, Minh-Tuan, et al.
Published: (2026)
TrafficVLM: A Controllable Visual Language Model for Traffic Video Captioning
by: Dinh, Quang Minh, et al.
Published: (2024)
by: Dinh, Quang Minh, et al.
Published: (2024)
NAYER: Noisy Layer Data Generation for Efficient and Effective Data-free Knowledge Distillation
by: Tran, Minh-Tuan, et al.
Published: (2023)
by: Tran, Minh-Tuan, et al.
Published: (2023)
TaleForge: Interactive Multimodal System for Personalized Story Creation
by: Nguyen, Minh-Loi, et al.
Published: (2025)
by: Nguyen, Minh-Loi, et al.
Published: (2025)
Robust Deepfake Detection: Mitigating Spatial Attention Drift via Calibrated Complementary Ensembles
by: Le-Phan, Minh-Khoa, et al.
Published: (2026)
by: Le-Phan, Minh-Khoa, et al.
Published: (2026)
Online 3D Multi-Camera Perception through Robust 2D Tracking and Depth-based Late Aggregation
by: Le, Vu-Minh, et al.
Published: (2025)
by: Le, Vu-Minh, et al.
Published: (2025)
TF-SASM: Training-free Spatial-aware Sparse Memory for Multi-object Tracking
by: Nguyen-Quang, Thuc, et al.
Published: (2024)
by: Nguyen-Quang, Thuc, et al.
Published: (2024)
A2VIS: Amodal-Aware Approach to Video Instance Segmentation
by: Tran, Minh, et al.
Published: (2024)
by: Tran, Minh, et al.
Published: (2024)
Event-Enriched Image Analysis Grand Challenge at ACM Multimedia 2025
by: Tran, Thien-Phuc, et al.
Published: (2025)
by: Tran, Thien-Phuc, et al.
Published: (2025)
Enhancing Domain Adaptation through Prompt Gradient Alignment
by: Phan, Hoang, et al.
Published: (2024)
by: Phan, Hoang, et al.
Published: (2024)
Phantasia: Context-Adaptive Backdoors in Vision Language Models
by: Tran, Nam Duong, et al.
Published: (2026)
by: Tran, Nam Duong, et al.
Published: (2026)
OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
by: Nguyen, Hieu, et al.
Published: (2025)
by: Nguyen, Hieu, et al.
Published: (2025)
MambaU-Lite: A Lightweight Model based on Mamba and Integrated Channel-Spatial Attention for Skin Lesion Segmentation
by: Nguyen, Thi-Nhu-Quynh, et al.
Published: (2024)
by: Nguyen, Thi-Nhu-Quynh, et al.
Published: (2024)
PointSplat: Efficient Geometry-Driven Pruning and Transformer Refinement for 3D Gaussian Splatting
by: Tran, Anh Thuan, et al.
Published: (2026)
by: Tran, Anh Thuan, et al.
Published: (2026)
Unpaired Image Dehazing via Kolmogorov-Arnold Transformation of Latent Features
by: Tran, Le-Anh
Published: (2025)
by: Tran, Le-Anh
Published: (2025)
TimeRefine: Temporal Grounding with Time Refining Video LLM
by: Wang, Xizi, et al.
Published: (2024)
by: Wang, Xizi, et al.
Published: (2024)
Automated Image Recognition Framework
by: Nguyen, Quang-Binh, et al.
Published: (2025)
by: Nguyen, Quang-Binh, et al.
Published: (2025)
GenFlow: Interactive Modular System for Image Generation
by: Nguyen, Duc-Hung, et al.
Published: (2025)
by: Nguyen, Duc-Hung, et al.
Published: (2025)
Amodal Instance Segmentation with Diffusion Shape Prior Estimation
by: Tran, Minh, et al.
Published: (2024)
by: Tran, Minh, et al.
Published: (2024)
Directed-Tokens: A Robust Multi-Modality Alignment Approach to Large Language-Vision Models
by: Truong, Thanh-Dat, et al.
Published: (2025)
by: Truong, Thanh-Dat, et al.
Published: (2025)
Interactive Interface For Semantic Segmentation Dataset Synthesis
by: Tran, Ngoc-Do, et al.
Published: (2025)
by: Tran, Ngoc-Do, et al.
Published: (2025)
LLandMark: A Multi-Agent Framework for Landmark-Aware Multimodal Interactive Video Retrieval
by: Phung, Minh-Chi, et al.
Published: (2026)
by: Phung, Minh-Chi, et al.
Published: (2026)
Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic Alignment
by: Song, Yizhi, et al.
Published: (2024)
by: Song, Yizhi, et al.
Published: (2024)
LiftRefine: Progressively Refined View Synthesis from 3D Lifting with Volume-Triplane Representations
by: Do, Tung, et al.
Published: (2024)
by: Do, Tung, et al.
Published: (2024)
GUNNEL: Guided Mixup Augmentation and Multi-Model Fusion for Aquatic Animal Segmentation
by: Le, Minh-Quan, et al.
Published: (2021)
by: Le, Minh-Quan, et al.
Published: (2021)
RelWitness: Open-Vocabulary 3D Scene Graph Generation with Visual-Geometric Relation Witnesses
by: Nguyen, Minh Anh, et al.
Published: (2026)
by: Nguyen, Minh Anh, et al.
Published: (2026)
On the Evaluation and Refinement of Vision-Language Instruction Tuning Datasets
by: Liao, Ning, et al.
Published: (2023)
by: Liao, Ning, et al.
Published: (2023)
RefineStyle: Dynamic Convolution Refinement for StyleGAN
by: Xia, Siwei, et al.
Published: (2024)
by: Xia, Siwei, et al.
Published: (2024)
CamoFA: A Learnable Fourier-based Augmentation for Camouflage Segmentation
by: Le, Minh-Quan, et al.
Published: (2023)
by: Le, Minh-Quan, et al.
Published: (2023)
Similar Items
-
Towards Agentic AI for Multimodal-Guided Video Object Segmentation
by: Tran, Tuyen, et al.
Published: (2025) -
Unified Framework with Consistency across Modalities for Human Activity Recognition
by: Tran, Tuyen, et al.
Published: (2024) -
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
by: Le, Quang-Hung, et al.
Published: (2024) -
The 2nd Solution for LSVOS Challenge RVOS Track: Spatial-temporal Refinement for Consistent Semantic Segmentation
by: Tran, Tuyen
Published: (2024) -
SADL: An Effective In-Context Learning Method for Compositional Visual QA
by: Dang, Long Hoang, et al.
Published: (2024)