TF-SASM: Training-free Spatial-aware Sparse Memory for Multi-object Tracking
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Nguyen-Quang, Thuc, Tran, Minh-Triet |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
TaleForge: Interactive Multimodal System for Personalized Story Creation
par: Nguyen, Minh-Loi, et autres
Publié: (2025)
par: Nguyen, Minh-Loi, et autres
Publié: (2025)
OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
par: Nguyen, Hieu, et autres
Publié: (2025)
par: Nguyen, Hieu, et autres
Publié: (2025)
ShowFlow: From Robust Single Concept to Condition-Free Multi-Concept Generation
par: Hoang, Trong-Vu, et autres
Publié: (2025)
par: Hoang, Trong-Vu, et autres
Publié: (2025)
Robust Deepfake Detection: Mitigating Spatial Attention Drift via Calibrated Complementary Ensembles
par: Le-Phan, Minh-Khoa, et autres
Publié: (2026)
par: Le-Phan, Minh-Khoa, et autres
Publié: (2026)
SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
par: Nguyen, Trong-Thuan, et autres
Publié: (2025)
par: Nguyen, Trong-Thuan, et autres
Publié: (2025)
Automated Image Recognition Framework
par: Nguyen, Quang-Binh, et autres
Publié: (2025)
par: Nguyen, Quang-Binh, et autres
Publié: (2025)
GUNNEL: Guided Mixup Augmentation and Multi-Model Fusion for Aquatic Animal Segmentation
par: Le, Minh-Quan, et autres
Publié: (2021)
par: Le, Minh-Quan, et autres
Publié: (2021)
FaR: Enhancing Multi-Concept Text-to-Image Diffusion via Concept Fusion and Localized Refinement
par: Tran, Gia-Nghia, et autres
Publié: (2025)
par: Tran, Gia-Nghia, et autres
Publié: (2025)
Interactive Interface For Semantic Segmentation Dataset Synthesis
par: Tran, Ngoc-Do, et autres
Publié: (2025)
par: Tran, Ngoc-Do, et autres
Publié: (2025)
SATURN: Autoregressive Image Generation Guided by Scene Graphs
par: Vo, Thanh-Nhan, et autres
Publié: (2025)
par: Vo, Thanh-Nhan, et autres
Publié: (2025)
VENUS: Visual Editing with Noise Inversion Using Scene Graphs
par: Vo, Thanh-Nhan, et autres
Publié: (2026)
par: Vo, Thanh-Nhan, et autres
Publié: (2026)
CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
par: Nguyen, Quang-Binh, et autres
Publié: (2025)
par: Nguyen, Quang-Binh, et autres
Publié: (2025)
SAMURAI: Shape-Aware Multimodal Retrieval for 3D Object Identification
par: Vo, Dinh-Khoi, et autres
Publié: (2025)
par: Vo, Dinh-Khoi, et autres
Publié: (2025)
GenFlow: Interactive Modular System for Image Generation
par: Nguyen, Duc-Hung, et autres
Publié: (2025)
par: Nguyen, Duc-Hung, et autres
Publié: (2025)
EditSleuth: A Dataset of Grounded Reasoning Chains for Image-Edit Forensics
par: Nguyen, Van-Loc, et autres
Publié: (2026)
par: Nguyen, Van-Loc, et autres
Publié: (2026)
EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
par: Vo, Dinh-Khoi, et autres
Publié: (2025)
par: Vo, Dinh-Khoi, et autres
Publié: (2025)
THYME: Temporal Hierarchical-Cyclic Interactivity Modeling for Video Scene Graphs in Aerial Footage
par: Nguyen, Trong-Thuan, et autres
Publié: (2025)
par: Nguyen, Trong-Thuan, et autres
Publié: (2025)
GenKOL: Modular Generative AI Framework For Scalable Virtual KOL Generation
par: To, Tan-Hiep, et autres
Publié: (2025)
par: To, Tan-Hiep, et autres
Publié: (2025)
CamoFA: A Learnable Fourier-based Augmentation for Camouflage Segmentation
par: Le, Minh-Quan, et autres
Publié: (2023)
par: Le, Minh-Quan, et autres
Publié: (2023)
EDGER: EDge-Guided with HEatmap Refinement for Generalizable Image Forgery Localization
par: Le-Phan, Minh-Khoa, et autres
Publié: (2026)
par: Le-Phan, Minh-Khoa, et autres
Publié: (2026)
ARtVista: Gateway To Empower Anyone Into Artist
par: Hoang, Trong-Vu, et autres
Publié: (2024)
par: Hoang, Trong-Vu, et autres
Publié: (2024)
Event-Enriched Image Analysis Grand Challenge at ACM Multimedia 2025
par: Tran, Thien-Phuc, et autres
Publié: (2025)
par: Tran, Thien-Phuc, et autres
Publié: (2025)
SimGraph: A Unified Framework for Scene Graph-Based Image Generation and Editing
par: Vo, Thanh-Nhan, et autres
Publié: (2026)
par: Vo, Thanh-Nhan, et autres
Publié: (2026)
MaskDiff: Modeling Mask Distribution with Diffusion Probabilistic Model for Few-Shot Instance Segmentation
par: Le, Minh-Quan, et autres
Publié: (2023)
par: Le, Minh-Quan, et autres
Publié: (2023)
PANDORA: Pixel-wise Attention Dissolution and Latent Guidance for Zero-Shot Object Removal
par: Vo, Dinh-Khoi, et autres
Publié: (2026)
par: Vo, Dinh-Khoi, et autres
Publié: (2026)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
par: Anh, Duy Le Dinh, et autres
Publié: (2024)
par: Anh, Duy Le Dinh, et autres
Publié: (2024)
Transferable Multi-Bit Watermarking Across Frozen Diffusion Models via Latent Consistency Bridges
par: Nguyen-Le, Hong-Hanh, et autres
Publié: (2026)
par: Nguyen-Le, Hong-Hanh, et autres
Publié: (2026)
Any3DIS: Class-Agnostic 3D Instance Segmentation by 2D Mask Tracking
par: Nguyen, Phuc, et autres
Publié: (2024)
par: Nguyen, Phuc, et autres
Publié: (2024)
Integrating Image Features with Convolutional Sequence-to-sequence Network for Multilingual Visual Question Answering
par: Thai, Triet Minh, et autres
Publié: (2023)
par: Thai, Triet Minh, et autres
Publié: (2023)
VisionGuard: Synergistic Framework for Helmet Violation Detection
par: Nguyen, Lam-Huy, et autres
Publié: (2025)
par: Nguyen, Lam-Huy, et autres
Publié: (2025)
ReCap: Event-Aware Image Captioning with Article Retrieval and Semantic Gaussian Normalization
par: Nguyen, Thinh-Phuc, et autres
Publié: (2025)
par: Nguyen, Thinh-Phuc, et autres
Publié: (2025)
SAM-EG: Segment Anything Model with Egde Guidance framework for efficient Polyp Segmentation
par: Trinh, Quoc-Huy, et autres
Publié: (2024)
par: Trinh, Quoc-Huy, et autres
Publié: (2024)
CPAM: Context-Preserving Adaptive Manipulation for Zero-Shot Real Image Editing
par: Vo, Dinh-Khoi, et autres
Publié: (2025)
par: Vo, Dinh-Khoi, et autres
Publié: (2025)
SDPA++: A General Framework for Self-Supervised Denoising with Patch Aggregation
par: Nguyen, Huy Minh Nhat, et autres
Publié: (2025)
par: Nguyen, Huy Minh Nhat, et autres
Publié: (2025)
MambaU-Lite: A Lightweight Model based on Mamba and Integrated Channel-Spatial Attention for Skin Lesion Segmentation
par: Nguyen, Thi-Nhu-Quynh, et autres
Publié: (2024)
par: Nguyen, Thi-Nhu-Quynh, et autres
Publié: (2024)
Shape2Animal: Creative Animal Generation from Natural Silhouettes
par: Tran, Quoc-Duy, et autres
Publié: (2025)
par: Tran, Quoc-Duy, et autres
Publié: (2025)
Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
par: Tran, Tuyen, et autres
Publié: (2025)
par: Tran, Tuyen, et autres
Publié: (2025)
Driver Attention Tracking and Analysis
par: Nguyen, Dat Viet Thanh, et autres
Publié: (2024)
par: Nguyen, Dat Viet Thanh, et autres
Publié: (2024)
NeIn: Telling What You Don't Want
par: Bui, Nhat-Tan, et autres
Publié: (2024)
par: Bui, Nhat-Tan, et autres
Publié: (2024)
KiseKloset for Fashion Retrieval and Recommendation
par: Phan-Nguyen, Thanh-Tung, et autres
Publié: (2025)
par: Phan-Nguyen, Thanh-Tung, et autres
Publié: (2025)
Documents similaires
-
TaleForge: Interactive Multimodal System for Personalized Story Creation
par: Nguyen, Minh-Loi, et autres
Publié: (2025) -
OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
par: Nguyen, Hieu, et autres
Publié: (2025) -
ShowFlow: From Robust Single Concept to Condition-Free Multi-Concept Generation
par: Hoang, Trong-Vu, et autres
Publié: (2025) -
Robust Deepfake Detection: Mitigating Spatial Attention Drift via Calibrated Complementary Ensembles
par: Le-Phan, Minh-Khoa, et autres
Publié: (2026) -
SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
par: Nguyen, Trong-Thuan, et autres
Publié: (2025)