Unified Framework with Consistency across Modalities for Human Activity Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tran, Tuyen, Le, Thao Minh, Tran, Hung, Tran, Truyen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Agentic AI for Multimodal-Guided Video Object Segmentation
von: Tran, Tuyen, et al.
Veröffentlicht: (2025)
von: Tran, Tuyen, et al.
Veröffentlicht: (2025)
Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
von: Tran, Tuyen, et al.
Veröffentlicht: (2025)
von: Tran, Tuyen, et al.
Veröffentlicht: (2025)
The 2nd Solution for LSVOS Challenge RVOS Track: Spatial-temporal Refinement for Consistent Semantic Segmentation
von: Tran, Tuyen
Veröffentlicht: (2024)
von: Tran, Tuyen
Veröffentlicht: (2024)
SADL: An Effective In-Context Learning Method for Compositional Visual QA
von: Dang, Long Hoang, et al.
Veröffentlicht: (2024)
von: Dang, Long Hoang, et al.
Veröffentlicht: (2024)
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)
Finding the Trigger: Causal Abductive Reasoning on Video Events
von: Le, Thao Minh, et al.
Veröffentlicht: (2025)
von: Le, Thao Minh, et al.
Veröffentlicht: (2025)
FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation
von: Le, Minh Khoa, et al.
Veröffentlicht: (2026)
von: Le, Minh Khoa, et al.
Veröffentlicht: (2026)
Automated Image Recognition Framework
von: Nguyen, Quang-Binh, et al.
Veröffentlicht: (2025)
von: Nguyen, Quang-Binh, et al.
Veröffentlicht: (2025)
Online 3D Multi-Camera Perception through Robust 2D Tracking and Depth-based Late Aggregation
von: Le, Vu-Minh, et al.
Veröffentlicht: (2025)
von: Le, Vu-Minh, et al.
Veröffentlicht: (2025)
NAYER: Noisy Layer Data Generation for Efficient and Effective Data-free Knowledge Distillation
von: Tran, Minh-Tuan, et al.
Veröffentlicht: (2023)
von: Tran, Minh-Tuan, et al.
Veröffentlicht: (2023)
Reference-Based Post-OCR Processing with LLM for Precise Diacritic Text in Historical Document Recognition
von: Do, Thao, et al.
Veröffentlicht: (2024)
von: Do, Thao, et al.
Veröffentlicht: (2024)
SignBart -- New approach with the skeleton sequence for Isolated Sign language Recognition
von: Nguyen, Tinh, et al.
Veröffentlicht: (2025)
von: Nguyen, Tinh, et al.
Veröffentlicht: (2025)
Test-Time Instance-Specific Parameter Composition: A New Paradigm for Adaptive Generative Modeling
von: Tran, Minh-Tuan, et al.
Veröffentlicht: (2026)
von: Tran, Minh-Tuan, et al.
Veröffentlicht: (2026)
Interactive Interface For Semantic Segmentation Dataset Synthesis
von: Tran, Ngoc-Do, et al.
Veröffentlicht: (2025)
von: Tran, Ngoc-Do, et al.
Veröffentlicht: (2025)
Unpaired Image Dehazing via Kolmogorov-Arnold Transformation of Latent Features
von: Tran, Le-Anh
Veröffentlicht: (2025)
von: Tran, Le-Anh
Veröffentlicht: (2025)
Robust Deepfake Detection: Mitigating Spatial Attention Drift via Calibrated Complementary Ensembles
von: Le-Phan, Minh-Khoa, et al.
Veröffentlicht: (2026)
von: Le-Phan, Minh-Khoa, et al.
Veröffentlicht: (2026)
EDGER: EDge-Guided with HEatmap Refinement for Generalizable Image Forgery Localization
von: Le-Phan, Minh-Khoa, et al.
Veröffentlicht: (2026)
von: Le-Phan, Minh-Khoa, et al.
Veröffentlicht: (2026)
TriaGS: Differentiable Triangulation-Guided Geometric Consistency for 3D Gaussian Splatting
von: Tran, Quan, et al.
Veröffentlicht: (2025)
von: Tran, Quan, et al.
Veröffentlicht: (2025)
Amodal Instance Segmentation with Diffusion Shape Prior Estimation
von: Tran, Minh, et al.
Veröffentlicht: (2024)
von: Tran, Minh, et al.
Veröffentlicht: (2024)
GenFlow: Interactive Modular System for Image Generation
von: Nguyen, Duc-Hung, et al.
Veröffentlicht: (2025)
von: Nguyen, Duc-Hung, et al.
Veröffentlicht: (2025)
Stratified Domain Adaptation: A Progressive Self-Training Approach for Scene Text Recognition
von: Le, Kha Nhat, et al.
Veröffentlicht: (2024)
von: Le, Kha Nhat, et al.
Veröffentlicht: (2024)
Discrete Facial Encoding: : A Framework for Data-driven Facial Display Discovery
von: Tran, Minh, et al.
Veröffentlicht: (2025)
von: Tran, Minh, et al.
Veröffentlicht: (2025)
Hypergraph Laplacian Eigenmaps and Face Recognition Problems
von: Tran, Loc Hoang
Veröffentlicht: (2024)
von: Tran, Loc Hoang
Veröffentlicht: (2024)
VSRM: A Robust Mamba-Based Framework for Video Super-Resolution
von: Tran, Dinh Phu, et al.
Veröffentlicht: (2025)
von: Tran, Dinh Phu, et al.
Veröffentlicht: (2025)
TF-SASM: Training-free Spatial-aware Sparse Memory for Multi-object Tracking
von: Nguyen-Quang, Thuc, et al.
Veröffentlicht: (2024)
von: Nguyen-Quang, Thuc, et al.
Veröffentlicht: (2024)
LLandMark: A Multi-Agent Framework for Landmark-Aware Multimodal Interactive Video Retrieval
von: Phung, Minh-Chi, et al.
Veröffentlicht: (2026)
von: Phung, Minh-Chi, et al.
Veröffentlicht: (2026)
CarcassFormer: An End-to-end Transformer-based Framework for Simultaneous Localization, Segmentation and Classification of Poultry Carcass Defect
von: Tran, Minh, et al.
Veröffentlicht: (2024)
von: Tran, Minh, et al.
Veröffentlicht: (2024)
ReasonVQA: A Multi-hop Reasoning Benchmark with Structural Knowledge for Visual Question Answering
von: Tran, Duong T., et al.
Veröffentlicht: (2025)
von: Tran, Duong T., et al.
Veröffentlicht: (2025)
Event-Enriched Image Analysis Grand Challenge at ACM Multimedia 2025
von: Tran, Thien-Phuc, et al.
Veröffentlicht: (2025)
von: Tran, Thien-Phuc, et al.
Veröffentlicht: (2025)
Directed-Tokens: A Robust Multi-Modality Alignment Approach to Large Language-Vision Models
von: Truong, Thanh-Dat, et al.
Veröffentlicht: (2025)
von: Truong, Thanh-Dat, et al.
Veröffentlicht: (2025)
A Distributed Multi-Modal Sensing Approach for Human Activity Recognition in Real-Time Human-Robot Collaboration
von: Belcamino, Valerio, et al.
Veröffentlicht: (2026)
von: Belcamino, Valerio, et al.
Veröffentlicht: (2026)
SimGraph: A Unified Framework for Scene Graph-Based Image Generation and Editing
von: Vo, Thanh-Nhan, et al.
Veröffentlicht: (2026)
von: Vo, Thanh-Nhan, et al.
Veröffentlicht: (2026)
Adaptive federated learning for ship detection across diverse satellite imagery sources
von: La, Tran-Vu, et al.
Veröffentlicht: (2025)
von: La, Tran-Vu, et al.
Veröffentlicht: (2025)
Low-Resource Heuristics for Bahnaric Optical Character Recognition Improvement
von: Tran, Phat, et al.
Veröffentlicht: (2026)
von: Tran, Phat, et al.
Veröffentlicht: (2026)
Blur2Blur: Blur Conversion for Unsupervised Image Deblurring on Unknown Domains
von: Pham, Bang-Dang, et al.
Veröffentlicht: (2024)
von: Pham, Bang-Dang, et al.
Veröffentlicht: (2024)
OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
von: Nguyen, Hieu, et al.
Veröffentlicht: (2025)
von: Nguyen, Hieu, et al.
Veröffentlicht: (2025)
Enhancing Domain Adaptation through Prompt Gradient Alignment
von: Phan, Hoang, et al.
Veröffentlicht: (2024)
von: Phan, Hoang, et al.
Veröffentlicht: (2024)
HENASY: Learning to Assemble Scene-Entities for Egocentric Video-Language Model
von: Vo, Khoa, et al.
Veröffentlicht: (2024)
von: Vo, Khoa, et al.
Veröffentlicht: (2024)
A2VIS: Amodal-Aware Approach to Video Instance Segmentation
von: Tran, Minh, et al.
Veröffentlicht: (2024)
von: Tran, Minh, et al.
Veröffentlicht: (2024)
DualFit: A Two-Stage Virtual Try-On via Warping and Synthesis
von: Tran, Minh, et al.
Veröffentlicht: (2025)
von: Tran, Minh, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards Agentic AI for Multimodal-Guided Video Object Segmentation
von: Tran, Tuyen, et al.
Veröffentlicht: (2025) -
Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
von: Tran, Tuyen, et al.
Veröffentlicht: (2025) -
The 2nd Solution for LSVOS Challenge RVOS Track: Spatial-temporal Refinement for Consistent Semantic Segmentation
von: Tran, Tuyen
Veröffentlicht: (2024) -
SADL: An Effective In-Context Learning Method for Compositional Visual QA
von: Dang, Long Hoang, et al.
Veröffentlicht: (2024) -
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)