Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Thong Thanh, Wu, Xiaobao, Bin, Yi, Nguyen, Cong-Duy T, Ng, See-Kiong, Luu, Anh Tuan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Scale Contrastive Learning for Video Temporal Grounding
by: Nguyen, Thong Thanh, et al.
Published: (2024)
by: Nguyen, Thong Thanh, et al.
Published: (2024)
DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding
by: Nguyen, Thong, et al.
Published: (2023)
by: Nguyen, Thong, et al.
Published: (2023)
Encoding and Controlling Global Semantics for Long-form Video Question Answering
by: Nguyen, Thong Thanh, et al.
Published: (2024)
by: Nguyen, Thong Thanh, et al.
Published: (2024)
MAMA: Meta-optimized Angular Margin Contrastive Framework for Video-Language Representation Learning
by: Nguyen, Thong, et al.
Published: (2024)
by: Nguyen, Thong, et al.
Published: (2024)
Topic Modeling as Multi-Objective Contrastive Optimization
by: Nguyen, Thong, et al.
Published: (2024)
by: Nguyen, Thong, et al.
Published: (2024)
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
by: Nguyen, Thong, et al.
Published: (2025)
by: Nguyen, Thong, et al.
Published: (2025)
Vision-and-Language Pretraining
by: Nguyen, Thong, et al.
Published: (2022)
by: Nguyen, Thong, et al.
Published: (2022)
READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling
by: Nguyen, Thong, et al.
Published: (2023)
by: Nguyen, Thong, et al.
Published: (2023)
KDMCSE: Knowledge Distillation Multimodal Sentence Embeddings with Adaptive Angular margin Contrastive Learning
by: Nguyen, Cong-Duy, et al.
Published: (2024)
by: Nguyen, Cong-Duy, et al.
Published: (2024)
Enhancing Multimodal Entity Linking with Jaccard Distance-based Conditional Contrastive Learning and Contextual Visual Augmentation
by: Nguyen, Cong-Duy, et al.
Published: (2025)
by: Nguyen, Cong-Duy, et al.
Published: (2025)
CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base
by: Nguyen, Cong-Duy, et al.
Published: (2025)
by: Nguyen, Cong-Duy, et al.
Published: (2025)
Adaptive Contrastive Learning on Multimodal Transformer for Review Helpfulness Predictions
by: Nguyen, Thong, et al.
Published: (2022)
by: Nguyen, Thong, et al.
Published: (2022)
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
by: Cao, Tri, et al.
Published: (2026)
by: Cao, Tri, et al.
Published: (2026)
A Survey on Neural Topic Models: Methods, Applications, and Challenges
by: Wu, Xiaobao, et al.
Published: (2024)
by: Wu, Xiaobao, et al.
Published: (2024)
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
by: Le, Khoi, et al.
Published: (2026)
by: Le, Khoi, et al.
Published: (2026)
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives
by: Nguyen, Thong, et al.
Published: (2024)
by: Nguyen, Thong, et al.
Published: (2024)
Modeling Dynamic Topics in Chain-Free Fashion by Evolution-Tracking Contrastive Learning and Unassociated Word Exclusion
by: Wu, Xiaobao, et al.
Published: (2024)
by: Wu, Xiaobao, et al.
Published: (2024)
Learning Uncertainty from Sequential Internal Dispersion in Large Language Models
by: Srey, Ponhvoan, et al.
Published: (2026)
by: Srey, Ponhvoan, et al.
Published: (2026)
Curriculum Demonstration Selection for In-Context Learning
by: Vu, Duc Anh, et al.
Published: (2024)
by: Vu, Duc Anh, et al.
Published: (2024)
Don't Read Everything: A Curvature-Conditioned Query for Linear Attention
by: Le, Dong, et al.
Published: (2026)
by: Le, Dong, et al.
Published: (2026)
Depth-aware Panoptic Segmentation
by: Nguyen, Tuan, et al.
Published: (2024)
by: Nguyen, Tuan, et al.
Published: (2024)
Video Understanding: Through A Temporal Lens
by: Nguyen, Thong Thanh
Published: (2026)
by: Nguyen, Thong Thanh
Published: (2026)
Expand BERT Representation with Visual Information via Grounded Language Learning with Multimodal Partial Alignment
by: Nguyen, Cong-Duy, et al.
Published: (2023)
by: Nguyen, Cong-Duy, et al.
Published: (2023)
Gradient-Boosted Decision Tree for Listwise Context Model in Multimodal Review Helpfulness Prediction
by: Nguyen, Thong, et al.
Published: (2023)
by: Nguyen, Thong, et al.
Published: (2023)
On the Affinity, Rationality, and Diversity of Hierarchical Topic Modeling
by: Wu, Xiaobao, et al.
Published: (2024)
by: Wu, Xiaobao, et al.
Published: (2024)
More Bias, Less Bias: BiasPrompting for Enhanced Multiple-Choice Question Answering
by: Vu, Duc Anh, et al.
Published: (2025)
by: Vu, Duc Anh, et al.
Published: (2025)
HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video Understanding
by: Nguyen, Trong-Thuan, et al.
Published: (2023)
by: Nguyen, Trong-Thuan, et al.
Published: (2023)
CodeArena: A Collective Evaluation Platform for LLM Code Generation
by: Du, Mingzhe, et al.
Published: (2025)
by: Du, Mingzhe, et al.
Published: (2025)
Unified Interactive Multimodal Moment Retrieval via Cascaded Embedding-Reranking and Temporal-Aware Score Fusion
by: Thanh, Toan Le Ngo, et al.
Published: (2025)
by: Thanh, Toan Le Ngo, et al.
Published: (2025)
Mercury: A Code Efficiency Benchmark for Code Large Language Models
by: Du, Mingzhe, et al.
Published: (2024)
by: Du, Mingzhe, et al.
Published: (2024)
Panoptic Scene Graph Generation with Semantics-Prototype Learning
by: Li, Li, et al.
Published: (2023)
by: Li, Li, et al.
Published: (2023)
HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation
by: Nguyen, Trong-Thuan, et al.
Published: (2024)
by: Nguyen, Trong-Thuan, et al.
Published: (2024)
FASTopic: Pretrained Transformer is a Fast, Adaptive, Stable, and Transferable Topic Model
by: Wu, Xiaobao, et al.
Published: (2024)
by: Wu, Xiaobao, et al.
Published: (2024)
InfoCTM: A Mutual Information Maximization Perspective of Cross-Lingual Topic Modeling
by: Wu, Xiaobao, et al.
Published: (2023)
by: Wu, Xiaobao, et al.
Published: (2023)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
by: Nguyen, Nhi Ngoc-Yen, et al.
Published: (2026)
by: Nguyen, Nhi Ngoc-Yen, et al.
Published: (2026)
THYME: Temporal Hierarchical-Cyclic Interactivity Modeling for Video Scene Graphs in Aerial Footage
by: Nguyen, Trong-Thuan, et al.
Published: (2025)
by: Nguyen, Trong-Thuan, et al.
Published: (2025)
Spatio-Temporal Foundation Models: Vision, Challenges, and Opportunities
by: Goodge, Adam, et al.
Published: (2025)
by: Goodge, Adam, et al.
Published: (2025)
Adaptive Fusion Network with Temporal-Ranked and Motion-Intensity Dynamic Images for Micro-expression Recognition
by: Man, Thi Bich Phuong, et al.
Published: (2025)
by: Man, Thi Bich Phuong, et al.
Published: (2025)
SemRoDe: Macro Adversarial Training to Learn Representations That are Robust to Word-Level Attacks
by: Formento, Brian, et al.
Published: (2024)
by: Formento, Brian, et al.
Published: (2024)
SATURN: Autoregressive Image Generation Guided by Scene Graphs
by: Vo, Thanh-Nhan, et al.
Published: (2025)
by: Vo, Thanh-Nhan, et al.
Published: (2025)
Similar Items
-
Multi-Scale Contrastive Learning for Video Temporal Grounding
by: Nguyen, Thong Thanh, et al.
Published: (2024) -
DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding
by: Nguyen, Thong, et al.
Published: (2023) -
Encoding and Controlling Global Semantics for Long-form Video Question Answering
by: Nguyen, Thong Thanh, et al.
Published: (2024) -
MAMA: Meta-optimized Angular Margin Contrastive Framework for Video-Language Representation Learning
by: Nguyen, Thong, et al.
Published: (2024) -
Topic Modeling as Multi-Objective Contrastive Optimization
by: Nguyen, Thong, et al.
Published: (2024)