Encoding and Controlling Global Semantics for Long-form Video Question Answering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nguyen, Thong Thanh, Hu, Zhiyuan, Wu, Xiaobao, Nguyen, Cong-Duy T, Ng, See-Kiong, Luu, Anh Tuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Scale Contrastive Learning for Video Temporal Grounding
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
von: Nguyen, Thong, et al.
Veröffentlicht: (2025)
von: Nguyen, Thong, et al.
Veröffentlicht: (2025)
READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
MAMA: Meta-optimized Angular Margin Contrastive Framework for Video-Language Representation Learning
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
Vision-and-Language Pretraining
von: Nguyen, Thong, et al.
Veröffentlicht: (2022)
von: Nguyen, Thong, et al.
Veröffentlicht: (2022)
Topic Modeling as Multi-Objective Contrastive Optimization
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base
von: Nguyen, Cong-Duy, et al.
Veröffentlicht: (2025)
von: Nguyen, Cong-Duy, et al.
Veröffentlicht: (2025)
KDMCSE: Knowledge Distillation Multimodal Sentence Embeddings with Adaptive Angular margin Contrastive Learning
von: Nguyen, Cong-Duy, et al.
Veröffentlicht: (2024)
von: Nguyen, Cong-Duy, et al.
Veröffentlicht: (2024)
More Bias, Less Bias: BiasPrompting for Enhanced Multiple-Choice Question Answering
von: Vu, Duc Anh, et al.
Veröffentlicht: (2025)
von: Vu, Duc Anh, et al.
Veröffentlicht: (2025)
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
von: Cao, Tri, et al.
Veröffentlicht: (2026)
von: Cao, Tri, et al.
Veröffentlicht: (2026)
Enhancing Multimodal Entity Linking with Jaccard Distance-based Conditional Contrastive Learning and Contextual Visual Augmentation
von: Nguyen, Cong-Duy, et al.
Veröffentlicht: (2025)
von: Nguyen, Cong-Duy, et al.
Veröffentlicht: (2025)
A Survey on Neural Topic Models: Methods, Applications, and Challenges
von: Wu, Xiaobao, et al.
Veröffentlicht: (2024)
von: Wu, Xiaobao, et al.
Veröffentlicht: (2024)
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
von: Le, Khoi, et al.
Veröffentlicht: (2026)
von: Le, Khoi, et al.
Veröffentlicht: (2026)
Adaptive Contrastive Learning on Multimodal Transformer for Review Helpfulness Predictions
von: Nguyen, Thong, et al.
Veröffentlicht: (2022)
von: Nguyen, Thong, et al.
Veröffentlicht: (2022)
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
Enriching and Controlling Global Semantics for Text Summarization
von: Nguyen, Thong, et al.
Veröffentlicht: (2021)
von: Nguyen, Thong, et al.
Veröffentlicht: (2021)
Don't Read Everything: A Curvature-Conditioned Query for Linear Attention
von: Le, Dong, et al.
Veröffentlicht: (2026)
von: Le, Dong, et al.
Veröffentlicht: (2026)
Learning Uncertainty from Sequential Internal Dispersion in Large Language Models
von: Srey, Ponhvoan, et al.
Veröffentlicht: (2026)
von: Srey, Ponhvoan, et al.
Veröffentlicht: (2026)
Video Understanding: Through A Temporal Lens
von: Nguyen, Thong Thanh
Veröffentlicht: (2026)
von: Nguyen, Thong Thanh
Veröffentlicht: (2026)
Gradient-Boosted Decision Tree for Listwise Context Model in Multimodal Review Helpfulness Prediction
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
On the Affinity, Rationality, and Diversity of Hierarchical Topic Modeling
von: Wu, Xiaobao, et al.
Veröffentlicht: (2024)
von: Wu, Xiaobao, et al.
Veröffentlicht: (2024)
Curriculum Demonstration Selection for In-Context Learning
von: Vu, Duc Anh, et al.
Veröffentlicht: (2024)
von: Vu, Duc Anh, et al.
Veröffentlicht: (2024)
Expand BERT Representation with Visual Information via Grounded Language Learning with Multimodal Partial Alignment
von: Nguyen, Cong-Duy, et al.
Veröffentlicht: (2023)
von: Nguyen, Cong-Duy, et al.
Veröffentlicht: (2023)
Modeling Dynamic Topics in Chain-Free Fashion by Evolution-Tracking Contrastive Learning and Unassociated Word Exclusion
von: Wu, Xiaobao, et al.
Veröffentlicht: (2024)
von: Wu, Xiaobao, et al.
Veröffentlicht: (2024)
ViConsFormer: Constituting Meaningful Phrases of Scene Texts using Transformer-based Method in Vietnamese Text-based Visual Question Answering
von: Nguyen, Nghia Hieu, et al.
Veröffentlicht: (2024)
von: Nguyen, Nghia Hieu, et al.
Veröffentlicht: (2024)
FASTopic: Pretrained Transformer is a Fast, Adaptive, Stable, and Transferable Topic Model
von: Wu, Xiaobao, et al.
Veröffentlicht: (2024)
von: Wu, Xiaobao, et al.
Veröffentlicht: (2024)
InfoCTM: A Mutual Information Maximization Perspective of Cross-Lingual Topic Modeling
von: Wu, Xiaobao, et al.
Veröffentlicht: (2023)
von: Wu, Xiaobao, et al.
Veröffentlicht: (2023)
Semi-Supervised Semantic Segmentation using Redesigned Self-Training for White Blood Cells
von: Luu, Vinh Quoc, et al.
Veröffentlicht: (2024)
von: Luu, Vinh Quoc, et al.
Veröffentlicht: (2024)
MRAG: A Modular Retrieval Framework for Time-Sensitive Question Answering
von: Siyue, Zhang, et al.
Veröffentlicht: (2024)
von: Siyue, Zhang, et al.
Veröffentlicht: (2024)
MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering
von: Nguyen, Hai-Dang, et al.
Veröffentlicht: (2025)
von: Nguyen, Hai-Dang, et al.
Veröffentlicht: (2025)
VietMEAgent: Culturally-Aware Few-Shot Multimodal Explanation for Vietnamese Visual Question Answering
von: Nguyen, Hai-Dang, et al.
Veröffentlicht: (2025)
von: Nguyen, Hai-Dang, et al.
Veröffentlicht: (2025)
Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering
von: Fernando, Basura, et al.
Veröffentlicht: (2025)
von: Fernando, Basura, et al.
Veröffentlicht: (2025)
AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering
von: Tuong, Nguyen Anh, et al.
Veröffentlicht: (2026)
von: Tuong, Nguyen Anh, et al.
Veröffentlicht: (2026)
Mercury: A Code Efficiency Benchmark for Code Large Language Models
von: Du, Mingzhe, et al.
Veröffentlicht: (2024)
von: Du, Mingzhe, et al.
Veröffentlicht: (2024)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
Towards Reliable Truth-Aligned Uncertainty Estimation in Large Language Models
von: Srey, Ponhvoan, et al.
Veröffentlicht: (2026)
von: Srey, Ponhvoan, et al.
Veröffentlicht: (2026)
Towards Signboard-Oriented Visual Question Answering: ViSignVQA Dataset, Method and Benchmark
von: Nguyen, Hieu Minh, et al.
Veröffentlicht: (2025)
von: Nguyen, Hieu Minh, et al.
Veröffentlicht: (2025)
Unified Interactive Multimodal Moment Retrieval via Cascaded Embedding-Reranking and Temporal-Aware Score Fusion
von: Thanh, Toan Le Ngo, et al.
Veröffentlicht: (2025)
von: Thanh, Toan Le Ngo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multi-Scale Contrastive Learning for Video Temporal Grounding
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024) -
Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024) -
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
von: Nguyen, Thong, et al.
Veröffentlicht: (2025) -
READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling
von: Nguyen, Thong, et al.
Veröffentlicht: (2023) -
DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)