TI-JEPA: An Innovative Energy-based Joint Embedding Strategy for Text-Image Multimodal Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Vo, Khang H. N., Nguyen, Duc P. T., Nguyen, Thong, Quan, Tho T. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ViConsFormer: Constituting Meaningful Phrases of Scene Texts using Transformer-based Method in Vietnamese Text-based Visual Question Answering
by: Nguyen, Nghia Hieu, et al.
Published: (2024)
by: Nguyen, Nghia Hieu, et al.
Published: (2024)
Enriching and Controlling Global Semantics for Text Summarization
by: Nguyen, Thong, et al.
Published: (2021)
by: Nguyen, Thong, et al.
Published: (2021)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
by: Nguyen, Nhi Ngoc-Yen, et al.
Published: (2026)
by: Nguyen, Nhi Ngoc-Yen, et al.
Published: (2026)
Expand BERT Representation with Visual Information via Grounded Language Learning with Multimodal Partial Alignment
by: Nguyen, Cong-Duy, et al.
Published: (2023)
by: Nguyen, Cong-Duy, et al.
Published: (2023)
CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base
by: Nguyen, Cong-Duy, et al.
Published: (2025)
by: Nguyen, Cong-Duy, et al.
Published: (2025)
Towards Cultural Bridge by Bahnaric-Vietnamese Translation Using Transfer Learning of Sequence-To-Sequence Pre-training Language Model
by: Dat, Phan Tran Minh, et al.
Published: (2025)
by: Dat, Phan Tran Minh, et al.
Published: (2025)
URAG: Implementing a Unified Hybrid RAG for Precise Answers in University Admission Chatbots -- A Case Study at HCMUT
by: Nguyen, Long, et al.
Published: (2025)
by: Nguyen, Long, et al.
Published: (2025)
JEPA-T: Joint-Embedding Predictive Architecture with Text Fusion for Image Generation
by: Wan, Siheng, et al.
Published: (2025)
by: Wan, Siheng, et al.
Published: (2025)
ViHERMES: A Graph-Grounded Multihop Question Answering Benchmark and System for Vietnamese Healthcare Regulations
by: Nguyen, Long S. T., et al.
Published: (2026)
by: Nguyen, Long S. T., et al.
Published: (2026)
ViConBERT: Context-Gloss Aligned Vietnamese Word Embedding for Polysemous and Sense-Aware Representations
by: Huynh, Khang T., et al.
Published: (2025)
by: Huynh, Khang T., et al.
Published: (2025)
KDMCSE: Knowledge Distillation Multimodal Sentence Embeddings with Adaptive Angular margin Contrastive Learning
by: Nguyen, Cong-Duy, et al.
Published: (2024)
by: Nguyen, Cong-Duy, et al.
Published: (2024)
DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding
by: Nguyen, Thong, et al.
Published: (2023)
by: Nguyen, Thong, et al.
Published: (2023)
Leveraging Sentence-oriented Augmentation and Transformer-Based Architecture for Vietnamese-Bahnaric Translation
by: Nguyen, Tan Sang, et al.
Published: (2026)
by: Nguyen, Tan Sang, et al.
Published: (2026)
Image-Text Out-Of-Context Detection Using Synthetic Multimodal Misinformation
by: Shalabi, Fatma, et al.
Published: (2024)
by: Shalabi, Fatma, et al.
Published: (2024)
VM14K: First Vietnamese Medical Benchmark
by: Nguyen, Thong, et al.
Published: (2025)
by: Nguyen, Thong, et al.
Published: (2025)
Bridging LLMs and Symbolic Reasoning in Educational QA Systems: Insights from the XAI Challenge at IJCNN 2025
by: Nguyen, Long S. T., et al.
Published: (2025)
by: Nguyen, Long S. T., et al.
Published: (2025)
ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text Retrieval with Optimal Transport
by: Tran, Quoc-Khang, et al.
Published: (2026)
by: Tran, Quoc-Khang, et al.
Published: (2026)
RAPID: Retrieval-Augmented Parallel Inference Drafting for Text-Based Video Event Retrieval
by: Nguyen, Long, et al.
Published: (2025)
by: Nguyen, Long, et al.
Published: (2025)
BERT-based model for Vietnamese Fact Verification Dataset
by: Tran, Bao, et al.
Published: (2025)
by: Tran, Bao, et al.
Published: (2025)
JEPA4Rec: Learning Effective Language Representations for Sequential Recommendation via Joint Embedding Predictive Architecture
by: Nguyen, Minh-Anh, et al.
Published: (2025)
by: Nguyen, Minh-Anh, et al.
Published: (2025)
Topic-aware Causal Intervention for Counterfactual Detection
by: Nguyen, Thong, et al.
Published: (2024)
by: Nguyen, Thong, et al.
Published: (2024)
ViTextVQA: A Large-Scale Visual Question Answering Dataset and a Novel Multimodal Feature Fusion Method for Vietnamese Text Comprehension in Images
by: Van Nguyen, Quan, et al.
Published: (2024)
by: Van Nguyen, Quan, et al.
Published: (2024)
KGAlign: Joint Semantic-Structural Knowledge Encoding for Multimodal Fake News Detection
by: La, Tuan-Vinh, et al.
Published: (2025)
by: La, Tuan-Vinh, et al.
Published: (2025)
Cross-Data Knowledge Graph Construction for LLM-enabled Educational Question-Answering System: A Case Study at HCMUT
by: Bui, Tuan, et al.
Published: (2024)
by: Bui, Tuan, et al.
Published: (2024)
Low-Resource Heuristics for Bahnaric Optical Character Recognition Improvement
by: Tran, Phat, et al.
Published: (2026)
by: Tran, Phat, et al.
Published: (2026)
More Bias, Less Bias: BiasPrompting for Enhanced Multiple-Choice Question Answering
by: Vu, Duc Anh, et al.
Published: (2025)
by: Vu, Duc Anh, et al.
Published: (2025)
Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model
by: Huang, Haoyang, et al.
Published: (2025)
by: Huang, Haoyang, et al.
Published: (2025)
VN-MTEB: Vietnamese Massive Text Embedding Benchmark
by: Pham, Loc, et al.
Published: (2025)
by: Pham, Loc, et al.
Published: (2025)
READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling
by: Nguyen, Thong, et al.
Published: (2023)
by: Nguyen, Thong, et al.
Published: (2023)
GloCOM: A Short Text Neural Topic Model via Global Clustering Context
by: Nguyen, Quang Duc, et al.
Published: (2024)
by: Nguyen, Quang Duc, et al.
Published: (2024)
Vintern-1B: An Efficient Multimodal Large Language Model for Vietnamese
by: Doan, Khang T., et al.
Published: (2024)
by: Doan, Khang T., et al.
Published: (2024)
VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
by: Chen, Delong, et al.
Published: (2025)
by: Chen, Delong, et al.
Published: (2025)
DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
by: He, Xiangteng, et al.
Published: (2025)
by: He, Xiangteng, et al.
Published: (2025)
Mitigating Multimodal Hallucinations via Gradient-based Self-Reflection
by: Wang, Shan, et al.
Published: (2025)
by: Wang, Shan, et al.
Published: (2025)
A Vietnamese Dataset for Text Segmentation and Multiple Choices Reading Comprehension
by: Hai, Toan Nguyen, et al.
Published: (2025)
by: Hai, Toan Nguyen, et al.
Published: (2025)
Crossing Linguistic Horizons: Finetuning and Comprehensive Evaluation of Vietnamese Large Language Models
by: Truong, Sang T., et al.
Published: (2024)
by: Truong, Sang T., et al.
Published: (2024)
DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Factorized Discrete Flow Matching
by: Nguyen, Ngoc-Son, et al.
Published: (2025)
by: Nguyen, Ngoc-Son, et al.
Published: (2025)
MAMA: Meta-optimized Angular Margin Contrastive Framework for Video-Language Representation Learning
by: Nguyen, Thong, et al.
Published: (2024)
by: Nguyen, Thong, et al.
Published: (2024)
Video Understanding: Through A Temporal Lens
by: Nguyen, Thong Thanh
Published: (2026)
by: Nguyen, Thong Thanh
Published: (2026)
Coreference Resolution for Vietnamese Narrative Texts
by: Tran, Hieu-Dai, et al.
Published: (2025)
by: Tran, Hieu-Dai, et al.
Published: (2025)
Similar Items
-
ViConsFormer: Constituting Meaningful Phrases of Scene Texts using Transformer-based Method in Vietnamese Text-based Visual Question Answering
by: Nguyen, Nghia Hieu, et al.
Published: (2024) -
Enriching and Controlling Global Semantics for Text Summarization
by: Nguyen, Thong, et al.
Published: (2021) -
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
by: Nguyen, Nhi Ngoc-Yen, et al.
Published: (2026) -
Expand BERT Representation with Visual Information via Grounded Language Learning with Multimodal Partial Alignment
by: Nguyen, Cong-Duy, et al.
Published: (2023) -
CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base
by: Nguyen, Cong-Duy, et al.
Published: (2025)