Selective Off-Policy Reference Tuning with Plan Guidance
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Le, Duc Anh, Nguyen, Tien-Phat, Nguyen, Thien Huu, Van, Linh Ngo, Le, Trung |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Preserving Generalization of Language models in Few-shot Continual Relation Extraction
par: Tran, Quyen, et autres
Publié: (2024)
par: Tran, Quyen, et autres
Publié: (2024)
Zero-shot Cross-lingual Transfer Learning with Multiple Source and Target Languages for Information Extraction: Language Selection and Adversarial Training
par: Ngo, Nghia Trung, et autres
Publié: (2024)
par: Ngo, Nghia Trung, et autres
Publié: (2024)
Realistic Evaluation of Toxicity in Large Language Models
par: Luong, Tinh Son, et autres
Publié: (2024)
par: Luong, Tinh Son, et autres
Publié: (2024)
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching
par: Nguyen, Truong, et autres
Publié: (2026)
par: Nguyen, Truong, et autres
Publié: (2026)
mSCoRe: a $M$ultilingual and Scalable Benchmark for $S$kill-based $Co$mmonsense $Re$asoning
par: Ngo, Nghia Trung, et autres
Publié: (2025)
par: Ngo, Nghia Trung, et autres
Publié: (2025)
Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering
par: Ngo, Nghia Trung, et autres
Publié: (2024)
par: Ngo, Nghia Trung, et autres
Publié: (2024)
RAG-IT: Retrieval-Augmented Instruction Tuning for Automated Financial Analysis -- A Case Study for the Semiconductor Sector
par: To, Hai-Thien, et autres
Publié: (2024)
par: To, Hai-Thien, et autres
Publié: (2024)
XTRA: Cross-Lingual Topic Modeling with Topic and Representation Alignments
par: Nguyen, Tien Phat, et autres
Publié: (2025)
par: Nguyen, Tien Phat, et autres
Publié: (2025)
Spectral Flattening Is All Muon Needs: How Orthogonalization Controls Learning Rate and Convergence
par: Nguyen, Tien-Phat, et autres
Publié: (2026)
par: Nguyen, Tien-Phat, et autres
Publié: (2026)
Unified Interactive Multimodal Moment Retrieval via Cascaded Embedding-Reranking and Temporal-Aware Score Fusion
par: Thanh, Toan Le Ngo, et autres
Publié: (2025)
par: Thanh, Toan Le Ngo, et autres
Publié: (2025)
Few-Shot, No Problem: Descriptive Continual Relation Extraction
par: Thanh, Nguyen Xuan, et autres
Publié: (2025)
par: Thanh, Nguyen Xuan, et autres
Publié: (2025)
BSO: Safety Alignment Is Density Ratio Matching
par: Nguyen, Tien-Phat, et autres
Publié: (2026)
par: Nguyen, Tien-Phat, et autres
Publié: (2026)
GloCTM: Cross-Lingual Topic Modeling via a Global Context Space
par: Phat, Nguyen Tien, et autres
Publié: (2026)
par: Phat, Nguyen Tien, et autres
Publié: (2026)
DmC: Nearest Neighbor Guidance Diffusion Model for Offline Cross-domain Reinforcement Learning
par: Van, Linh Le Pham, et autres
Publié: (2025)
par: Van, Linh Le Pham, et autres
Publié: (2025)
ITSELF: Attention Guided Fine-Grained Alignment for Vision-Language Retrieval
par: Nguyen, Tien-Huy, et autres
Publié: (2026)
par: Nguyen, Tien-Huy, et autres
Publié: (2026)
CoT2Align: Cross-Chain of Thought Distillation via Optimal Transport Alignment for Language Models with Different Tokenizers
par: Le, Anh Duc, et autres
Publié: (2025)
par: Le, Anh Duc, et autres
Publié: (2025)
Few-shot Continual Relation Extraction via Open Information Extraction
par: Nguyen, Thiem, et autres
Publié: (2025)
par: Nguyen, Thiem, et autres
Publié: (2025)
Policy Learning for Off-Dynamics RL with Deficient Support
par: Van, Linh Le Pham, et autres
Publié: (2024)
par: Van, Linh Le Pham, et autres
Publié: (2024)
GloCOM: A Short Text Neural Topic Model via Global Clustering Context
par: Nguyen, Quang Duc, et autres
Publié: (2024)
par: Nguyen, Quang Duc, et autres
Publié: (2024)
Secure and Efficient UAV-Based Face Detection via Homomorphic Encryption and Edge Computing
par: Van Duc, Nguyen, et autres
Publié: (2025)
par: Van Duc, Nguyen, et autres
Publié: (2025)
Adaptive Prompting for Continual Relation Extraction: A Within-Task Variance Perspective
par: Le, Minh, et autres
Publié: (2024)
par: Le, Minh, et autres
Publié: (2024)
MADTempo: An Interactive System for Multi-Event Temporal Video Retrieval with Query Augmentation
par: Vu, Huu-An, et autres
Publié: (2025)
par: Vu, Huu-An, et autres
Publié: (2025)
NeuroMax: Enhancing Neural Topic Modeling via Maximizing Mutual Information and Group Topic Regularization
par: Pham, Duy-Tung, et autres
Publié: (2024)
par: Pham, Duy-Tung, et autres
Publié: (2024)
Divide and Refine: Enhancing Multimodal Representation and Explainability for Emotion Recognition in Conversation
par: Mai, Anh-Tuan, et autres
Publié: (2026)
par: Mai, Anh-Tuan, et autres
Publié: (2026)
Taipan: Efficient and Expressive State Space Language Models with Selective Attention
par: Van Nguyen, Chien, et autres
Publié: (2024)
par: Van Nguyen, Chien, et autres
Publié: (2024)
CausalPlan: Empowering Efficient LLM Multi-Agent Collaboration Through Causality-Driven Planning
par: Nguyen, Minh Hoang, et autres
Publié: (2025)
par: Nguyen, Minh Hoang, et autres
Publié: (2025)
SwiftBrush v2: Make Your One-step Diffusion Model Better Than Its Teacher
par: Dao, Trung, et autres
Publié: (2024)
par: Dao, Trung, et autres
Publié: (2024)
Multi-modal Adaptive Mixture of Experts for Cold-start Recommendation
par: Nguyen, Van-Khang, et autres
Publié: (2025)
par: Nguyen, Van-Khang, et autres
Publié: (2025)
SuperRAG: Beyond RAG with Layout-Aware Graph Modeling
par: Yang, Jeff, et autres
Publié: (2025)
par: Yang, Jeff, et autres
Publié: (2025)
Beyond Low-rank Decomposition: A Shortcut Approach for Efficient On-Device Learning
par: Nguyen, Le-Trung, et autres
Publié: (2025)
par: Nguyen, Le-Trung, et autres
Publié: (2025)
SurveyG: A Multi-Agent LLM Framework with Hierarchical Citation Graph for Automated Survey Generation
par: Nguye, Minh-Anh, et autres
Publié: (2025)
par: Nguye, Minh-Anh, et autres
Publié: (2025)
Reasoning Planning for Language Models
par: Nguyen, Bao, et autres
Publié: (2025)
par: Nguyen, Bao, et autres
Publié: (2025)
A Concept is More Than a Word: Diversified Unlearning in Text-to-Image Diffusion Models
par: Pham, Duc Hao, et autres
Publié: (2026)
par: Pham, Duc Hao, et autres
Publié: (2026)
Whisper based Cross-Lingual Phoneme Recognition between Vietnamese and English
par: Minh, Nguyen Huu Nhat, et autres
Publié: (2025)
par: Minh, Nguyen Huu Nhat, et autres
Publié: (2025)
xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models
par: Luong, Phung Duc, et autres
Publié: (2025)
par: Luong, Phung Duc, et autres
Publié: (2025)
Spectral Text Fusion: A Frequency-Aware Approach to Multimodal Time-Series Forecasting
par: Nguyen, Huu Hiep, et autres
Publié: (2026)
par: Nguyen, Huu Hiep, et autres
Publié: (2026)
Sharpness-Guided Group Relative Policy Optimization via Probability Shaping
par: Le, Tue, et autres
Publié: (2025)
par: Le, Tue, et autres
Publié: (2025)
BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
par: Le, Huy, et autres
Publié: (2025)
par: Le, Huy, et autres
Publié: (2025)
Auto-Prompting with Retrieval Guidance for Frame Detection in Logistics
par: Duc, Do Minh, et autres
Publié: (2025)
par: Duc, Do Minh, et autres
Publié: (2025)
Low-Field Magnetic Resonance Image Quality Enhancement using a Conditional Flow Matching Model
par: Nguyen, Huu Tien, et autres
Publié: (2025)
par: Nguyen, Huu Tien, et autres
Publié: (2025)
Documents similaires
-
Preserving Generalization of Language models in Few-shot Continual Relation Extraction
par: Tran, Quyen, et autres
Publié: (2024) -
Zero-shot Cross-lingual Transfer Learning with Multiple Source and Target Languages for Information Extraction: Language Selection and Adversarial Training
par: Ngo, Nghia Trung, et autres
Publié: (2024) -
Realistic Evaluation of Toxicity in Large Language Models
par: Luong, Tinh Son, et autres
Publié: (2024) -
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching
par: Nguyen, Truong, et autres
Publié: (2026) -
mSCoRe: a $M$ultilingual and Scalable Benchmark for $S$kill-based $Co$mmonsense $Re$asoning
par: Ngo, Nghia Trung, et autres
Publié: (2025)