Reference-Free Reinforcement Learning Fine-Tuning for MT: A Seq2Seq Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Garcia-Estrada, Ernesto, Escolano, Carlos, Fonallosa, José A. R. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Seq2Seq2Seq: Lossless Data Compression via Discrete Latent Transformers and Reinforcement Learning
by: Khodabandeh, Mahdi, et al.
Published: (2026)
by: Khodabandeh, Mahdi, et al.
Published: (2026)
Exploiting the Potential of Seq2Seq Models as Robust Few-Shot Learners
by: Lee, Jihyeon, et al.
Published: (2023)
by: Lee, Jihyeon, et al.
Published: (2023)
Leveraging Graph Structure in Seq2Seq Models for Knowledge Graph Link Prediction
by: Phuc, Luu Huu, et al.
Published: (2026)
by: Phuc, Luu Huu, et al.
Published: (2026)
SeqPO-SiMT: Sequential Policy Optimization for Simultaneous Machine Translation
by: Xu, Ting, et al.
Published: (2025)
by: Xu, Ting, et al.
Published: (2025)
Improving Bangla Linguistics: Advanced LSTM, Bi-LSTM, and Seq2Seq Models for Translating Sylheti to Modern Bangla
by: Das, Sourav Kumar, et al.
Published: (2025)
by: Das, Sourav Kumar, et al.
Published: (2025)
OptiSeq: Ordering Examples On-The-Fly for In-Context Learning
by: Bhope, Rahul Atul, et al.
Published: (2025)
by: Bhope, Rahul Atul, et al.
Published: (2025)
SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters
by: Yang, Yan, et al.
Published: (2024)
by: Yang, Yan, et al.
Published: (2024)
SeqPE: Transformer with Sequential Position Encoding
by: Li, Huayang, et al.
Published: (2025)
by: Li, Huayang, et al.
Published: (2025)
SA-DiffuSeq: Addressing Computational and Scalability Challenges in Long-Document Generation with Sparse Attention
by: Christoforos, Alexandros, et al.
Published: (2025)
by: Christoforos, Alexandros, et al.
Published: (2025)
Why Reinforcement Fine-Tuning Enables MLLMs Preserve Prior Knowledge Better: A Data Perspective
by: Zhang, Zhihao, et al.
Published: (2025)
by: Zhang, Zhihao, et al.
Published: (2025)
Supervised Fine-Tuning as Inverse Reinforcement Learning
by: Sun, Hao
Published: (2024)
by: Sun, Hao
Published: (2024)
Prior Prompt Engineering for Reinforcement Fine-Tuning
by: Taveekitworachai, Pittawat, et al.
Published: (2025)
by: Taveekitworachai, Pittawat, et al.
Published: (2025)
LLMs for Explainable Business Decision-Making: A Reinforcement Learning Fine-Tuning Approach
by: Cheng, Xiang, et al.
Published: (2025)
by: Cheng, Xiang, et al.
Published: (2025)
CantonMT: Cantonese to English NMT Platform with Fine-Tuned Models Using Synthetic Back-Translation Data
by: Hong, Kung Yin, et al.
Published: (2024)
by: Hong, Kung Yin, et al.
Published: (2024)
EHR-SeqSQL : A Sequential Text-to-SQL Dataset For Interactively Exploring Electronic Health Records
by: Ryu, Jaehee, et al.
Published: (2024)
by: Ryu, Jaehee, et al.
Published: (2024)
UniGenCoder: Merging Seq2Seq and Seq2Tree Paradigms for Unified Code Generation
by: Shao, Liangying, et al.
Published: (2025)
by: Shao, Liangying, et al.
Published: (2025)
Bidirectional Awareness Induction in Autoregressive Seq2Seq Models
by: Hu, Jia Cheng, et al.
Published: (2024)
by: Hu, Jia Cheng, et al.
Published: (2024)
DPI: Exploiting Parameter Heterogeneity for Interference-Free Fine-Tuning
by: Liu, Xiaoyu, et al.
Published: (2026)
by: Liu, Xiaoyu, et al.
Published: (2026)
Supervised Fine-Tuning versus Reinforcement Learning: A Study of Post-Training Methods for Large Language Models
by: Jiang, Haitao, et al.
Published: (2026)
by: Jiang, Haitao, et al.
Published: (2026)
Knowledge-Guided Biomarker Identification for Label-Free Single-Cell RNA-Seq Data: A Reinforcement Learning Perspective
by: Xiao, Meng, et al.
Published: (2025)
by: Xiao, Meng, et al.
Published: (2025)
Efficient Differentially Private Fine-Tuning of LLMs via Reinforcement Learning
by: Khadangi, Afshin, et al.
Published: (2025)
by: Khadangi, Afshin, et al.
Published: (2025)
SPARK: Stepwise Process-Aware Rewards for Reference-Free Reinforcement Learning
by: Rahman, Salman, et al.
Published: (2025)
by: Rahman, Salman, et al.
Published: (2025)
CARFT: Boosting LLM Reasoning via Contrastive Learning with Annotated Chain-of-Thought-based Reinforced Fine-Tuning
by: Zhu, Wenqiao, et al.
Published: (2025)
by: Zhu, Wenqiao, et al.
Published: (2025)
Fact4ac at the Financial Misinformation Detection Challenge Task: Reference-Free Financial Misinformation Detection via Fine-Tuning and Few-Shot Prompting of Large Language Models
by: Hoang, Cuong, et al.
Published: (2026)
by: Hoang, Cuong, et al.
Published: (2026)
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
by: Huang, Zeyu, et al.
Published: (2025)
by: Huang, Zeyu, et al.
Published: (2025)
Parameter-Efficient Fine-Tuning for Medical Text Summarization: A Comparative Study of Lora, Prompt Tuning, and Full Fine-Tuning
by: Shernazarov, Ulugbek, et al.
Published: (2026)
by: Shernazarov, Ulugbek, et al.
Published: (2026)
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
by: Qu, Yuxiao, et al.
Published: (2025)
by: Qu, Yuxiao, et al.
Published: (2025)
EMORL: Ensemble Multi-Objective Reinforcement Learning for Efficient and Flexible LLM Fine-Tuning
by: Kong, Lingxiao, et al.
Published: (2025)
by: Kong, Lingxiao, et al.
Published: (2025)
Public Transit Arrival Prediction: a Seq2Seq RNN Approach
by: Bhutani, Nancy, et al.
Published: (2022)
by: Bhutani, Nancy, et al.
Published: (2022)
Don't Fine-Tune, Decode: Syntax Error-Free Tool Use via Constrained Decoding
by: Zhang, Kexun, et al.
Published: (2023)
by: Zhang, Kexun, et al.
Published: (2023)
American Sign Language to Text Translation using Transformer and Seq2Seq with LSTM
by: Putra, Gregorius Guntur Sunardi, et al.
Published: (2024)
by: Putra, Gregorius Guntur Sunardi, et al.
Published: (2024)
Active Learning-Guided Seq2Seq Variational Autoencoder for Multi-target Inhibitor Generation
by: Vilalta-Mor, Júlia, et al.
Published: (2025)
by: Vilalta-Mor, Júlia, et al.
Published: (2025)
P2LHAP:Wearable sensor-based human activity recognition, segmentation and forecast through Patch-to-Label Seq2Seq Transformer
by: Li, Shuangjian, et al.
Published: (2024)
by: Li, Shuangjian, et al.
Published: (2024)
Natural Language Fine-Tuning
by: Liu, Jia, et al.
Published: (2024)
by: Liu, Jia, et al.
Published: (2024)
ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization
by: Bae, Suyoung, et al.
Published: (2026)
by: Bae, Suyoung, et al.
Published: (2026)
MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues
by: Bai, Ge, et al.
Published: (2024)
by: Bai, Ge, et al.
Published: (2024)
Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models
by: Zhu, Wenhui, et al.
Published: (2025)
by: Zhu, Wenhui, et al.
Published: (2025)
Supervised Fine-Tuning or In-Context Learning? Evaluating LLMs for Clinical NER
by: Baroian, Andrei
Published: (2025)
by: Baroian, Andrei
Published: (2025)
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
by: Fu, Yuqian, et al.
Published: (2025)
by: Fu, Yuqian, et al.
Published: (2025)
Tuning-Free Personalized Alignment via Trial-Error-Explain In-Context Learning
by: Cho, Hyundong, et al.
Published: (2025)
by: Cho, Hyundong, et al.
Published: (2025)
Similar Items
-
Seq2Seq2Seq: Lossless Data Compression via Discrete Latent Transformers and Reinforcement Learning
by: Khodabandeh, Mahdi, et al.
Published: (2026) -
Exploiting the Potential of Seq2Seq Models as Robust Few-Shot Learners
by: Lee, Jihyeon, et al.
Published: (2023) -
Leveraging Graph Structure in Seq2Seq Models for Knowledge Graph Link Prediction
by: Phuc, Luu Huu, et al.
Published: (2026) -
SeqPO-SiMT: Sequential Policy Optimization for Simultaneous Machine Translation
by: Xu, Ting, et al.
Published: (2025) -
Improving Bangla Linguistics: Advanced LSTM, Bi-LSTM, and Seq2Seq Models for Translating Sylheti to Modern Bangla
by: Das, Sourav Kumar, et al.
Published: (2025)