Salvato in:
| Autori principali: | Mąka, Paweł, Semerci, Yusuf Can, Scholtes, Jan, Spanakis, Gerasimos |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2509.14031 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
di: Mąka, Paweł, et al.
Pubblicazione: (2024)
di: Mąka, Paweł, et al.
Pubblicazione: (2024)
Sequence Shortening for Context-Aware Machine Translation
di: Mąka, Paweł, et al.
Pubblicazione: (2024)
di: Mąka, Paweł, et al.
Pubblicazione: (2024)
Cross-Modal Robustness Transfer (CMRT): Training Robust Speech Translation Models Using Adversarial Text
di: Issam, Abderrahmane, et al.
Pubblicazione: (2026)
di: Issam, Abderrahmane, et al.
Pubblicazione: (2026)
Fixed and Adaptive Simultaneous Machine Translation Strategies Using Adapters
di: Issam, Abderrahmane, et al.
Pubblicazione: (2024)
di: Issam, Abderrahmane, et al.
Pubblicazione: (2024)
DTW-Align: Bridging the Modality Gap in End-to-End Speech Translation with Dynamic Time Warping Alignment
di: Issam, Abderrahmane, et al.
Pubblicazione: (2025)
di: Issam, Abderrahmane, et al.
Pubblicazione: (2025)
Language Models as Artificial Learners: Investigating Crosslinguistic Influence
di: Issam, Abderrahmane, et al.
Pubblicazione: (2026)
di: Issam, Abderrahmane, et al.
Pubblicazione: (2026)
A Representation Level Analysis of NMT Model Robustness to Grammatical Errors
di: Issam, Abderrahmane, et al.
Pubblicazione: (2025)
di: Issam, Abderrahmane, et al.
Pubblicazione: (2025)
Navigating WebAI: Training Agents to Complete Web Tasks with Large Language Models and Reinforcement Learning
di: Thil, Lucas-Andreï, et al.
Pubblicazione: (2024)
di: Thil, Lucas-Andreï, et al.
Pubblicazione: (2024)
More Compute Is What You Need
di: Guo, Zhen
Pubblicazione: (2024)
di: Guo, Zhen
Pubblicazione: (2024)
DIDS: Domain Impact-aware Data Sampling for Large Language Model Training
di: Shi, Weijie, et al.
Pubblicazione: (2025)
di: Shi, Weijie, et al.
Pubblicazione: (2025)
Did You Forget What I Asked? Prospective Memory Failures in Large Language Models
di: Mittal, Avni
Pubblicazione: (2026)
di: Mittal, Avni
Pubblicazione: (2026)
What Would You Ask When You First Saw $a^2+b^2=c^2$? Evaluating LLM on Curiosity-Driven Questioning
di: Javaji, Shashidhar Reddy, et al.
Pubblicazione: (2024)
di: Javaji, Shashidhar Reddy, et al.
Pubblicazione: (2024)
Automated Multi-Language to English Machine Translation Using Generative Pre-Trained Transformers
di: Pelofske, Elijah, et al.
Pubblicazione: (2024)
di: Pelofske, Elijah, et al.
Pubblicazione: (2024)
Forget What You Know about LLMs Evaluations -- LLMs are Like a Chameleon
di: Cohen-Inger, Nurit, et al.
Pubblicazione: (2025)
di: Cohen-Inger, Nurit, et al.
Pubblicazione: (2025)
Planning and Editing What You Retrieve for Enhanced Tool Learning
di: Huang, Tenghao, et al.
Pubblicazione: (2024)
di: Huang, Tenghao, et al.
Pubblicazione: (2024)
Synthetic Data RL: Task Definition Is All You Need
di: Guo, Yiduo, et al.
Pubblicazione: (2025)
di: Guo, Yiduo, et al.
Pubblicazione: (2025)
Regurgitative Training: The Value of Real Data in Training Large Language Models
di: Zhang, Jinghui, et al.
Pubblicazione: (2024)
di: Zhang, Jinghui, et al.
Pubblicazione: (2024)
You Only Train Once: Differentiable Subset Selection for Omics Data
di: Chopard, Daphné, et al.
Pubblicazione: (2025)
di: Chopard, Daphné, et al.
Pubblicazione: (2025)
More Agents Is All You Need
di: Li, Junyou, et al.
Pubblicazione: (2024)
di: Li, Junyou, et al.
Pubblicazione: (2024)
Mitigating Reversal Curse in Large Language Models via Semantic-aware Permutation Training
di: Guo, Qingyan, et al.
Pubblicazione: (2024)
di: Guo, Qingyan, et al.
Pubblicazione: (2024)
ToolACE-R: Model-aware Iterative Training and Adaptive Refinement for Tool Learning
di: Zeng, Xingshan, et al.
Pubblicazione: (2025)
di: Zeng, Xingshan, et al.
Pubblicazione: (2025)
APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
di: Guan, Ziyi, et al.
Pubblicazione: (2024)
di: Guan, Ziyi, et al.
Pubblicazione: (2024)
Many-to-English Machine Translation Tools, Data, and Pretrained Models
di: Gowda, Thamme, et al.
Pubblicazione: (2021)
di: Gowda, Thamme, et al.
Pubblicazione: (2021)
You Can Generate It Again: Data-to-Text Generation with Verification and Correction Prompting
di: Ren, Xuan, et al.
Pubblicazione: (2023)
di: Ren, Xuan, et al.
Pubblicazione: (2023)
Reasoning with Sampling: Your Base Model is Smarter Than You Think
di: Karan, Aayush, et al.
Pubblicazione: (2025)
di: Karan, Aayush, et al.
Pubblicazione: (2025)
Should You Use Your Large Language Model to Explore or Exploit?
di: Harris, Keegan, et al.
Pubblicazione: (2025)
di: Harris, Keegan, et al.
Pubblicazione: (2025)
Tensor Product Attention Is All You Need
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
LoRA Is Slower Than You Think
di: Ko, Seokmin
Pubblicazione: (2025)
di: Ko, Seokmin
Pubblicazione: (2025)
Attention Smoothing Is All You Need For Unlearning
di: Zade, Saleh Zare, et al.
Pubblicazione: (2026)
di: Zade, Saleh Zare, et al.
Pubblicazione: (2026)
Fast Training Dataset Attribution via In-Context Learning
di: Fotouhi, Milad, et al.
Pubblicazione: (2024)
di: Fotouhi, Milad, et al.
Pubblicazione: (2024)
Data Efficacy for Language Model Training
di: Dai, Yalun, et al.
Pubblicazione: (2025)
di: Dai, Yalun, et al.
Pubblicazione: (2025)
Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training
di: Kesgin, H. Toprak, et al.
Pubblicazione: (2024)
di: Kesgin, H. Toprak, et al.
Pubblicazione: (2024)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
di: Xiao, Chaojun, et al.
Pubblicazione: (2024)
di: Xiao, Chaojun, et al.
Pubblicazione: (2024)
Guidance is All You Need: Temperature-Guided Reasoning in Large Language Models
di: Gomaa, Eyad, et al.
Pubblicazione: (2024)
di: Gomaa, Eyad, et al.
Pubblicazione: (2024)
SynthDST: Synthetic Data is All You Need for Few-Shot Dialog State Tracking
di: Kulkarni, Atharva, et al.
Pubblicazione: (2024)
di: Kulkarni, Atharva, et al.
Pubblicazione: (2024)
SAEs Are Good for Steering -- If You Select the Right Features
di: Arad, Dana, et al.
Pubblicazione: (2025)
di: Arad, Dana, et al.
Pubblicazione: (2025)
Training and Evaluating Language Models with Template-based Data Generation
di: Zhang, Yifan
Pubblicazione: (2024)
di: Zhang, Yifan
Pubblicazione: (2024)
Does Training on Synthetic Data Make Models Less Robust?
di: Zhang, Lingze, et al.
Pubblicazione: (2025)
di: Zhang, Lingze, et al.
Pubblicazione: (2025)
Is Child-Directed Speech Effective Training Data for Language Models?
di: Feng, Steven Y., et al.
Pubblicazione: (2024)
di: Feng, Steven Y., et al.
Pubblicazione: (2024)
Optimisation Is Not What You Need
di: Ibias, Alfredo
Pubblicazione: (2025)
di: Ibias, Alfredo
Pubblicazione: (2025)
Documenti analoghi
-
Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
di: Mąka, Paweł, et al.
Pubblicazione: (2024) -
Sequence Shortening for Context-Aware Machine Translation
di: Mąka, Paweł, et al.
Pubblicazione: (2024) -
Cross-Modal Robustness Transfer (CMRT): Training Robust Speech Translation Models Using Adversarial Text
di: Issam, Abderrahmane, et al.
Pubblicazione: (2026) -
Fixed and Adaptive Simultaneous Machine Translation Strategies Using Adapters
di: Issam, Abderrahmane, et al.
Pubblicazione: (2024) -
DTW-Align: Bridging the Modality Gap in End-to-End Speech Translation with Dynamic Time Warping Alignment
di: Issam, Abderrahmane, et al.
Pubblicazione: (2025)