Guardado en:
| Autores principales: | Mickus, Timothee, Vázquez, Raúl, Attieh, Joseph |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2404.17918 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Isotropy, Clusters, and Classifiers
por: Mickus, Timothee, et al.
Publicado: (2024)
por: Mickus, Timothee, et al.
Publicado: (2024)
Life Cycle-Aware Evaluation of Knowledge Distillation for Machine Translation: Environmental Impact and Translation Quality Trade-offs
por: Attieh, Joseph, et al.
Publicado: (2026)
por: Attieh, Joseph, et al.
Publicado: (2026)
KD4MT: A Survey of Knowledge Distillation for Machine Translation
por: de Gibert, Ona, et al.
Publicado: (2026)
por: de Gibert, Ona, et al.
Publicado: (2026)
MAMMOTH: Massively Multilingual Modular Open Translation @ Helsinki
por: Mickus, Timothee, et al.
Publicado: (2024)
por: Mickus, Timothee, et al.
Publicado: (2024)
Your Model is Overconfident, and Other Lies We Tell Ourselves
por: Mickus, Timothee, et al.
Publicado: (2025)
por: Mickus, Timothee, et al.
Publicado: (2025)
Can Machine Translation Bridge Multilingual Pretraining and Cross-lingual Transfer Learning?
por: Ji, Shaoxiong, et al.
Publicado: (2024)
por: Ji, Shaoxiong, et al.
Publicado: (2024)
A Comparison of Language Modeling and Translation as Multilingual Pretraining Objectives
por: Li, Zihao, et al.
Publicado: (2024)
por: Li, Zihao, et al.
Publicado: (2024)
Adapting Definition Modeling for New Languages: A Case Study on Belarusian
por: Kazakouskaya, Daniela, et al.
Publicado: (2025)
por: Kazakouskaya, Daniela, et al.
Publicado: (2025)
Language Models Learn Universal Representations of Numbers and Here's Why You Should Care
por: Štefánik, Michal, et al.
Publicado: (2025)
por: Štefánik, Michal, et al.
Publicado: (2025)
SemEval-2024 Shared Task 6: SHROOM, a Shared-task on Hallucinations and Related Observable Overgeneration Mistakes
por: Mickus, Timothee, et al.
Publicado: (2024)
por: Mickus, Timothee, et al.
Publicado: (2024)
Domain-specific or Uncertainty-aware models: Does it really make a difference for biomedical text classification?
por: Sinha, Aman, et al.
Publicado: (2024)
por: Sinha, Aman, et al.
Publicado: (2024)
Scaling Low-Resource MT via Synthetic Data Generation with LLMs
por: de Gibert, Ona, et al.
Publicado: (2025)
por: de Gibert, Ona, et al.
Publicado: (2025)
SemEval-2025 Task 3: Mu-SHROOM, the Multilingual Shared Task on Hallucinations and Related Observable Overgeneration Mistakes
por: Vázquez, Raúl, et al.
Publicado: (2025)
por: Vázquez, Raúl, et al.
Publicado: (2025)
Pre-trained Language Models Learn Remarkably Accurate Representations of Numbers
por: Kadlčík, Marek, et al.
Publicado: (2025)
por: Kadlčík, Marek, et al.
Publicado: (2025)
AXOLOTL'24 Shared Task on Multilingual Explainable Semantic Change Modeling
por: Fedorova, Mariia, et al.
Publicado: (2024)
por: Fedorova, Mariia, et al.
Publicado: (2024)
Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
por: Štefánik, Michal, et al.
Publicado: (2025)
por: Štefánik, Michal, et al.
Publicado: (2025)
Sell More, Play Less: Benchmarking LLM Realistic Selling Skill
por: Su, Xuanbo, et al.
Publicado: (2026)
por: Su, Xuanbo, et al.
Publicado: (2026)
Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention
por: Guo, Zhenyu, et al.
Publicado: (2025)
por: Guo, Zhenyu, et al.
Publicado: (2025)
Confabulations from ACL Publications (CAP): A Dataset for Scientific Hallucination Detection
por: Gamba, Federica, et al.
Publicado: (2025)
por: Gamba, Federica, et al.
Publicado: (2025)
Improving Language Transfer Capability of Decoder-only Architecture in Multilingual Neural Machine Translation
por: Qu, Zhi, et al.
Publicado: (2024)
por: Qu, Zhi, et al.
Publicado: (2024)
BridG MT: Enhancing LLMs' Machine Translation Capabilities with Sentence Bridging and Gradual MT
por: Choi, Seung-Woo, et al.
Publicado: (2024)
por: Choi, Seung-Woo, et al.
Publicado: (2024)
What Have We Achieved on Non-autoregressive Translation?
por: Li, Yafu, et al.
Publicado: (2024)
por: Li, Yafu, et al.
Publicado: (2024)
Tower+: Bridging Generality and Translation Specialization in Multilingual LLMs
por: Rei, Ricardo, et al.
Publicado: (2025)
por: Rei, Ricardo, et al.
Publicado: (2025)
What is it for a Machine Learning Model to Have a Capability?
por: Harding, Jacqueline, et al.
Publicado: (2024)
por: Harding, Jacqueline, et al.
Publicado: (2024)
GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models
por: Luo, Hengyu, et al.
Publicado: (2025)
por: Luo, Hengyu, et al.
Publicado: (2025)
Mention Attention for Pronoun Translation
por: Tang, Gongbo, et al.
Publicado: (2024)
por: Tang, Gongbo, et al.
Publicado: (2024)
You Didn't Have to Say It like That: Subliminal Learning from Faithful Paraphrases
por: Gisler, Isaia, et al.
Publicado: (2026)
por: Gisler, Isaia, et al.
Publicado: (2026)
Adding Multimodal Capabilities to a Text-only Translation Model
por: Vijayan, Vipin, et al.
Publicado: (2024)
por: Vijayan, Vipin, et al.
Publicado: (2024)
You Need Better Attention Priors
por: Litman, Elon, et al.
Publicado: (2026)
por: Litman, Elon, et al.
Publicado: (2026)
A Novel Paradigm Boosting Translation Capabilities of Large Language Models
por: Guo, Jiaxin, et al.
Publicado: (2024)
por: Guo, Jiaxin, et al.
Publicado: (2024)
PROST-LLM: Progressively Enhancing the Speech-to-Speech Translation Capability in LLMs
por: Xu, Jing, et al.
Publicado: (2026)
por: Xu, Jing, et al.
Publicado: (2026)
Are We Paying Attention to Her? Investigating Gender Disambiguation and Attention in Machine Translation
por: Manna, Chiara, et al.
Publicado: (2025)
por: Manna, Chiara, et al.
Publicado: (2025)
Proverbs Run in Pairs: Evaluating Proverb Translation Capability of Large Language Model
por: Wang, Minghan, et al.
Publicado: (2025)
por: Wang, Minghan, et al.
Publicado: (2025)
Unlocking Reasoning Capability on Machine Translation in Large Language Models
por: Rajaee, Sara, et al.
Publicado: (2026)
por: Rajaee, Sara, et al.
Publicado: (2026)
How Does Sequence Modeling Architecture Influence Base Capabilities of Pre-trained Language Models? Exploring Key Architecture Design Principles to Avoid Base Capabilities Degradation
por: Lu, Xin, et al.
Publicado: (2025)
por: Lu, Xin, et al.
Publicado: (2025)
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
por: Ko, Hyunwoo, et al.
Publicado: (2025)
por: Ko, Hyunwoo, et al.
Publicado: (2025)
Beyond Isolated Capabilities: Bridging Long CoT Reasoning and Long-Context Understanding
por: Wang, Yifei
Publicado: (2025)
por: Wang, Yifei
Publicado: (2025)
You Have Thirteen Hours in Which to Solve the Labyrinth: Enhancing AI Game Masters with Function Calling
por: Song, Jaewoo, et al.
Publicado: (2024)
por: Song, Jaewoo, et al.
Publicado: (2024)
Your Multimodal Speech Model Says I Have a Face for Radio
por: Nachesa, Maya K., et al.
Publicado: (2026)
por: Nachesa, Maya K., et al.
Publicado: (2026)
We Know I Know You Know; Choreographic Programming With Multicast and Multiply Located Values
por: Bates, Mako, et al.
Publicado: (2024)
por: Bates, Mako, et al.
Publicado: (2024)
Ejemplares similares
-
Isotropy, Clusters, and Classifiers
por: Mickus, Timothee, et al.
Publicado: (2024) -
Life Cycle-Aware Evaluation of Knowledge Distillation for Machine Translation: Environmental Impact and Translation Quality Trade-offs
por: Attieh, Joseph, et al.
Publicado: (2026) -
KD4MT: A Survey of Knowledge Distillation for Machine Translation
por: de Gibert, Ona, et al.
Publicado: (2026) -
MAMMOTH: Massively Multilingual Modular Open Translation @ Helsinki
por: Mickus, Timothee, et al.
Publicado: (2024) -
Your Model is Overconfident, and Other Lies We Tell Ourselves
por: Mickus, Timothee, et al.
Publicado: (2025)