Improving Rare Word Translation With Dictionaries and Attention Masking
Fuente:
arXiv
Salvato in:
| Autori principali: | Sible, Kenneth J., Chiang, David |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Using Source-Side Confidence Estimation for Reliable Translation into Unfamiliar Languages
di: Sible, Kenneth J., et al.
Pubblicazione: (2025)
di: Sible, Kenneth J., et al.
Pubblicazione: (2025)
Masked Hard-Attention Transformers Recognize Exactly the Star-Free Languages
di: Yang, Andy, et al.
Pubblicazione: (2023)
di: Yang, Andy, et al.
Pubblicazione: (2023)
DICTDIS: Dictionary Constrained Disambiguation for Improved NMT
di: Maheshwari, Ayush, et al.
Pubblicazione: (2022)
di: Maheshwari, Ayush, et al.
Pubblicazione: (2022)
Simulating Hard Attention Using Soft Attention
di: Yang, Andy, et al.
Pubblicazione: (2024)
di: Yang, Andy, et al.
Pubblicazione: (2024)
SEA: Sparse Linear Attention with Estimated Attention Mask
di: Lee, Heejun, et al.
Pubblicazione: (2023)
di: Lee, Heejun, et al.
Pubblicazione: (2023)
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
di: Wang, Junxuan, et al.
Pubblicazione: (2025)
di: Wang, Junxuan, et al.
Pubblicazione: (2025)
Nostra Domina at EvaLatin 2024: Improving Latin Polarity Detection through Data Augmentation
di: Bothwell, Stephen, et al.
Pubblicazione: (2024)
di: Bothwell, Stephen, et al.
Pubblicazione: (2024)
Improving Word Translation via Two-Stage Contrastive Learning
di: Li, Yaoyiran, et al.
Pubblicazione: (2022)
di: Li, Yaoyiran, et al.
Pubblicazione: (2022)
Features that Make a Difference: Leveraging Gradients for Improved Dictionary Learning
di: Olmo, Jeffrey, et al.
Pubblicazione: (2024)
di: Olmo, Jeffrey, et al.
Pubblicazione: (2024)
The Rotary Position Embedding May Cause Dimension Inefficiency in Attention Heads for Long-Distance Retrieval
di: Chiang, Ting-Rui, et al.
Pubblicazione: (2025)
di: Chiang, Ting-Rui, et al.
Pubblicazione: (2025)
Faster Transformer Decoding: N-gram Masked Self-Attention
di: Chelba, Ciprian, et al.
Pubblicazione: (2020)
di: Chelba, Ciprian, et al.
Pubblicazione: (2020)
Pay Less Attention to Function Words for Free Robustness of Vision-Language Models
di: Tian, Qiwei, et al.
Pubblicazione: (2025)
di: Tian, Qiwei, et al.
Pubblicazione: (2025)
Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More
di: Zhuang, Xialie, et al.
Pubblicazione: (2025)
di: Zhuang, Xialie, et al.
Pubblicazione: (2025)
Edeflip: Supervised Word Translation between English and Yoruba
di: Abioye, Ikeoluwa, et al.
Pubblicazione: (2025)
di: Abioye, Ikeoluwa, et al.
Pubblicazione: (2025)
Trainable Dynamic Mask Sparse Attention
di: Shi, Jingze, et al.
Pubblicazione: (2025)
di: Shi, Jingze, et al.
Pubblicazione: (2025)
Efficiently Dispatching Flash Attention For Partially Filled Attention Masks
di: Sharma, Agniv, et al.
Pubblicazione: (2024)
di: Sharma, Agniv, et al.
Pubblicazione: (2024)
Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
di: Zhussip, Magauiya, et al.
Pubblicazione: (2025)
di: Zhussip, Magauiya, et al.
Pubblicazione: (2025)
Beyond Shared Vocabulary: Increasing Representational Word Similarities across Languages for Multilingual Machine Translation
di: Wu, Di, et al.
Pubblicazione: (2023)
di: Wu, Di, et al.
Pubblicazione: (2023)
Towards Understanding the Word Sensitivity of Attention Layers: A Study via Random Features
di: Bombari, Simone, et al.
Pubblicazione: (2024)
di: Bombari, Simone, et al.
Pubblicazione: (2024)
Simultaneous Masking, Not Prompting Optimization: A Paradigm Shift in Fine-tuning LLMs for Simultaneous Translation
di: Raffel, Matthew, et al.
Pubblicazione: (2024)
di: Raffel, Matthew, et al.
Pubblicazione: (2024)
On Translating Technical Terminology: A Translation Workflow for Machine-Translated Acronyms
di: Yue, Richard, et al.
Pubblicazione: (2024)
di: Yue, Richard, et al.
Pubblicazione: (2024)
An Improved Deep Learning Model for Word Embeddings Based Clustering for Large Text Datasets
di: Sutrakar, Vijay Kumar, et al.
Pubblicazione: (2025)
di: Sutrakar, Vijay Kumar, et al.
Pubblicazione: (2025)
Attention Sinks in Massively Multilingual Neural Machine Translation:Discovery, Analysis, and Mitigation
di: Mutisya, Hillary, et al.
Pubblicazione: (2026)
di: Mutisya, Hillary, et al.
Pubblicazione: (2026)
Improving Transformers with Dynamically Composable Multi-Head Attention
di: Xiao, Da, et al.
Pubblicazione: (2024)
di: Xiao, Da, et al.
Pubblicazione: (2024)
Self-Augmented In-Context Learning for Unsupervised Word Translation
di: Li, Yaoyiran, et al.
Pubblicazione: (2024)
di: Li, Yaoyiran, et al.
Pubblicazione: (2024)
Improving Text Style Transfer using Masked Diffusion Language Models with Inference-time Scaling
di: Padole, Tejomay Kishor, et al.
Pubblicazione: (2025)
di: Padole, Tejomay Kishor, et al.
Pubblicazione: (2025)
Geometry-Adaptive Explainer for Faithful Dictionary-Based Interpretability under Distribution Shift
di: Lim, Sungjun, et al.
Pubblicazione: (2026)
di: Lim, Sungjun, et al.
Pubblicazione: (2026)
LogicDiff: Logic-Guided Denoising Improves Zero-Shot Reasoning in Masked Diffusion Language Models
di: Aman, Shaik
Pubblicazione: (2026)
di: Aman, Shaik
Pubblicazione: (2026)
Masked Thought: Simply Masking Partial Reasoning Steps Can Improve Mathematical Reasoning Learning of Language Models
di: Chen, Changyu, et al.
Pubblicazione: (2024)
di: Chen, Changyu, et al.
Pubblicazione: (2024)
OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention
di: Zhou, Chenyu, et al.
Pubblicazione: (2026)
di: Zhou, Chenyu, et al.
Pubblicazione: (2026)
Selective Attention Improves Transformer
di: Leviathan, Yaniv, et al.
Pubblicazione: (2024)
di: Leviathan, Yaniv, et al.
Pubblicazione: (2024)
SWE2: SubWord Enriched and Significant Word Emphasized Framework for Hate Speech Detection
di: Mou, Guanyi, et al.
Pubblicazione: (2024)
di: Mou, Guanyi, et al.
Pubblicazione: (2024)
Transformers in Uniform TC$^0$
di: Chiang, David
Pubblicazione: (2024)
di: Chiang, David
Pubblicazione: (2024)
Wikipedia is Not a Dictionary, Delete! Text Classification as a Proxy for Analysing Wiki Deletion Discussions
di: Borkakoty, Hsuvas, et al.
Pubblicazione: (2025)
di: Borkakoty, Hsuvas, et al.
Pubblicazione: (2025)
Counting Like Transformers: Compiling Temporal Counting Logic Into Softmax Transformers
di: Yang, Andy, et al.
Pubblicazione: (2024)
di: Yang, Andy, et al.
Pubblicazione: (2024)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
di: Gu, Zhuohan, et al.
Pubblicazione: (2024)
di: Gu, Zhuohan, et al.
Pubblicazione: (2024)
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
di: Karvonen, Adam, et al.
Pubblicazione: (2024)
di: Karvonen, Adam, et al.
Pubblicazione: (2024)
Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders
di: Zheng, Carolina, et al.
Pubblicazione: (2025)
di: Zheng, Carolina, et al.
Pubblicazione: (2025)
Task-Informed Anti-Curriculum by Masking Improves Downstream Performance on Text
di: Jarca, Andrei, et al.
Pubblicazione: (2025)
di: Jarca, Andrei, et al.
Pubblicazione: (2025)
Low-Cost Generation and Evaluation of Dictionary Example Sentences
di: Cai, Bill, et al.
Pubblicazione: (2024)
di: Cai, Bill, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Using Source-Side Confidence Estimation for Reliable Translation into Unfamiliar Languages
di: Sible, Kenneth J., et al.
Pubblicazione: (2025) -
Masked Hard-Attention Transformers Recognize Exactly the Star-Free Languages
di: Yang, Andy, et al.
Pubblicazione: (2023) -
DICTDIS: Dictionary Constrained Disambiguation for Improved NMT
di: Maheshwari, Ayush, et al.
Pubblicazione: (2022) -
Simulating Hard Attention Using Soft Attention
di: Yang, Andy, et al.
Pubblicazione: (2024) -
SEA: Sparse Linear Attention with Estimated Attention Mask
di: Lee, Heejun, et al.
Pubblicazione: (2023)