Efficient transformer with reinforced position embedding for language models
Fuente:
arXiv
Salvato in:
| Autori principali: | Hsiao, Yen-Che, Dutta, Abhishek |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Towards Autonomous Agents: Adaptive-planning, Reasoning, and Acting in Language Models
di: Dutta, Abhishek, et al.
Pubblicazione: (2024)
di: Dutta, Abhishek, et al.
Pubblicazione: (2024)
Unveiling Reasoning Thresholds in Language Models: Scaling, Fine-Tuning, and Interpretability through Attention Maps
di: Hsiao, Yen-Che, et al.
Pubblicazione: (2025)
di: Hsiao, Yen-Che, et al.
Pubblicazione: (2025)
Hybrid Coordinate Descent for Efficient Neural Network Learning Using Line Search and Gradient Descent
di: Hsiao, Yen-Che, et al.
Pubblicazione: (2024)
di: Hsiao, Yen-Che, et al.
Pubblicazione: (2024)
Adaptive Reasoning and Acting in Medical Language Agents
di: Dutta, Abhishek, et al.
Pubblicazione: (2024)
di: Dutta, Abhishek, et al.
Pubblicazione: (2024)
Enhanced Arabic-language cyberbullying detection: deep embedding and transformer (BERT) approaches
di: Aljohani, Ebtesam Jaber, et al.
Pubblicazione: (2025)
di: Aljohani, Ebtesam Jaber, et al.
Pubblicazione: (2025)
Why transformers are obviously good models of language
di: Hill, Felix
Pubblicazione: (2024)
di: Hill, Felix
Pubblicazione: (2024)
Interpreting the linear structure of vision-language model embedding spaces
di: Papadimitriou, Isabel, et al.
Pubblicazione: (2025)
di: Papadimitriou, Isabel, et al.
Pubblicazione: (2025)
Prompt reinforcing for long-term planning of large language models
di: Lin, Hsien-Chin, et al.
Pubblicazione: (2025)
di: Lin, Hsien-Chin, et al.
Pubblicazione: (2025)
Vocabulary embeddings organize linguistic structure early in language model training
di: Papadimitriou, Isabel, et al.
Pubblicazione: (2025)
di: Papadimitriou, Isabel, et al.
Pubblicazione: (2025)
A semantic embedding space based on large language models for modelling human beliefs
di: Lee, Byunghwee, et al.
Pubblicazione: (2024)
di: Lee, Byunghwee, et al.
Pubblicazione: (2024)
NLD-LLM: A systematic framework for evaluating small language transformer models on natural language description
di: Jelodar, Hamed, et al.
Pubblicazione: (2025)
di: Jelodar, Hamed, et al.
Pubblicazione: (2025)
Derivation of Back-propagation for Graph Convolutional Networks using Matrix Calculus and its Application to Explainable Artificial Intelligence
di: Hsiao, Yen-Che, et al.
Pubblicazione: (2024)
di: Hsiao, Yen-Che, et al.
Pubblicazione: (2024)
Lost without translation -- Can transformer (language models) understand mood states?
di: Shivaprakash, Prakrithi, et al.
Pubblicazione: (2025)
di: Shivaprakash, Prakrithi, et al.
Pubblicazione: (2025)
Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
di: Raposo, David, et al.
Pubblicazione: (2024)
di: Raposo, David, et al.
Pubblicazione: (2024)
Survey on reinforcement learning for language processing
di: Uc-Cetina, Victor, et al.
Pubblicazione: (2021)
di: Uc-Cetina, Victor, et al.
Pubblicazione: (2021)
Humans and transformer LMs: Abstraction drives language learning
di: Jian, Jasper, et al.
Pubblicazione: (2026)
di: Jian, Jasper, et al.
Pubblicazione: (2026)
Physical models realizing the transformer architecture of large language models
di: Chen, Zeqian
Pubblicazione: (2025)
di: Chen, Zeqian
Pubblicazione: (2025)
Comparison of different Unique hard attention transformer models by the formal languages they can recognize
di: Ryvkin, Leonid
Pubblicazione: (2025)
di: Ryvkin, Leonid
Pubblicazione: (2025)
A comparative study of transformer-based embeddings for topic coherence
di: Ding, Alex, et al.
Pubblicazione: (2026)
di: Ding, Alex, et al.
Pubblicazione: (2026)
Human-like fleeting memory improves language learning but impairs reading time prediction in transformer language models
di: Thamma, Abishek, et al.
Pubblicazione: (2025)
di: Thamma, Abishek, et al.
Pubblicazione: (2025)
Spatio-temporal transformer to support automatic sign language translation
di: Ruiz, Christian, et al.
Pubblicazione: (2025)
di: Ruiz, Christian, et al.
Pubblicazione: (2025)
Multilingual acoustic word embeddings for zero-resource languages
di: Jacobs, Christiaan
Pubblicazione: (2024)
di: Jacobs, Christiaan
Pubblicazione: (2024)
Med-gte-hybrid: A contextual embedding transformer model for extracting actionable information from clinical texts
di: Kumar, Aditya, et al.
Pubblicazione: (2025)
di: Kumar, Aditya, et al.
Pubblicazione: (2025)
A path to natural language through tokenisation and transformers
di: Berman, David S., et al.
Pubblicazione: (2026)
di: Berman, David S., et al.
Pubblicazione: (2026)
Efficient argument classification with compact language models and ChatGPT-4 refinements
di: Pietron, Marcin, et al.
Pubblicazione: (2024)
di: Pietron, Marcin, et al.
Pubblicazione: (2024)
'Neural howlround' in large language models: a self-reinforcing bias phenomenon, and a dynamic attenuation solution
di: Drake, Seth
Pubblicazione: (2025)
di: Drake, Seth
Pubblicazione: (2025)
Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
di: Albadarneh, Israa A., et al.
Pubblicazione: (2025)
di: Albadarneh, Israa A., et al.
Pubblicazione: (2025)
Can large language models understand uncommon meanings of common words?
di: Wu, Jinyang, et al.
Pubblicazione: (2024)
di: Wu, Jinyang, et al.
Pubblicazione: (2024)
A multitask transformer to sign language translation using motion gesture primitives
di: López, Fredy Alejandro Mendoza, et al.
Pubblicazione: (2025)
di: López, Fredy Alejandro Mendoza, et al.
Pubblicazione: (2025)
Large language models have learned to use language
di: Lupyan, Gary
Pubblicazione: (2025)
di: Lupyan, Gary
Pubblicazione: (2025)
Tiny language models
di: Gross, Ronit D., et al.
Pubblicazione: (2025)
di: Gross, Ronit D., et al.
Pubblicazione: (2025)
Detecting out-of-distribution text using topological features of transformer-based language models
di: Pollano, Andres, et al.
Pubblicazione: (2023)
di: Pollano, Andres, et al.
Pubblicazione: (2023)
Can we teach language models to gloss endangered languages?
di: Ginn, Michael, et al.
Pubblicazione: (2024)
di: Ginn, Michael, et al.
Pubblicazione: (2024)
Retrieval augmentation of large language models for lay language generation
di: Guo, Yue, et al.
Pubblicazione: (2022)
di: Guo, Yue, et al.
Pubblicazione: (2022)
Do large language models resemble humans in language use?
di: Cai, Zhenguang G., et al.
Pubblicazione: (2023)
di: Cai, Zhenguang G., et al.
Pubblicazione: (2023)
Studies with impossible languages falsify LMs as models of human language
di: Bowers, Jeffrey S., et al.
Pubblicazione: (2025)
di: Bowers, Jeffrey S., et al.
Pubblicazione: (2025)
Generalist embedding models are better at short-context clinical semantic search than specialized embedding models
di: Excoffier, Jean-Baptiste, et al.
Pubblicazione: (2024)
di: Excoffier, Jean-Baptiste, et al.
Pubblicazione: (2024)
Why do language models perform worse for morphologically complex languages?
di: Arnett, Catherine, et al.
Pubblicazione: (2024)
di: Arnett, Catherine, et al.
Pubblicazione: (2024)
Prompting language influences diagnostic reasoning and accuracy of large language models
di: Bazoge, Adrien, et al.
Pubblicazione: (2026)
di: Bazoge, Adrien, et al.
Pubblicazione: (2026)
A systematic review of geospatial location embedding approaches in large language models: A path to spatial AI systems
di: Tucker, Sean
Pubblicazione: (2024)
di: Tucker, Sean
Pubblicazione: (2024)
Documenti analoghi
-
Towards Autonomous Agents: Adaptive-planning, Reasoning, and Acting in Language Models
di: Dutta, Abhishek, et al.
Pubblicazione: (2024) -
Unveiling Reasoning Thresholds in Language Models: Scaling, Fine-Tuning, and Interpretability through Attention Maps
di: Hsiao, Yen-Che, et al.
Pubblicazione: (2025) -
Hybrid Coordinate Descent for Efficient Neural Network Learning Using Line Search and Gradient Descent
di: Hsiao, Yen-Che, et al.
Pubblicazione: (2024) -
Adaptive Reasoning and Acting in Medical Language Agents
di: Dutta, Abhishek, et al.
Pubblicazione: (2024) -
Enhanced Arabic-language cyberbullying detection: deep embedding and transformer (BERT) approaches
di: Aljohani, Ebtesam Jaber, et al.
Pubblicazione: (2025)