Better Embeddings with Coupled Adam
Fuente:
arXiv
Saved in:
| Main Authors: | Stollenwerk, Felix, Stollenwerk, Tobias |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Mathematical Relationship Between Layer Normalization and Dynamic Activation Functions
by: Stollenwerk, Felix
Published: (2025)
by: Stollenwerk, Felix
Published: (2025)
Output Embedding Centering for Stable LLM Pretraining
by: Stollenwerk, Felix, et al.
Published: (2026)
by: Stollenwerk, Felix, et al.
Published: (2026)
SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training
by: Huang, Tianjin, et al.
Published: (2025)
by: Huang, Tianjin, et al.
Published: (2025)
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
by: Leemann, Tobias, et al.
Published: (2024)
by: Leemann, Tobias, et al.
Published: (2024)
Better LLM Reasoning via Dual-Play
by: Zhang, Zhengxin, et al.
Published: (2025)
by: Zhang, Zhengxin, et al.
Published: (2025)
Lifelong Knowledge Editing requires Better Regularization
by: Gupta, Akshat, et al.
Published: (2025)
by: Gupta, Akshat, et al.
Published: (2025)
Better Alignment with Instruction Back-and-Forth Translation
by: Nguyen, Thao, et al.
Published: (2024)
by: Nguyen, Thao, et al.
Published: (2024)
Verbal Process Supervision Elicits Better Coding Agents
by: Chen, Hao-Yuan, et al.
Published: (2025)
by: Chen, Hao-Yuan, et al.
Published: (2025)
Does Biomedical Training Lead to Better Medical Performance?
by: Dada, Amin, et al.
Published: (2024)
by: Dada, Amin, et al.
Published: (2024)
Instruction-tuned Language Models are Better Knowledge Learners
by: Jiang, Zhengbao, et al.
Published: (2024)
by: Jiang, Zhengbao, et al.
Published: (2024)
Better Estimation of the Kullback--Leibler Divergence Between Language Models
by: Amini, Afra, et al.
Published: (2025)
by: Amini, Afra, et al.
Published: (2025)
R.I.P.: Better Models by Survival of the Fittest Prompts
by: Yu, Ping, et al.
Published: (2025)
by: Yu, Ping, et al.
Published: (2025)
ProgRM: Build Better GUI Agents with Progress Rewards
by: Zhang, Danyang, et al.
Published: (2025)
by: Zhang, Danyang, et al.
Published: (2025)
POSS: Position Specialist Generates Better Draft for Speculative Decoding
by: Huang, Langlin, et al.
Published: (2025)
by: Huang, Langlin, et al.
Published: (2025)
When Does Multimodality Lead to Better Time Series Forecasting?
by: Zhang, Xiyuan, et al.
Published: (2025)
by: Zhang, Xiyuan, et al.
Published: (2025)
Neural Parameter Search for Slimmer Fine-Tuned Models and Better Transfer
by: Du, Guodong, et al.
Published: (2025)
by: Du, Guodong, et al.
Published: (2025)
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
by: Sung, Yi-Lin, et al.
Published: (2025)
by: Sung, Yi-Lin, et al.
Published: (2025)
Improving Large Models with Small models: Lower Costs and Better Performance
by: Chen, Dong, et al.
Published: (2024)
by: Chen, Dong, et al.
Published: (2024)
Teach Better or Show Smarter? On Instructions and Exemplars in Automatic Prompt Optimization
by: Wan, Xingchen, et al.
Published: (2024)
by: Wan, Xingchen, et al.
Published: (2024)
Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation
by: Vu, Tu, et al.
Published: (2024)
by: Vu, Tu, et al.
Published: (2024)
ctELM: Decoding and Manipulating Embeddings of Clinical Trials with Embedding Language Models
by: Ondov, Brian, et al.
Published: (2026)
by: Ondov, Brian, et al.
Published: (2026)
Language Models can Self-Improve at State-Value Estimation for Better Search
by: Mendes, Ethan, et al.
Published: (2025)
by: Mendes, Ethan, et al.
Published: (2025)
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
by: Zhang, Jiazheng, et al.
Published: (2025)
by: Zhang, Jiazheng, et al.
Published: (2025)
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
Are You Sure? Rank Them Again: Repeated Ranking For Better Preference Datasets
by: Devine, Peter
Published: (2024)
by: Devine, Peter
Published: (2024)
Fine-Tuning and Prompt Optimization: Two Great Steps that Work Better Together
by: Soylu, Dilara, et al.
Published: (2024)
by: Soylu, Dilara, et al.
Published: (2024)
QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
by: Tseng, Albert, et al.
Published: (2024)
by: Tseng, Albert, et al.
Published: (2024)
Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves
by: Deng, Yihe, et al.
Published: (2023)
by: Deng, Yihe, et al.
Published: (2023)
BPO: Staying Close to the Behavior LLM Creates Better Online LLM Alignment
by: Xu, Wenda, et al.
Published: (2024)
by: Xu, Wenda, et al.
Published: (2024)
MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language Modeling
by: Limisiewicz, Tomasz, et al.
Published: (2024)
by: Limisiewicz, Tomasz, et al.
Published: (2024)
Understanding LLM Embeddings for Regression
by: Tang, Eric, et al.
Published: (2024)
by: Tang, Eric, et al.
Published: (2024)
ELF: Embedded Language Flows
by: Hu, Keya, et al.
Published: (2026)
by: Hu, Keya, et al.
Published: (2026)
RobustSentEmbed: Robust Sentence Embeddings Using Adversarial Self-Supervised Contrastive Learning
by: Asl, Javad Rafiei, et al.
Published: (2024)
by: Asl, Javad Rafiei, et al.
Published: (2024)
GPT-SW3: An Autoregressive Language Model for the Nordic Languages
by: Ekgren, Ariel, et al.
Published: (2023)
by: Ekgren, Ariel, et al.
Published: (2023)
Enhancing Antibiotic Stewardship using a Natural Language Approach for Better Feature Representation
by: Lee, Simon A., et al.
Published: (2024)
by: Lee, Simon A., et al.
Published: (2024)
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation
by: He, Zhenyu, et al.
Published: (2024)
by: He, Zhenyu, et al.
Published: (2024)
Hakim: Farsi Text Embedding Model
by: Sarmadi, Mehran, et al.
Published: (2025)
by: Sarmadi, Mehran, et al.
Published: (2025)
Evaluating Embedding Frameworks for Scientific Domain
by: Ahmed, Nouman, et al.
Published: (2025)
by: Ahmed, Nouman, et al.
Published: (2025)
Counterfactual Reasoning with Knowledge Graph Embeddings
by: Zellinger, Lena, et al.
Published: (2024)
by: Zellinger, Lena, et al.
Published: (2024)
Word Embeddings Are Steers for Language Models
by: Han, Chi, et al.
Published: (2023)
by: Han, Chi, et al.
Published: (2023)
Similar Items
-
On the Mathematical Relationship Between Layer Normalization and Dynamic Activation Functions
by: Stollenwerk, Felix
Published: (2025) -
Output Embedding Centering for Stable LLM Pretraining
by: Stollenwerk, Felix, et al.
Published: (2026) -
SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training
by: Huang, Tianjin, et al.
Published: (2025) -
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
by: Leemann, Tobias, et al.
Published: (2024) -
Better LLM Reasoning via Dual-Play
by: Zhang, Zhengxin, et al.
Published: (2025)