Gespeichert in:
| 1. Verfasser: | Aggarwal, Arpit |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2405.04585 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RePo: Language Models with Context Re-Positioning
von: Li, Huayang, et al.
Veröffentlicht: (2025)
von: Li, Huayang, et al.
Veröffentlicht: (2025)
SeqPE: Transformer with Sequential Position Encoding
von: Li, Huayang, et al.
Veröffentlicht: (2025)
von: Li, Huayang, et al.
Veröffentlicht: (2025)
HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation
von: Chen, Yuhan, et al.
Veröffentlicht: (2024)
von: Chen, Yuhan, et al.
Veröffentlicht: (2024)
Group Representational Position Encoding
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
Polynomial Composition Activations: Unleashing the Dynamics of Large Language Models
von: Zhuo, Zhijian, et al.
Veröffentlicht: (2024)
von: Zhuo, Zhijian, et al.
Veröffentlicht: (2024)
Hierarchical Orthogonal Residual Spread for Precise Massive Editing in Large Language Models
von: Gu, Xiaojie, et al.
Veröffentlicht: (2026)
von: Gu, Xiaojie, et al.
Veröffentlicht: (2026)
Fourier Head: Helping Large Language Models Learn Complex Probability Distributions
von: Gillman, Nate, et al.
Veröffentlicht: (2024)
von: Gillman, Nate, et al.
Veröffentlicht: (2024)
Rotary Positional Embeddings as Phase Modulation: Theoretical Bounds on the RoPE Base for Long-Context Transformers
von: Liu, Feilong
Veröffentlicht: (2026)
von: Liu, Feilong
Veröffentlicht: (2026)
Position Engineering: Boosting Large Language Models through Positional Information Manipulation
von: He, Zhiyuan, et al.
Veröffentlicht: (2024)
von: He, Zhiyuan, et al.
Veröffentlicht: (2024)
RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably
von: Du, Yufeng, et al.
Veröffentlicht: (2026)
von: Du, Yufeng, et al.
Veröffentlicht: (2026)
ORION: Teaching Language Models to Reason Efficiently in the Language of Thought
von: Tanmay, Kumar, et al.
Veröffentlicht: (2025)
von: Tanmay, Kumar, et al.
Veröffentlicht: (2025)
Extracting and Encoding: Leveraging Large Language Models and Medical Knowledge to Enhance Radiological Text Representation
von: Messina, Pablo, et al.
Veröffentlicht: (2024)
von: Messina, Pablo, et al.
Veröffentlicht: (2024)
Group-Aware Reinforcement Learning for Output Diversity in Large Language Models
von: Anschel, Oron, et al.
Veröffentlicht: (2025)
von: Anschel, Oron, et al.
Veröffentlicht: (2025)
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
HoPE: Hyperbolic Rotary Positional Encoding for Stable Long-Range Dependency Modeling in Large Language Models
von: Dai, Chang, et al.
Veröffentlicht: (2025)
von: Dai, Chang, et al.
Veröffentlicht: (2025)
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation
von: He, Zhenyu, et al.
Veröffentlicht: (2024)
von: He, Zhenyu, et al.
Veröffentlicht: (2024)
CoPE: Clipped RoPE as A Scalable Free Lunch for Long Context LLMs
von: Li, Haoran, et al.
Veröffentlicht: (2026)
von: Li, Haoran, et al.
Veröffentlicht: (2026)
MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language Modeling
von: Limisiewicz, Tomasz, et al.
Veröffentlicht: (2024)
von: Limisiewicz, Tomasz, et al.
Veröffentlicht: (2024)
Needle in the Haystack for Memory Based Large Language Models
von: Nelson, Elliot, et al.
Veröffentlicht: (2024)
von: Nelson, Elliot, et al.
Veröffentlicht: (2024)
PoTPTQ: A Two-step Power-of-Two Post-training for LLMs
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
Sequential Large Language Model-Based Hyper-parameter Optimization
von: Mahammadli, Kanan, et al.
Veröffentlicht: (2024)
von: Mahammadli, Kanan, et al.
Veröffentlicht: (2024)
Self-Supervised Position Debiasing for Large Language Models
von: Liu, Zhongkun, et al.
Veröffentlicht: (2024)
von: Liu, Zhongkun, et al.
Veröffentlicht: (2024)
Struc-EMB: The Potential of Structure-Aware Encoding in Language Embeddings
von: Liu, Shikun, et al.
Veröffentlicht: (2025)
von: Liu, Shikun, et al.
Veröffentlicht: (2025)
PoPreRo: A New Dataset for Popularity Prediction of Romanian Reddit Posts
von: Rogoz, Ana-Cristina, et al.
Veröffentlicht: (2024)
von: Rogoz, Ana-Cristina, et al.
Veröffentlicht: (2024)
Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents
von: Kirchhof, Michael, et al.
Veröffentlicht: (2025)
von: Kirchhof, Michael, et al.
Veröffentlicht: (2025)
From Construction to Injection: Edit-Based Fingerprints for Large Language Models
von: Li, Yue, et al.
Veröffentlicht: (2025)
von: Li, Yue, et al.
Veröffentlicht: (2025)
Demystifying the Slash Pattern in Attention: The Role of RoPE
von: Cheng, Yuan, et al.
Veröffentlicht: (2026)
von: Cheng, Yuan, et al.
Veröffentlicht: (2026)
Unraveling LoRA Interference: Orthogonal Subspaces for Robust Model Merging
von: Zhang, Haobo, et al.
Veröffentlicht: (2025)
von: Zhang, Haobo, et al.
Veröffentlicht: (2025)
Method-Based Reasoning for Large Language Models: Extraction, Reuse, and Continuous Improvement
von: Su, Hong
Veröffentlicht: (2025)
von: Su, Hong
Veröffentlicht: (2025)
PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models
von: Xu, Yinggan, et al.
Veröffentlicht: (2025)
von: Xu, Yinggan, et al.
Veröffentlicht: (2025)
CASCADE: Case-Based Continual Adaptation for Large Language Models During Deployment
von: Guo, Siyuan, et al.
Veröffentlicht: (2026)
von: Guo, Siyuan, et al.
Veröffentlicht: (2026)
Large Language Model Pruning
von: Huang, Hanjuan, et al.
Veröffentlicht: (2024)
von: Huang, Hanjuan, et al.
Veröffentlicht: (2024)
Causality for Large Language Models
von: Wu, Anpeng, et al.
Veröffentlicht: (2024)
von: Wu, Anpeng, et al.
Veröffentlicht: (2024)
Large Language Model Unlearning
von: Yao, Yuanshun, et al.
Veröffentlicht: (2023)
von: Yao, Yuanshun, et al.
Veröffentlicht: (2023)
Foundations of Large Language Models
von: Xiao, Tong, et al.
Veröffentlicht: (2025)
von: Xiao, Tong, et al.
Veröffentlicht: (2025)
Large Language Models as Optimizers
von: Yang, Chengrun, et al.
Veröffentlicht: (2023)
von: Yang, Chengrun, et al.
Veröffentlicht: (2023)
Orthogonal Finetuning for Direct Preference Optimization
von: Yang, Chenxu, et al.
Veröffentlicht: (2024)
von: Yang, Chenxu, et al.
Veröffentlicht: (2024)
The Geometries of Truth Are Orthogonal Across Tasks
von: Azizian, Waiss, et al.
Veröffentlicht: (2025)
von: Azizian, Waiss, et al.
Veröffentlicht: (2025)
BAPO: Base-Anchored Preference Optimization for Overcoming Forgetting in Large Language Models Personalization
von: Lee, Gihun, et al.
Veröffentlicht: (2024)
von: Lee, Gihun, et al.
Veröffentlicht: (2024)
SEE: Sememe Entanglement Encoding for Transformer-bases Models Compression
von: Zhang, Jing, et al.
Veröffentlicht: (2024)
von: Zhang, Jing, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
RePo: Language Models with Context Re-Positioning
von: Li, Huayang, et al.
Veröffentlicht: (2025) -
SeqPE: Transformer with Sequential Position Encoding
von: Li, Huayang, et al.
Veröffentlicht: (2025) -
HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation
von: Chen, Yuhan, et al.
Veröffentlicht: (2024) -
Group Representational Position Encoding
von: Zhang, Yifan, et al.
Veröffentlicht: (2025) -
Polynomial Composition Activations: Unleashing the Dynamics of Large Language Models
von: Zhuo, Zhijian, et al.
Veröffentlicht: (2024)