Gespeichert in:
| Hauptverfasser: | Yang, Yijun, Huang, Zeyu, Zhu, Wenhao, Qiu, Zihan, Yuan, Fei, Pan, Jeff Z., Titov, Ivan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2506.02921 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Post-hoc Reward Calibration: A Case Study on Length Bias
von: Huang, Zeyu, et al.
Veröffentlicht: (2024)
von: Huang, Zeyu, et al.
Veröffentlicht: (2024)
Layerwise Recurrent Router for Mixture-of-Experts
von: Qiu, Zihan, et al.
Veröffentlicht: (2024)
von: Qiu, Zihan, et al.
Veröffentlicht: (2024)
Masking in Multi-hop QA: An Analysis of How Language Models Perform with Context Permutation
von: Huang, Wenyu, et al.
Veröffentlicht: (2025)
von: Huang, Wenyu, et al.
Veröffentlicht: (2025)
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
von: Huang, Zeyu, et al.
Veröffentlicht: (2025)
von: Huang, Zeyu, et al.
Veröffentlicht: (2025)
Controlling What You Share: Assessing Language Model Adherence to Privacy Preferences
von: Ramírez, Guillem, et al.
Veröffentlicht: (2025)
von: Ramírez, Guillem, et al.
Veröffentlicht: (2025)
Evaluating and Improving Graph to Text Generation with Large Language Models
von: He, Jie, et al.
Veröffentlicht: (2025)
von: He, Jie, et al.
Veröffentlicht: (2025)
Generalizing From Short to Long: Effective Data Synthesis for Long-Context Instruction Tuning
von: Zhu, Wenhao, et al.
Veröffentlicht: (2025)
von: Zhu, Wenhao, et al.
Veröffentlicht: (2025)
Unlearning Traces the Influential Training Data of Language Models
von: Isonuma, Masaru, et al.
Veröffentlicht: (2024)
von: Isonuma, Masaru, et al.
Veröffentlicht: (2024)
Training-Free Long-Context Scaling of Large Language Models
von: An, Chenxin, et al.
Veröffentlicht: (2024)
von: An, Chenxin, et al.
Veröffentlicht: (2024)
Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
von: Qiu, Zihan, et al.
Veröffentlicht: (2025)
von: Qiu, Zihan, et al.
Veröffentlicht: (2025)
CCF: A Context Compression Framework for Efficient Long-Sequence Language Modeling
von: Li, Wenhao, et al.
Veröffentlicht: (2025)
von: Li, Wenhao, et al.
Veröffentlicht: (2025)
Generalisation First, Memorisation Second? Memorisation Localisation for Natural Language Classification Tasks
von: Dankers, Verna, et al.
Veröffentlicht: (2024)
von: Dankers, Verna, et al.
Veröffentlicht: (2024)
A Closer Look into Mixture-of-Experts in Large Language Models
von: Lo, Ka Man, et al.
Veröffentlicht: (2024)
von: Lo, Ka Man, et al.
Veröffentlicht: (2024)
Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection
von: Ramírez, Guillem, et al.
Veröffentlicht: (2024)
von: Ramírez, Guillem, et al.
Veröffentlicht: (2024)
Prompting Large Language Models with Knowledge Graphs for Question Answering Involving Long-tail Facts
von: Huang, Wenyu, et al.
Veröffentlicht: (2024)
von: Huang, Wenyu, et al.
Veröffentlicht: (2024)
Mitigating Copy Bias in In-Context Learning through Neuron Pruning
von: Ali, Ameen, et al.
Veröffentlicht: (2024)
von: Ali, Ameen, et al.
Veröffentlicht: (2024)
Empirical Study on Updating Key-Value Memories in Transformer Feed-forward Layers
von: Qiu, Zihan, et al.
Veröffentlicht: (2024)
von: Qiu, Zihan, et al.
Veröffentlicht: (2024)
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
von: Huang, Xu, et al.
Veröffentlicht: (2025)
von: Huang, Xu, et al.
Veröffentlicht: (2025)
UniArk: Improving Generalisation and Consistency for Factual Knowledge Extraction through Debiasing
von: Yang, Yijun, et al.
Veröffentlicht: (2024)
von: Yang, Yijun, et al.
Veröffentlicht: (2024)
LongEmbed: Extending Embedding Models for Long Context Retrieval
von: Zhu, Dawei, et al.
Veröffentlicht: (2024)
von: Zhu, Dawei, et al.
Veröffentlicht: (2024)
A Comprehensive Survey on Long Context Language Modeling
von: Liu, Jiaheng, et al.
Veröffentlicht: (2025)
von: Liu, Jiaheng, et al.
Veröffentlicht: (2025)
DZ-TDPO: Non-Destructive Temporal Alignment for Mutable State Tracking in Long-Context Dialogue
von: Liao, Yijun
Veröffentlicht: (2025)
von: Liao, Yijun
Veröffentlicht: (2025)
Shared Doubt: Zero-shot Cross-Lingual Confidence Estimation for Language Models
von: Kyriakou, Athina, et al.
Veröffentlicht: (2026)
von: Kyriakou, Athina, et al.
Veröffentlicht: (2026)
Enhancing Long Document Long Form Summarisation with Self-Planning
von: Du, Xiaotang, et al.
Veröffentlicht: (2025)
von: Du, Xiaotang, et al.
Veröffentlicht: (2025)
Anthropomimetic Uncertainty: What Verbalized Uncertainty in Language Models is Missing
von: Ulmer, Dennis, et al.
Veröffentlicht: (2025)
von: Ulmer, Dennis, et al.
Veröffentlicht: (2025)
Detecting and Pruning Prominent but Detrimental Neurons in Large Language Models
von: Ali, Ameen, et al.
Veröffentlicht: (2025)
von: Ali, Ameen, et al.
Veröffentlicht: (2025)
Cache & Distil: Optimising API Calls to Large Language Models
von: Ramírez, Guillem, et al.
Veröffentlicht: (2023)
von: Ramírez, Guillem, et al.
Veröffentlicht: (2023)
SPELL: Self-Play Reinforcement Learning for Evolving Long-Context Language Models
von: Yang, Ziyi, et al.
Veröffentlicht: (2025)
von: Yang, Ziyi, et al.
Veröffentlicht: (2025)
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
von: Qiu, Zihan, et al.
Veröffentlicht: (2025)
von: Qiu, Zihan, et al.
Veröffentlicht: (2025)
RULER: What's the Real Context Size of Your Long-Context Language Models?
von: Hsieh, Cheng-Ping, et al.
Veröffentlicht: (2024)
von: Hsieh, Cheng-Ping, et al.
Veröffentlicht: (2024)
CLongEval: A Chinese Benchmark for Evaluating Long-Context Large Language Models
von: Qiu, Zexuan, et al.
Veröffentlicht: (2024)
von: Qiu, Zexuan, et al.
Veröffentlicht: (2024)
Thus Spake Long-Context Large Language Model
von: Liu, Xiaoran, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoran, et al.
Veröffentlicht: (2025)
SIP: Injecting a Structural Inductive Bias into a Seq2Seq Model by Simulation
von: Lindemann, Matthias, et al.
Veröffentlicht: (2023)
von: Lindemann, Matthias, et al.
Veröffentlicht: (2023)
MindMerger: Efficient Boosting LLM Reasoning in non-English Languages
von: Huang, Zixian, et al.
Veröffentlicht: (2024)
von: Huang, Zixian, et al.
Veröffentlicht: (2024)
Finding Culture-Sensitive Neurons in Vision-Language Models
von: Zhao, Xiutian, et al.
Veröffentlicht: (2025)
von: Zhao, Xiutian, et al.
Veröffentlicht: (2025)
Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervision
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
M-Wanda: Improving One-Shot Pruning for Multilingual LLMs
von: Choenni, Rochelle, et al.
Veröffentlicht: (2025)
von: Choenni, Rochelle, et al.
Veröffentlicht: (2025)
A Controlled Study on Long Context Extension and Generalization in LLMs
von: Lu, Yi, et al.
Veröffentlicht: (2024)
von: Lu, Yi, et al.
Veröffentlicht: (2024)
AgentLongBench: A Controllable Long Benchmark For Long-Contexts Agents via Environment Rollouts
von: Fang, Shicheng, et al.
Veröffentlicht: (2026)
von: Fang, Shicheng, et al.
Veröffentlicht: (2026)
Evaluating and Safeguarding the Adversarial Robustness of Retrieval-Based In-Context Learning
von: Yu, Simon, et al.
Veröffentlicht: (2024)
von: Yu, Simon, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Post-hoc Reward Calibration: A Case Study on Length Bias
von: Huang, Zeyu, et al.
Veröffentlicht: (2024) -
Layerwise Recurrent Router for Mixture-of-Experts
von: Qiu, Zihan, et al.
Veröffentlicht: (2024) -
Masking in Multi-hop QA: An Analysis of How Language Models Perform with Context Permutation
von: Huang, Wenyu, et al.
Veröffentlicht: (2025) -
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
von: Huang, Zeyu, et al.
Veröffentlicht: (2025) -
Controlling What You Share: Assessing Language Model Adherence to Privacy Preferences
von: Ramírez, Guillem, et al.
Veröffentlicht: (2025)