Identifying and Analyzing Performance-Critical Tokens in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bai, Yu, Huang, Heyan, Piano, Cesare Spinoso-Di, Rondeau, Marc-Antoine, Chen, Sanxing, Gao, Yang, Cheung, Jackie Chi Kit |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CItruS: Chunked Instruction-aware State Eviction for Long Sequence Modeling
von: Bai, Yu, et al.
Veröffentlicht: (2024)
von: Bai, Yu, et al.
Veröffentlicht: (2024)
$(RSA)^2$: A Rhetorical-Strategy-Aware Rational Speech Act Framework for Figurative Language Understanding
von: Piano, Cesare Spinoso-Di, et al.
Veröffentlicht: (2025)
von: Piano, Cesare Spinoso-Di, et al.
Veröffentlicht: (2025)
Does This Summary Answer My Question? Modeling Query-Focused Summary Readers with Rational Speech Acts
von: Piano, Cesare Spinoso-Di, et al.
Veröffentlicht: (2024)
von: Piano, Cesare Spinoso-Di, et al.
Veröffentlicht: (2024)
Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning
von: Flores, Lorenzo Jaime Yu, et al.
Veröffentlicht: (2026)
von: Flores, Lorenzo Jaime Yu, et al.
Veröffentlicht: (2026)
Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMs
von: Cheng, Ziling, et al.
Veröffentlicht: (2025)
von: Cheng, Ziling, et al.
Veröffentlicht: (2025)
Testing the Assumptions of Active Learning for Translation Tasks with Few Samples
von: Flores, Lorenzo Jaime Yu, et al.
Veröffentlicht: (2026)
von: Flores, Lorenzo Jaime Yu, et al.
Veröffentlicht: (2026)
Solving the Challenge Set without Solving the Task: On Winograd Schemas as a Test of Pronominal Coreference Resolution
von: Porada, Ian, et al.
Veröffentlicht: (2024)
von: Porada, Ian, et al.
Veröffentlicht: (2024)
PreSumm: Predicting Summarization Performance Without Summarizing
von: Koniaev, Steven, et al.
Veröffentlicht: (2025)
von: Koniaev, Steven, et al.
Veröffentlicht: (2025)
Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations
von: Yu, Lei, et al.
Veröffentlicht: (2024)
von: Yu, Lei, et al.
Veröffentlicht: (2024)
A Controlled Reevaluation of Coreference Resolution Models
von: Porada, Ian, et al.
Veröffentlicht: (2024)
von: Porada, Ian, et al.
Veröffentlicht: (2024)
Improving the Calibration of Confidence Scores in Text Generation Using the Output Distribution's Characteristics
von: Flores, Lorenzo Jaime Yu, et al.
Veröffentlicht: (2025)
von: Flores, Lorenzo Jaime Yu, et al.
Veröffentlicht: (2025)
Can Vision Language Models Be Adaptive in Mathematics Education? A Learner Model-based Rubric Study
von: Gao, Jie, et al.
Veröffentlicht: (2026)
von: Gao, Jie, et al.
Veröffentlicht: (2026)
$\texttt{COSMIC}$: Mutual Information for Task-Agnostic Summarization Evaluation
von: Darrin, Maxime, et al.
Veröffentlicht: (2024)
von: Darrin, Maxime, et al.
Veröffentlicht: (2024)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2025)
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2025)
To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External Contexts
von: Huang, Yukun, et al.
Veröffentlicht: (2024)
von: Huang, Yukun, et al.
Veröffentlicht: (2024)
CriticEval: Evaluating Large Language Model as Critic
von: Lan, Tian, et al.
Veröffentlicht: (2024)
von: Lan, Tian, et al.
Veröffentlicht: (2024)
Challenges to Evaluating the Generalization of Coreference Resolution Models: A Measurement Modeling Perspective
von: Porada, Ian, et al.
Veröffentlicht: (2023)
von: Porada, Ian, et al.
Veröffentlicht: (2023)
Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation
von: Cheng, Ziling, et al.
Veröffentlicht: (2025)
von: Cheng, Ziling, et al.
Veröffentlicht: (2025)
ChatShop: Interactive Information Seeking with Language Agents
von: Chen, Sanxing, et al.
Veröffentlicht: (2024)
von: Chen, Sanxing, et al.
Veröffentlicht: (2024)
Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
von: Huang, Yukun, et al.
Veröffentlicht: (2025)
von: Huang, Yukun, et al.
Veröffentlicht: (2025)
Real-time Factuality Assessment from Adversarial Feedback
von: Chen, Sanxing, et al.
Veröffentlicht: (2024)
von: Chen, Sanxing, et al.
Veröffentlicht: (2024)
How Far Can In-Context Alignment Go? Exploring the State of In-Context Alignment
von: Huang, Heyan, et al.
Veröffentlicht: (2024)
von: Huang, Heyan, et al.
Veröffentlicht: (2024)
ECBD: Evidence-Centered Benchmark Design for NLP
von: Liu, Yu Lu, et al.
Veröffentlicht: (2024)
von: Liu, Yu Lu, et al.
Veröffentlicht: (2024)
Analyzing The Language of Visual Tokens
von: Chan, David M., et al.
Veröffentlicht: (2024)
von: Chan, David M., et al.
Veröffentlicht: (2024)
FlashBack:Efficient Retrieval-Augmented Language Modeling for Long Context Inference
von: Liu, Runheng, et al.
Veröffentlicht: (2024)
von: Liu, Runheng, et al.
Veröffentlicht: (2024)
Error Diversity Matters: An Error-Resistant Ensemble Method for Unsupervised Dependency Parsing
von: Shayegh, Behzad, et al.
Veröffentlicht: (2024)
von: Shayegh, Behzad, et al.
Veröffentlicht: (2024)
How Teachers Can Use Large Language Models and Bloom's Taxonomy to Create Educational Quizzes
von: Elkins, Sabina, et al.
Veröffentlicht: (2024)
von: Elkins, Sabina, et al.
Veröffentlicht: (2024)
Beyond Exact Match: Semantically Reassessing Event Extraction by Large Language Models
von: Lu, Yi-Fan, et al.
Veröffentlicht: (2024)
von: Lu, Yi-Fan, et al.
Veröffentlicht: (2024)
Debate, Reflect, and Distill: Multi-Agent Feedback with Tree-Structured Preference Optimization for Efficient Language Model Enhancement
von: Zhou, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Zhou, Xiaofeng, et al.
Veröffentlicht: (2025)
Building Knowledge-Grounded Dialogue Systems with Graph-Based Semantic Modeling
von: Yang, Yizhe, et al.
Veröffentlicht: (2022)
von: Yang, Yizhe, et al.
Veröffentlicht: (2022)
Word Matters: What Influences Domain Adaptation in Summarization?
von: Li, Yinghao, et al.
Veröffentlicht: (2024)
von: Li, Yinghao, et al.
Veröffentlicht: (2024)
Performance Evaluation of Tokenizers in Large Language Models for the Assamese Language
von: Tamang, Sagar, et al.
Veröffentlicht: (2024)
von: Tamang, Sagar, et al.
Veröffentlicht: (2024)
EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios
von: Xu, Bin, et al.
Veröffentlicht: (2025)
von: Xu, Bin, et al.
Veröffentlicht: (2025)
Analyzing the Performance of Large Language Models on Code Summarization
von: Haldar, Rajarshi, et al.
Veröffentlicht: (2024)
von: Haldar, Rajarshi, et al.
Veröffentlicht: (2024)
Enhancing Large Language Model Reasoning via Selective Critical Token Fine-Tuning
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2025)
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2025)
Can Large Language Models Identify Authorship?
von: Huang, Baixiang, et al.
Veröffentlicht: (2024)
von: Huang, Baixiang, et al.
Veröffentlicht: (2024)
When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training
von: Chen, Sanxing, et al.
Veröffentlicht: (2025)
von: Chen, Sanxing, et al.
Veröffentlicht: (2025)
Incorporating Self-Rewriting into Large Language Model Reasoning Reinforcement
von: Yao, Jiashu, et al.
Veröffentlicht: (2025)
von: Yao, Jiashu, et al.
Veröffentlicht: (2025)
From Tokens to Materials: Leveraging Language Models for Scientific Discovery
von: Wan, Yuwei, et al.
Veröffentlicht: (2024)
von: Wan, Yuwei, et al.
Veröffentlicht: (2024)
Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages
von: Tamang, S., et al.
Veröffentlicht: (2024)
von: Tamang, S., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CItruS: Chunked Instruction-aware State Eviction for Long Sequence Modeling
von: Bai, Yu, et al.
Veröffentlicht: (2024) -
$(RSA)^2$: A Rhetorical-Strategy-Aware Rational Speech Act Framework for Figurative Language Understanding
von: Piano, Cesare Spinoso-Di, et al.
Veröffentlicht: (2025) -
Does This Summary Answer My Question? Modeling Query-Focused Summary Readers with Rational Speech Acts
von: Piano, Cesare Spinoso-Di, et al.
Veröffentlicht: (2024) -
Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning
von: Flores, Lorenzo Jaime Yu, et al.
Veröffentlicht: (2026) -
Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMs
von: Cheng, Ziling, et al.
Veröffentlicht: (2025)