What Matters to an LLM? Behavioral and Computational Evidences from Summarization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Yongxin, Wu, Changshun, Mulhem, Philippe, Schwab, Didier, Peyrard, Maxime |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TempPerturb-Eval: On the Joint Effects of Internal Temperature and External Perturbations in RAG Robustness
von: Zhou, Yongxin, et al.
Veröffentlicht: (2025)
von: Zhou, Yongxin, et al.
Veröffentlicht: (2025)
What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search
von: Zhang, Xinhao, et al.
Veröffentlicht: (2026)
von: Zhang, Xinhao, et al.
Veröffentlicht: (2026)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
von: Méloux, Maxime, et al.
Veröffentlicht: (2025)
von: Méloux, Maxime, et al.
Veröffentlicht: (2025)
What Really Controls Temporal Reasoning in Large Language Models: Tokenisation or Representation of Time?
von: Bhatia, Gagan, et al.
Veröffentlicht: (2026)
von: Bhatia, Gagan, et al.
Veröffentlicht: (2026)
Date Fragments: A Hidden Bottleneck of Tokenization for Temporal Reasoning
von: Bhatia, Gagan, et al.
Veröffentlicht: (2025)
von: Bhatia, Gagan, et al.
Veröffentlicht: (2025)
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
von: Méloux, Maxime, et al.
Veröffentlicht: (2025)
von: Méloux, Maxime, et al.
Veröffentlicht: (2025)
Agentic AI: The Era of Semantic Decoding
von: Peyrard, Maxime, et al.
Veröffentlicht: (2024)
von: Peyrard, Maxime, et al.
Veröffentlicht: (2024)
PSentScore: Evaluating Sentiment Polarity in Dialogue Summarization
von: Zhou, Yongxin, et al.
Veröffentlicht: (2023)
von: Zhou, Yongxin, et al.
Veröffentlicht: (2023)
Word Matters: What Influences Domain Adaptation in Summarization?
von: Li, Yinghao, et al.
Veröffentlicht: (2024)
von: Li, Yinghao, et al.
Veröffentlicht: (2024)
Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning
von: Geng, Saibo, et al.
Veröffentlicht: (2023)
von: Geng, Saibo, et al.
Veröffentlicht: (2023)
Can GPT models Follow Human Summarization Guidelines? A Study for Targeted Communication Goals
von: Zhou, Yongxin, et al.
Veröffentlicht: (2023)
von: Zhou, Yongxin, et al.
Veröffentlicht: (2023)
$\texttt{COSMIC}$: Mutual Information for Task-Agnostic Summarization Evaluation
von: Darrin, Maxime, et al.
Veröffentlicht: (2024)
von: Darrin, Maxime, et al.
Veröffentlicht: (2024)
Understanding LLM Behavior in Multi-Target Cross-Lingual Summarization
von: Ryu, Sangwon, et al.
Veröffentlicht: (2026)
von: Ryu, Sangwon, et al.
Veröffentlicht: (2026)
REFINER: Reasoning Feedback on Intermediate Representations
von: Paul, Debjit, et al.
Veröffentlicht: (2023)
von: Paul, Debjit, et al.
Veröffentlicht: (2023)
Analyzing LLM Behavior in Dialogue Summarization: Unveiling Circumstantial Hallucination Trends
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
Reassessing Graph Linearization for Sequence-to-sequence AMR Parsing: On the Advantages and Limitations of Triple-Based Encoding
von: Kang, Jeongwoo, et al.
Veröffentlicht: (2025)
von: Kang, Jeongwoo, et al.
Veröffentlicht: (2025)
UMLS-KGI-BERT: Data-Centric Knowledge Integration in Transformers for Biomedical Entity Recognition
von: Mannion, Aidan, et al.
Veröffentlicht: (2023)
von: Mannion, Aidan, et al.
Veröffentlicht: (2023)
Should Cross-Lingual AMR Parsing go Meta? An Empirical Assessment of Meta-Learning and Joint Learning AMR Parsing
von: Kang, Jeongwoo, et al.
Veröffentlicht: (2024)
von: Kang, Jeongwoo, et al.
Veröffentlicht: (2024)
Unlearning What Matters: Token-Level Attribution for Precise Language Model Unlearning
von: Wu, Jiawei, et al.
Veröffentlicht: (2026)
von: Wu, Jiawei, et al.
Veröffentlicht: (2026)
Multi-LLM Text Summarization
von: Fang, Jiangnan, et al.
Veröffentlicht: (2024)
von: Fang, Jiangnan, et al.
Veröffentlicht: (2024)
Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking
von: Roitero, Kevin, et al.
Veröffentlicht: (2025)
von: Roitero, Kevin, et al.
Veröffentlicht: (2025)
SummExecEdit: A Factual Consistency Benchmark in Summarization with Executable Edits
von: Thorat, Onkar, et al.
Veröffentlicht: (2024)
von: Thorat, Onkar, et al.
Veröffentlicht: (2024)
What Are They Talking About? A Benchmark of Knowledge-Grounded Discussion Summarization
von: Zhou, Weixiao, et al.
Veröffentlicht: (2025)
von: Zhou, Weixiao, et al.
Veröffentlicht: (2025)
GLIMPSE: Pragmatically Informative Multi-Document Summarization for Scholarly Reviews
von: Darrin, Maxime, et al.
Veröffentlicht: (2024)
von: Darrin, Maxime, et al.
Veröffentlicht: (2024)
Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question Coverage
von: Xie, Kaige, et al.
Veröffentlicht: (2024)
von: Xie, Kaige, et al.
Veröffentlicht: (2024)
Query-Focused Extractive Summarization for Sentiment Explanation
von: Moubtahij, Ahmed, et al.
Veröffentlicht: (2025)
von: Moubtahij, Ahmed, et al.
Veröffentlicht: (2025)
zip2zip: Inference-Time Adaptive Tokenization via Online Compression
von: Geng, Saibo, et al.
Veröffentlicht: (2025)
von: Geng, Saibo, et al.
Veröffentlicht: (2025)
Understanding LLM Reasoning for Abstractive Summarization
von: Yuan, Haohan, et al.
Veröffentlicht: (2025)
von: Yuan, Haohan, et al.
Veröffentlicht: (2025)
Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News Articles
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2023)
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2023)
Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities
von: Zhang, Xiangxu, et al.
Veröffentlicht: (2026)
von: Zhang, Xiangxu, et al.
Veröffentlicht: (2026)
Learning from Self Critique and Refinement for Faithful LLM Summarization
von: Hu, Ting-Yao, et al.
Veröffentlicht: (2025)
von: Hu, Ting-Yao, et al.
Veröffentlicht: (2025)
Learning to Summarize from LLM-generated Feedback
von: Song, Hwanjun, et al.
Veröffentlicht: (2024)
von: Song, Hwanjun, et al.
Veröffentlicht: (2024)
References Matter: Investigating the Impact of Reference Set Variation on Summarization Evaluation
von: Casola, Silvia, et al.
Veröffentlicht: (2025)
von: Casola, Silvia, et al.
Veröffentlicht: (2025)
ArgCMV: An Argument Summarization Benchmark for the LLM-era
von: Gurjar, Omkar, et al.
Veröffentlicht: (2025)
von: Gurjar, Omkar, et al.
Veröffentlicht: (2025)
Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior Data
von: Lu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Lu, Yuxuan, et al.
Veröffentlicht: (2025)
A Computational Framework for Behavioral Assessment of LLM Therapists
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2024)
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2024)
FactPICO: Factuality Evaluation for Plain Language Summarization of Medical Evidence
von: Joseph, Sebastian Antony, et al.
Veröffentlicht: (2024)
von: Joseph, Sebastian Antony, et al.
Veröffentlicht: (2024)
Learning to Summarize by Learning to Quiz: Adversarial Agentic Collaboration for Long Document Summarization
von: Wang, Weixuan, et al.
Veröffentlicht: (2025)
von: Wang, Weixuan, et al.
Veröffentlicht: (2025)
Measuring What Matters -- or What's Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors
von: Walsh, Cole, et al.
Veröffentlicht: (2026)
von: Walsh, Cole, et al.
Veröffentlicht: (2026)
Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence
von: Mo, Kaijie, et al.
Veröffentlicht: (2026)
von: Mo, Kaijie, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
TempPerturb-Eval: On the Joint Effects of Internal Temperature and External Perturbations in RAG Robustness
von: Zhou, Yongxin, et al.
Veröffentlicht: (2025) -
What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search
von: Zhang, Xinhao, et al.
Veröffentlicht: (2026) -
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
von: Méloux, Maxime, et al.
Veröffentlicht: (2025) -
What Really Controls Temporal Reasoning in Large Language Models: Tokenisation or Representation of Time?
von: Bhatia, Gagan, et al.
Veröffentlicht: (2026) -
Date Fragments: A Hidden Bottleneck of Tokenization for Temporal Reasoning
von: Bhatia, Gagan, et al.
Veröffentlicht: (2025)