Large Language Models Decide Early and Explain Later
Fuente:
arXiv
Salvato in:
| Autori principali: | Datta, Ayan, Zhao, Zhixue, Verma, Bhuvanesh, Mamidi, Radhika, Marreddy, Mounika, Mehler, Alexander |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Early Encoding to Late Suppression: Interpreting LLMs on Character Counting Tasks
di: Datta, Ayan, et al.
Pubblicazione: (2026)
di: Datta, Ayan, et al.
Pubblicazione: (2026)
IndicSentEval: How Effectively do Multilingual Transformer Models encode Linguistic Properties for Indic Languages?
di: Aravapalli, Akhilesh, et al.
Pubblicazione: (2024)
di: Aravapalli, Akhilesh, et al.
Pubblicazione: (2024)
SemEval-2024 Task 8: Weighted Layer Averaging RoBERTa for Black-Box Machine-Generated Text Detection
di: Datta, Ayan, et al.
Pubblicazione: (2024)
di: Datta, Ayan, et al.
Pubblicazione: (2024)
Analyzing Biases in Political Dialogue: Tagging U.S. Presidential Debates with an Extended DAMSL Framework
di: Prahallad, Lavanya, et al.
Pubblicazione: (2025)
di: Prahallad, Lavanya, et al.
Pubblicazione: (2025)
Significance of Chain of Thought in Gender Bias Mitigation for English-Dravidian Machine Translation
di: Prahallad, Lavanya, et al.
Pubblicazione: (2024)
di: Prahallad, Lavanya, et al.
Pubblicazione: (2024)
Enhancing Hate Speech Detection on Social Media: A Comparative Analysis of Machine Learning Models and Text Transformation Approaches
di: Mishra, Saurabh, et al.
Pubblicazione: (2026)
di: Mishra, Saurabh, et al.
Pubblicazione: (2026)
Zero-Shot Multi-task Hallucination Detection
di: Bhamidipati, Patanjali, et al.
Pubblicazione: (2024)
di: Bhamidipati, Patanjali, et al.
Pubblicazione: (2024)
Can Confidence Estimates Decide When Chain-of-Thought Is Necessary for LLMs?
di: Lewis-Lim, Samuel, et al.
Pubblicazione: (2025)
di: Lewis-Lim, Samuel, et al.
Pubblicazione: (2025)
DFKI-NLP at SemEval-2024 Task 2: Towards Robust LLMs Using Data Perturbations and MinMax Training
di: Verma, Bhuvanesh, et al.
Pubblicazione: (2024)
di: Verma, Bhuvanesh, et al.
Pubblicazione: (2024)
USDC: A Dataset of $\underline{U}$ser $\underline{S}$tance and $\underline{D}$ogmatism in Long $\underline{C}$onversations
di: Marreddy, Mounika, et al.
Pubblicazione: (2024)
di: Marreddy, Mounika, et al.
Pubblicazione: (2024)
MMTM: Tri-Modal Topic Modeling for Long-Form Video via Similarity-Gated Fusion
di: Abusaleh, Ali, et al.
Pubblicazione: (2026)
di: Abusaleh, Ali, et al.
Pubblicazione: (2026)
Comparing Explanation Faithfulness between Multilingual and Monolingual Fine-tuned Language Models
di: Zhao, Zhixue, et al.
Pubblicazione: (2024)
di: Zhao, Zhixue, et al.
Pubblicazione: (2024)
Investigating Hallucinations in Pruned Large Language Models for Abstractive Summarization
di: Chrysostomou, George, et al.
Pubblicazione: (2023)
di: Chrysostomou, George, et al.
Pubblicazione: (2023)
Position: Editing Large Language Models Poses Serious Safety Risks
di: Youssef, Paul, et al.
Pubblicazione: (2025)
di: Youssef, Paul, et al.
Pubblicazione: (2025)
ReAGent: A Model-agnostic Feature Attribution Method for Generative Language Models
di: Zhao, Zhixue, et al.
Pubblicazione: (2024)
di: Zhao, Zhixue, et al.
Pubblicazione: (2024)
Mast Kalandar at SemEval-2024 Task 8: On the Trail of Textual Origins: RoBERTa-BiLSTM Approach to Detect AI-Generated Text
di: Bafna, Jainit Sushil, et al.
Pubblicazione: (2024)
di: Bafna, Jainit Sushil, et al.
Pubblicazione: (2024)
Exploring Vision Language Models for Multimodal and Multilingual Stance Detection
di: Vasilakes, Jake, et al.
Pubblicazione: (2025)
di: Vasilakes, Jake, et al.
Pubblicazione: (2025)
Multi-modal brain encoding models for multi-modal stimuli
di: Oota, Subba Reddy, et al.
Pubblicazione: (2025)
di: Oota, Subba Reddy, et al.
Pubblicazione: (2025)
Compression Laws for Large Language Models
di: Sengupta, Ayan, et al.
Pubblicazione: (2025)
di: Sengupta, Ayan, et al.
Pubblicazione: (2025)
It's All About In-Context Learning! Teaching Extremely Low-Resource Languages to LLMs
di: Li, Yue, et al.
Pubblicazione: (2025)
di: Li, Yue, et al.
Pubblicazione: (2025)
The Spatial Semantics of Iconic Gesture
di: Lücking, Andy, et al.
Pubblicazione: (2024)
di: Lücking, Andy, et al.
Pubblicazione: (2024)
The Art of Scaling Test-Time Compute for Large Language Models
di: Agarwal, Aradhye, et al.
Pubblicazione: (2025)
di: Agarwal, Aradhye, et al.
Pubblicazione: (2025)
Label Set Optimization via Activation Distribution Kurtosis for Zero-shot Classification with Generative Models
di: Li, Yue, et al.
Pubblicazione: (2024)
di: Li, Yue, et al.
Pubblicazione: (2024)
Has this Fact been Edited? Detecting Knowledge Edits in Language Models
di: Youssef, Paul, et al.
Pubblicazione: (2024)
di: Youssef, Paul, et al.
Pubblicazione: (2024)
First Finish Search: Efficient Test-Time Scaling in Large Language Models
di: Agarwal, Aradhye, et al.
Pubblicazione: (2025)
di: Agarwal, Aradhye, et al.
Pubblicazione: (2025)
Pretraining Exposure Explains Popularity Judgments in Large Language Models
di: Mozafari, Jamshid, et al.
Pubblicazione: (2026)
di: Mozafari, Jamshid, et al.
Pubblicazione: (2026)
Racing Thoughts: Explaining Contextualization Errors in Large Language Models
di: Lepori, Michael A., et al.
Pubblicazione: (2024)
di: Lepori, Michael A., et al.
Pubblicazione: (2024)
Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions
di: Chen, Hanjie, et al.
Pubblicazione: (2024)
di: Chen, Hanjie, et al.
Pubblicazione: (2024)
Explanation Generation for Contradiction Reconciliation with LLMs
di: Chan, Jason, et al.
Pubblicazione: (2026)
di: Chan, Jason, et al.
Pubblicazione: (2026)
Position: Logical Soundness is not a Reliable Criterion for Neurosymbolic Fact-Checking with LLMs
di: Chan, Jason, et al.
Pubblicazione: (2026)
di: Chan, Jason, et al.
Pubblicazione: (2026)
RULEBREAKERS: Challenging LLMs at the Crossroads between Formal Logic and Human-like Reasoning
di: Chan, Jason, et al.
Pubblicazione: (2024)
di: Chan, Jason, et al.
Pubblicazione: (2024)
Position: On the Methodological Pitfalls of Evaluating Base LLMs for Reasoning
di: Chan, Jason, et al.
Pubblicazione: (2025)
di: Chan, Jason, et al.
Pubblicazione: (2025)
Syntactic Language Change in English and German: Metrics, Parsers, and Convergences
di: Chen, Yanran, et al.
Pubblicazione: (2024)
di: Chen, Yanran, et al.
Pubblicazione: (2024)
Mitigating Content Effects on Reasoning in Language Models through Fine-Grained Activation Steering
di: Valentino, Marco, et al.
Pubblicazione: (2025)
di: Valentino, Marco, et al.
Pubblicazione: (2025)
Efficient Pruning of Text-to-Image Models: Insights from Pruning Stable Diffusion
di: Ramesh, Samarth N, et al.
Pubblicazione: (2024)
di: Ramesh, Samarth N, et al.
Pubblicazione: (2024)
MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
di: Huang, Yue, et al.
Pubblicazione: (2023)
di: Huang, Yue, et al.
Pubblicazione: (2023)
Step-by-Step Unmasking for Parameter-Efficient Fine-tuning of Large Language Models
di: Agarwal, Aradhye, et al.
Pubblicazione: (2024)
di: Agarwal, Aradhye, et al.
Pubblicazione: (2024)
Explaining Large Language Models with gSMILE
di: Dehghani, Zeinab, et al.
Pubblicazione: (2025)
di: Dehghani, Zeinab, et al.
Pubblicazione: (2025)
Incorporating Attribution Importance for Improving Faithfulness Metrics
di: Zhao, Zhixue, et al.
Pubblicazione: (2023)
di: Zhao, Zhixue, et al.
Pubblicazione: (2023)
Contextual Compression in Retrieval-Augmented Generation for Large Language Models: A Survey
di: Verma, Sourav
Pubblicazione: (2024)
di: Verma, Sourav
Pubblicazione: (2024)
Documenti analoghi
-
From Early Encoding to Late Suppression: Interpreting LLMs on Character Counting Tasks
di: Datta, Ayan, et al.
Pubblicazione: (2026) -
IndicSentEval: How Effectively do Multilingual Transformer Models encode Linguistic Properties for Indic Languages?
di: Aravapalli, Akhilesh, et al.
Pubblicazione: (2024) -
SemEval-2024 Task 8: Weighted Layer Averaging RoBERTa for Black-Box Machine-Generated Text Detection
di: Datta, Ayan, et al.
Pubblicazione: (2024) -
Analyzing Biases in Political Dialogue: Tagging U.S. Presidential Debates with an Extended DAMSL Framework
di: Prahallad, Lavanya, et al.
Pubblicazione: (2025) -
Significance of Chain of Thought in Gender Bias Mitigation for English-Dravidian Machine Translation
di: Prahallad, Lavanya, et al.
Pubblicazione: (2024)