From Early Encoding to Late Suppression: Interpreting LLMs on Character Counting Tasks
Fuente:
arXiv
Guardado en:
| Autores principales: | Datta, Ayan, Marreddy, Mounika, Mehler, Alexander, Zhao, Zhixue, Mamidi, Radhika |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Large Language Models Decide Early and Explain Later
por: Datta, Ayan, et al.
Publicado: (2026)
por: Datta, Ayan, et al.
Publicado: (2026)
SemEval-2024 Task 8: Weighted Layer Averaging RoBERTa for Black-Box Machine-Generated Text Detection
por: Datta, Ayan, et al.
Publicado: (2024)
por: Datta, Ayan, et al.
Publicado: (2024)
IndicSentEval: How Effectively do Multilingual Transformer Models encode Linguistic Properties for Indic Languages?
por: Aravapalli, Akhilesh, et al.
Publicado: (2024)
por: Aravapalli, Akhilesh, et al.
Publicado: (2024)
Analyzing Biases in Political Dialogue: Tagging U.S. Presidential Debates with an Extended DAMSL Framework
por: Prahallad, Lavanya, et al.
Publicado: (2025)
por: Prahallad, Lavanya, et al.
Publicado: (2025)
Significance of Chain of Thought in Gender Bias Mitigation for English-Dravidian Machine Translation
por: Prahallad, Lavanya, et al.
Publicado: (2024)
por: Prahallad, Lavanya, et al.
Publicado: (2024)
Zero-Shot Multi-task Hallucination Detection
por: Bhamidipati, Patanjali, et al.
Publicado: (2024)
por: Bhamidipati, Patanjali, et al.
Publicado: (2024)
Enhancing Hate Speech Detection on Social Media: A Comparative Analysis of Machine Learning Models and Text Transformation Approaches
por: Mishra, Saurabh, et al.
Publicado: (2026)
por: Mishra, Saurabh, et al.
Publicado: (2026)
Mast Kalandar at SemEval-2024 Task 8: On the Trail of Textual Origins: RoBERTa-BiLSTM Approach to Detect AI-Generated Text
por: Bafna, Jainit Sushil, et al.
Publicado: (2024)
por: Bafna, Jainit Sushil, et al.
Publicado: (2024)
USDC: A Dataset of $\underline{U}$ser $\underline{S}$tance and $\underline{D}$ogmatism in Long $\underline{C}$onversations
por: Marreddy, Mounika, et al.
Publicado: (2024)
por: Marreddy, Mounika, et al.
Publicado: (2024)
Explanation Generation for Contradiction Reconciliation with LLMs
por: Chan, Jason, et al.
Publicado: (2026)
por: Chan, Jason, et al.
Publicado: (2026)
Position: On the Methodological Pitfalls of Evaluating Base LLMs for Reasoning
por: Chan, Jason, et al.
Publicado: (2025)
por: Chan, Jason, et al.
Publicado: (2025)
Position: Logical Soundness is not a Reliable Criterion for Neurosymbolic Fact-Checking with LLMs
por: Chan, Jason, et al.
Publicado: (2026)
por: Chan, Jason, et al.
Publicado: (2026)
RULEBREAKERS: Challenging LLMs at the Crossroads between Formal Logic and Human-like Reasoning
por: Chan, Jason, et al.
Publicado: (2024)
por: Chan, Jason, et al.
Publicado: (2024)
It's All About In-Context Learning! Teaching Extremely Low-Resource Languages to LLMs
por: Li, Yue, et al.
Publicado: (2025)
por: Li, Yue, et al.
Publicado: (2025)
Tracing and Reversing Edits in LLMs
por: Youssef, Paul, et al.
Publicado: (2025)
por: Youssef, Paul, et al.
Publicado: (2025)
How to Make LLMs Forget: On Reversing In-Context Knowledge Edits
por: Youssef, Paul, et al.
Publicado: (2024)
por: Youssef, Paul, et al.
Publicado: (2024)
Multi-modal brain encoding models for multi-modal stimuli
por: Oota, Subba Reddy, et al.
Publicado: (2025)
por: Oota, Subba Reddy, et al.
Publicado: (2025)
Can Confidence Estimates Decide When Chain-of-Thought Is Necessary for LLMs?
por: Lewis-Lim, Samuel, et al.
Publicado: (2025)
por: Lewis-Lim, Samuel, et al.
Publicado: (2025)
The Spatial Semantics of Iconic Gesture
por: Lücking, Andy, et al.
Publicado: (2024)
por: Lücking, Andy, et al.
Publicado: (2024)
Comparing Explanation Faithfulness between Multilingual and Monolingual Fine-tuned Language Models
por: Zhao, Zhixue, et al.
Publicado: (2024)
por: Zhao, Zhixue, et al.
Publicado: (2024)
Late-to-Early Training: LET LLMs Learn Earlier, So Faster and Better
por: Zhao, Ji, et al.
Publicado: (2026)
por: Zhao, Ji, et al.
Publicado: (2026)
Every Character Counts: From Vulnerability to Defense in Phishing Detection
por: Chiper, Maria, et al.
Publicado: (2025)
por: Chiper, Maria, et al.
Publicado: (2025)
Survey-to-Behavior: Downstream Alignment of Human Values in LLMs via Survey Questions
por: Nie, Shangrui, et al.
Publicado: (2025)
por: Nie, Shangrui, et al.
Publicado: (2025)
LLMs Encode Harmfulness and Refusal Separately
por: Zhao, Jiachen, et al.
Publicado: (2025)
por: Zhao, Jiachen, et al.
Publicado: (2025)
Label Set Optimization via Activation Distribution Kurtosis for Zero-shot Classification with Generative Models
por: Li, Yue, et al.
Publicado: (2024)
por: Li, Yue, et al.
Publicado: (2024)
Mechanistic Interpretability of Large-Scale Counting in LLMs through a System-2 Strategy
por: Hasani, Hosein, et al.
Publicado: (2026)
por: Hasani, Hosein, et al.
Publicado: (2026)
Incorporating Attribution Importance for Improving Faithfulness Metrics
por: Zhao, Zhixue, et al.
Publicado: (2023)
por: Zhao, Zhixue, et al.
Publicado: (2023)
ReAGent: A Model-agnostic Feature Attribution Method for Generative Language Models
por: Zhao, Zhixue, et al.
Publicado: (2024)
por: Zhao, Zhixue, et al.
Publicado: (2024)
Exploring the Limits of Model Compression in LLMs: A Knowledge Distillation Study on QA Tasks
por: Datta, Joyeeta, et al.
Publicado: (2025)
por: Datta, Joyeeta, et al.
Publicado: (2025)
Exploring Vision Language Models for Multimodal and Multilingual Stance Detection
por: Vasilakes, Jake, et al.
Publicado: (2025)
por: Vasilakes, Jake, et al.
Publicado: (2025)
Position: Enough of Scaling LLMs! Lets Focus on Downscaling
por: Goel, Yash, et al.
Publicado: (2025)
por: Goel, Yash, et al.
Publicado: (2025)
Understanding the Ability of LLMs to Handle Character-Level Perturbation
por: Zhuo, Anyuan, et al.
Publicado: (2025)
por: Zhuo, Anyuan, et al.
Publicado: (2025)
SCRum-9: Multilingual Stance Classification over Rumours on Social Media
por: Li, Yue, et al.
Publicado: (2025)
por: Li, Yue, et al.
Publicado: (2025)
Contextual Position Encoding: Learning to Count What's Important
por: Golovneva, Olga, et al.
Publicado: (2024)
por: Golovneva, Olga, et al.
Publicado: (2024)
Enhancing Character-Level Understanding in LLMs through Token Internal Structure Learning
por: Xu, Zhu, et al.
Publicado: (2024)
por: Xu, Zhu, et al.
Publicado: (2024)
Do LLMs Provide Consistent Answers to Health-Related Questions across Languages?
por: Schlicht, Ipek Baris, et al.
Publicado: (2025)
por: Schlicht, Ipek Baris, et al.
Publicado: (2025)
Value-Guided KV Compression for LLMs via Approximated CUR Decomposition
por: Sengupta, Ayan, et al.
Publicado: (2025)
por: Sengupta, Ayan, et al.
Publicado: (2025)
LLMs Encode How Difficult Problems Are
por: Lugoloobi, William, et al.
Publicado: (2025)
por: Lugoloobi, William, et al.
Publicado: (2025)
CharBench: Evaluating the Role of Tokenization in Character-Level Tasks
por: Uzan, Omri, et al.
Publicado: (2025)
por: Uzan, Omri, et al.
Publicado: (2025)
Efficient Pruning of Text-to-Image Models: Insights from Pruning Stable Diffusion
por: Ramesh, Samarth N, et al.
Publicado: (2024)
por: Ramesh, Samarth N, et al.
Publicado: (2024)
Ejemplares similares
-
Large Language Models Decide Early and Explain Later
por: Datta, Ayan, et al.
Publicado: (2026) -
SemEval-2024 Task 8: Weighted Layer Averaging RoBERTa for Black-Box Machine-Generated Text Detection
por: Datta, Ayan, et al.
Publicado: (2024) -
IndicSentEval: How Effectively do Multilingual Transformer Models encode Linguistic Properties for Indic Languages?
por: Aravapalli, Akhilesh, et al.
Publicado: (2024) -
Analyzing Biases in Political Dialogue: Tagging U.S. Presidential Debates with an Extended DAMSL Framework
por: Prahallad, Lavanya, et al.
Publicado: (2025) -
Significance of Chain of Thought in Gender Bias Mitigation for English-Dravidian Machine Translation
por: Prahallad, Lavanya, et al.
Publicado: (2024)