Extracting Paragraphs from LLM Token Activations
Fuente:
arXiv
Saved in:
| Main Authors: | Pochinkov, Nicholas, Benoit, Angelo, Agarwal, Lovkush, Majid, Zainab Ali, Ter-Minassian, Lucile |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Democratizing AI Governance: Balancing Expertise and Public Participation
by: Ter-Minassian, Lucile
Published: (2025)
by: Ter-Minassian, Lucile
Published: (2025)
Dissecting Language Models: Machine Unlearning via Selective Pruning
by: Pochinkov, Nicholas, et al.
Published: (2024)
by: Pochinkov, Nicholas, et al.
Published: (2024)
Beyond Tokens in Language Models: Interpreting Activations through Text Genre Chunks
by: Benito-Rodriguez, Éloïse, et al.
Published: (2025)
by: Benito-Rodriguez, Éloïse, et al.
Published: (2025)
Investigating Neuron Ablation in Attention Heads: The Case for Peak Activation Centering
by: Pochinkov, Nicholas, et al.
Published: (2024)
by: Pochinkov, Nicholas, et al.
Published: (2024)
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
by: Veuthey, Jaime Raldua, et al.
Published: (2025)
by: Veuthey, Jaime Raldua, et al.
Published: (2025)
ParaScopes: What do Language Models Activations Encode About Future Text?
by: Pochinkov, Nicky, et al.
Published: (2025)
by: Pochinkov, Nicky, et al.
Published: (2025)
QuaLLM: An LLM-based Framework to Extract Quantitative Insights from Online Forums
by: Rao, Varun Nagaraj, et al.
Published: (2024)
by: Rao, Varun Nagaraj, et al.
Published: (2024)
LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
by: Zebaze, Armel, et al.
Published: (2025)
by: Zebaze, Armel, et al.
Published: (2025)
Localizing Paragraph Memorization in Language Models
by: Stoehr, Niklas, et al.
Published: (2024)
by: Stoehr, Niklas, et al.
Published: (2024)
Query-driven Relevant Paragraph Extraction from Legal Judgments
by: Santosh, T. Y. S. S, et al.
Published: (2024)
by: Santosh, T. Y. S. S, et al.
Published: (2024)
PLANNER: Generating Diversified Paragraph via Latent Language Diffusion Model
by: Zhang, Yizhe, et al.
Published: (2023)
by: Zhang, Yizhe, et al.
Published: (2023)
Paragraph Segmentation Revisited: Towards a Standard Task for Structuring Speech
by: Retkowski, Fabian, et al.
Published: (2025)
by: Retkowski, Fabian, et al.
Published: (2025)
Explainable AI for survival analysis: a median-SHAP approach
by: Ter-Minassian, Lucile, et al.
Published: (2024)
by: Ter-Minassian, Lucile, et al.
Published: (2024)
A Paragraph-level Multi-task Learning Model for Scientific Fact-Verification
by: Li, Xiangci, et al.
Published: (2020)
by: Li, Xiangci, et al.
Published: (2020)
X-PARADE: Cross-Lingual Textual Entailment and Information Divergence across Paragraphs
by: Rodriguez, Juan Diego, et al.
Published: (2023)
by: Rodriguez, Juan Diego, et al.
Published: (2023)
ParaRev: Building a dataset for Scientific Paragraph Revision annotated with revision instruction
by: Jourdan, Léane, et al.
Published: (2025)
by: Jourdan, Léane, et al.
Published: (2025)
Modularity in Transformers: Investigating Neuron Separability & Specialization
by: Pochinkov, Nicholas, et al.
Published: (2024)
by: Pochinkov, Nicholas, et al.
Published: (2024)
A Multi-lingual Dataset of Classified Paragraphs from Open Access Scientific Publications
by: Jeangirard, Eric
Published: (2025)
by: Jeangirard, Eric
Published: (2025)
Hierarchical Bias-Driven Stratification for Interpretable Causal Effect Estimation
by: Ter-Minassian, Lucile, et al.
Published: (2024)
by: Ter-Minassian, Lucile, et al.
Published: (2024)
Is merging worth it? Securely evaluating the information gain for causal dataset acquisition
by: Fawkes, Jake, et al.
Published: (2024)
by: Fawkes, Jake, et al.
Published: (2024)
Token-free Models for Sarcasm Detection
by: Mamtani, Sumit, et al.
Published: (2025)
by: Mamtani, Sumit, et al.
Published: (2025)
Do LLMs Understand Your Translations? Evaluating Paragraph-level MT with Question Answering
by: Fernandes, Patrick, et al.
Published: (2025)
by: Fernandes, Patrick, et al.
Published: (2025)
Multilingual Non-Factoid Question Answering with Answer Paragraph Selection
by: Mishra, Ritwik, et al.
Published: (2024)
by: Mishra, Ritwik, et al.
Published: (2024)
SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness
by: Huo, Jiahao, et al.
Published: (2026)
by: Huo, Jiahao, et al.
Published: (2026)
MUTANT: A Recipe for Multilingual Tokenizer Design
by: Rana, Souvik, et al.
Published: (2025)
by: Rana, Souvik, et al.
Published: (2025)
Advancing Topic Segmentation and Outline Generation in Chinese Texts: The Paragraph-level Topic Representation, Corpus, and Benchmark
by: Jiang, Feng, et al.
Published: (2023)
by: Jiang, Feng, et al.
Published: (2023)
Extracting Unlearned Information from LLMs with Activation Steering
by: Seyitoğlu, Atakan, et al.
Published: (2024)
by: Seyitoğlu, Atakan, et al.
Published: (2024)
Implementing Long Text Style Transfer with LLMs through Dual-Layered Sentence and Paragraph Structure Extraction and Mapping
by: Wu, Yusen, et al.
Published: (2025)
by: Wu, Yusen, et al.
Published: (2025)
Pointer-Guided Pre-Training: Infusing Large Language Models with Paragraph-Level Contextual Awareness
by: Hillebrand, Lars, et al.
Published: (2024)
by: Hillebrand, Lars, et al.
Published: (2024)
COSMIC: Generalized Refusal Direction Identification in LLM Activations
by: Siu, Vincent, et al.
Published: (2025)
by: Siu, Vincent, et al.
Published: (2025)
Tokenization Matters: Improving Zero-Shot NER for Indic Languages
by: Pattnayak, Priyaranjan, et al.
Published: (2025)
by: Pattnayak, Priyaranjan, et al.
Published: (2025)
Extracting Prompts by Inverting LLM Outputs
by: Zhang, Collin, et al.
Published: (2024)
by: Zhang, Collin, et al.
Published: (2024)
Explicit Learning and the LLM in Machine Translation
by: Marmonier, Malik, et al.
Published: (2025)
by: Marmonier, Malik, et al.
Published: (2025)
HAMburger: Accelerating LLM Inference via Token Smashing
by: Liu, Jingyu, et al.
Published: (2025)
by: Liu, Jingyu, et al.
Published: (2025)
EXECUTE: A Multilingual Benchmark for LLM Token Understanding
by: Edman, Lukas, et al.
Published: (2025)
by: Edman, Lukas, et al.
Published: (2025)
Token-Level Marginalization for Multi-Label LLM Classifiers
by: Praharaj, Anjaneya, et al.
Published: (2025)
by: Praharaj, Anjaneya, et al.
Published: (2025)
Subword Tokenization Strategies for Kurdish Word Embeddings
by: Salehi, Ali, et al.
Published: (2025)
by: Salehi, Ali, et al.
Published: (2025)
ELLIS Alicante at CQs-Gen 2025: Winning the critical thinking questions shared task: LLM-based question generation and selection
by: Favero, Lucile, et al.
Published: (2025)
by: Favero, Lucile, et al.
Published: (2025)
Extending Token Computation for LLM Reasoning
by: Liao, Bingli, et al.
Published: (2024)
by: Liao, Bingli, et al.
Published: (2024)
Disentangling meaning from language in LLM-based machine translation
by: Lasnier, Théo, et al.
Published: (2026)
by: Lasnier, Théo, et al.
Published: (2026)
Similar Items
-
Democratizing AI Governance: Balancing Expertise and Public Participation
by: Ter-Minassian, Lucile
Published: (2025) -
Dissecting Language Models: Machine Unlearning via Selective Pruning
by: Pochinkov, Nicholas, et al.
Published: (2024) -
Beyond Tokens in Language Models: Interpreting Activations through Text Genre Chunks
by: Benito-Rodriguez, Éloïse, et al.
Published: (2025) -
Investigating Neuron Ablation in Attention Heads: The Case for Peak Activation Centering
by: Pochinkov, Nicholas, et al.
Published: (2024) -
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
by: Veuthey, Jaime Raldua, et al.
Published: (2025)