LIME-LLM: Probing Models with Fluent Counterfactuals, Not Broken Text
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Mihaila, George, Polat, Suleyman Olcay, Nemkova, Poli, Sharma, Himanshu, Urs, Namratha V., Albert, Mark V. |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Comparing LLM Text Annotation Skills: A Study on Human Rights Violations in Social Media Data
par: Nemkova, Poli Apollinaire, et autres
Publié: (2025)
par: Nemkova, Poli Apollinaire, et autres
Publié: (2025)
Synthetic Adaptive Guided Embeddings (SAGE): A Novel Knowledge Distillation Method
par: Polat, Suleyman Olcay, et autres
Publié: (2025)
par: Polat, Suleyman Olcay, et autres
Publié: (2025)
Cross-Lingual Stability and Bias in Instruction-Tuned Language Models for Humanitarian NLP
par: Nemkova, Poli, et autres
Publié: (2025)
par: Nemkova, Poli, et autres
Publié: (2025)
Do Large Language Models Know Conflict? Investigating Parametric vs. Non-Parametric Knowledge of LLMs for Conflict Forecasting
par: Nemkova, Apollinaire Poli, et autres
Publié: (2025)
par: Nemkova, Apollinaire Poli, et autres
Publié: (2025)
Fluent dreaming for language models
par: Thompson, T. Ben, et autres
Publié: (2024)
par: Thompson, T. Ben, et autres
Publié: (2024)
Fluent but Unfeeling: The Emotional Blind Spots of Language Models
par: Shu, Bangzhao, et autres
Publié: (2025)
par: Shu, Bangzhao, et autres
Publié: (2025)
LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings
par: Sztwiertnia, Sebastian, et autres
Publié: (2025)
par: Sztwiertnia, Sebastian, et autres
Publié: (2025)
FLRT: Fluent Student-Teacher Redteaming
par: Thompson, T. Ben, et autres
Publié: (2024)
par: Thompson, T. Ben, et autres
Publié: (2024)
The Effect of Model Size on LLM Post-hoc Explainability via LIME
par: Heyen, Henning, et autres
Publié: (2024)
par: Heyen, Henning, et autres
Publié: (2024)
Towards Automated Situation Awareness: A RAG-Based Framework for Peacebuilding Reports
par: Nemkova, Poli A., et autres
Publié: (2025)
par: Nemkova, Poli A., et autres
Publié: (2025)
Whose Good, Whose Place? The Moral Geography of Agentic AI for Social Good
par: Nemkova, Poli, et autres
Publié: (2026)
par: Nemkova, Poli, et autres
Publié: (2026)
A Comparative Analysis of Counterfactual Explanation Methods for Text Classifiers
par: McAleese, Stephen, et autres
Publié: (2024)
par: McAleese, Stephen, et autres
Publié: (2024)
Counterfactual Probing for Hallucination Detection and Mitigation in Large Language Models
par: Feng, Yijun
Publié: (2025)
par: Feng, Yijun
Publié: (2025)
Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages
par: Samuel, David, et autres
Publié: (2025)
par: Samuel, David, et autres
Publié: (2025)
LLM as a Broken Telephone: Iterative Generation Distorts Information
par: Mohamed, Amr, et autres
Publié: (2025)
par: Mohamed, Amr, et autres
Publié: (2025)
Organic Data-Driven Approach for Turkish Grammatical Error Correction and LLMs
par: Ersoy, Asım, et autres
Publié: (2024)
par: Ersoy, Asım, et autres
Publié: (2024)
CEval: A Benchmark for Evaluating Counterfactual Text Generation
par: Nguyen, Van Bach, et autres
Publié: (2024)
par: Nguyen, Van Bach, et autres
Publié: (2024)
SHROOM-INDElab at SemEval-2024 Task 6: Zero- and Few-Shot LLM-Based Classification for Hallucination Detection
par: Allen, Bradley P., et autres
Publié: (2024)
par: Allen, Bradley P., et autres
Publié: (2024)
Counterfactual Simulatability of LLM Explanations for Generation Tasks
par: Limpijankit, Marvin, et autres
Publié: (2025)
par: Limpijankit, Marvin, et autres
Publié: (2025)
AI Generated Text Detection
par: Alikhanov, Adilkhan, et autres
Publié: (2026)
par: Alikhanov, Adilkhan, et autres
Publié: (2026)
Multi-Aspect Controllable Text Generation with Disentangled Counterfactual Augmentation
par: Liu, Yi, et autres
Publié: (2024)
par: Liu, Yi, et autres
Publié: (2024)
Reasoning Elicitation in Language Models via Counterfactual Feedback
par: Hüyük, Alihan, et autres
Publié: (2024)
par: Hüyük, Alihan, et autres
Publié: (2024)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
par: Agarwal, Dhruv, et autres
Publié: (2025)
par: Agarwal, Dhruv, et autres
Publié: (2025)
Fixing the Broken Compass: Diagnosing and Improving Inference-Time Reward Modeling
par: Li, Jiachun, et autres
Publié: (2025)
par: Li, Jiachun, et autres
Publié: (2025)
Equal Access, Unequal Interaction: A Counterfactual Audit of LLM Fairness
par: Amiri-Margavi, Alireza, et autres
Publié: (2026)
par: Amiri-Margavi, Alireza, et autres
Publié: (2026)
Can Large Language Models put 2 and 2 together? Probing for Entailed Arithmetical Relationships
par: Panas, D., et autres
Publié: (2024)
par: Panas, D., et autres
Publié: (2024)
COMMENTATOR: A Code-mixed Multilingual Text Annotation Framework
par: Sheth, Rajvee, et autres
Publié: (2024)
par: Sheth, Rajvee, et autres
Publié: (2024)
Presentations are not always linear! GNN meets LLM for Document-to-Presentation Transformation with Attribution
par: Maheshwari, Himanshu, et autres
Publié: (2024)
par: Maheshwari, Himanshu, et autres
Publié: (2024)
Aligning Large Language Models with Counterfactual DPO
par: Butcher, Bradley
Publié: (2024)
par: Butcher, Bradley
Publié: (2024)
Probing the Capacity of Language Model Agents to Operationalize Disparate Experiential Context Despite Distraction
par: George, Sonny, et autres
Publié: (2024)
par: George, Sonny, et autres
Publié: (2024)
An LLM Maturity Model for Reliable and Transparent Text-to-Query
par: Yu, Lei, et autres
Publié: (2024)
par: Yu, Lei, et autres
Publié: (2024)
Thinking Fast, Thinking Wrong: Intuitiveness Modulates LLM Counterfactual Reasoning in Policy Evaluation
par: He, Yanjie
Publié: (2026)
par: He, Yanjie
Publié: (2026)
From Graph Retrieval to Schema Realization: Counterfactual Validation for Text-to-SPARQL over Heterogeneous Knowledge Graphs
par: Zhao, Yang, et autres
Publié: (2025)
par: Zhao, Yang, et autres
Publié: (2025)
A Human-in-the-Loop Approach to Improving Cross-Text Prosody Transfer
par: Maurya, Himanshu, et autres
Publié: (2024)
par: Maurya, Himanshu, et autres
Publié: (2024)
RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals
par: Reber, David, et autres
Publié: (2024)
par: Reber, David, et autres
Publié: (2024)
CLOMO: Counterfactual Logical Modification with Large Language Models
par: Huang, Yinya, et autres
Publié: (2023)
par: Huang, Yinya, et autres
Publié: (2023)
Reverse Probing: Supervised Token-level Uncertainty Quantification for Large Language Models in Clinical Text
par: Xiao, Bushi, et autres
Publié: (2026)
par: Xiao, Bushi, et autres
Publié: (2026)
Zero-shot LLM-guided Counterfactual Generation: A Case Study on NLP Model Evaluation
par: Bhattacharjee, Amrita, et autres
Publié: (2024)
par: Bhattacharjee, Amrita, et autres
Publié: (2024)
Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation
par: Wang, Qianli, et autres
Publié: (2026)
par: Wang, Qianli, et autres
Publié: (2026)
Cross-lingual Editing in Multilingual Language Models
par: Beniwal, Himanshu, et autres
Publié: (2024)
par: Beniwal, Himanshu, et autres
Publié: (2024)
Documents similaires
-
Comparing LLM Text Annotation Skills: A Study on Human Rights Violations in Social Media Data
par: Nemkova, Poli Apollinaire, et autres
Publié: (2025) -
Synthetic Adaptive Guided Embeddings (SAGE): A Novel Knowledge Distillation Method
par: Polat, Suleyman Olcay, et autres
Publié: (2025) -
Cross-Lingual Stability and Bias in Instruction-Tuned Language Models for Humanitarian NLP
par: Nemkova, Poli, et autres
Publié: (2025) -
Do Large Language Models Know Conflict? Investigating Parametric vs. Non-Parametric Knowledge of LLMs for Conflict Forecasting
par: Nemkova, Apollinaire Poli, et autres
Publié: (2025) -
Fluent dreaming for language models
par: Thompson, T. Ben, et autres
Publié: (2024)