Benchmarking Debiasing Methods for LLM-based Parameter Estimates
Fuente:
arXiv
Saved in:
| Main Authors: | de Pieuchon, Nicolas Audinet, Daoud, Adel, Jerzak, Connor T., Johansson, Moa, Johansson, Richard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Large Language Models (or Humans) Disentangle Text?
by: de Pieuchon, Nicolas Audinet, et al.
Published: (2024)
by: de Pieuchon, Nicolas Audinet, et al.
Published: (2024)
Detecting and Mitigating Treatment Leakage in Text-Based Causal Inference: Distillation and Sensitivity Analysis
by: Daoud, Adel, et al.
Published: (2025)
by: Daoud, Adel, et al.
Published: (2025)
Remote Auditing: Design-based Tests of Randomization, Selection, and Missingness with Broadly Accessible Satellite Imagery
by: Jerzak, Connor T., et al.
Published: (2025)
by: Jerzak, Connor T., et al.
Published: (2025)
Fact Recall, Heuristics or Pure Guesswork? Precise Interpretations of Language Models for Fact Completion
by: Saynova, Denitsa, et al.
Published: (2024)
by: Saynova, Denitsa, et al.
Published: (2024)
Debiasing Machine Learning Predictions for Causal Inference Without Additional Ground Truth Data: "One Map, Many Trials" in Satellite-Driven Poverty Analysis
by: Pettersson, Markus B., et al.
Published: (2025)
by: Pettersson, Markus B., et al.
Published: (2025)
A Scoping Review of Earth Observation and Machine Learning for Causal Inference: Implications for the Geography of Poverty
by: Sakamoto, Kazuki, et al.
Published: (2024)
by: Sakamoto, Kazuki, et al.
Published: (2024)
What Happens to a Dataset Transformed by a Projection-based Concept Removal Method?
by: Johansson, Richard
Published: (2024)
by: Johansson, Richard
Published: (2024)
Reasoning in Transformers -- Mitigating Spurious Correlations and Reasoning Shortcuts
by: Enström, Daniel, et al.
Published: (2024)
by: Enström, Daniel, et al.
Published: (2024)
PACE: Procedural Abstractions for Communicating Efficiently
by: Thomas, Jonathan D., et al.
Published: (2024)
by: Thomas, Jonathan D., et al.
Published: (2024)
Specify What? Enhancing Neural Specification Synthesis by Symbolic Methods
by: Granberry, George, et al.
Published: (2024)
by: Granberry, George, et al.
Published: (2024)
Platonic Representations for Poverty Mapping: Unified Vision-Language Codes or Agent-Induced Novelty?
by: Murugaboopathy, Satiyabooshan, et al.
Published: (2025)
by: Murugaboopathy, Satiyabooshan, et al.
Published: (2025)
Chinese vs. World Bank Development Projects: Insights from Earth Observation and Computer Vision on Wealth Gains in Africa, 2002-2013
by: Daoud, Adel, et al.
Published: (2025)
by: Daoud, Adel, et al.
Published: (2025)
Effect Heterogeneity with Earth Observation in Randomized Controlled Trials: Exploring the Role of Data, Model, and Evaluation Metric Choice
by: Jerzak, Connor T., et al.
Published: (2024)
by: Jerzak, Connor T., et al.
Published: (2024)
Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs
by: Balter, Samuel G., et al.
Published: (2026)
by: Balter, Samuel G., et al.
Published: (2026)
How Well Do Large Language Models Disambiguate Swedish Words?
by: Johansson, Richard
Published: (2024)
by: Johansson, Richard
Published: (2024)
Identifying Non-Replicable Social Science Studies with Language Models
by: Saynova, Denitsa, et al.
Published: (2025)
by: Saynova, Denitsa, et al.
Published: (2025)
Learning Efficient Recursive Numeral Systems via Reinforcement Learning
by: Silvi, Andrea, et al.
Published: (2024)
by: Silvi, Andrea, et al.
Published: (2024)
Optimizing Multi-Scale Representations to Detect Effect Heterogeneity Using Earth Observation and Computer Vision: Applications to Two Anti-Poverty RCTs
by: Zhu, Fucheng Warren, et al.
Published: (2024)
by: Zhu, Fucheng Warren, et al.
Published: (2024)
Recursive numeral systems are highly regular and easy to process
by: Prasertsom, Ponrawee, et al.
Published: (2025)
by: Prasertsom, Ponrawee, et al.
Published: (2025)
Deciphering the Interplay of Parametric and Non-parametric Memory in Retrieval-augmented Language Models
by: Farahani, Mehrdad, et al.
Published: (2024)
by: Farahani, Mehrdad, et al.
Published: (2024)
Evaluating the relationship between regularity and learnability in recursive numeral systems using Reinforcement Learning
by: Silvi, Andrea, et al.
Published: (2026)
by: Silvi, Andrea, et al.
Published: (2026)
To Copy or Not to Copy: Copying Is Easier to Induce Than Recall
by: Farahani, Mehrdad, et al.
Published: (2026)
by: Farahani, Mehrdad, et al.
Published: (2026)
Queryable LoRA: Instruction-Regularized Routing Over Shared Low-Rank Update Atoms
by: Vaidya, Omatharv Bharat, et al.
Published: (2026)
by: Vaidya, Omatharv Bharat, et al.
Published: (2026)
Fine Tuning Methods for Low-resource Languages
by: Bakkenes, Tim, et al.
Published: (2025)
by: Bakkenes, Tim, et al.
Published: (2025)
Twitch: Learning Abstractions for Equational Theorem Proving
by: Axelrod, Guy, et al.
Published: (2026)
by: Axelrod, Guy, et al.
Published: (2026)
On the Military Applications of Large Language Models
by: Johansson, Satu, et al.
Published: (2025)
by: Johansson, Satu, et al.
Published: (2025)
Logical Negation Augmenting and Debiasing for Prompt-based Methods
by: Li, Yitian, et al.
Published: (2024)
by: Li, Yitian, et al.
Published: (2024)
CUB: Benchmarking Context Utilisation Techniques for Language Models
by: Hagström, Lovisa, et al.
Published: (2025)
by: Hagström, Lovisa, et al.
Published: (2025)
Benchmarking LLM-based Relevance Judgment Methods
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs
by: Restrepo, David, et al.
Published: (2024)
by: Restrepo, David, et al.
Published: (2024)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
by: Yang, Bo, et al.
Published: (2026)
by: Yang, Bo, et al.
Published: (2026)
A Multi-LLM Debiasing Framework
by: Owens, Deonna M., et al.
Published: (2024)
by: Owens, Deonna M., et al.
Published: (2024)
Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models?
by: Modica, Luca, et al.
Published: (2026)
by: Modica, Luca, et al.
Published: (2026)
Bring Your Own Knowledge: A Survey of Methods for LLM Knowledge Expansion
by: Wang, Mingyang, et al.
Published: (2025)
by: Wang, Mingyang, et al.
Published: (2025)
PEFT-Bench: A Parameter-Efficient Fine-Tuning Methods Benchmark
by: Belanec, Robert, et al.
Published: (2025)
by: Belanec, Robert, et al.
Published: (2025)
Rethinking Prompt-based Debiasing in Large Language Models
by: Yang, Xinyi, et al.
Published: (2025)
by: Yang, Xinyi, et al.
Published: (2025)
Learning Approximate and Exact Numeral Systems via Reinforcement Learning
by: Carlsson, Emil, et al.
Published: (2021)
by: Carlsson, Emil, et al.
Published: (2021)
Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
by: Zhou, Hongli, et al.
Published: (2026)
by: Zhou, Hongli, et al.
Published: (2026)
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
by: Habba, Eliya, et al.
Published: (2026)
by: Habba, Eliya, et al.
Published: (2026)
Beyond Interaction Effects: Two Logics for Studying Population Inequalities
by: Daoud, Adel
Published: (2025)
by: Daoud, Adel
Published: (2025)
Similar Items
-
Can Large Language Models (or Humans) Disentangle Text?
by: de Pieuchon, Nicolas Audinet, et al.
Published: (2024) -
Detecting and Mitigating Treatment Leakage in Text-Based Causal Inference: Distillation and Sensitivity Analysis
by: Daoud, Adel, et al.
Published: (2025) -
Remote Auditing: Design-based Tests of Randomization, Selection, and Missingness with Broadly Accessible Satellite Imagery
by: Jerzak, Connor T., et al.
Published: (2025) -
Fact Recall, Heuristics or Pure Guesswork? Precise Interpretations of Language Models for Fact Completion
by: Saynova, Denitsa, et al.
Published: (2024) -
Debiasing Machine Learning Predictions for Causal Inference Without Additional Ground Truth Data: "One Map, Many Trials" in Satellite-Driven Poverty Analysis
by: Pettersson, Markus B., et al.
Published: (2025)