Causal Estimation of Tokenisation Bias
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lesci, Pietro, Meister, Clara, Hofmann, Thomas, Vlachos, Andreas, Pimentel, Tiago |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Causal Estimation of Memorisation Profiles
von: Lesci, Pietro, et al.
Veröffentlicht: (2024)
von: Lesci, Pietro, et al.
Veröffentlicht: (2024)
AnchorAL: Computationally Efficient Active Learning for Large and Imbalanced Datasets
von: Lesci, Pietro, et al.
Veröffentlicht: (2024)
von: Lesci, Pietro, et al.
Veröffentlicht: (2024)
Tokenisation via Convex Relaxations
von: Tempus, Jan, et al.
Veröffentlicht: (2026)
von: Tempus, Jan, et al.
Veröffentlicht: (2026)
Tokenisation over Bounded Alphabets is Hard
von: Kastreva, Violeta, et al.
Veröffentlicht: (2025)
von: Kastreva, Violeta, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Decoding with Minimum Bayes Risk
von: Daheim, Nico, et al.
Veröffentlicht: (2025)
von: Daheim, Nico, et al.
Veröffentlicht: (2025)
An LLM Feature-based Framework for Dialogue Constructiveness Assessment
von: Zhou, Lexin, et al.
Veröffentlicht: (2024)
von: Zhou, Lexin, et al.
Veröffentlicht: (2024)
What Language is This? Ask Your Tokenizer
von: Meister, Clara, et al.
Veröffentlicht: (2026)
von: Meister, Clara, et al.
Veröffentlicht: (2026)
Locally Typical Sampling
von: Meister, Clara, et al.
Veröffentlicht: (2022)
von: Meister, Clara, et al.
Veröffentlicht: (2022)
The Role of Ambiguity in Error Prediction via Uncertainty Quantification
von: Staliūnaitė, Ieva Raminta, et al.
Veröffentlicht: (2026)
von: Staliūnaitė, Ieva Raminta, et al.
Veröffentlicht: (2026)
Ev2R: Evaluating Evidence Retrieval in Automated Fact-Checking
von: Akhtar, Mubashara, et al.
Veröffentlicht: (2024)
von: Akhtar, Mubashara, et al.
Veröffentlicht: (2024)
Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
PRECISE: Reducing the Bias of LLM Evaluations Using Prediction-Powered Ranking Estimation
von: Divekar, Abhishek, et al.
Veröffentlicht: (2026)
von: Divekar, Abhishek, et al.
Veröffentlicht: (2026)
Explicit Word Density Estimation for Language Modelling
von: Andonov, Jovan, et al.
Veröffentlicht: (2024)
von: Andonov, Jovan, et al.
Veröffentlicht: (2024)
Text Rationalization for Robust Causal Effect Estimation
von: Zhang, Lijinghua, et al.
Veröffentlicht: (2025)
von: Zhang, Lijinghua, et al.
Veröffentlicht: (2025)
RCT Rejection Sampling for Causal Estimation Evaluation
von: Keith, Katherine A., et al.
Veröffentlicht: (2023)
von: Keith, Katherine A., et al.
Veröffentlicht: (2023)
How Causal Abstraction Underpins Computational Explanation
von: Geiger, Atticus, et al.
Veröffentlicht: (2025)
von: Geiger, Atticus, et al.
Veröffentlicht: (2025)
Tokenisation is NP-Complete
von: Whittington, Philip, et al.
Veröffentlicht: (2024)
von: Whittington, Philip, et al.
Veröffentlicht: (2024)
Relative Bias: A Comparative Framework for Quantifying Bias in LLMs
von: Arbabi, Alireza, et al.
Veröffentlicht: (2025)
von: Arbabi, Alireza, et al.
Veröffentlicht: (2025)
A Hitchhiker's Guide to Scaling Law Estimation
von: Choshen, Leshem, et al.
Veröffentlicht: (2024)
von: Choshen, Leshem, et al.
Veröffentlicht: (2024)
More Thinking, More Bias: Length-Driven Position Bias in Reasoning Models
von: Wang, Xiao
Veröffentlicht: (2026)
von: Wang, Xiao
Veröffentlicht: (2026)
CausalARC: Abstract Reasoning with Causal World Models
von: Maasch, Jacqueline, et al.
Veröffentlicht: (2025)
von: Maasch, Jacqueline, et al.
Veröffentlicht: (2025)
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
von: Huang, Jing, et al.
Veröffentlicht: (2025)
von: Huang, Jing, et al.
Veröffentlicht: (2025)
Token Homogenization under Positional Bias
von: Yusupov, Viacheslav, et al.
Veröffentlicht: (2025)
von: Yusupov, Viacheslav, et al.
Veröffentlicht: (2025)
The Impact of Inference Acceleration on Bias of LLMs
von: Kirsten, Elisabeth, et al.
Veröffentlicht: (2024)
von: Kirsten, Elisabeth, et al.
Veröffentlicht: (2024)
CausalVLBench: Benchmarking Visual Causal Reasoning in Large Vision-Language Models
von: Komanduri, Aneesh, et al.
Veröffentlicht: (2025)
von: Komanduri, Aneesh, et al.
Veröffentlicht: (2025)
ByteSpan: Information-Driven Subword Tokenisation
von: Goriely, Zébulon, et al.
Veröffentlicht: (2025)
von: Goriely, Zébulon, et al.
Veröffentlicht: (2025)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
LLM4Causal: Democratized Causal Tools for Everyone via Large Language Model
von: Jiang, Haitao, et al.
Veröffentlicht: (2023)
von: Jiang, Haitao, et al.
Veröffentlicht: (2023)
Self-Exploring Language Models for Explainable Link Forecasting on Temporal Graphs via Reinforcement Learning
von: Ding, Zifeng, et al.
Veröffentlicht: (2025)
von: Ding, Zifeng, et al.
Veröffentlicht: (2025)
On the Inductive Bias of Stacking Towards Improving Reasoning
von: Saunshi, Nikunj, et al.
Veröffentlicht: (2024)
von: Saunshi, Nikunj, et al.
Veröffentlicht: (2024)
Understanding and Mitigating Tokenization Bias in Language Models
von: Phan, Buu, et al.
Veröffentlicht: (2024)
von: Phan, Buu, et al.
Veröffentlicht: (2024)
Causal Evaluation of Language Models
von: Chen, Sirui, et al.
Veröffentlicht: (2024)
von: Chen, Sirui, et al.
Veröffentlicht: (2024)
Causality for Large Language Models
von: Wu, Anpeng, et al.
Veröffentlicht: (2024)
von: Wu, Anpeng, et al.
Veröffentlicht: (2024)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2023)
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2023)
How to Compute the Probability of a Word
von: Pimentel, Tiago, et al.
Veröffentlicht: (2024)
von: Pimentel, Tiago, et al.
Veröffentlicht: (2024)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
von: Yang, Jinming, et al.
Veröffentlicht: (2026)
von: Yang, Jinming, et al.
Veröffentlicht: (2026)
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2026)
Mitigating Extrinsic Gender Bias for Bangla Classification Tasks
von: Joy, Sajib Kumar Saha, et al.
Veröffentlicht: (2024)
von: Joy, Sajib Kumar Saha, et al.
Veröffentlicht: (2024)
Tokenised Flow Matching for Hierarchical Simulation Based Inference
von: Charles, Giovanni, et al.
Veröffentlicht: (2026)
von: Charles, Giovanni, et al.
Veröffentlicht: (2026)
Causally-Enhanced Reinforcement Policy Optimization
von: Wang, Xiangqi, et al.
Veröffentlicht: (2025)
von: Wang, Xiangqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Causal Estimation of Memorisation Profiles
von: Lesci, Pietro, et al.
Veröffentlicht: (2024) -
AnchorAL: Computationally Efficient Active Learning for Large and Imbalanced Datasets
von: Lesci, Pietro, et al.
Veröffentlicht: (2024) -
Tokenisation via Convex Relaxations
von: Tempus, Jan, et al.
Veröffentlicht: (2026) -
Tokenisation over Bounded Alphabets is Hard
von: Kastreva, Violeta, et al.
Veröffentlicht: (2025) -
Uncertainty-Aware Decoding with Minimum Bayes Risk
von: Daheim, Nico, et al.
Veröffentlicht: (2025)