Using Embedding Models to Improve Probabilistic Race Prediction
Fuente:
arXiv
Guardado en:
| Autores principales: | Dasanaike, Noah, Imai, Kosuke |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Large Language Models Naively Recover Ethnicity from Individual Records
por: Dasanaike, Noah
Publicado: (2026)
por: Dasanaike, Noah
Publicado: (2026)
EnsembleLink: Accurate Record Linkage Without Training Data
por: Dasanaike, Noah
Publicado: (2026)
por: Dasanaike, Noah
Publicado: (2026)
Causal Representation Learning with Generative Artificial Intelligence: Application to Texts as Treatments
por: Imai, Kosuke, et al.
Publicado: (2024)
por: Imai, Kosuke, et al.
Publicado: (2024)
Out-of-the-Box Conditional Text Embeddings from Large Language Models
por: Yamada, Kosuke, et al.
Publicado: (2025)
por: Yamada, Kosuke, et al.
Publicado: (2025)
Hallucination Detection: A Probabilistic Framework Using Embeddings Distance Analysis
por: Ricco, Emanuele, et al.
Publicado: (2025)
por: Ricco, Emanuele, et al.
Publicado: (2025)
Improving Text Embeddings for Smaller Language Models Using Contrastive Fine-tuning
por: Ukarapol, Trapoom, et al.
Publicado: (2024)
por: Ukarapol, Trapoom, et al.
Publicado: (2024)
On the Influence of Gender and Race in Romantic Relationship Prediction from Large Language Models
por: Sancheti, Abhilasha, et al.
Publicado: (2024)
por: Sancheti, Abhilasha, et al.
Publicado: (2024)
Hierarchical Text Classification Using Black Box Large Language Models
por: Yoshimura, Kosuke, et al.
Publicado: (2025)
por: Yoshimura, Kosuke, et al.
Publicado: (2025)
Improving the Accuracy and Efficiency of Legal Document Tagging with Large Language Models and Instruction Prompts
por: Johnson, Emily, et al.
Publicado: (2025)
por: Johnson, Emily, et al.
Publicado: (2025)
SiLVERScore: Semantically-Aware Embeddings for Sign Language Generation Evaluation
por: Imai, Saki, et al.
Publicado: (2025)
por: Imai, Saki, et al.
Publicado: (2025)
Third-Party Language Model Performance Prediction from Instruction
por: Nadkarni, Rahul, et al.
Publicado: (2024)
por: Nadkarni, Rahul, et al.
Publicado: (2024)
Repetition Improves Language Model Embeddings
por: Springer, Jacob Mitchell, et al.
Publicado: (2024)
por: Springer, Jacob Mitchell, et al.
Publicado: (2024)
Vec2Summ: Text Summarization via Probabilistic Sentence Embeddings
por: Li, Mao, et al.
Publicado: (2025)
por: Li, Mao, et al.
Publicado: (2025)
Improving Probabilistic Models in Text Classification via Active Learning
por: Bosley, Mitchell, et al.
Publicado: (2022)
por: Bosley, Mitchell, et al.
Publicado: (2022)
Probabilistic Aggregation and Targeted Embedding Optimization for Collective Moral Reasoning in Large Language Models
por: Yuan, Chenchen, et al.
Publicado: (2025)
por: Yuan, Chenchen, et al.
Publicado: (2025)
Improving Text Embeddings with Large Language Models
por: Wang, Liang, et al.
Publicado: (2023)
por: Wang, Liang, et al.
Publicado: (2023)
ATEB: Evaluating and Improving Advanced NLP Tasks for Text Embedding Models
por: Han, Simeng, et al.
Publicado: (2025)
por: Han, Simeng, et al.
Publicado: (2025)
Structured Context Recomposition for Large Language Models Using Probabilistic Layer Realignment
por: Teel, Jonathan, et al.
Publicado: (2025)
por: Teel, Jonathan, et al.
Publicado: (2025)
Automated Essay Scoring Using Grammatical Variety and Errors with Multi-Task Learning and Item Response Theory
por: Doi, Kosuke, et al.
Publicado: (2024)
por: Doi, Kosuke, et al.
Publicado: (2024)
Enhanced Data Race Prediction Through Modular Reasoning
por: Ang, Zhendong, et al.
Publicado: (2025)
por: Ang, Zhendong, et al.
Publicado: (2025)
Racing Thoughts: Explaining Contextualization Errors in Large Language Models
por: Lepori, Michael A., et al.
Publicado: (2024)
por: Lepori, Michael A., et al.
Publicado: (2024)
Estimating Racial Disparities When Race is Not Observed
por: McCartan, Cory, et al.
Publicado: (2023)
por: McCartan, Cory, et al.
Publicado: (2023)
Improving Quotation Attribution with Fictional Character Embeddings
por: Michel, Gaspard, et al.
Publicado: (2024)
por: Michel, Gaspard, et al.
Publicado: (2024)
Demographic Attributes Prediction from Speech Using WavLM Embeddings
por: Yang, Yuchen, et al.
Publicado: (2025)
por: Yang, Yuchen, et al.
Publicado: (2025)
Transfer Learning for Text Diffusion Models
por: Han, Kehang, et al.
Publicado: (2024)
por: Han, Kehang, et al.
Publicado: (2024)
Improving General Text Embedding Model: Tackling Task Conflict and Data Imbalance through Model Merging
por: Li, Mingxin, et al.
Publicado: (2024)
por: Li, Mingxin, et al.
Publicado: (2024)
Initialization of Large Language Models via Reparameterization to Mitigate Loss Spikes
por: Nishida, Kosuke, et al.
Publicado: (2024)
por: Nishida, Kosuke, et al.
Publicado: (2024)
MAGNET: Improving the Multilingual Fairness of Language Models with Adaptive Gradient-Based Tokenization
por: Ahia, Orevaoghene, et al.
Publicado: (2024)
por: Ahia, Orevaoghene, et al.
Publicado: (2024)
Evaluating $n$-Gram Novelty of Language Models Using Rusty-DAWG
por: Merrill, William, et al.
Publicado: (2024)
por: Merrill, William, et al.
Publicado: (2024)
Improving Self Consistency in LLMs through Probabilistic Tokenization
por: Sathe, Ashutosh, et al.
Publicado: (2024)
por: Sathe, Ashutosh, et al.
Publicado: (2024)
Improving Sentence Embeddings with Automatic Generation of Training Data Using Few-shot Examples
por: Sato, Soma, et al.
Publicado: (2024)
por: Sato, Soma, et al.
Publicado: (2024)
SemPA: Improving Sentence Embeddings of Large Language Models through Semantic Preference Alignment
por: Chen, Ziyang, et al.
Publicado: (2026)
por: Chen, Ziyang, et al.
Publicado: (2026)
Sequence Repetition Enhances Token Embeddings and Improves Sequence Labeling with Decoder-only Language Models
por: Kukić, Matija Luka, et al.
Publicado: (2026)
por: Kukić, Matija Luka, et al.
Publicado: (2026)
Marathon: A Race Through the Realm of Long Context with Large Language Models
por: Zhang, Lei, et al.
Publicado: (2023)
por: Zhang, Lei, et al.
Publicado: (2023)
Race, Ethnicity and Their Implication on Bias in Large Language Models
por: Hu, Shiyue, et al.
Publicado: (2026)
por: Hu, Shiyue, et al.
Publicado: (2026)
$M^3$ Scaling Law: Optimizing Multi-Epoch, Multi-Lingual, and Multi-Stage Training for Low-Resource Language Models
por: Akimoto, Kosuke, et al.
Publicado: (2024)
por: Akimoto, Kosuke, et al.
Publicado: (2024)
Improving Embedding Accuracy for Document Retrieval Using Entity Relationship Maps and Model-Aware Contrastive Sampling
por: Aviss, Thea
Publicado: (2024)
por: Aviss, Thea
Publicado: (2024)
Leveraging Large Language Models for Career Mobility Analysis: A Study of Gender, Race, and Job Change Using U.S. Online Resume Profiles
por: Achananuparp, Palakorn, et al.
Publicado: (2025)
por: Achananuparp, Palakorn, et al.
Publicado: (2025)
Context-level Language Modeling by Learning Predictive Context Embeddings
por: Dai, Beiya, et al.
Publicado: (2025)
por: Dai, Beiya, et al.
Publicado: (2025)
Partial Colexifications Improve Concept Embeddings
por: Rubehn, Arne, et al.
Publicado: (2025)
por: Rubehn, Arne, et al.
Publicado: (2025)
Ejemplares similares
-
Large Language Models Naively Recover Ethnicity from Individual Records
por: Dasanaike, Noah
Publicado: (2026) -
EnsembleLink: Accurate Record Linkage Without Training Data
por: Dasanaike, Noah
Publicado: (2026) -
Causal Representation Learning with Generative Artificial Intelligence: Application to Texts as Treatments
por: Imai, Kosuke, et al.
Publicado: (2024) -
Out-of-the-Box Conditional Text Embeddings from Large Language Models
por: Yamada, Kosuke, et al.
Publicado: (2025) -
Hallucination Detection: A Probabilistic Framework Using Embeddings Distance Analysis
por: Ricco, Emanuele, et al.
Publicado: (2025)