Improving the Distributional Alignment of LLMs using Supervision
Fuente:
arXiv
Guardado en:
| Autores principales: | Kambhatla, Gauri, Gautam, Sanjana, Zhang, Angela, Liu, Alex, Srinivasan, Ravi, Li, Junyi Jessy, Lease, Matthew |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Promoting Constructive Deliberation: Reframing for Receptiveness
por: Kambhatla, Gauri, et al.
Publicado: (2024)
por: Kambhatla, Gauri, et al.
Publicado: (2024)
Measuring Lexical Diversity of Synthetic Data Generated through Fine-Grained Persona Prompting
por: Kambhatla, Gauri, et al.
Publicado: (2025)
por: Kambhatla, Gauri, et al.
Publicado: (2025)
Wrapper Boxes: Faithful Attribution of Model Predictions to Training Data
por: Su, Yiheng, et al.
Publicado: (2023)
por: Su, Yiheng, et al.
Publicado: (2023)
Robo-Instruct: Simulator-Augmented Instruction Alignment For Finetuning Code LLMs
por: Hu, Zichao, et al.
Publicado: (2024)
por: Hu, Zichao, et al.
Publicado: (2024)
How Researchers Navigate Accountability, Transparency, and Trust When Using AI Tools in Early-Stage Research: A Think-Aloud Study
por: Gautam, Sanjana, et al.
Publicado: (2026)
por: Gautam, Sanjana, et al.
Publicado: (2026)
LLMs Lean on Priors, Not Programming Language Semantics
por: Thimmaiah, Aditya, et al.
Publicado: (2025)
por: Thimmaiah, Aditya, et al.
Publicado: (2025)
Which course? Discourse! Teaching Discourse and Generation in the Era of LLMs
por: Li, Junyi Jessy, et al.
Publicado: (2026)
por: Li, Junyi Jessy, et al.
Publicado: (2026)
CREATE: Testing LLMs for Associative Creativity
por: Wadhwa, Manya, et al.
Publicado: (2026)
por: Wadhwa, Manya, et al.
Publicado: (2026)
Blind Spots and Biases: Exploring the Role of Annotator Cognitive Biases in NLP
por: Gautam, Sanjana, et al.
Publicado: (2024)
por: Gautam, Sanjana, et al.
Publicado: (2024)
LLM REgression with a Latent Iterative State Head
por: Su, Yiheng, et al.
Publicado: (2026)
por: Su, Yiheng, et al.
Publicado: (2026)
Benchmark Transparency: Measuring the Impact of Data on Evaluation
por: Kovatchev, Venelin, et al.
Publicado: (2024)
por: Kovatchev, Venelin, et al.
Publicado: (2024)
Language Models (Mostly) Do Not Consider Emotion Triggers When Predicting Emotion
por: Singh, Smriti, et al.
Publicado: (2023)
por: Singh, Smriti, et al.
Publicado: (2023)
Diverse, but Divisive: LLMs Can Exaggerate Gender Differences in Opinion Related to Harms of Misinformation
por: Neumann, Terrence, et al.
Publicado: (2024)
por: Neumann, Terrence, et al.
Publicado: (2024)
Is It JUST Semantics? A Case Study of Discourse Particle Understanding in LLMs
por: Sheffield, William, et al.
Publicado: (2025)
por: Sheffield, William, et al.
Publicado: (2025)
Behavioral Analysis of Information Salience in Large Language Models
por: Trienes, Jan, et al.
Publicado: (2025)
por: Trienes, Jan, et al.
Publicado: (2025)
WUGNECTIVES: Novel Entity Inferences of Language Models from Discourse Connectives
por: Brubaker, Daniel, et al.
Publicado: (2025)
por: Brubaker, Daniel, et al.
Publicado: (2025)
Strategic Dialogue Assessment: The Crooked Path to Innocence
por: Zheng, Anshun Asher, et al.
Publicado: (2025)
por: Zheng, Anshun Asher, et al.
Publicado: (2025)
Using Natural Language Explanations to Rescale Human Judgments
por: Wadhwa, Manya, et al.
Publicado: (2023)
por: Wadhwa, Manya, et al.
Publicado: (2023)
Learning to Refine with Fine-Grained Natural Language Feedback
por: Wadhwa, Manya, et al.
Publicado: (2024)
por: Wadhwa, Manya, et al.
Publicado: (2024)
Detection and Measurement of Syntactic Templates in Generated Text
por: Shaib, Chantal, et al.
Publicado: (2024)
por: Shaib, Chantal, et al.
Publicado: (2024)
Counterfactual Probing for the Influence of Affect and Specificity on Intergroup Bias
por: Govindarajan, Venkata S, et al.
Publicado: (2023)
por: Govindarajan, Venkata S, et al.
Publicado: (2023)
Chord Embeddings: Analyzing What They Capture and Their Role for Next Chord Prediction and Artist Attribute Prediction
por: Lahnala, Allison, et al.
Publicado: (2021)
por: Lahnala, Allison, et al.
Publicado: (2021)
Do they mean 'us'? Interpreting Referring Expressions in Intergroup Bias
por: Govindarajan, Venkata S, et al.
Publicado: (2024)
por: Govindarajan, Venkata S, et al.
Publicado: (2024)
Caught in the Web of Words: Do LLMs Fall for Spin in Medical Literature?
por: Yun, Hye Sun, et al.
Publicado: (2025)
por: Yun, Hye Sun, et al.
Publicado: (2025)
Who Owns Creativity and Who Does the Work? Trade-offs in LLM-Supported Research Ideation
por: Liu, Houjiang, et al.
Publicado: (2026)
por: Liu, Houjiang, et al.
Publicado: (2026)
Help! Need Advice on Identifying Advice
por: Govindarajan, Venkata Subrahmanyan, et al.
Publicado: (2020)
por: Govindarajan, Venkata Subrahmanyan, et al.
Publicado: (2020)
Multimodal QUD: Inquisitive Questions from Scientific Figures
por: Wu, Yating, et al.
Publicado: (2026)
por: Wu, Yating, et al.
Publicado: (2026)
Which questions should I answer? Salience Prediction of Inquisitive Questions
por: Wu, Yating, et al.
Publicado: (2024)
por: Wu, Yating, et al.
Publicado: (2024)
SPRI: Aligning Large Language Models with Context-Situated Principles
por: Zhan, Hongli, et al.
Publicado: (2025)
por: Zhan, Hongli, et al.
Publicado: (2025)
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
por: Wadhwa, Manya, et al.
Publicado: (2025)
por: Wadhwa, Manya, et al.
Publicado: (2025)
Human-centered NLP Fact-checking: Co-Designing with Fact-checkers using Matchmaking for AI
por: Liu, Houjiang, et al.
Publicado: (2023)
por: Liu, Houjiang, et al.
Publicado: (2023)
Finding Pareto Trade-offs in Fair and Accurate Detection of Toxic Speech
por: Gupta, Soumyajit, et al.
Publicado: (2022)
por: Gupta, Soumyajit, et al.
Publicado: (2022)
Self-Alignment: Improving Alignment of Cultural Values in LLMs via In-Context Learning
por: Choenni, Rochelle, et al.
Publicado: (2024)
por: Choenni, Rochelle, et al.
Publicado: (2024)
QUDsim: Quantifying Discourse Similarities in LLM-Generated Text
por: Namuduri, Ramya, et al.
Publicado: (2025)
por: Namuduri, Ramya, et al.
Publicado: (2025)
How people talk about each other: Modeling Generalized Intergroup Bias and Emotion
por: Govindarajan, Venkata S, et al.
Publicado: (2022)
por: Govindarajan, Venkata S, et al.
Publicado: (2022)
Improving Task Diversity in Label Efficient Supervised Finetuning of LLMs
por: Arabelly, Abhinav, et al.
Publicado: (2025)
por: Arabelly, Abhinav, et al.
Publicado: (2025)
LABELING COPILOT: A Deep Research Agent for Automated Data Curation in Computer Vision
por: Ganguly, Debargha, et al.
Publicado: (2025)
por: Ganguly, Debargha, et al.
Publicado: (2025)
Large Language Models are Capable of Offering Cognitive Reappraisal, if Guided
por: Zhan, Hongli, et al.
Publicado: (2024)
por: Zhan, Hongli, et al.
Publicado: (2024)
Challenges in Adapting Multilingual LLMs to Low-Resource Languages using LoRA PEFT Tuning
por: Khade, Omkar, et al.
Publicado: (2024)
por: Khade, Omkar, et al.
Publicado: (2024)
Large Language Models Produce Responses Perceived to be Empathic
por: Lee, Yoon Kyung, et al.
Publicado: (2024)
por: Lee, Yoon Kyung, et al.
Publicado: (2024)
Ejemplares similares
-
Promoting Constructive Deliberation: Reframing for Receptiveness
por: Kambhatla, Gauri, et al.
Publicado: (2024) -
Measuring Lexical Diversity of Synthetic Data Generated through Fine-Grained Persona Prompting
por: Kambhatla, Gauri, et al.
Publicado: (2025) -
Wrapper Boxes: Faithful Attribution of Model Predictions to Training Data
por: Su, Yiheng, et al.
Publicado: (2023) -
Robo-Instruct: Simulator-Augmented Instruction Alignment For Finetuning Code LLMs
por: Hu, Zichao, et al.
Publicado: (2024) -
How Researchers Navigate Accountability, Transparency, and Trust When Using AI Tools in Early-Stage Research: A Think-Aloud Study
por: Gautam, Sanjana, et al.
Publicado: (2026)