Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Jain, Shomik, Lanchantin, Jack, Nickel, Maximilian, Ross, Candace, Ullrich, Karen, Wilson, Ashia, Watson-Daniels, Jamelle |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
As an AI Language Model, "Yes I Would Recommend Calling the Police": Norm Inconsistency in LLM Decision-Making
by: Jain, Shomik, et al.
Published: (2024)
by: Jain, Shomik, et al.
Published: (2024)
Algorithmic Fairness and Color-blind Racism: Navigating the Intersection
by: Watson-Daniels, Jamelle
Published: (2024)
by: Watson-Daniels, Jamelle
Published: (2024)
Scarce Resource Allocations That Rely On Machine Learning Should Be Randomized
by: Jain, Shomik, et al.
Published: (2024)
by: Jain, Shomik, et al.
Published: (2024)
Allocation Multiplicity: Evaluating the Promises of the Rashomon Set
by: Jain, Shomik, et al.
Published: (2025)
by: Jain, Shomik, et al.
Published: (2025)
Automating Transparency Mechanisms in the Judicial System Using LLMs: Opportunities and Challenges
by: Shastri, Ishana, et al.
Published: (2024)
by: Shastri, Ishana, et al.
Published: (2024)
Algorithmic Pluralism: A Structural Approach To Equal Opportunity
by: Jain, Shomik, et al.
Published: (2023)
by: Jain, Shomik, et al.
Published: (2023)
Interaction Context Often Increases Sycophancy in LLMs
by: Jain, Shomik, et al.
Published: (2025)
by: Jain, Shomik, et al.
Published: (2025)
Homogeneous Algorithms Can Reduce Competition in Personalized Pricing
by: Jo, Nathanael, et al.
Published: (2025)
by: Jo, Nathanael, et al.
Published: (2025)
Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space
by: Verma, Gaurav, et al.
Published: (2024)
by: Verma, Gaurav, et al.
Published: (2024)
SCRuB: Social Concept Reasoning under Rubric-Based Evaluation
by: Watson-Daniels, Jamelle, et al.
Published: (2026)
by: Watson-Daniels, Jamelle, et al.
Published: (2026)
DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models
by: Teotia, Revant, et al.
Published: (2025)
by: Teotia, Revant, et al.
Published: (2025)
Position: AI Evaluations Should be Grounded on a Theory of Capability
by: Jo, Nathanael, et al.
Published: (2025)
by: Jo, Nathanael, et al.
Published: (2025)
Representative Ranking for Deliberation in the Public Sphere
by: Revel, Manon, et al.
Published: (2025)
by: Revel, Manon, et al.
Published: (2025)
Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages
by: Ovalle, Anaelia, et al.
Published: (2025)
by: Ovalle, Anaelia, et al.
Published: (2025)
Mean-field underdamped Langevin dynamics and its spacetime discretization
by: Fu, Qiang, et al.
Published: (2023)
by: Fu, Qiang, et al.
Published: (2023)
Facebook Political Ads And Accountability: Outside Groups Are Most Negative, Especially When Hiding Donors
by: Jain, Shomik, et al.
Published: (2020)
by: Jain, Shomik, et al.
Published: (2020)
Semivalue-based data valuation is arbitrary and gameable
by: Diehl, Hannah, et al.
Published: (2025)
by: Diehl, Hannah, et al.
Published: (2025)
A Course Shared Task on Evaluating LLM Output for Clinical Questions
by: Hou, Yufang, et al.
Published: (2024)
by: Hou, Yufang, et al.
Published: (2024)
High-accuracy sampling from constrained spaces with the Metropolis-adjusted Preconditioned Langevin Algorithm
by: Srinivasan, Vishwak, et al.
Published: (2024)
by: Srinivasan, Vishwak, et al.
Published: (2024)
Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
by: Szymanski, Annalisa, et al.
Published: (2024)
by: Szymanski, Annalisa, et al.
Published: (2024)
OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
by: Aggarwal, Pranjal, et al.
Published: (2025)
by: Aggarwal, Pranjal, et al.
Published: (2025)
Sense and Sensitivity: Evaluating the simulation of social dynamics via Large Language Models
by: Ju, Da, et al.
Published: (2024)
by: Ju, Da, et al.
Published: (2024)
What makes a good metric? Evaluating automatic metrics for text-to-image consistency
by: Ross, Candace, et al.
Published: (2024)
by: Ross, Candace, et al.
Published: (2024)
Fast sampling from constrained spaces using the Metropolis-adjusted Mirror Langevin algorithm
by: Srinivasan, Vishwak, et al.
Published: (2023)
by: Srinivasan, Vishwak, et al.
Published: (2023)
Decoupling Task-Solving and Output Formatting in LLM Generation
by: Deng, Haikang, et al.
Published: (2025)
by: Deng, Haikang, et al.
Published: (2025)
UCD: Unlearning in LLMs via Contrastive Decoding
by: Suriyakumar, Vinith M., et al.
Published: (2025)
by: Suriyakumar, Vinith M., et al.
Published: (2025)
DiffusionWorldViewer: Exposing and Broadening the Worldview Reflected by Generative Text-to-Image Models
by: De Simone, Zoe, et al.
Published: (2023)
by: De Simone, Zoe, et al.
Published: (2023)
Towards Geographic Inclusion in the Evaluation of Text-to-Image Models
by: Hall, Melissa, et al.
Published: (2024)
by: Hall, Melissa, et al.
Published: (2024)
A Unified Taxonomy-Guided Instruction Tuning Framework for Entity Set Expansion and Taxonomy Expansion
by: Shen, Yanzhen, et al.
Published: (2024)
by: Shen, Yanzhen, et al.
Published: (2024)
LLM Output Detectability and Task Performance Can be Jointly Optimized
by: Saito, Koshiro, et al.
Published: (2026)
by: Saito, Koshiro, et al.
Published: (2026)
LLM Pretraining with Continuous Concepts
by: Tack, Jihoon, et al.
Published: (2025)
by: Tack, Jihoon, et al.
Published: (2025)
DIG In: Evaluating Disparities in Image Generations with Indicators for Geographic Diversity
by: Hall, Melissa, et al.
Published: (2023)
by: Hall, Melissa, et al.
Published: (2023)
LITE: LLM-Impelled efficient Taxonomy Evaluation
by: Zhang, Lin, et al.
Published: (2025)
by: Zhang, Lin, et al.
Published: (2025)
NaturalThoughts: Selecting and Distilling Reasoning Traces for General Reasoning Tasks
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
STED and Consistency Scoring: A Framework for Evaluating LLM Structured Output Reliability
by: Wang, Guanghui, et al.
Published: (2025)
by: Wang, Guanghui, et al.
Published: (2025)
Alignment has a Fantasia Problem
by: Jo, Nathanael, et al.
Published: (2026)
by: Jo, Nathanael, et al.
Published: (2026)
Adaptive Decoding via Latent Preference Optimization
by: Dhuliawala, Shehzaad, et al.
Published: (2024)
by: Dhuliawala, Shehzaad, et al.
Published: (2024)
Diverse Preference Optimization
by: Lanchantin, Jack, et al.
Published: (2025)
by: Lanchantin, Jack, et al.
Published: (2025)
Creo: From One-Shot Image Generation to Progressive, Co-Creative Ideation
by: De Simone, Zoe, et al.
Published: (2026)
by: De Simone, Zoe, et al.
Published: (2026)
Perspective Dial: Measuring Perspective of Text and Guiding LLM Outputs
by: Kim, Taejin, et al.
Published: (2025)
by: Kim, Taejin, et al.
Published: (2025)
Similar Items
-
As an AI Language Model, "Yes I Would Recommend Calling the Police": Norm Inconsistency in LLM Decision-Making
by: Jain, Shomik, et al.
Published: (2024) -
Algorithmic Fairness and Color-blind Racism: Navigating the Intersection
by: Watson-Daniels, Jamelle
Published: (2024) -
Scarce Resource Allocations That Rely On Machine Learning Should Be Randomized
by: Jain, Shomik, et al.
Published: (2024) -
Allocation Multiplicity: Evaluating the Promises of the Rashomon Set
by: Jain, Shomik, et al.
Published: (2025) -
Automating Transparency Mechanisms in the Judicial System Using LLMs: Opportunities and Challenges
by: Shastri, Ishana, et al.
Published: (2024)