Saved in:
| Main Authors: | Goethals, Sofie, Rhue, Lauren |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2412.10281 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating LLMs for Gender Disparities in Notable Persons
by: Rhue, Lauren, et al.
Published: (2024)
by: Rhue, Lauren, et al.
Published: (2024)
Prompt-Counterfactual Explanations for Generative AI System Behavior
by: Goethals, Sofie, et al.
Published: (2026)
by: Goethals, Sofie, et al.
Published: (2026)
Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices
by: Reusens, Manon, et al.
Published: (2026)
by: Reusens, Manon, et al.
Published: (2026)
Cash or Comfort? How LLMs Value Your Inconvenience
by: Cedro, Mateusz, et al.
Published: (2025)
by: Cedro, Mateusz, et al.
Published: (2025)
The Basic B*** Effect: The Use of LLM-based Agents Reduces the Distinctiveness and Diversity of People's Choices
by: Matz, Sandra C., et al.
Published: (2025)
by: Matz, Sandra C., et al.
Published: (2025)
Reranking individuals: The effect of fair classification within-groups
by: Goethals, Sofie, et al.
Published: (2024)
by: Goethals, Sofie, et al.
Published: (2024)
Resource-constrained Fairness
by: Goethals, Sofie, et al.
Published: (2024)
by: Goethals, Sofie, et al.
Published: (2024)
Manson superstar
Published: (2006)
Published: (2006)
When One LLM Drools, Multi-LLM Collaboration Rules
by: Feng, Shangbin, et al.
Published: (2025)
by: Feng, Shangbin, et al.
Published: (2025)
DiscoTrack: A Multilingual LLM Benchmark for Discourse Tracking
by: Bu, Lanni, et al.
Published: (2025)
by: Bu, Lanni, et al.
Published: (2025)
A Customer Journey in the Land of Oz: Leveraging the Wizard of Oz Technique to Model Emotions in Customer Service Interactions
by: Labat, Sofie, et al.
Published: (2025)
by: Labat, Sofie, et al.
Published: (2025)
LLMBridge: An LLM Pipeline for End-to-end Referential Bridging Resolution in English
by: Levine, Lauren, et al.
Published: (2026)
by: Levine, Lauren, et al.
Published: (2026)
Brevity is the soul of sustainability: Characterizing LLM response lengths
by: Poddar, Soham, et al.
Published: (2025)
by: Poddar, Soham, et al.
Published: (2025)
LCO: LLM-based Constraint Optimization for Safer Agentic LLMs in Real-world Tasks
by: Wan, Jiayong, et al.
Published: (2026)
by: Wan, Jiayong, et al.
Published: (2026)
ReZero: Enhancing LLM search ability by trying one-more-time
by: Dao, Alan, et al.
Published: (2025)
by: Dao, Alan, et al.
Published: (2025)
One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
by: Shen, Chengyu, et al.
Published: (2026)
by: Shen, Chengyu, et al.
Published: (2026)
One Token to Fool LLM-as-a-Judge
by: Zhao, Yulai, et al.
Published: (2025)
by: Zhao, Yulai, et al.
Published: (2025)
LLM one-shot style transfer for Authorship Attribution and Verification
by: Miralles-González, Pablo, et al.
Published: (2025)
by: Miralles-González, Pablo, et al.
Published: (2025)
Researchers waste 80% of LLM annotation costs by classifying one text at a time
by: Pipal, Christian, et al.
Published: (2026)
by: Pipal, Christian, et al.
Published: (2026)
TRUEBench: Can LLM Response Meet Real-world Constraints as Productivity Assistant?
by: Park, Jiho, et al.
Published: (2025)
by: Park, Jiho, et al.
Published: (2025)
MultiZebraLogic: A Multilingual Logical Reasoning Benchmark
by: Bruun, Sofie Helene, et al.
Published: (2025)
by: Bruun, Sofie Helene, et al.
Published: (2025)
Beyond One Path: Evaluating and Enhancing Divergent Thinking in Interactive LLM Agents
by: Park, Jihyeong, et al.
Published: (2026)
by: Park, Jihyeong, et al.
Published: (2026)
One Language, Two Scripts: Probing Script-Invariance in LLM Concept Representations
by: Karne, Sripad
Published: (2026)
by: Karne, Sripad
Published: (2026)
OneLLM: One Framework to Align All Modalities with Language
by: Han, Jiaming, et al.
Published: (2023)
by: Han, Jiaming, et al.
Published: (2023)
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
by: He, Wei, et al.
Published: (2025)
by: He, Wei, et al.
Published: (2025)
StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
by: Chen, Yanxu, et al.
Published: (2025)
by: Chen, Yanxu, et al.
Published: (2025)
OneEval: Benchmarking LLM Knowledge-intensive Reasoning over Diverse Knowledge Bases
by: Chen, Yongrui, et al.
Published: (2025)
by: Chen, Yongrui, et al.
Published: (2025)
Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents
by: Yu, Jianxiang, et al.
Published: (2026)
by: Yu, Jianxiang, et al.
Published: (2026)
One LLM to Train Them All: Multi-Task Learning Framework for Fact-Checking
by: Larsson, Malin Astrid, et al.
Published: (2026)
by: Larsson, Malin Astrid, et al.
Published: (2026)
OneShield -- the Next Generation of LLM Guardrails
by: DeLuca, Chad, et al.
Published: (2025)
by: DeLuca, Chad, et al.
Published: (2025)
Bias in the Mirror: Are LLMs opinions robust to their own adversarial attacks ?
by: Rennard, Virgile, et al.
Published: (2024)
by: Rennard, Virgile, et al.
Published: (2024)
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
by: Chen, Guoxuan, et al.
Published: (2024)
by: Chen, Guoxuan, et al.
Published: (2024)
Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based Evaluation
by: Lee, Dongryeol, et al.
Published: (2024)
by: Lee, Dongryeol, et al.
Published: (2024)
One Persona, Many Cues, Different Results: How Sociodemographic Cues Impact LLM Personalization
by: Weeber, Franziska, et al.
Published: (2026)
by: Weeber, Franziska, et al.
Published: (2026)
One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment
by: Cai, Hongru, et al.
Published: (2026)
by: Cai, Hongru, et al.
Published: (2026)
One Size Fits None: Heuristic Collapse in LLM Investment Advice
by: Ross, Jillian, et al.
Published: (2026)
by: Ross, Jillian, et al.
Published: (2026)
DepressLLM: Interpretable domain-adapted language model for depression detection from real-world narratives
by: Moon, Sehwan, et al.
Published: (2025)
by: Moon, Sehwan, et al.
Published: (2025)
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set
by: Finkelstein, Mara, et al.
Published: (2024)
by: Finkelstein, Mara, et al.
Published: (2024)
One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety
by: Arif, Samee, et al.
Published: (2026)
by: Arif, Samee, et al.
Published: (2026)
CodeNav: Beyond tool-use to using real-world codebases with LLM agents
by: Gupta, Tanmay, et al.
Published: (2024)
by: Gupta, Tanmay, et al.
Published: (2024)
Similar Items
-
Evaluating LLMs for Gender Disparities in Notable Persons
by: Rhue, Lauren, et al.
Published: (2024) -
Prompt-Counterfactual Explanations for Generative AI System Behavior
by: Goethals, Sofie, et al.
Published: (2026) -
Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices
by: Reusens, Manon, et al.
Published: (2026) -
Cash or Comfort? How LLMs Value Your Inconvenience
by: Cedro, Mateusz, et al.
Published: (2025) -
The Basic B*** Effect: The Use of LLM-based Agents Reduces the Distinctiveness and Diversity of People's Choices
by: Matz, Sandra C., et al.
Published: (2025)