TRUEBench: Can LLM Response Meet Real-world Constraints as Productivity Assistant?
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Jiho, Song, Jongyoon, Choi, Minjin, Heo, Kyuho, Huh, Taehun, Kim, Ji Won |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pay What LLM Wants: Can LLM Simulate Economics Experiment with 522 Real-human Persona?
by: Choi, Junhyuk, et al.
Published: (2025)
by: Choi, Junhyuk, et al.
Published: (2025)
Self-HarmLLM: Can Large Language Model Harm Itself?
by: Kim, Heehwan, et al.
Published: (2025)
by: Kim, Heehwan, et al.
Published: (2025)
A Multifaceted Analysis of Negative Bias in Large Language Models through the Lens of Parametric Knowledge
by: Song, Jongyoon, et al.
Published: (2025)
by: Song, Jongyoon, et al.
Published: (2025)
Large Language Models are Skeptics: False Negative Problem of Input-conflicting Hallucination
by: Song, Jongyoon, et al.
Published: (2024)
by: Song, Jongyoon, et al.
Published: (2024)
R2-KG: General-Purpose Dual-Agent Framework for Reliable Reasoning on Knowledge Graphs
by: Jo, Sumin, et al.
Published: (2025)
by: Jo, Sumin, et al.
Published: (2025)
CCQA: Generating Question from Solution Can Improve Inference-Time Reasoning in SLMs
by: Kim, Jin Young, et al.
Published: (2025)
by: Kim, Jin Young, et al.
Published: (2025)
LLM Meets Scene Graph: Can Large Language Models Understand and Generate Scene Graphs? A Benchmark and Empirical Study
by: Yang, Dongil, et al.
Published: (2025)
by: Yang, Dongil, et al.
Published: (2025)
From Reading to Compressing: Exploring the Multi-document Reader for Prompt Compression
by: Choi, Eunseong, et al.
Published: (2024)
by: Choi, Eunseong, et al.
Published: (2024)
ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation
by: Kim, Jiho, et al.
Published: (2025)
by: Kim, Jiho, et al.
Published: (2025)
GRAM: Generative Recommendation via Semantic-aware Multi-granular Late Fusion
by: Lee, Sunkyung, et al.
Published: (2025)
by: Lee, Sunkyung, et al.
Published: (2025)
Integrated Framework for LLM Evaluation with Answer Generation
by: Lee, Sujeong, et al.
Published: (2025)
by: Lee, Sujeong, et al.
Published: (2025)
Can Separators Improve Chain-of-Thought Prompting?
by: Park, Yoonjeong, et al.
Published: (2024)
by: Park, Yoonjeong, et al.
Published: (2024)
Coding-Free and Privacy-Preserving Agentic Framework for Data-Driven Clinical Research
by: Kim, Taehun, et al.
Published: (2026)
by: Kim, Taehun, et al.
Published: (2026)
HuDEx: Integrating Hallucination Detection and Explainability for Enhancing the Reliability of LLM responses
by: Lee, Sujeong, et al.
Published: (2025)
by: Lee, Sujeong, et al.
Published: (2025)
SelectLLM: Can LLMs Select Important Instructions to Annotate?
by: Parkar, Ritik Sachin, et al.
Published: (2024)
by: Parkar, Ritik Sachin, et al.
Published: (2024)
YA-TA: Towards Personalized Question-Answering Teaching Assistants using Instructor-Student Dual Retrieval-augmented Knowledge Fusion
by: Yang, Dongil, et al.
Published: (2024)
by: Yang, Dongil, et al.
Published: (2024)
Offline Policy Evaluation of Multi-Turn LLM Health Coaching with Real Users
by: Ozolcer, Melik, et al.
Published: (2025)
by: Ozolcer, Melik, et al.
Published: (2025)
Entity-level Factual Adaptiveness of Fine-tuning based Abstractive Summarization Models
by: Song, Jongyoon, et al.
Published: (2024)
by: Song, Jongyoon, et al.
Published: (2024)
Can AI Assistants Know What They Don't Know?
by: Cheng, Qinyuan, et al.
Published: (2024)
by: Cheng, Qinyuan, et al.
Published: (2024)
Re-Ex: Revising after Explanation Reduces the Factual Errors in LLM Responses
by: Kim, Juyeon, et al.
Published: (2024)
by: Kim, Juyeon, et al.
Published: (2024)
Can David Beat Goliath? On Multi-Hop Reasoning with Resource-Constrained Agents
by: Han, Hojae, et al.
Published: (2026)
by: Han, Hojae, et al.
Published: (2026)
Can Risk-taking AI-Assistants suitably represent entities
by: Mazyaki, Ali, et al.
Published: (2025)
by: Mazyaki, Ali, et al.
Published: (2025)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
by: Long, Xiang, et al.
Published: (2026)
by: Long, Xiang, et al.
Published: (2026)
Efficient semantic uncertainty quantification in language models via diversity-steered sampling
by: Park, Ji Won, et al.
Published: (2025)
by: Park, Ji Won, et al.
Published: (2025)
Contrastive Learning to Improve Retrieval for Real-world Fact Checking
by: Sriram, Aniruddh, et al.
Published: (2024)
by: Sriram, Aniruddh, et al.
Published: (2024)
Not All Personas Are Worth It: Culture-Reflective Persona Data Augmentation
by: Han, Ji-Eun, et al.
Published: (2025)
by: Han, Ji-Eun, et al.
Published: (2025)
Large Language Model Meets Constraint Propagation
by: Bonlarron, Alexandre, et al.
Published: (2025)
by: Bonlarron, Alexandre, et al.
Published: (2025)
MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
Can Multiple Responses from an LLM Reveal the Sources of Its Uncertainty?
by: Nan, Yang, et al.
Published: (2025)
by: Nan, Yang, et al.
Published: (2025)
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
by: He, Wei, et al.
Published: (2025)
by: He, Wei, et al.
Published: (2025)
CHILL at SemEval-2025 Task 2: You Can't Just Throw Entities and Hope -- Make Your LLM to Get Them Right
by: Lee, Jaebok, et al.
Published: (2025)
by: Lee, Jaebok, et al.
Published: (2025)
ATLAS: Constraints-Aware Multi-Agent Collaboration for Real-World Travel Planning
by: Choi, Jihye, et al.
Published: (2025)
by: Choi, Jihye, et al.
Published: (2025)
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
by: Jin, Jiho, et al.
Published: (2026)
by: Jin, Jiho, et al.
Published: (2026)
Correcting Negative Bias in Large Language Models through Negative Attention Score Alignment
by: Yu, Sangwon, et al.
Published: (2024)
by: Yu, Sangwon, et al.
Published: (2024)
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
by: Song, Seyoung, et al.
Published: (2025)
by: Song, Seyoung, et al.
Published: (2025)
Understanding LLM Development Through Longitudinal Study: Insights from the Open Ko-LLM Leaderboard
by: Park, Chanjun, et al.
Published: (2024)
by: Park, Chanjun, et al.
Published: (2024)
Human Psychometric Questionnaires Mischaracterize LLM Behavior
by: Song, Woojung, et al.
Published: (2025)
by: Song, Woojung, et al.
Published: (2025)
AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy
by: Schoenegger, Philipp, et al.
Published: (2024)
by: Schoenegger, Philipp, et al.
Published: (2024)
ELITR-Bench: A Meeting Assistant Benchmark for Long-Context Language Models
by: Thonet, Thibaut, et al.
Published: (2024)
by: Thonet, Thibaut, et al.
Published: (2024)
PatientSim: A Persona-Driven Simulator for Realistic Doctor-Patient Interactions
by: Kyung, Daeun, et al.
Published: (2025)
by: Kyung, Daeun, et al.
Published: (2025)
Similar Items
-
Pay What LLM Wants: Can LLM Simulate Economics Experiment with 522 Real-human Persona?
by: Choi, Junhyuk, et al.
Published: (2025) -
Self-HarmLLM: Can Large Language Model Harm Itself?
by: Kim, Heehwan, et al.
Published: (2025) -
A Multifaceted Analysis of Negative Bias in Large Language Models through the Lens of Parametric Knowledge
by: Song, Jongyoon, et al.
Published: (2025) -
Large Language Models are Skeptics: False Negative Problem of Input-conflicting Hallucination
by: Song, Jongyoon, et al.
Published: (2024) -
R2-KG: General-Purpose Dual-Agent Framework for Reliable Reasoning on Knowledge Graphs
by: Jo, Sumin, et al.
Published: (2025)