A Survey of Useful LLM Evaluation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Peng, Ji-Lun, Cheng, Sijia, Diau, Egil, Shih, Yung-Yu, Chen, Po-Heng, Lin, Yen-Ting, Chen, Yun-Nung |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Measuring Taiwanese Mandarin Language Understanding
par: Chen, Po-Heng, et autres
Publié: (2024)
par: Chen, Po-Heng, et autres
Publié: (2024)
Rethinking Role-Playing Evaluation: Anonymous Benchmarking and a Systematic Study of Personality Effects
par: Peng, Ji-Lun, et autres
Publié: (2026)
par: Peng, Ji-Lun, et autres
Publié: (2026)
Exploring Personality-Aware Interactions in Salesperson Dialogue Agents
par: Cheng, Sijia, et autres
Publié: (2025)
par: Cheng, Sijia, et autres
Publié: (2025)
The Cognitive Foundations of Economic Exchange: A Modular Framework Grounded in Behavioral Evidence
par: Diau, Egil
Publié: (2025)
par: Diau, Egil
Publié: (2025)
Finance as Extended Biology: Reciprocity as the Cognitive Substrate of Financial Behavior
par: Diau, Egil
Publié: (2025)
par: Diau, Egil
Publié: (2025)
Reciprocity as the Foundational Substrate of Society: How Reciprocal Dynamics Scale into Social Systems
par: Diau, Egil
Publié: (2025)
par: Diau, Egil
Publié: (2025)
LLMs are Biased Evaluators But Not Biased for Retrieval Augmented Generation
par: Chen, Yen-Shan, et autres
Publié: (2024)
par: Chen, Yen-Shan, et autres
Publié: (2024)
Rehearsing Answers to Probable Questions with Perspective-Taking
par: Shih, Yung-Yu, et autres
Publié: (2024)
par: Shih, Yung-Yu, et autres
Publié: (2024)
LLM Inference Enhanced by External Knowledge: A Survey
par: Lin, Yu-Hsuan, et autres
Publié: (2025)
par: Lin, Yu-Hsuan, et autres
Publié: (2025)
Creativity in LLM-based Multi-Agent Systems: A Survey
par: Lin, Yi-Cheng, et autres
Publié: (2025)
par: Lin, Yi-Cheng, et autres
Publié: (2025)
TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
par: Chen, Yen-Shan, et autres
Publié: (2026)
par: Chen, Yen-Shan, et autres
Publié: (2026)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
par: Chen, Yen-Shan, et autres
Publié: (2026)
par: Chen, Yen-Shan, et autres
Publié: (2026)
I Need Help! Evaluating LLM's Ability to Ask for Users' Support: A Case Study on Text-to-SQL Generation
par: Wu, Cheng-Kuang, et autres
Publié: (2024)
par: Wu, Cheng-Kuang, et autres
Publié: (2024)
None of the Above, Less of the Right: Parallel Patterns between Humans and LLMs on Multi-Choice Questions Answering
par: Tam, Zhi Rui, et autres
Publié: (2025)
par: Tam, Zhi Rui, et autres
Publié: (2025)
VisTW: Benchmarking Vision-Language Models for Traditional Chinese in Taiwan
par: Tam, Zhi Rui, et autres
Publié: (2025)
par: Tam, Zhi Rui, et autres
Publié: (2025)
MedVoiceBias: A Controlled Study of Audio LLM Behavior in Clinical Decision-Making
par: Tam, Zhi Rui, et autres
Publié: (2025)
par: Tam, Zhi Rui, et autres
Publié: (2025)
Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors
par: Chen, Yen-Shan, et autres
Publié: (2025)
par: Chen, Yen-Shan, et autres
Publié: (2025)
Two Tales of Persona in LLMs: A Survey of Role-Playing and Personalization
par: Tseng, Yu-Min, et autres
Publié: (2024)
par: Tseng, Yu-Min, et autres
Publié: (2024)
Injecting Salesperson's Dialogue Strategies in Large Language Models with Chain-of-Thought Reasoning
par: Chang, Wen-Yu, et autres
Publié: (2024)
par: Chang, Wen-Yu, et autres
Publié: (2024)
StreamBench: Towards Benchmarking Continuous Improvement of Language Agents
par: Wu, Cheng-Kuang, et autres
Publié: (2024)
par: Wu, Cheng-Kuang, et autres
Publié: (2024)
The Role of Exploration Modules in Small Language Models for Knowledge Graph Question Answering
par: Cheng, Yi-Jie, et autres
Publié: (2025)
par: Cheng, Yi-Jie, et autres
Publié: (2025)
Mitigating Forgetting in LLM Fine-Tuning via Low-Perplexity Token Learning
par: Wu, Chao-Chung, et autres
Publié: (2025)
par: Wu, Chao-Chung, et autres
Publié: (2025)
Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models
par: Tam, Zhi Rui, et autres
Publié: (2024)
par: Tam, Zhi Rui, et autres
Publié: (2024)
Visualizing Dialogues: Enhancing Image Selection through Dialogue Understanding with Large Language Models
par: Kao, Chang-Sheng, et autres
Publié: (2024)
par: Kao, Chang-Sheng, et autres
Publié: (2024)
Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models?
par: Tam, Zhi Rui, et autres
Publié: (2025)
par: Tam, Zhi Rui, et autres
Publié: (2025)
Answer, Refuse, or Guess? Investigating Risk-Aware Decision Making in Language Models
par: Wu, Cheng-Kuang, et autres
Publié: (2025)
par: Wu, Cheng-Kuang, et autres
Publié: (2025)
Compound AI Systems Optimization: A Survey of Methods, Challenges, and Future Directions
par: Lee, Yu-Ang, et autres
Publié: (2025)
par: Lee, Yu-Ang, et autres
Publié: (2025)
A Survey of Generative Information Retrieval
par: Kuo, Tzu-Lin, et autres
Publié: (2024)
par: Kuo, Tzu-Lin, et autres
Publié: (2024)
InstUPR : Instruction-based Unsupervised Passage Reranking with Large Language Models
par: Huang, Chao-Wei, et autres
Publié: (2024)
par: Huang, Chao-Wei, et autres
Publié: (2024)
PairDistill: Pairwise Relevance Distillation for Dense Retrieval
par: Huang, Chao-Wei, et autres
Publié: (2024)
par: Huang, Chao-Wei, et autres
Publié: (2024)
Balancing Knowledge Delivery and Emotional Comfort in Healthcare Conversational Systems
par: Tsai, Shang-Chi, et autres
Publié: (2025)
par: Tsai, Shang-Chi, et autres
Publié: (2025)
FactAlign: Long-form Factuality Alignment of Large Language Models
par: Huang, Chao-Wei, et autres
Publié: (2024)
par: Huang, Chao-Wei, et autres
Publié: (2024)
Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning
par: Chung, Ho-Lam, et autres
Publié: (2025)
par: Chung, Ho-Lam, et autres
Publié: (2025)
RADAR: Retrieval-Augmented Detector with Adversarial Refinement for Robust Fake News Detection
par: Ma, Song-Duo, et autres
Publié: (2026)
par: Ma, Song-Duo, et autres
Publié: (2026)
DogeRM: Equipping Reward Models with Domain Knowledge through Model Merging
par: Lin, Tzu-Han, et autres
Publié: (2024)
par: Lin, Tzu-Han, et autres
Publié: (2024)
AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning
par: Lin, Tzu-Han, et autres
Publié: (2025)
par: Lin, Tzu-Han, et autres
Publié: (2025)
From Simulation to Strategy: Automating Personalized Interaction Planning for Conversational Agents
par: Chang, Wen-Yu, et autres
Publié: (2025)
par: Chang, Wen-Yu, et autres
Publié: (2025)
Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
par: Cao, Yixin, et autres
Publié: (2025)
par: Cao, Yixin, et autres
Publié: (2025)
MIDI-LLM: Adapting Large Language Models for Text-to-MIDI Music Generation
par: Wu, Shih-Lun, et autres
Publié: (2025)
par: Wu, Shih-Lun, et autres
Publié: (2025)
Multi-expert Prompting Improves Reliability, Safety, and Usefulness of Large Language Models
par: Long, Do Xuan, et autres
Publié: (2024)
par: Long, Do Xuan, et autres
Publié: (2024)
Documents similaires
-
Measuring Taiwanese Mandarin Language Understanding
par: Chen, Po-Heng, et autres
Publié: (2024) -
Rethinking Role-Playing Evaluation: Anonymous Benchmarking and a Systematic Study of Personality Effects
par: Peng, Ji-Lun, et autres
Publié: (2026) -
Exploring Personality-Aware Interactions in Salesperson Dialogue Agents
par: Cheng, Sijia, et autres
Publié: (2025) -
The Cognitive Foundations of Economic Exchange: A Modular Framework Grounded in Behavioral Evidence
par: Diau, Egil
Publié: (2025) -
Finance as Extended Biology: Reciprocity as the Cognitive Substrate of Financial Behavior
par: Diau, Egil
Publié: (2025)