A Survey of Useful LLM Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Peng, Ji-Lun, Cheng, Sijia, Diau, Egil, Shih, Yung-Yu, Chen, Po-Heng, Lin, Yen-Ting, Chen, Yun-Nung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Measuring Taiwanese Mandarin Language Understanding
di: Chen, Po-Heng, et al.
Pubblicazione: (2024)
di: Chen, Po-Heng, et al.
Pubblicazione: (2024)
Rethinking Role-Playing Evaluation: Anonymous Benchmarking and a Systematic Study of Personality Effects
di: Peng, Ji-Lun, et al.
Pubblicazione: (2026)
di: Peng, Ji-Lun, et al.
Pubblicazione: (2026)
Exploring Personality-Aware Interactions in Salesperson Dialogue Agents
di: Cheng, Sijia, et al.
Pubblicazione: (2025)
di: Cheng, Sijia, et al.
Pubblicazione: (2025)
The Cognitive Foundations of Economic Exchange: A Modular Framework Grounded in Behavioral Evidence
di: Diau, Egil
Pubblicazione: (2025)
di: Diau, Egil
Pubblicazione: (2025)
Finance as Extended Biology: Reciprocity as the Cognitive Substrate of Financial Behavior
di: Diau, Egil
Pubblicazione: (2025)
di: Diau, Egil
Pubblicazione: (2025)
Reciprocity as the Foundational Substrate of Society: How Reciprocal Dynamics Scale into Social Systems
di: Diau, Egil
Pubblicazione: (2025)
di: Diau, Egil
Pubblicazione: (2025)
LLMs are Biased Evaluators But Not Biased for Retrieval Augmented Generation
di: Chen, Yen-Shan, et al.
Pubblicazione: (2024)
di: Chen, Yen-Shan, et al.
Pubblicazione: (2024)
Rehearsing Answers to Probable Questions with Perspective-Taking
di: Shih, Yung-Yu, et al.
Pubblicazione: (2024)
di: Shih, Yung-Yu, et al.
Pubblicazione: (2024)
LLM Inference Enhanced by External Knowledge: A Survey
di: Lin, Yu-Hsuan, et al.
Pubblicazione: (2025)
di: Lin, Yu-Hsuan, et al.
Pubblicazione: (2025)
Creativity in LLM-based Multi-Agent Systems: A Survey
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
di: Chen, Yen-Shan, et al.
Pubblicazione: (2026)
di: Chen, Yen-Shan, et al.
Pubblicazione: (2026)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
di: Chen, Yen-Shan, et al.
Pubblicazione: (2026)
di: Chen, Yen-Shan, et al.
Pubblicazione: (2026)
I Need Help! Evaluating LLM's Ability to Ask for Users' Support: A Case Study on Text-to-SQL Generation
di: Wu, Cheng-Kuang, et al.
Pubblicazione: (2024)
di: Wu, Cheng-Kuang, et al.
Pubblicazione: (2024)
None of the Above, Less of the Right: Parallel Patterns between Humans and LLMs on Multi-Choice Questions Answering
di: Tam, Zhi Rui, et al.
Pubblicazione: (2025)
di: Tam, Zhi Rui, et al.
Pubblicazione: (2025)
VisTW: Benchmarking Vision-Language Models for Traditional Chinese in Taiwan
di: Tam, Zhi Rui, et al.
Pubblicazione: (2025)
di: Tam, Zhi Rui, et al.
Pubblicazione: (2025)
MedVoiceBias: A Controlled Study of Audio LLM Behavior in Clinical Decision-Making
di: Tam, Zhi Rui, et al.
Pubblicazione: (2025)
di: Tam, Zhi Rui, et al.
Pubblicazione: (2025)
Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors
di: Chen, Yen-Shan, et al.
Pubblicazione: (2025)
di: Chen, Yen-Shan, et al.
Pubblicazione: (2025)
Two Tales of Persona in LLMs: A Survey of Role-Playing and Personalization
di: Tseng, Yu-Min, et al.
Pubblicazione: (2024)
di: Tseng, Yu-Min, et al.
Pubblicazione: (2024)
Injecting Salesperson's Dialogue Strategies in Large Language Models with Chain-of-Thought Reasoning
di: Chang, Wen-Yu, et al.
Pubblicazione: (2024)
di: Chang, Wen-Yu, et al.
Pubblicazione: (2024)
StreamBench: Towards Benchmarking Continuous Improvement of Language Agents
di: Wu, Cheng-Kuang, et al.
Pubblicazione: (2024)
di: Wu, Cheng-Kuang, et al.
Pubblicazione: (2024)
The Role of Exploration Modules in Small Language Models for Knowledge Graph Question Answering
di: Cheng, Yi-Jie, et al.
Pubblicazione: (2025)
di: Cheng, Yi-Jie, et al.
Pubblicazione: (2025)
Mitigating Forgetting in LLM Fine-Tuning via Low-Perplexity Token Learning
di: Wu, Chao-Chung, et al.
Pubblicazione: (2025)
di: Wu, Chao-Chung, et al.
Pubblicazione: (2025)
Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models
di: Tam, Zhi Rui, et al.
Pubblicazione: (2024)
di: Tam, Zhi Rui, et al.
Pubblicazione: (2024)
Visualizing Dialogues: Enhancing Image Selection through Dialogue Understanding with Large Language Models
di: Kao, Chang-Sheng, et al.
Pubblicazione: (2024)
di: Kao, Chang-Sheng, et al.
Pubblicazione: (2024)
Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models?
di: Tam, Zhi Rui, et al.
Pubblicazione: (2025)
di: Tam, Zhi Rui, et al.
Pubblicazione: (2025)
Answer, Refuse, or Guess? Investigating Risk-Aware Decision Making in Language Models
di: Wu, Cheng-Kuang, et al.
Pubblicazione: (2025)
di: Wu, Cheng-Kuang, et al.
Pubblicazione: (2025)
Compound AI Systems Optimization: A Survey of Methods, Challenges, and Future Directions
di: Lee, Yu-Ang, et al.
Pubblicazione: (2025)
di: Lee, Yu-Ang, et al.
Pubblicazione: (2025)
A Survey of Generative Information Retrieval
di: Kuo, Tzu-Lin, et al.
Pubblicazione: (2024)
di: Kuo, Tzu-Lin, et al.
Pubblicazione: (2024)
InstUPR : Instruction-based Unsupervised Passage Reranking with Large Language Models
di: Huang, Chao-Wei, et al.
Pubblicazione: (2024)
di: Huang, Chao-Wei, et al.
Pubblicazione: (2024)
PairDistill: Pairwise Relevance Distillation for Dense Retrieval
di: Huang, Chao-Wei, et al.
Pubblicazione: (2024)
di: Huang, Chao-Wei, et al.
Pubblicazione: (2024)
Balancing Knowledge Delivery and Emotional Comfort in Healthcare Conversational Systems
di: Tsai, Shang-Chi, et al.
Pubblicazione: (2025)
di: Tsai, Shang-Chi, et al.
Pubblicazione: (2025)
FactAlign: Long-form Factuality Alignment of Large Language Models
di: Huang, Chao-Wei, et al.
Pubblicazione: (2024)
di: Huang, Chao-Wei, et al.
Pubblicazione: (2024)
Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning
di: Chung, Ho-Lam, et al.
Pubblicazione: (2025)
di: Chung, Ho-Lam, et al.
Pubblicazione: (2025)
RADAR: Retrieval-Augmented Detector with Adversarial Refinement for Robust Fake News Detection
di: Ma, Song-Duo, et al.
Pubblicazione: (2026)
di: Ma, Song-Duo, et al.
Pubblicazione: (2026)
DogeRM: Equipping Reward Models with Domain Knowledge through Model Merging
di: Lin, Tzu-Han, et al.
Pubblicazione: (2024)
di: Lin, Tzu-Han, et al.
Pubblicazione: (2024)
AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning
di: Lin, Tzu-Han, et al.
Pubblicazione: (2025)
di: Lin, Tzu-Han, et al.
Pubblicazione: (2025)
From Simulation to Strategy: Automating Personalized Interaction Planning for Conversational Agents
di: Chang, Wen-Yu, et al.
Pubblicazione: (2025)
di: Chang, Wen-Yu, et al.
Pubblicazione: (2025)
Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
di: Cao, Yixin, et al.
Pubblicazione: (2025)
di: Cao, Yixin, et al.
Pubblicazione: (2025)
MIDI-LLM: Adapting Large Language Models for Text-to-MIDI Music Generation
di: Wu, Shih-Lun, et al.
Pubblicazione: (2025)
di: Wu, Shih-Lun, et al.
Pubblicazione: (2025)
Multi-expert Prompting Improves Reliability, Safety, and Usefulness of Large Language Models
di: Long, Do Xuan, et al.
Pubblicazione: (2024)
di: Long, Do Xuan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Measuring Taiwanese Mandarin Language Understanding
di: Chen, Po-Heng, et al.
Pubblicazione: (2024) -
Rethinking Role-Playing Evaluation: Anonymous Benchmarking and a Systematic Study of Personality Effects
di: Peng, Ji-Lun, et al.
Pubblicazione: (2026) -
Exploring Personality-Aware Interactions in Salesperson Dialogue Agents
di: Cheng, Sijia, et al.
Pubblicazione: (2025) -
The Cognitive Foundations of Economic Exchange: A Modular Framework Grounded in Behavioral Evidence
di: Diau, Egil
Pubblicazione: (2025) -
Finance as Extended Biology: Reciprocity as the Cognitive Substrate of Financial Behavior
di: Diau, Egil
Pubblicazione: (2025)