SkillVerse : Assessing and Enhancing LLMs with Tree Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Tian, Yufei, Sun, Jiao, Peng, Nanyun, Zhang, Zizhao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Detecting Machine-Generated Long-Form Content with Latent-Space Variables
por: Tian, Yufei, et al.
Publicado: (2024)
por: Tian, Yufei, et al.
Publicado: (2024)
Rethinking Creativity Evaluation: A Critical Analysis of Existing Creativity Evaluations
por: Lu, Li-Chun, et al.
Publicado: (2025)
por: Lu, Li-Chun, et al.
Publicado: (2025)
Evaluating Human Alignment and Model Faithfulness of LLM Rationale
por: Fayyaz, Mohsen, et al.
Publicado: (2024)
por: Fayyaz, Mohsen, et al.
Publicado: (2024)
Multimodal Cultural Safety: Evaluation Framework and Alignment Strategies
por: Qiu, Haoyi, et al.
Publicado: (2025)
por: Qiu, Haoyi, et al.
Publicado: (2025)
REFFLY: Melody-Constrained Lyrics Editing Model
por: Zhao, Songyan, et al.
Publicado: (2024)
por: Zhao, Songyan, et al.
Publicado: (2024)
PhonologyBench: Evaluating Phonological Skills of Large Language Models
por: Suvarna, Ashima, et al.
Publicado: (2024)
por: Suvarna, Ashima, et al.
Publicado: (2024)
Are Akpans Trick or Treat: Unveiling Helpful Biases in Assistant Systems
por: Sun, Jiao, et al.
Publicado: (2022)
por: Sun, Jiao, et al.
Publicado: (2022)
AMRFact: Enhancing Summarization Factuality Evaluation with AMR-Driven Negative Samples Generation
por: Qiu, Haoyi, et al.
Publicado: (2023)
por: Qiu, Haoyi, et al.
Publicado: (2023)
The Unreasonable Effectiveness of Model Merging for Cross-Lingual Transfer in LLMs
por: Bandarkar, Lucas, et al.
Publicado: (2025)
por: Bandarkar, Lucas, et al.
Publicado: (2025)
Are Large Language Models Capable of Generating Human-Level Narratives?
por: Tian, Yufei, et al.
Publicado: (2024)
por: Tian, Yufei, et al.
Publicado: (2024)
Open-Domain Text Evaluation via Contrastive Distribution Methods
por: Lu, Sidi, et al.
Publicado: (2023)
por: Lu, Sidi, et al.
Publicado: (2023)
Scientific Discourse Tagging for Evidence Extraction
por: Li, Xiangci, et al.
Publicado: (2019)
por: Li, Xiangci, et al.
Publicado: (2019)
A Paragraph-level Multi-task Learning Model for Scientific Fact-Verification
por: Li, Xiangci, et al.
Publicado: (2020)
por: Li, Xiangci, et al.
Publicado: (2020)
IQ Test for LLMs: An Evaluation Framework for Uncovering Core Skills in LLMs
por: Maimon, Aviya, et al.
Publicado: (2025)
por: Maimon, Aviya, et al.
Publicado: (2025)
Extracting Small Translation Specialists from LLMs by Aggressively Pruning Experts
por: Martin, Liu O., et al.
Publicado: (2026)
por: Martin, Liu O., et al.
Publicado: (2026)
Vulnerability of LLMs to Vertically Aligned Text Manipulations
por: Li, Zhecheng, et al.
Publicado: (2024)
por: Li, Zhecheng, et al.
Publicado: (2024)
VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models
por: Qiu, Haoyi, et al.
Publicado: (2024)
por: Qiu, Haoyi, et al.
Publicado: (2024)
Con-ReCall: Detecting Pre-training Data in LLMs via Contrastive Decoding
por: Wang, Cheng, et al.
Publicado: (2024)
por: Wang, Cheng, et al.
Publicado: (2024)
CoKe: Customizable Fine-Grained Story Evaluation via Chain-of-Keyword Rationalization
por: Joshi, Brihi, et al.
Publicado: (2025)
por: Joshi, Brihi, et al.
Publicado: (2025)
Structured Outputs Enable General-Purpose LLMs to be Medical Experts
por: Guo, Guangfu, et al.
Publicado: (2025)
por: Guo, Guangfu, et al.
Publicado: (2025)
Decoupling Task-Solving and Output Formatting in LLM Generation
por: Deng, Haikang, et al.
Publicado: (2025)
por: Deng, Haikang, et al.
Publicado: (2025)
RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment
por: Yang, Kevin, et al.
Publicado: (2023)
por: Yang, Kevin, et al.
Publicado: (2023)
Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs
por: Doddapaneni, Sumanth, et al.
Publicado: (2024)
por: Doddapaneni, Sumanth, et al.
Publicado: (2024)
Guiding Through Complexity: What Makes Good Supervision for Hard Math Reasoning Tasks?
por: He, Xuan, et al.
Publicado: (2024)
por: He, Xuan, et al.
Publicado: (2024)
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
por: Yerukola, Akhila, et al.
Publicado: (2025)
por: Yerukola, Akhila, et al.
Publicado: (2025)
OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving
por: Zhang, Xinyu, et al.
Publicado: (2026)
por: Zhang, Xinyu, et al.
Publicado: (2026)
TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language Models
por: Zhang, Junyi, et al.
Publicado: (2025)
por: Zhang, Junyi, et al.
Publicado: (2025)
ALI-Agent: Assessing LLMs' Alignment with Human Values via Agent-based Evaluation
por: Zheng, Jingnan, et al.
Publicado: (2024)
por: Zheng, Jingnan, et al.
Publicado: (2024)
MMPersuade: A Dataset and Evaluation Framework for Multimodal Persuasion
por: Qiu, Haoyi, et al.
Publicado: (2025)
por: Qiu, Haoyi, et al.
Publicado: (2025)
Enhancing LLM Character-Level Manipulation via Divide and Conquer
por: Xiong, Zhen, et al.
Publicado: (2025)
por: Xiong, Zhen, et al.
Publicado: (2025)
MacGyver: Are Large Language Models Creative Problem Solvers?
por: Tian, Yufei, et al.
Publicado: (2023)
por: Tian, Yufei, et al.
Publicado: (2023)
Skill-LLM: Repurposing General-Purpose LLMs for Skill Extraction
por: Herandi, Amirhossein, et al.
Publicado: (2024)
por: Herandi, Amirhossein, et al.
Publicado: (2024)
Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation
por: Limkonchotiwat, Peerat, et al.
Publicado: (2025)
por: Limkonchotiwat, Peerat, et al.
Publicado: (2025)
Learning Action Conditions from Instructional Manuals for Instruction Understanding
por: Wu, Te-Lin, et al.
Publicado: (2022)
por: Wu, Te-Lin, et al.
Publicado: (2022)
Steering MoE LLMs via Expert (De)Activation
por: Fayyaz, Mohsen, et al.
Publicado: (2025)
por: Fayyaz, Mohsen, et al.
Publicado: (2025)
LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMs
por: Guo, Pei-Fu, et al.
Publicado: (2025)
por: Guo, Pei-Fu, et al.
Publicado: (2025)
Preference Learning Unlocks LLMs' Psycho-Counseling Skills
por: Zhang, Mian, et al.
Publicado: (2025)
por: Zhang, Mian, et al.
Publicado: (2025)
EmoVerse: Exploring Multimodal Large Language Models for Sentiment and Emotion Understanding
por: Li, Ao, et al.
Publicado: (2024)
por: Li, Ao, et al.
Publicado: (2024)
Evaluating Cultural and Social Awareness of LLM Web Agents
por: Qiu, Haoyi, et al.
Publicado: (2024)
por: Qiu, Haoyi, et al.
Publicado: (2024)
Grading the Unspoken: Evaluating Tacit Reasoning in Quantum Field Theory and String Theory with LLMs
por: Yu, Xingyang, et al.
Publicado: (2026)
por: Yu, Xingyang, et al.
Publicado: (2026)
Ejemplares similares
-
Detecting Machine-Generated Long-Form Content with Latent-Space Variables
por: Tian, Yufei, et al.
Publicado: (2024) -
Rethinking Creativity Evaluation: A Critical Analysis of Existing Creativity Evaluations
por: Lu, Li-Chun, et al.
Publicado: (2025) -
Evaluating Human Alignment and Model Faithfulness of LLM Rationale
por: Fayyaz, Mohsen, et al.
Publicado: (2024) -
Multimodal Cultural Safety: Evaluation Framework and Alignment Strategies
por: Qiu, Haoyi, et al.
Publicado: (2025) -
REFFLY: Melody-Constrained Lyrics Editing Model
por: Zhao, Songyan, et al.
Publicado: (2024)