Gespeichert in:
| Hauptverfasser: | Hu, Jinpeng, Dong, Tengteng, Gang, Luo, Ma, Hui, Zou, Peng, Sun, Xiao, Guo, Dan, Yang, Xun, Wang, Meng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2407.05721 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Psyche-R1: Towards Reliable Psychological LLMs through Unified Empathy, Expertise, and Reasoning
von: Dai, Chongyuan, et al.
Veröffentlicht: (2025)
von: Dai, Chongyuan, et al.
Veröffentlicht: (2025)
Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors
von: Li, Jia, et al.
Veröffentlicht: (2025)
von: Li, Jia, et al.
Veröffentlicht: (2025)
AgentMental: An Interactive Multi-Agent Framework for Explainable and Adaptive Mental Health Assessment
von: Hu, Jinpeng, et al.
Veröffentlicht: (2025)
von: Hu, Jinpeng, et al.
Veröffentlicht: (2025)
Think-Augmented Function Calling: Improving LLM Parameter Accuracy Through Embedded Reasoning
von: Wei, Lei, et al.
Veröffentlicht: (2026)
von: Wei, Lei, et al.
Veröffentlicht: (2026)
In-Context Examples Matter: Improving Emotion Recognition in Conversation with Instruction Tuning
von: Ma, Hui, et al.
Veröffentlicht: (2025)
von: Ma, Hui, et al.
Veröffentlicht: (2025)
Unveiling LLM Evaluation Focused on Metrics: Challenges and Solutions
von: Hu, Taojun, et al.
Veröffentlicht: (2024)
von: Hu, Taojun, et al.
Veröffentlicht: (2024)
Understanding Layer Significance in LLM Alignment
von: Shi, Guangyuan, et al.
Veröffentlicht: (2024)
von: Shi, Guangyuan, et al.
Veröffentlicht: (2024)
TypyBench: Evaluating LLM Type Inference for Untyped Python Repositories
von: Dong, Honghua, et al.
Veröffentlicht: (2025)
von: Dong, Honghua, et al.
Veröffentlicht: (2025)
DOCBENCH: A Benchmark for Evaluating LLM-based Document Reading Systems
von: Zou, Anni, et al.
Veröffentlicht: (2024)
von: Zou, Anni, et al.
Veröffentlicht: (2024)
Ψ-Arena: Interactive Assessment and Optimization of LLM-based Psychological Counselors with Tripartite Feedback
von: Zhu, Shijing, et al.
Veröffentlicht: (2025)
von: Zhu, Shijing, et al.
Veröffentlicht: (2025)
CSCE: Boosting LLM Reasoning by Simultaneous Enhancing of Causal Significance and Consistency
von: Wang, Kangsheng, et al.
Veröffentlicht: (2024)
von: Wang, Kangsheng, et al.
Veröffentlicht: (2024)
Lost in the Mix: Evaluating LLM Understanding of Code-Switched Text
von: Mohamed, Amr, et al.
Veröffentlicht: (2025)
von: Mohamed, Amr, et al.
Veröffentlicht: (2025)
LLM Hallucination Detection: HSAD
von: Li, JinXin, et al.
Veröffentlicht: (2025)
von: Li, JinXin, et al.
Veröffentlicht: (2025)
ScreenLLM: Stateful Screen Schema for Efficient Action Understanding and Prediction
von: Jin, Yiqiao, et al.
Veröffentlicht: (2025)
von: Jin, Yiqiao, et al.
Veröffentlicht: (2025)
LLM-based NLG Evaluation: Current Status and Challenges
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
Enhancing LLM Reasoning with Multi-Path Collaborative Reactive and Reflection agents
von: He, Chengbo, et al.
Veröffentlicht: (2024)
von: He, Chengbo, et al.
Veröffentlicht: (2024)
MedDialBench: Benchmarking LLM Diagnostic Robustness under Parametric Adversarial Patient Behaviors
von: Luo, Xiaotian, et al.
Veröffentlicht: (2026)
von: Luo, Xiaotian, et al.
Veröffentlicht: (2026)
Explaining Length Bias in LLM-Based Preference Evaluations
von: Hu, Zhengyu, et al.
Veröffentlicht: (2024)
von: Hu, Zhengyu, et al.
Veröffentlicht: (2024)
LLM-Guided Strategy Synthesis for Scalable Equality Saturation
von: Yin, Chenyun, et al.
Veröffentlicht: (2026)
von: Yin, Chenyun, et al.
Veröffentlicht: (2026)
Evaluating Human Alignment and Model Faithfulness of LLM Rationale
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2024)
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2024)
Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual Understanding
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
Code Fingerprints: Disentangled Attribution of LLM-Generated Code
von: Guo, Jiaxun, et al.
Veröffentlicht: (2026)
von: Guo, Jiaxun, et al.
Veröffentlicht: (2026)
HEART-Bench: Do LLM Agents Exhibit Human-like Psychology?
von: Peng, Weihan, et al.
Veröffentlicht: (2026)
von: Peng, Weihan, et al.
Veröffentlicht: (2026)
Benchmarking LLM Guardrails in Handling Multilingual Toxicity
von: Yang, Yahan, et al.
Veröffentlicht: (2024)
von: Yang, Yahan, et al.
Veröffentlicht: (2024)
WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction
von: Zhang, Binbin, et al.
Veröffentlicht: (2025)
von: Zhang, Binbin, et al.
Veröffentlicht: (2025)
Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
von: Cao, Yixin, et al.
Veröffentlicht: (2025)
von: Cao, Yixin, et al.
Veröffentlicht: (2025)
Are LLM-based Evaluators Confusing NLG Quality Criteria?
von: Hu, Xinyu, et al.
Veröffentlicht: (2024)
von: Hu, Xinyu, et al.
Veröffentlicht: (2024)
Citation-Enhanced Generation for LLM-based Chatbots
von: Li, Weitao, et al.
Veröffentlicht: (2024)
von: Li, Weitao, et al.
Veröffentlicht: (2024)
Skill-Conditioned Gated Self-Distillation for LLM Reasoning
von: Huang, Jiazhen, et al.
Veröffentlicht: (2026)
von: Huang, Jiazhen, et al.
Veröffentlicht: (2026)
LLM-A*: Large Language Model Enhanced Incremental Heuristic Search on Path Planning
von: Meng, Silin, et al.
Veröffentlicht: (2024)
von: Meng, Silin, et al.
Veröffentlicht: (2024)
LLM-MedQA: Enhancing Medical Question Answering through Case Studies in Large Language Models
von: Yang, Hang, et al.
Veröffentlicht: (2024)
von: Yang, Hang, et al.
Veröffentlicht: (2024)
DuanzAI: Slang-Enhanced LLM with Prompt for Humor Understanding
von: Rohn, Yesian
Veröffentlicht: (2024)
von: Rohn, Yesian
Veröffentlicht: (2024)
BoRP: Bootstrapped Regression Probing for Scalable and Human-Aligned LLM Evaluation
von: Sun, Peng, et al.
Veröffentlicht: (2026)
von: Sun, Peng, et al.
Veröffentlicht: (2026)
Understanding LLM Embeddings for Regression
von: Tang, Eric, et al.
Veröffentlicht: (2024)
von: Tang, Eric, et al.
Veröffentlicht: (2024)
IDGen: Item Discrimination Induced Prompt Generation for LLM Evaluation
von: Lin, Fan, et al.
Veröffentlicht: (2024)
von: Lin, Fan, et al.
Veröffentlicht: (2024)
Exploring LLM Multi-Agents for ICD Coding
von: Li, Rumeng, et al.
Veröffentlicht: (2024)
von: Li, Rumeng, et al.
Veröffentlicht: (2024)
PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
von: Yang, Dongjie, et al.
Veröffentlicht: (2024)
von: Yang, Dongjie, et al.
Veröffentlicht: (2024)
MIRAI: Evaluating LLM Agents for Event Forecasting
von: Ye, Chenchen, et al.
Veröffentlicht: (2024)
von: Ye, Chenchen, et al.
Veröffentlicht: (2024)
HuggingGraph: Understanding the Supply Chain of LLM Ecosystem
von: Rahman, Mohammad Shahedur, et al.
Veröffentlicht: (2025)
von: Rahman, Mohammad Shahedur, et al.
Veröffentlicht: (2025)
HTAA: Enhancing LLM Planning via Hybrid Toolset Agentization & Adaptation
von: Huang, Chengrui, et al.
Veröffentlicht: (2026)
von: Huang, Chengrui, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Psyche-R1: Towards Reliable Psychological LLMs through Unified Empathy, Expertise, and Reasoning
von: Dai, Chongyuan, et al.
Veröffentlicht: (2025) -
Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors
von: Li, Jia, et al.
Veröffentlicht: (2025) -
AgentMental: An Interactive Multi-Agent Framework for Explainable and Adaptive Mental Health Assessment
von: Hu, Jinpeng, et al.
Veröffentlicht: (2025) -
Think-Augmented Function Calling: Improving LLM Parameter Accuracy Through Embedded Reasoning
von: Wei, Lei, et al.
Veröffentlicht: (2026) -
In-Context Examples Matter: Improving Emotion Recognition in Conversation with Instruction Tuning
von: Ma, Hui, et al.
Veröffentlicht: (2025)