FedEval-LLM: Federated Evaluation of Large Language Models on Downstream Tasks with Collective Wisdom
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Yuanqin, Kang, Yan, Fan, Lixin, Yang, Qiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FedCoT: Federated Chain-of-Thought Distillation for Large Language Models
von: Fan, Tao, et al.
Veröffentlicht: (2024)
von: Fan, Tao, et al.
Veröffentlicht: (2024)
FedMKT: Federated Mutual Knowledge Transfer for Large and Small Language Models
von: Fan, Tao, et al.
Veröffentlicht: (2024)
von: Fan, Tao, et al.
Veröffentlicht: (2024)
PPC-GPT: Federated Task-Specific Compression of Large Language Models via Pruning and Chain-of-Thought Distillation
von: Fan, Tao, et al.
Veröffentlicht: (2025)
von: Fan, Tao, et al.
Veröffentlicht: (2025)
Federated Co-tuning Framework for Large and Small Language Models
von: Fan, Tao, et al.
Veröffentlicht: (2024)
von: Fan, Tao, et al.
Veröffentlicht: (2024)
MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures
von: Ni, Jinjie, et al.
Veröffentlicht: (2024)
von: Ni, Jinjie, et al.
Veröffentlicht: (2024)
Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake Analysis
von: Chen, Kai, et al.
Veröffentlicht: (2023)
von: Chen, Kai, et al.
Veröffentlicht: (2023)
DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
von: Chung, Tsz Ting, et al.
Veröffentlicht: (2025)
von: Chung, Tsz Ting, et al.
Veröffentlicht: (2025)
Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks
von: Li, Miaomiao, et al.
Veröffentlicht: (2025)
von: Li, Miaomiao, et al.
Veröffentlicht: (2025)
VocabTailor: Dynamic Vocabulary Selection for Downstream Tasks in Small Language Models
von: Zhang, Hanling, et al.
Veröffentlicht: (2025)
von: Zhang, Hanling, et al.
Veröffentlicht: (2025)
AdEval: Alignment-based Dynamic Evaluation to Mitigate Data Contamination in Large Language Models
von: Fan, Yang
Veröffentlicht: (2025)
von: Fan, Yang
Veröffentlicht: (2025)
FedLLM-Bench: Realistic Benchmarks for Federated Learning of Large Language Models
von: Ye, Rui, et al.
Veröffentlicht: (2024)
von: Ye, Rui, et al.
Veröffentlicht: (2024)
SecureBoost+: Large Scale and High-Performance Vertical Federated Gradient Boosting Decision Tree
von: Fan, Tao, et al.
Veröffentlicht: (2021)
von: Fan, Tao, et al.
Veröffentlicht: (2021)
The Shape of Wisdom: Decision Trajectories in Language Models
von: Rana, Shailesh
Veröffentlicht: (2026)
von: Rana, Shailesh
Veröffentlicht: (2026)
Revisiting the Scaling Properties of Downstream Metrics in Large Language Model Training
von: Krajewski, Jakub, et al.
Veröffentlicht: (2025)
von: Krajewski, Jakub, et al.
Veröffentlicht: (2025)
StructEval: Deepen and Broaden Large Language Model Assessment via Structured Evaluation
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
Grounding Foundation Models through Federated Transfer Learning: A General Framework
von: Kang, Yan, et al.
Veröffentlicht: (2023)
von: Kang, Yan, et al.
Veröffentlicht: (2023)
FedProxy: Federated Fine-Tuning of LLMs via Proxy SLMs and Heterogeneity-Aware Fusion
von: Fan, Tao, et al.
Veröffentlicht: (2026)
von: Fan, Tao, et al.
Veröffentlicht: (2026)
FedTracker: Furnishing Ownership Verification and Traceability for Federated Learning Model
von: Shao, Shuo, et al.
Veröffentlicht: (2022)
von: Shao, Shuo, et al.
Veröffentlicht: (2022)
CityBench: Evaluating the Capabilities of Large Language Models for Urban Tasks
von: Feng, Jie, et al.
Veröffentlicht: (2024)
von: Feng, Jie, et al.
Veröffentlicht: (2024)
DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
A Meta-learning Framework for Tuning Parameters of Protection Mechanisms in Trustworthy Federated Learning
von: Zhang, Xiaojin, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaojin, et al.
Veröffentlicht: (2023)
ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks
von: Yu, Xiaodong, et al.
Veröffentlicht: (2023)
von: Yu, Xiaodong, et al.
Veröffentlicht: (2023)
Learning to Rewrite Prompts for Bootstrapping LLMs on Downstream Tasks
von: Zhou, Qinhao, et al.
Veröffentlicht: (2025)
von: Zhou, Qinhao, et al.
Veröffentlicht: (2025)
ChemEval: A Comprehensive Multi-Level Chemical Evaluation for Large Language Models
von: Huang, Yuqing, et al.
Veröffentlicht: (2024)
von: Huang, Yuqing, et al.
Veröffentlicht: (2024)
Large-Small Model Collaborative Framework for Federated Continual Learning
von: Yu, Hao, et al.
Veröffentlicht: (2025)
von: Yu, Hao, et al.
Veröffentlicht: (2025)
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models
von: Nguyen, Trong-Hieu, et al.
Veröffentlicht: (2024)
von: Nguyen, Trong-Hieu, et al.
Veröffentlicht: (2024)
Hierarchical Federated Unlearning for Large Language Models
von: Zhong, Yisheng, et al.
Veröffentlicht: (2025)
von: Zhong, Yisheng, et al.
Veröffentlicht: (2025)
QualEval: Qualitative Evaluation for Model Improvement
von: Murahari, Vishvak, et al.
Veröffentlicht: (2023)
von: Murahari, Vishvak, et al.
Veröffentlicht: (2023)
Federated Data-Efficient Instruction Tuning for Large Language Models
von: Qin, Zhen, et al.
Veröffentlicht: (2024)
von: Qin, Zhen, et al.
Veröffentlicht: (2024)
CriticEval: Evaluating Large Language Model as Critic
von: Lan, Tian, et al.
Veröffentlicht: (2024)
von: Lan, Tian, et al.
Veröffentlicht: (2024)
Task-Informed Anti-Curriculum by Masking Improves Downstream Performance on Text
von: Jarca, Andrei, et al.
Veröffentlicht: (2025)
von: Jarca, Andrei, et al.
Veröffentlicht: (2025)
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd
von: Liu, Shang, et al.
Veröffentlicht: (2024)
von: Liu, Shang, et al.
Veröffentlicht: (2024)
NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes
von: Fan, Lizhou, et al.
Veröffentlicht: (2023)
von: Fan, Lizhou, et al.
Veröffentlicht: (2023)
ZKP-FedEval: Verifiable and Privacy-Preserving Federated Evaluation using Zero-Knowledge Proofs
von: Commey, Daniel, et al.
Veröffentlicht: (2025)
von: Commey, Daniel, et al.
Veröffentlicht: (2025)
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
von: Adib, Shefayat E Shams, et al.
Veröffentlicht: (2026)
von: Adib, Shefayat E Shams, et al.
Veröffentlicht: (2026)
HypoEval: Hypothesis-Guided Evaluation for Natural Language Generation
von: Li, Mingxuan, et al.
Veröffentlicht: (2025)
von: Li, Mingxuan, et al.
Veröffentlicht: (2025)
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
von: Sansford, Hannah, et al.
Veröffentlicht: (2024)
von: Sansford, Hannah, et al.
Veröffentlicht: (2024)
LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
von: Ren, Huimin, et al.
Veröffentlicht: (2025)
von: Ren, Huimin, et al.
Veröffentlicht: (2025)
LogicTree: Structured Proof Exploration for Coherent and Rigorous Logical Reasoning with Large Language Models
von: He, Kang, et al.
Veröffentlicht: (2025)
von: He, Kang, et al.
Veröffentlicht: (2025)
Measuring all the noises of LLM Evals
von: Wang, Sida
Veröffentlicht: (2025)
von: Wang, Sida
Veröffentlicht: (2025)
Ähnliche Einträge
-
FedCoT: Federated Chain-of-Thought Distillation for Large Language Models
von: Fan, Tao, et al.
Veröffentlicht: (2024) -
FedMKT: Federated Mutual Knowledge Transfer for Large and Small Language Models
von: Fan, Tao, et al.
Veröffentlicht: (2024) -
PPC-GPT: Federated Task-Specific Compression of Large Language Models via Pruning and Chain-of-Thought Distillation
von: Fan, Tao, et al.
Veröffentlicht: (2025) -
Federated Co-tuning Framework for Large and Small Language Models
von: Fan, Tao, et al.
Veröffentlicht: (2024) -
MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures
von: Ni, Jinjie, et al.
Veröffentlicht: (2024)