Can Many-Shot In-Context Learning Help LLMs as Evaluators? A Preliminary Empirical Study
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Mingyang, Zheng, Mao, Luo, Xuan, Pan, Yue |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GRP: Goal-Reversed Prompting for Zero-Shot Evaluation with LLMs
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
Counting-Stars: A Multi-evidence, Position-aware, and Scalable Benchmark for Evaluating Long-Context Large Language Models
von: Song, Mingyang, et al.
Veröffentlicht: (2024)
von: Song, Mingyang, et al.
Veröffentlicht: (2024)
FastCuRL: Curriculum Reinforcement Learning with Stage-wise Context Scaling for Efficient Training R1-like Reasoning Models
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMs
von: Kim, Sangyeop, et al.
Veröffentlicht: (2025)
von: Kim, Sangyeop, et al.
Veröffentlicht: (2025)
Large Language Models as Zero-Shot Keyphrase Extractors: A Preliminary Empirical Study
von: Song, Mingyang, et al.
Veröffentlicht: (2023)
von: Song, Mingyang, et al.
Veröffentlicht: (2023)
A Preliminary Empirical Study on Prompt-based Unsupervised Keyphrase Extraction
von: Song, Mingyang, et al.
Veröffentlicht: (2024)
von: Song, Mingyang, et al.
Veröffentlicht: (2024)
On Many-Shot In-Context Learning for Long-Context Evaluation
von: Zou, Kaijian, et al.
Veröffentlicht: (2024)
von: Zou, Kaijian, et al.
Veröffentlicht: (2024)
An Empirical Study of Many-Shot In-Context Learning for Machine Translation of Low-Resource Languages
von: Lu, Yinhan, et al.
Veröffentlicht: (2026)
von: Lu, Yinhan, et al.
Veröffentlicht: (2026)
CodeDelegator: Mitigating Context Pollution via Role Separation in Code-as-Action Agents
von: Fei, Tianxiang, et al.
Veröffentlicht: (2026)
von: Fei, Tianxiang, et al.
Veröffentlicht: (2026)
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
A Survey of Query Optimization in Large Language Models
von: Song, Mingyang, et al.
Veröffentlicht: (2024)
von: Song, Mingyang, et al.
Veröffentlicht: (2024)
Many-Shot In-Context Learning
von: Agarwal, Rishabh, et al.
Veröffentlicht: (2024)
von: Agarwal, Rishabh, et al.
Veröffentlicht: (2024)
Beyond the Illusion of Consensus: From Surface Heuristics to Knowledge-Grounded Evaluation in LLM-as-a-Judge
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
Model Merging in the Era of Large Language Models: Methods, Applications, and Future Directions
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
Can LLMs Help Create Grammar?: Automating Grammar Creation for Endangered Languages with In-Context Learning
von: Spencer, Piyapath T, et al.
Veröffentlicht: (2024)
von: Spencer, Piyapath T, et al.
Veröffentlicht: (2024)
A Survey of On-Policy Distillation for Large Language Models
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness
von: Luo, Xuan, et al.
Veröffentlicht: (2025)
von: Luo, Xuan, et al.
Veröffentlicht: (2025)
Distilling Many-Shot In-Context Learning into a Cheat Sheet
von: Honda, Ukyo, et al.
Veröffentlicht: (2025)
von: Honda, Ukyo, et al.
Veröffentlicht: (2025)
Can LLMs Learn by Teaching for Better Reasoning? A Preliminary Study
von: Ning, Xuefei, et al.
Veröffentlicht: (2024)
von: Ning, Xuefei, et al.
Veröffentlicht: (2024)
Many-Shot In-Context Learning for Molecular Inverse Design
von: Moayedpour, Saeed, et al.
Veröffentlicht: (2024)
von: Moayedpour, Saeed, et al.
Veröffentlicht: (2024)
TAT-R1: Terminology-Aware Translation with Reinforcement Learning and Word Alignment
von: Li, Zheng, et al.
Veröffentlicht: (2025)
von: Li, Zheng, et al.
Veröffentlicht: (2025)
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
von: Xiong, Miao, et al.
Veröffentlicht: (2023)
von: Xiong, Miao, et al.
Veröffentlicht: (2023)
Selecting Demonstrations for Many-Shot In-Context Learning via Gradient Matching
von: Zhang, Jianfei, et al.
Veröffentlicht: (2025)
von: Zhang, Jianfei, et al.
Veröffentlicht: (2025)
Efficient Many-Shot In-Context Learning with Dynamic Block-Sparse Attention
von: Xiao, Emily, et al.
Veröffentlicht: (2025)
von: Xiao, Emily, et al.
Veröffentlicht: (2025)
Towards Compute-Optimal Many-Shot In-Context Learning
von: Golchin, Shahriar, et al.
Veröffentlicht: (2025)
von: Golchin, Shahriar, et al.
Veröffentlicht: (2025)
PRISM: Probability Reallocation with In-Span Masking for Knowledge-Sensitive Alignment
von: Xu, Chenning, et al.
Veröffentlicht: (2026)
von: Xu, Chenning, et al.
Veröffentlicht: (2026)
Many-Shot In-Context Learning in Multimodal Foundation Models
von: Jiang, Yixing, et al.
Veröffentlicht: (2024)
von: Jiang, Yixing, et al.
Veröffentlicht: (2024)
MiMoTable: A Multi-scale Spreadsheet Benchmark with Meta Operations for Table Reasoning
von: Li, Zheng, et al.
Veröffentlicht: (2024)
von: Li, Zheng, et al.
Veröffentlicht: (2024)
PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script Generation
von: Xu, Chenning, et al.
Veröffentlicht: (2026)
von: Xu, Chenning, et al.
Veröffentlicht: (2026)
Can Models Help Us Create Better Models? Evaluating LLMs as Data Scientists
von: Pietruszka, Michał, et al.
Veröffentlicht: (2024)
von: Pietruszka, Michał, et al.
Veröffentlicht: (2024)
Is Context Helpful for Chat Translation Evaluation?
von: Agrawal, Sweta, et al.
Veröffentlicht: (2024)
von: Agrawal, Sweta, et al.
Veröffentlicht: (2024)
HardMTBench: Stress-Testing Chinese-English Translation on Knowledge-Intensive Domains
von: Li, Zheng, et al.
Veröffentlicht: (2026)
von: Li, Zheng, et al.
Veröffentlicht: (2026)
Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
von: Du, Yexing, et al.
Veröffentlicht: (2024)
von: Du, Yexing, et al.
Veröffentlicht: (2024)
An Empirical Study of Many-to-Many Summarization with Large Language Models
von: Wang, Jiaan, et al.
Veröffentlicht: (2025)
von: Wang, Jiaan, et al.
Veröffentlicht: (2025)
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?
von: Yan, Kai, et al.
Veröffentlicht: (2025)
von: Yan, Kai, et al.
Veröffentlicht: (2025)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
von: Huang, Brandon, et al.
Veröffentlicht: (2024)
von: Huang, Brandon, et al.
Veröffentlicht: (2024)
Can Hallucinations Help? Boosting LLMs for Drug Discovery
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench
von: Valmeekam, Karthik, et al.
Veröffentlicht: (2024)
von: Valmeekam, Karthik, et al.
Veröffentlicht: (2024)
Structured Semantic Information Helps Retrieve Better Examples for In-Context Learning Applied to Few-Shot Relation Extraction
von: Chakma, Aunabil, et al.
Veröffentlicht: (2026)
von: Chakma, Aunabil, et al.
Veröffentlicht: (2026)
Scaling Laws for Many-Shot In-Context Learning with Self-Generated Annotations
von: Gu, Zhengyao, et al.
Veröffentlicht: (2025)
von: Gu, Zhengyao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GRP: Goal-Reversed Prompting for Zero-Shot Evaluation with LLMs
von: Song, Mingyang, et al.
Veröffentlicht: (2025) -
Counting-Stars: A Multi-evidence, Position-aware, and Scalable Benchmark for Evaluating Long-Context Large Language Models
von: Song, Mingyang, et al.
Veröffentlicht: (2024) -
FastCuRL: Curriculum Reinforcement Learning with Stage-wise Context Scaling for Efficient Training R1-like Reasoning Models
von: Song, Mingyang, et al.
Veröffentlicht: (2025) -
What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMs
von: Kim, Sangyeop, et al.
Veröffentlicht: (2025) -
Large Language Models as Zero-Shot Keyphrase Extractors: A Preliminary Empirical Study
von: Song, Mingyang, et al.
Veröffentlicht: (2023)