GRP: Goal-Reversed Prompting for Zero-Shot Evaluation with LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Mingyang, Zheng, Mao, Luo, Xuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Many-Shot In-Context Learning Help LLMs as Evaluators? A Preliminary Empirical Study
by: Song, Mingyang, et al.
Published: (2024)
by: Song, Mingyang, et al.
Published: (2024)
Counting-Stars: A Multi-evidence, Position-aware, and Scalable Benchmark for Evaluating Long-Context Large Language Models
by: Song, Mingyang, et al.
Published: (2024)
by: Song, Mingyang, et al.
Published: (2024)
Beyond the Illusion of Consensus: From Surface Heuristics to Knowledge-Grounded Evaluation in LLM-as-a-Judge
by: Song, Mingyang, et al.
Published: (2026)
by: Song, Mingyang, et al.
Published: (2026)
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
by: Song, Mingyang, et al.
Published: (2025)
by: Song, Mingyang, et al.
Published: (2025)
Model Merging in the Era of Large Language Models: Methods, Applications, and Future Directions
by: Song, Mingyang, et al.
Published: (2026)
by: Song, Mingyang, et al.
Published: (2026)
A Survey of Query Optimization in Large Language Models
by: Song, Mingyang, et al.
Published: (2024)
by: Song, Mingyang, et al.
Published: (2024)
A Survey of On-Policy Distillation for Large Language Models
by: Song, Mingyang, et al.
Published: (2026)
by: Song, Mingyang, et al.
Published: (2026)
FastCuRL: Curriculum Reinforcement Learning with Stage-wise Context Scaling for Efficient Training R1-like Reasoning Models
by: Song, Mingyang, et al.
Published: (2025)
by: Song, Mingyang, et al.
Published: (2025)
SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation
by: Yang, Wenjie, et al.
Published: (2025)
by: Yang, Wenjie, et al.
Published: (2025)
PRISM: Probability Reallocation with In-Span Masking for Knowledge-Sensitive Alignment
by: Xu, Chenning, et al.
Published: (2026)
by: Xu, Chenning, et al.
Published: (2026)
Towards Hierarchical Multi-Agent Workflows for Zero-Shot Prompt Optimization
by: Liu, Yuchi, et al.
Published: (2024)
by: Liu, Yuchi, et al.
Published: (2024)
Nsanku: Evaluating Zero-Shot Translation Performance of LLMs for Ghanaian Languages
by: Moore, Stephen E., et al.
Published: (2026)
by: Moore, Stephen E., et al.
Published: (2026)
Large Language Models as Zero-Shot Keyphrase Extractors: A Preliminary Empirical Study
by: Song, Mingyang, et al.
Published: (2023)
by: Song, Mingyang, et al.
Published: (2023)
On Zero-Shot Counterspeech Generation by LLMs
by: Saha, Punyajoy, et al.
Published: (2024)
by: Saha, Punyajoy, et al.
Published: (2024)
HARGPT: Are LLMs Zero-Shot Human Activity Recognizers?
by: Ji, Sijie, et al.
Published: (2024)
by: Ji, Sijie, et al.
Published: (2024)
TAT-R1: Terminology-Aware Translation with Reinforcement Learning and Word Alignment
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
HardMTBench: Stress-Testing Chinese-English Translation on Knowledge-Intensive Domains
by: Li, Zheng, et al.
Published: (2026)
by: Li, Zheng, et al.
Published: (2026)
MiMoTable: A Multi-scale Spreadsheet Benchmark with Meta Operations for Table Reasoning
by: Li, Zheng, et al.
Published: (2024)
by: Li, Zheng, et al.
Published: (2024)
PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script Generation
by: Xu, Chenning, et al.
Published: (2026)
by: Xu, Chenning, et al.
Published: (2026)
Which Words Matter Most in Zero-Shot Prompts?
by: Sadr, Nikta Gohari, et al.
Published: (2025)
by: Sadr, Nikta Gohari, et al.
Published: (2025)
Better Zero-Shot Reasoning with Role-Play Prompting
by: Kong, Aobo, et al.
Published: (2023)
by: Kong, Aobo, et al.
Published: (2023)
Are LLMs Good Zero-Shot Fallacy Classifiers?
by: Pan, Fengjun, et al.
Published: (2024)
by: Pan, Fengjun, et al.
Published: (2024)
Instances Need More Care: Rewriting Prompts for Instances with LLMs in the Loop Yields Better Zero-Shot Performance
by: Srivastava, Saurabh, et al.
Published: (2023)
by: Srivastava, Saurabh, et al.
Published: (2023)
Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs
by: Sekar, Anirudh, et al.
Published: (2026)
by: Sekar, Anirudh, et al.
Published: (2026)
Beyond the Next Token: Towards Prompt-Robust Zero-Shot Classification via Efficient Multi-Token Prediction
by: Qian, Junlang, et al.
Published: (2025)
by: Qian, Junlang, et al.
Published: (2025)
Better Benchmarking LLMs for Zero-Shot Dependency Parsing
by: Ezquerro, Ana, et al.
Published: (2025)
by: Ezquerro, Ana, et al.
Published: (2025)
Zero-Shot Belief: A Hard Problem for LLMs
by: Murzaku, John, et al.
Published: (2025)
by: Murzaku, John, et al.
Published: (2025)
Towards Zero-Shot, Controllable Dialog Planning with LLMs
by: Väth, Dirk, et al.
Published: (2024)
by: Väth, Dirk, et al.
Published: (2024)
LLMs Are Zero-Shot Context-Aware Simultaneous Translators
by: Koshkin, Roman, et al.
Published: (2024)
by: Koshkin, Roman, et al.
Published: (2024)
A Preliminary Empirical Study on Prompt-based Unsupervised Keyphrase Extraction
by: Song, Mingyang, et al.
Published: (2024)
by: Song, Mingyang, et al.
Published: (2024)
HY-MT1.5 Technical Report
by: Zheng, Mao, et al.
Published: (2025)
by: Zheng, Mao, et al.
Published: (2025)
Taxonomy-Guided Zero-Shot Recommendations with LLMs
by: Liang, Yueqing, et al.
Published: (2024)
by: Liang, Yueqing, et al.
Published: (2024)
OmniVox: Zero-Shot Emotion Recognition with Omni-LLMs
by: Murzaku, John, et al.
Published: (2025)
by: Murzaku, John, et al.
Published: (2025)
Towards Reliable Latent Knowledge Estimation in LLMs: Zero-Prompt Many-Shot Based Factual Knowledge Extraction
by: Wu, Qinyuan, et al.
Published: (2024)
by: Wu, Qinyuan, et al.
Published: (2024)
Evaluating Compact LLMs for Zero-Shot Iberian Language Tasks on End-User Devices
by: Seller, Luís Couto, et al.
Published: (2025)
by: Seller, Luís Couto, et al.
Published: (2025)
English Prompts are Better for NLI-based Zero-Shot Emotion Classification than Target-Language Prompts
by: Bareiß, Patrick, et al.
Published: (2024)
by: Bareiß, Patrick, et al.
Published: (2024)
MAGIC-Enhanced Keyword Prompting for Zero-Shot Audio Captioning with CLIP Models
by: Govindarajan, Vijay, et al.
Published: (2025)
by: Govindarajan, Vijay, et al.
Published: (2025)
Exploring Zero-Shot ACSA with Unified Meaning Representation in Chain-of-Thought Prompting
by: Ventirozos, Filippos, et al.
Published: (2025)
by: Ventirozos, Filippos, et al.
Published: (2025)
Enabling Natural Zero-Shot Prompting on Encoder Models via Statement-Tuning
by: Elshabrawy, Ahmed, et al.
Published: (2024)
by: Elshabrawy, Ahmed, et al.
Published: (2024)
Navigating Prompt Complexity for Zero-Shot Classification: A Study of Large Language Models in Computational Social Science
by: Mu, Yida, et al.
Published: (2023)
by: Mu, Yida, et al.
Published: (2023)
Similar Items
-
Can Many-Shot In-Context Learning Help LLMs as Evaluators? A Preliminary Empirical Study
by: Song, Mingyang, et al.
Published: (2024) -
Counting-Stars: A Multi-evidence, Position-aware, and Scalable Benchmark for Evaluating Long-Context Large Language Models
by: Song, Mingyang, et al.
Published: (2024) -
Beyond the Illusion of Consensus: From Surface Heuristics to Knowledge-Grounded Evaluation in LLM-as-a-Judge
by: Song, Mingyang, et al.
Published: (2026) -
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
by: Song, Mingyang, et al.
Published: (2025) -
Model Merging in the Era of Large Language Models: Methods, Applications, and Future Directions
by: Song, Mingyang, et al.
Published: (2026)