One Token to Fool LLM-as-a-Judge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Yulai, Liu, Haolin, Yu, Dian, Kung, Sunyuan, Chen, Meijia, Mi, Haitao, Yu, Dong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
von: Yu, Dian, et al.
Veröffentlicht: (2025)
von: Yu, Dian, et al.
Veröffentlicht: (2025)
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
von: Liu, Haolin, et al.
Veröffentlicht: (2026)
von: Liu, Haolin, et al.
Veröffentlicht: (2026)
Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation
von: Zhou, Yujun, et al.
Veröffentlicht: (2025)
von: Zhou, Yujun, et al.
Veröffentlicht: (2025)
Scaling Synthetic Data Creation with 1,000,000,000 Personas
von: Ge, Tao, et al.
Veröffentlicht: (2024)
von: Ge, Tao, et al.
Veröffentlicht: (2024)
LiteSearch: Efficacious Tree Search for LLM
von: Wang, Ante, et al.
Veröffentlicht: (2024)
von: Wang, Ante, et al.
Veröffentlicht: (2024)
Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
von: Tian, Ye, et al.
Veröffentlicht: (2024)
von: Tian, Ye, et al.
Veröffentlicht: (2024)
Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
Stable and Efficient Single-Rollout RL for Multimodal Reasoning
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
von: Yue, Murong, et al.
Veröffentlicht: (2024)
von: Yue, Murong, et al.
Veröffentlicht: (2024)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
von: Dai, Runpeng, et al.
Veröffentlicht: (2025)
von: Dai, Runpeng, et al.
Veröffentlicht: (2025)
Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
von: Panaganti, Kishan, et al.
Veröffentlicht: (2026)
von: Panaganti, Kishan, et al.
Veröffentlicht: (2026)
Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
Dual-Uncertainty Guided Policy Learning for Multimodal Reasoning
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
von: Shi, Yucheng, et al.
Veröffentlicht: (2025)
von: Shi, Yucheng, et al.
Veröffentlicht: (2025)
R-Zero: Self-Evolving Reasoning LLM from Zero Data
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
Quantitative LLM Judges
von: Sahoo, Aishwarya, et al.
Veröffentlicht: (2025)
von: Sahoo, Aishwarya, et al.
Veröffentlicht: (2025)
Adding Conditional Control to Diffusion Models with Reinforcement Learning
von: Zhao, Yulai, et al.
Veröffentlicht: (2024)
von: Zhao, Yulai, et al.
Veröffentlicht: (2024)
RAFT: Realistic Attacks to Fool Text Detectors
von: Wang, James, et al.
Veröffentlicht: (2024)
von: Wang, James, et al.
Veröffentlicht: (2024)
Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
von: Zhang, Yuheng, et al.
Veröffentlicht: (2024)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2024)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
von: Xu, Ran, et al.
Veröffentlicht: (2025)
von: Xu, Ran, et al.
Veröffentlicht: (2025)
Mitigating Heterogeneous Token Overfitting in LLM Knowledge Editing
von: Liu, Tianci, et al.
Veröffentlicht: (2025)
von: Liu, Tianci, et al.
Veröffentlicht: (2025)
Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning
von: Liu, Hanbing, et al.
Veröffentlicht: (2025)
von: Liu, Hanbing, et al.
Veröffentlicht: (2025)
Guided Self-Evolving LLMs with Minimal Human Supervision
von: Yu, Wenhao, et al.
Veröffentlicht: (2025)
von: Yu, Wenhao, et al.
Veröffentlicht: (2025)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
Trust in One Round: Confidence Estimation for Large Language Models via Structural Signals
von: Yang, Pengyue, et al.
Veröffentlicht: (2026)
von: Yang, Pengyue, et al.
Veröffentlicht: (2026)
SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression
von: Wang, Xin, et al.
Veröffentlicht: (2024)
von: Wang, Xin, et al.
Veröffentlicht: (2024)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection
von: Wei, Zhipeng, et al.
Veröffentlicht: (2024)
von: Wei, Zhipeng, et al.
Veröffentlicht: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
From Token to Token Pair: Efficient Prompt Compression for Large Language Models in Clinical Prediction
von: Zhu, Mingcheng, et al.
Veröffentlicht: (2026)
von: Zhu, Mingcheng, et al.
Veröffentlicht: (2026)
Reinforcement Learning-based Knowledge Distillation with LLM-as-a-Judge
von: Shen, Yiyang, et al.
Veröffentlicht: (2026)
von: Shen, Yiyang, et al.
Veröffentlicht: (2026)
Token-Budget-Aware LLM Reasoning
von: Han, Tingxu, et al.
Veröffentlicht: (2024)
von: Han, Tingxu, et al.
Veröffentlicht: (2024)
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025)
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025)
One Pass Streaming Algorithm for Super Long Token Attention Approximation in Sublinear Space
von: Addanki, Raghav, et al.
Veröffentlicht: (2023)
von: Addanki, Raghav, et al.
Veröffentlicht: (2023)
How to Correctly Report LLM-as-a-Judge Evaluations
von: Lee, Chungpa, et al.
Veröffentlicht: (2025)
von: Lee, Chungpa, et al.
Veröffentlicht: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM Embeddings
von: Ding, Xueying, et al.
Veröffentlicht: (2025)
von: Ding, Xueying, et al.
Veröffentlicht: (2025)
Learning Rate Matters: Vanilla LoRA May Suffice for LLM Fine-tuning
von: Lee, Yu-Ang, et al.
Veröffentlicht: (2026)
von: Lee, Yu-Ang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
von: Yu, Dian, et al.
Veröffentlicht: (2025) -
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
von: Liu, Haolin, et al.
Veröffentlicht: (2026) -
Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation
von: Zhou, Yujun, et al.
Veröffentlicht: (2025) -
Scaling Synthetic Data Creation with 1,000,000,000 Personas
von: Ge, Tao, et al.
Veröffentlicht: (2024) -
LiteSearch: Efficacious Tree Search for LLM
von: Wang, Ante, et al.
Veröffentlicht: (2024)