RAVEL: Reasoning Agents for Validating and Evaluating LLM Text Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, Andrew Zhuoer, Wang, Cunxiang, Luo, Yu, Wen, Bosi, Wang, Yidong, Fan, Lin, Zhou, Yilin, Wang, Zikang, Yu, Wenbo, Wu, Lindong, Wang, Hongning, Huang, Minlie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RLAR: An Agentic Reward System for Multi-task Reinforcement Learning on Large Language Models
von: Feng, Andrew Zhuoer, et al.
Veröffentlicht: (2026)
von: Feng, Andrew Zhuoer, et al.
Veröffentlicht: (2026)
HoWToBench: Holistic Evaluation for LLM's Capability in Human-level Writing using Tree of Writing
von: Feng, Andrew Zhuoer, et al.
Veröffentlicht: (2026)
von: Feng, Andrew Zhuoer, et al.
Veröffentlicht: (2026)
IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation
von: Wen, Bosi, et al.
Veröffentlicht: (2025)
von: Wen, Bosi, et al.
Veröffentlicht: (2025)
IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation
von: Wen, Bosi, et al.
Veröffentlicht: (2026)
von: Wen, Bosi, et al.
Veröffentlicht: (2026)
HPSS: Heuristic Prompting Strategy Search for LLM Evaluators
von: Wen, Bosi, et al.
Veröffentlicht: (2025)
von: Wen, Bosi, et al.
Veröffentlicht: (2025)
UDA: Unsupervised Debiasing Alignment for Pair-wise LLM-as-a-Judge
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
TraceSIR: A Multi-Agent Framework for Structured Analysis and Reporting of Agentic Execution Traces
von: Yang, Shu-Xun, et al.
Veröffentlicht: (2026)
von: Yang, Shu-Xun, et al.
Veröffentlicht: (2026)
StepMathAgent: A Step-Wise Agent for Evaluating Mathematical Processes through Tree-of-Error
von: Yang, Shu-Xun, et al.
Veröffentlicht: (2025)
von: Yang, Shu-Xun, et al.
Veröffentlicht: (2025)
CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation
von: Ke, Pei, et al.
Veröffentlicht: (2023)
von: Ke, Pei, et al.
Veröffentlicht: (2023)
Agent-SafetyBench: Evaluating the Safety of LLM Agents
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
Unlocking Reasoning Potential in Large Langauge Models by Scaling Code-form Planning
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
LongSafety: Evaluating Long-Context Safety of Large Language Models
von: Lu, Yida, et al.
Veröffentlicht: (2025)
von: Lu, Yida, et al.
Veröffentlicht: (2025)
AMOR: A Recipe for Building Adaptable Modular Knowledge Agents Through Process Feedback
von: Guan, Jian, et al.
Veröffentlicht: (2024)
von: Guan, Jian, et al.
Veröffentlicht: (2024)
Think Socially via Cognitive Reasoning
von: Zhou, Jinfeng, et al.
Veröffentlicht: (2025)
von: Zhou, Jinfeng, et al.
Veröffentlicht: (2025)
Language Model Decoding as Direct Metrics Optimization
von: Ji, Haozhe, et al.
Veröffentlicht: (2023)
von: Ji, Haozhe, et al.
Veröffentlicht: (2023)
SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation
von: Liu, Xiaoze, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoze, et al.
Veröffentlicht: (2024)
Training Language Model to Critique for Better Refinement
von: Yu, Tianshu, et al.
Veröffentlicht: (2025)
von: Yu, Tianshu, et al.
Veröffentlicht: (2025)
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
Benchmarking Complex Instruction-Following with Multiple Constraints Composition
von: Wen, Bosi, et al.
Veröffentlicht: (2024)
von: Wen, Bosi, et al.
Veröffentlicht: (2024)
MAPS: Advancing Multi-Modal Reasoning in Expert-Level Physical Science
von: Zhu, Erle, et al.
Veröffentlicht: (2025)
von: Zhu, Erle, et al.
Veröffentlicht: (2025)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
von: Huang, Jing, et al.
Veröffentlicht: (2024)
von: Huang, Jing, et al.
Veröffentlicht: (2024)
PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
von: Wang, Yidong, et al.
Veröffentlicht: (2023)
von: Wang, Yidong, et al.
Veröffentlicht: (2023)
Learning Task Decomposition to Assist Humans in Competitive Programming
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
Reasoning on Multiple Needles In A Haystack
von: Wang, Yidong
Veröffentlicht: (2025)
von: Wang, Yidong
Veröffentlicht: (2025)
Trust-Region Adaptive Policy Optimization
von: Su, Mingyu, et al.
Veröffentlicht: (2025)
von: Su, Mingyu, et al.
Veröffentlicht: (2025)
Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints
von: Yang, Junxiao, et al.
Veröffentlicht: (2025)
von: Yang, Junxiao, et al.
Veröffentlicht: (2025)
Towards Efficient Exact Optimization of Language Model Alignment
von: Ji, Haozhe, et al.
Veröffentlicht: (2024)
von: Ji, Haozhe, et al.
Veröffentlicht: (2024)
Uncovering the Text Embedding in Text-to-Image Diffusion Models
von: Yu, Hu, et al.
Veröffentlicht: (2024)
von: Yu, Hu, et al.
Veröffentlicht: (2024)
AlignBench: Benchmarking Chinese Alignment of Large Language Models
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation
von: Tian, Yanzhi, et al.
Veröffentlicht: (2026)
von: Tian, Yanzhi, et al.
Veröffentlicht: (2026)
Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
von: Zhang, Zhexin, et al.
Veröffentlicht: (2023)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2023)
Grounding LLMs in Scientific Discovery via Embodied Actions
von: Zhang, Bo, et al.
Veröffentlicht: (2026)
von: Zhang, Bo, et al.
Veröffentlicht: (2026)
$R^3$: "This is My SQL, Are You With Me?" A Consensus-Based Multi-Agent System for Text-to-SQL Tasks
von: Xia, Hanchen, et al.
Veröffentlicht: (2024)
von: Xia, Hanchen, et al.
Veröffentlicht: (2024)
DVD: A Robust Method for Detecting Variant Contamination in Large Language Model Evaluation
von: Liang, Renzhao, et al.
Veröffentlicht: (2026)
von: Liang, Renzhao, et al.
Veröffentlicht: (2026)
LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models
von: Gui, Jiayi, et al.
Veröffentlicht: (2024)
von: Gui, Jiayi, et al.
Veröffentlicht: (2024)
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
von: Cui, Shiyao, et al.
Veröffentlicht: (2025)
von: Cui, Shiyao, et al.
Veröffentlicht: (2025)
VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning
von: Wang, Zikang, et al.
Veröffentlicht: (2025)
von: Wang, Zikang, et al.
Veröffentlicht: (2025)
When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity
von: Cui, Shiyao, et al.
Veröffentlicht: (2025)
von: Cui, Shiyao, et al.
Veröffentlicht: (2025)
Data-Efficient RLVR via Off-Policy Influence Guidance
von: Zhu, Erle, et al.
Veröffentlicht: (2025)
von: Zhu, Erle, et al.
Veröffentlicht: (2025)
Detector-Evasive LLM Paraphrasing via Constrained Policy Optimization
von: Wang, Mingyi, et al.
Veröffentlicht: (2026)
von: Wang, Mingyi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
RLAR: An Agentic Reward System for Multi-task Reinforcement Learning on Large Language Models
von: Feng, Andrew Zhuoer, et al.
Veröffentlicht: (2026) -
HoWToBench: Holistic Evaluation for LLM's Capability in Human-level Writing using Tree of Writing
von: Feng, Andrew Zhuoer, et al.
Veröffentlicht: (2026) -
IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation
von: Wen, Bosi, et al.
Veröffentlicht: (2025) -
IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation
von: Wen, Bosi, et al.
Veröffentlicht: (2026) -
HPSS: Heuristic Prompting Strategy Search for LLM Evaluators
von: Wen, Bosi, et al.
Veröffentlicht: (2025)