LLM-REVal: Can We Trust LLM Reviewers Yet?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Rui, Gu, Jia-Chen, Kung, Po-Nien, Xia, Heming, liu, Junfeng, Kong, Xiangwen, Sui, Zhifang, Peng, Nanyun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Decoupling Task-Solving and Output Formatting in LLM Generation
von: Deng, Haikang, et al.
Veröffentlicht: (2025)
von: Deng, Haikang, et al.
Veröffentlicht: (2025)
Adaptable Logical Control for Large Language Models
von: Zhang, Honghua, et al.
Veröffentlicht: (2024)
von: Zhang, Honghua, et al.
Veröffentlicht: (2024)
Can Large Multimodal Models Uncover Deep Semantics Behind Images?
von: Yang, Yixin, et al.
Veröffentlicht: (2024)
von: Yang, Yixin, et al.
Veröffentlicht: (2024)
Can We Trust LLM Detectors?
von: Sandhan, Jivnesh, et al.
Veröffentlicht: (2026)
von: Sandhan, Jivnesh, et al.
Veröffentlicht: (2026)
STAR: Boosting Low-Resource Information Extraction by Structure-to-Text Data Generation with Large Language Models
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2023)
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2023)
HauntAttack: When Attack Follows Reasoning as a Shadow
von: Ma, Jingyuan, et al.
Veröffentlicht: (2025)
von: Ma, Jingyuan, et al.
Veröffentlicht: (2025)
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?
von: Wang, Leyao, et al.
Veröffentlicht: (2026)
von: Wang, Leyao, et al.
Veröffentlicht: (2026)
When Can We Trust LLM Graders? Calibrating Confidence for Automated Assessment
von: Ferrer, Robinson, et al.
Veröffentlicht: (2026)
von: Ferrer, Robinson, et al.
Veröffentlicht: (2026)
Improving Event Definition Following For Zero-Shot Event Detection
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
Towards Harmonized Uncertainty Estimation for Large Language Models
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
How Far are LLMs from Being Our Digital Twins? A Benchmark for Persona-Based Behavior Chain Simulation
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
Evaluating Cultural and Social Awareness of LLM Web Agents
von: Qiu, Haoyi, et al.
Veröffentlicht: (2024)
von: Qiu, Haoyi, et al.
Veröffentlicht: (2024)
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
von: Li, Zheng, et al.
Veröffentlicht: (2025)
von: Li, Zheng, et al.
Veröffentlicht: (2025)
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
von: Atil, Berk, et al.
Veröffentlicht: (2025)
von: Atil, Berk, et al.
Veröffentlicht: (2025)
Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM Sycophancy
von: Feng, Zhaoxin, et al.
Veröffentlicht: (2026)
von: Feng, Zhaoxin, et al.
Veröffentlicht: (2026)
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
von: Schroeder, Kayla, et al.
Veröffentlicht: (2024)
von: Schroeder, Kayla, et al.
Veröffentlicht: (2024)
AMRFact: Enhancing Summarization Factuality Evaluation with AMR-Driven Negative Samples Generation
von: Qiu, Haoyi, et al.
Veröffentlicht: (2023)
von: Qiu, Haoyi, et al.
Veröffentlicht: (2023)
Taking a Deep Breath: Enhancing Language Modeling of Large Language Models with Sentinel Tokens
von: Luo, Weiyao, et al.
Veröffentlicht: (2024)
von: Luo, Weiyao, et al.
Veröffentlicht: (2024)
GenEARL: A Training-Free Generative Framework for Multimodal Event Argument Role Labeling
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
Evaluating Human Alignment and Model Faithfulness of LLM Rationale
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2024)
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2024)
A Probabilistic Inference Scaling Theory for LLM Self-Correction
von: Yang, Zhe, et al.
Veröffentlicht: (2025)
von: Yang, Zhe, et al.
Veröffentlicht: (2025)
Can We Trust a Black-box LLM? LLM Untrustworthy Boundary Detection via Bias-Diffusion and Multi-Agent Reinforcement Learning
von: Zhou, Xiaotian, et al.
Veröffentlicht: (2026)
von: Zhou, Xiaotian, et al.
Veröffentlicht: (2026)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
von: Badawi, Abeer, et al.
Veröffentlicht: (2025)
von: Badawi, Abeer, et al.
Veröffentlicht: (2025)
In Agents We Trust, but Who Do Agents Trust? Latent Source Preferences Steer LLM Generations
von: Khan, Mohammad Aflah, et al.
Veröffentlicht: (2026)
von: Khan, Mohammad Aflah, et al.
Veröffentlicht: (2026)
Beyond Single Frames: Can LMMs Comprehend Temporal and Contextual Narratives in Image Sequences?
von: Wang, Xiaochen, et al.
Veröffentlicht: (2025)
von: Wang, Xiaochen, et al.
Veröffentlicht: (2025)
CoLT: Reasoning with Chain of Latent Tool Calls
von: Zhu, Fangwei, et al.
Veröffentlicht: (2026)
von: Zhu, Fangwei, et al.
Veröffentlicht: (2026)
Tutorial Proposal: Speculative Decoding for Efficient LLM Inference
von: Xia, Heming, et al.
Veröffentlicht: (2025)
von: Xia, Heming, et al.
Veröffentlicht: (2025)
LLM-based Human Simulations Have Not Yet Been Reliable
von: Wang, Qian, et al.
Veröffentlicht: (2025)
von: Wang, Qian, et al.
Veröffentlicht: (2025)
Self-Routing RAG: Binding Selective Retrieval with Knowledge Verbalization
von: Wu, Di, et al.
Veröffentlicht: (2025)
von: Wu, Di, et al.
Veröffentlicht: (2025)
SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration
von: Xia, Heming, et al.
Veröffentlicht: (2024)
von: Xia, Heming, et al.
Veröffentlicht: (2024)
Can We Trust the Performance Evaluation of Uncertainty Estimation Methods in Text Summarization?
von: He, Jianfeng, et al.
Veröffentlicht: (2024)
von: He, Jianfeng, et al.
Veröffentlicht: (2024)
Are We There Yet? Revealing the Risks of Utilizing Large Language Models in Scholarly Peer Review
von: Ye, Rui, et al.
Veröffentlicht: (2024)
von: Ye, Rui, et al.
Veröffentlicht: (2024)
Multimodal Cultural Safety: Evaluation Framework and Alignment Strategies
von: Qiu, Haoyi, et al.
Veröffentlicht: (2025)
von: Qiu, Haoyi, et al.
Veröffentlicht: (2025)
Energy-Regularized Sequential Model Editing on Hyperspheres
von: Liu, Qingyuan, et al.
Veröffentlicht: (2025)
von: Liu, Qingyuan, et al.
Veröffentlicht: (2025)
Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding
von: Xia, Heming, et al.
Veröffentlicht: (2024)
von: Xia, Heming, et al.
Veröffentlicht: (2024)
HistLens: Mapping Idea Change across Concepts and Corpora
von: Jing, Yi, et al.
Veröffentlicht: (2026)
von: Jing, Yi, et al.
Veröffentlicht: (2026)
LLM Compression: How Far Can We Go in Balancing Size and Performance?
von: Sk, Sahil, et al.
Veröffentlicht: (2025)
von: Sk, Sahil, et al.
Veröffentlicht: (2025)
RealVul: Can We Detect Vulnerabilities in Web Applications with LLM?
von: Cao, Di, et al.
Veröffentlicht: (2024)
von: Cao, Di, et al.
Veröffentlicht: (2024)
Enhancing LLM Character-Level Manipulation via Divide and Conquer
von: Xiong, Zhen, et al.
Veröffentlicht: (2025)
von: Xiong, Zhen, et al.
Veröffentlicht: (2025)
Can Large Language Models Always Solve Easy Problems if They Can Solve Harder Ones?
von: Yang, Zhe, et al.
Veröffentlicht: (2024)
von: Yang, Zhe, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Decoupling Task-Solving and Output Formatting in LLM Generation
von: Deng, Haikang, et al.
Veröffentlicht: (2025) -
Adaptable Logical Control for Large Language Models
von: Zhang, Honghua, et al.
Veröffentlicht: (2024) -
Can Large Multimodal Models Uncover Deep Semantics Behind Images?
von: Yang, Yixin, et al.
Veröffentlicht: (2024) -
Can We Trust LLM Detectors?
von: Sandhan, Jivnesh, et al.
Veröffentlicht: (2026) -
STAR: Boosting Low-Resource Information Extraction by Structure-to-Text Data Generation with Large Language Models
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2023)