SAFETY-J: Evaluating Safety with Critique
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yixiu, Zheng, Yuxiang, Xia, Shijie, Li, Jiajun, Tu, Yi, Song, Chaoling, Liu, Pengfei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?
by: Huang, Zhen, et al.
Published: (2024)
by: Huang, Zhen, et al.
Published: (2024)
DIVE: Diversified Iterative Self-Improvement
by: Qin, Yiwei, et al.
Published: (2025)
by: Qin, Yiwei, et al.
Published: (2025)
Evaluating Mathematical Reasoning Beyond Accuracy
by: Xia, Shijie, et al.
Published: (2024)
by: Xia, Shijie, et al.
Published: (2024)
The Critique of Critique
by: Sun, Shichao, et al.
Published: (2024)
by: Sun, Shichao, et al.
Published: (2024)
Halu-J: Critique-Based Hallucination Judge
by: Wang, Binjie, et al.
Published: (2024)
by: Wang, Binjie, et al.
Published: (2024)
O1 Replication Journey: A Strategic Progress Report -- Part 1
by: Qin, Yiwei, et al.
Published: (2024)
by: Qin, Yiwei, et al.
Published: (2024)
Dancing with Critiques: Enhancing LLM Reasoning with Stepwise Natural Language Self-Critique
by: Li, Yansi, et al.
Published: (2025)
by: Li, Yansi, et al.
Published: (2025)
OlympicArena Medal Ranks: Who Is the Most Intelligent AI So Far?
by: Huang, Zhen, et al.
Published: (2024)
by: Huang, Zhen, et al.
Published: (2024)
Don't Take the Premise for Granted: Evaluating the Premise Critique Ability of Large Language Models
by: Li, Jinzhe, et al.
Published: (2025)
by: Li, Jinzhe, et al.
Published: (2025)
CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?
by: Zong, Qing, et al.
Published: (2025)
by: Zong, Qing, et al.
Published: (2025)
CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation
by: Ke, Pei, et al.
Published: (2023)
by: Ke, Pei, et al.
Published: (2023)
Cross-Cultural Expert-Level Art Critique Evaluation with Vision-Language Models
by: Yu, Haorui, et al.
Published: (2026)
by: Yu, Haorui, et al.
Published: (2026)
Generative AI Act II: Test Time Scaling Drives Cognition Engineering
by: Xia, Shijie, et al.
Published: (2025)
by: Xia, Shijie, et al.
Published: (2025)
Assessing LLMs in Art Contexts: Critique Generation and Theory of Mind Evaluation
by: Arita, Takaya, et al.
Published: (2025)
by: Arita, Takaya, et al.
Published: (2025)
Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
by: Liu, Xuannan, et al.
Published: (2025)
by: Liu, Xuannan, et al.
Published: (2025)
LIMO: Less is More for Reasoning
by: Ye, Yixin, et al.
Published: (2025)
by: Ye, Yixin, et al.
Published: (2025)
Rectify Evaluation Preference: Improving LLMs' Critique on Math Reasoning via Perplexity-aware Reinforcement Learning
by: Tian, Changyuan, et al.
Published: (2025)
by: Tian, Changyuan, et al.
Published: (2025)
RealCritic: Towards Effectiveness-Driven Evaluation of Language Model Critiques
by: Tang, Zhengyang, et al.
Published: (2025)
by: Tang, Zhengyang, et al.
Published: (2025)
Self-Critique Guided Iterative Reasoning for Multi-hop Question Answering
by: Chu, Zheng, et al.
Published: (2025)
by: Chu, Zheng, et al.
Published: (2025)
Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States
by: Xiao, Yang, et al.
Published: (2025)
by: Xiao, Yang, et al.
Published: (2025)
The Lighthouse of Language: Enhancing LLM Agents via Critique-Guided Improvement
by: Yang, Ruihan, et al.
Published: (2025)
by: Yang, Ruihan, et al.
Published: (2025)
Improving Model Factuality with Fine-grained Critique-based Evaluator
by: Xie, Yiqing, et al.
Published: (2024)
by: Xie, Yiqing, et al.
Published: (2024)
LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing
by: Du, Jiangshu, et al.
Published: (2024)
by: Du, Jiangshu, et al.
Published: (2024)
Critique-Coder: Enhancing Coder Models by Critique Reinforcement Learning
by: Ruan, Chi, et al.
Published: (2025)
by: Ruan, Chi, et al.
Published: (2025)
SparseRM: A Lightweight Preference Modeling with Sparse Autoencoder
by: Liu, Dengcan, et al.
Published: (2025)
by: Liu, Dengcan, et al.
Published: (2025)
How Far Are LLMs from Believable AI? A Benchmark for Evaluating the Believability of Human Behavior Simulation
by: Xiao, Yang, et al.
Published: (2023)
by: Xiao, Yang, et al.
Published: (2023)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
by: Liu, Yixin, et al.
Published: (2025)
by: Liu, Yixin, et al.
Published: (2025)
FRoG: Evaluating Fuzzy Reasoning of Generalized Quantifiers in Large Language Models
by: Li, Yiyuan, et al.
Published: (2024)
by: Li, Yiyuan, et al.
Published: (2024)
Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate
by: Wang, Yubo, et al.
Published: (2025)
by: Wang, Yubo, et al.
Published: (2025)
Digital Socrates: Evaluating LLMs through Explanation Critiques
by: Gu, Yuling, et al.
Published: (2023)
by: Gu, Yuling, et al.
Published: (2023)
Critique-RL: Training Language Models for Critiquing through Two-Stage Reinforcement Learning
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
A Unified Representation Underlying the Judgment of Large Language Models
by: Lu, Yi-Long, et al.
Published: (2025)
by: Lu, Yi-Long, et al.
Published: (2025)
SafetyBench: Evaluating the Safety of Large Language Models
by: Zhang, Zhexin, et al.
Published: (2023)
by: Zhang, Zhexin, et al.
Published: (2023)
ChiMed 2.0: Advancing Chinese Medical Dataset in Facilitating Large Language Modeling
by: Tian, Yuanhe, et al.
Published: (2025)
by: Tian, Yuanhe, et al.
Published: (2025)
LoPA: Scaling dLLM Inference via Lookahead Parallel Decoding
by: Xu, Chenkai, et al.
Published: (2025)
by: Xu, Chenkai, et al.
Published: (2025)
Auxiliary Metrics Help Decoding Skill Neurons in the Wild
by: Zhao, Yixiu, et al.
Published: (2025)
by: Zhao, Yixiu, et al.
Published: (2025)
MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critique
by: Zeng, Gailun, et al.
Published: (2025)
by: Zeng, Gailun, et al.
Published: (2025)
Paradox of De-identification: A Critique of HIPAA Safe Harbour in the Age of LLMs
by: Jiang, Lavender Y., et al.
Published: (2026)
by: Jiang, Lavender Y., et al.
Published: (2026)
MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety
by: Song, Jialin, et al.
Published: (2026)
by: Song, Jialin, et al.
Published: (2026)
Training Language Model to Critique for Better Refinement
by: Yu, Tianshu, et al.
Published: (2025)
by: Yu, Tianshu, et al.
Published: (2025)
Similar Items
-
O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?
by: Huang, Zhen, et al.
Published: (2024) -
DIVE: Diversified Iterative Self-Improvement
by: Qin, Yiwei, et al.
Published: (2025) -
Evaluating Mathematical Reasoning Beyond Accuracy
by: Xia, Shijie, et al.
Published: (2024) -
The Critique of Critique
by: Sun, Shichao, et al.
Published: (2024) -
Halu-J: Critique-Based Hallucination Judge
by: Wang, Binjie, et al.
Published: (2024)