Weak-to-Strong Generalization beyond Accuracy: a Pilot Study in Safety, Toxicity, and Legal Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Ye, Ruimeng, Xiao, Yang, Hui, Bo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Weak-to-Strong Generalization with Failure Trajectories: A Tree-based Approach to Elicit Optimal Policy in Strong Models
por: Ye, Ruimeng, et al.
Publicado: (2025)
por: Ye, Ruimeng, et al.
Publicado: (2025)
Theoretical Analysis of Weak-to-Strong Generalization
por: Lang, Hunter, et al.
Publicado: (2024)
por: Lang, Hunter, et al.
Publicado: (2024)
How Do Latent Reasoning Methods Perform Under Weak and Strong Supervision?
por: Cui, Yingqian, et al.
Publicado: (2026)
por: Cui, Yingqian, et al.
Publicado: (2026)
Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
por: Uzunoglu, Arda, et al.
Publicado: (2026)
por: Uzunoglu, Arda, et al.
Publicado: (2026)
Reasoning's Razor: Reasoning Improves Accuracy but Can Hurt Recall at Critical Operating Points in Safety and Hallucination Detection
por: Chegini, Atoosa, et al.
Publicado: (2025)
por: Chegini, Atoosa, et al.
Publicado: (2025)
Bayesian WeakS-to-Strong from Text Classification to Generation
por: Cui, Ziyun, et al.
Publicado: (2024)
por: Cui, Ziyun, et al.
Publicado: (2024)
EnsemW2S: Enhancing Weak-to-Strong Generalization with Large Language Model Ensembles
por: Agrawal, Aakriti, et al.
Publicado: (2025)
por: Agrawal, Aakriti, et al.
Publicado: (2025)
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
por: Liang, Xiao, et al.
Publicado: (2025)
por: Liang, Xiao, et al.
Publicado: (2025)
Lexical Hints of Accuracy in LLM Reasoning Chains
por: Vanhoyweghen, Arne, et al.
Publicado: (2025)
por: Vanhoyweghen, Arne, et al.
Publicado: (2025)
Scaling Reasoning Hop Exposes Weaknesses: Demystifying and Improving Hop Generalization in Large Language Models
por: Li, Zhaoyi, et al.
Publicado: (2026)
por: Li, Zhaoyi, et al.
Publicado: (2026)
DC-W2S: Dual-Consensus Weak-to-Strong Training for Reliable Process Reward Modeling in Biological Reasoning
por: Chan, Chi-Min, et al.
Publicado: (2026)
por: Chan, Chi-Min, et al.
Publicado: (2026)
Knowledge Graph-Assisted LLM Post-Training for Enhanced Legal Reasoning
por: Song, Dezhao, et al.
Publicado: (2026)
por: Song, Dezhao, et al.
Publicado: (2026)
Enhancing LLM Factual Accuracy with RAG to Counter Hallucinations: A Case Study on Domain-Specific Queries in Private Knowledge-Bases
por: Li, Jiarui, et al.
Publicado: (2024)
por: Li, Jiarui, et al.
Publicado: (2024)
Weak-to-Strong Elicitation via Mismatched Wrong Drafts
por: Deng, Wei
Publicado: (2026)
por: Deng, Wei
Publicado: (2026)
When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning
por: Barale, Claire, et al.
Publicado: (2025)
por: Barale, Claire, et al.
Publicado: (2025)
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
por: Le, Khoi, et al.
Publicado: (2026)
por: Le, Khoi, et al.
Publicado: (2026)
Weak-to-Strong Compositional Learning from Generative Models for Language-based Object Detection
por: Park, Kwanyong, et al.
Publicado: (2024)
por: Park, Kwanyong, et al.
Publicado: (2024)
Think-Augmented Function Calling: Improving LLM Parameter Accuracy Through Embedded Reasoning
por: Wei, Lei, et al.
Publicado: (2026)
por: Wei, Lei, et al.
Publicado: (2026)
On Giant's Shoulders: Effortless Weak to Strong by Dynamic Logits Fusion
por: Fan, Chenghao, et al.
Publicado: (2024)
por: Fan, Chenghao, et al.
Publicado: (2024)
The Amazing Agent Race: Strong Tool Users, Weak Navigators
por: Kim, Zae Myung, et al.
Publicado: (2026)
por: Kim, Zae Myung, et al.
Publicado: (2026)
NyayaMind- A Framework for Transparent Legal Reasoning and Judgment Prediction in the Indian Legal System
por: Shukla, Parjanya Aditya, et al.
Publicado: (2026)
por: Shukla, Parjanya Aditya, et al.
Publicado: (2026)
From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning
por: Huang, Yuzhen, et al.
Publicado: (2025)
por: Huang, Yuzhen, et al.
Publicado: (2025)
Safety Reasoning with Guidelines
por: Wang, Haoyu, et al.
Publicado: (2025)
por: Wang, Haoyu, et al.
Publicado: (2025)
AppellateGen: A Benchmark for Appellate Legal Judgment Generation
por: Yang, Hongkun, et al.
Publicado: (2026)
por: Yang, Hongkun, et al.
Publicado: (2026)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
por: Yang, Wenkai, et al.
Publicado: (2026)
por: Yang, Wenkai, et al.
Publicado: (2026)
Measuring and Mitigating Toxicity in Large Language Models: A Comprehensive Replication Study
por: Surana, Mokshit, et al.
Publicado: (2026)
por: Surana, Mokshit, et al.
Publicado: (2026)
Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models
por: Zhou, Zhanhui, et al.
Publicado: (2024)
por: Zhou, Zhanhui, et al.
Publicado: (2024)
MATATA: Weakly Supervised End-to-End MAthematical Tool-Augmented Reasoning for Tabular Applications
por: Vinayagame, Vishnou, et al.
Publicado: (2024)
por: Vinayagame, Vishnou, et al.
Publicado: (2024)
Offline Exploration-Aware Fine-Tuning for Long-Chain Mathematical Reasoning
por: Mu, Yongyu, et al.
Publicado: (2026)
por: Mu, Yongyu, et al.
Publicado: (2026)
IL-TUR: Benchmark for Indian Legal Text Understanding and Reasoning
por: Joshi, Abhinav, et al.
Publicado: (2024)
por: Joshi, Abhinav, et al.
Publicado: (2024)
Evaluating the Limits of Large Language Models in Multilingual Legal Reasoning
por: Ioannou, Antreas, et al.
Publicado: (2025)
por: Ioannou, Antreas, et al.
Publicado: (2025)
Zero-to-Strong Generalization: Eliciting Strong Capabilities of Large Language Models Iteratively without Gold Labels
por: Liu, Chaoqun, et al.
Publicado: (2024)
por: Liu, Chaoqun, et al.
Publicado: (2024)
LegalLens Shared Task 2024: Legal Violation Identification in Unstructured Text
por: Hagag, Ben, et al.
Publicado: (2024)
por: Hagag, Ben, et al.
Publicado: (2024)
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content
por: Stepanov, Ihor, et al.
Publicado: (2026)
por: Stepanov, Ihor, et al.
Publicado: (2026)
Dissecting Long-Chain-of-Thought Reasoning Models: An Empirical Study
por: Mu, Yongyu, et al.
Publicado: (2025)
por: Mu, Yongyu, et al.
Publicado: (2025)
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
por: Chen, Zixiang, et al.
Publicado: (2024)
por: Chen, Zixiang, et al.
Publicado: (2024)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
por: AlDahoul, Nouar, et al.
Publicado: (2025)
por: AlDahoul, Nouar, et al.
Publicado: (2025)
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis
por: Yang, Yushi, et al.
Publicado: (2024)
por: Yang, Yushi, et al.
Publicado: (2024)
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
por: Liu, Qihao, et al.
Publicado: (2025)
por: Liu, Qihao, et al.
Publicado: (2025)
ToVo: Toxicity Taxonomy via Voting
por: Luong, Tinh Son, et al.
Publicado: (2024)
por: Luong, Tinh Son, et al.
Publicado: (2024)
Ejemplares similares
-
Weak-to-Strong Generalization with Failure Trajectories: A Tree-based Approach to Elicit Optimal Policy in Strong Models
por: Ye, Ruimeng, et al.
Publicado: (2025) -
Theoretical Analysis of Weak-to-Strong Generalization
por: Lang, Hunter, et al.
Publicado: (2024) -
How Do Latent Reasoning Methods Perform Under Weak and Strong Supervision?
por: Cui, Yingqian, et al.
Publicado: (2026) -
Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
por: Uzunoglu, Arda, et al.
Publicado: (2026) -
Reasoning's Razor: Reasoning Improves Accuracy but Can Hurt Recall at Critical Operating Points in Safety and Hallucination Detection
por: Chegini, Atoosa, et al.
Publicado: (2025)