Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
Fuente:
arXiv
Saved in:
| Main Authors: | Uzunoglu, Arda, Zhang, Alvin, Khashabi, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
by: Uzunoglu, Arda, et al.
Published: (2025)
by: Uzunoglu, Arda, et al.
Published: (2025)
Theoretical Analysis of Weak-to-Strong Generalization
by: Lang, Hunter, et al.
Published: (2024)
by: Lang, Hunter, et al.
Published: (2024)
WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment
by: Ou, Jiefu, et al.
Published: (2024)
by: Ou, Jiefu, et al.
Published: (2024)
When to Trust the Cheap Check: Weak and Strong Verification for Reasoning
by: Kiyani, Shayan, et al.
Published: (2026)
by: Kiyani, Shayan, et al.
Published: (2026)
Bayesian WeakS-to-Strong from Text Classification to Generation
by: Cui, Ziyun, et al.
Published: (2024)
by: Cui, Ziyun, et al.
Published: (2024)
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
by: Le, Khoi, et al.
Published: (2026)
by: Le, Khoi, et al.
Published: (2026)
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
by: Pawelczyk, Martin, et al.
Published: (2024)
by: Pawelczyk, Martin, et al.
Published: (2024)
EnsemW2S: Enhancing Weak-to-Strong Generalization with Large Language Model Ensembles
by: Agrawal, Aakriti, et al.
Published: (2025)
by: Agrawal, Aakriti, et al.
Published: (2025)
Weak-to-Strong Generalization beyond Accuracy: a Pilot Study in Safety, Toxicity, and Legal Reasoning
by: Ye, Ruimeng, et al.
Published: (2024)
by: Ye, Ruimeng, et al.
Published: (2024)
Weak-to-Strong Compositional Learning from Generative Models for Language-based Object Detection
by: Park, Kwanyong, et al.
Published: (2024)
by: Park, Kwanyong, et al.
Published: (2024)
Weak-to-Strong Elicitation via Mismatched Wrong Drafts
by: Deng, Wei
Published: (2026)
by: Deng, Wei
Published: (2026)
Trust Region On-Policy Distillation
by: Xing, Xingrun, et al.
Published: (2026)
by: Xing, Xingrun, et al.
Published: (2026)
TRE: Encouraging Exploration in the Trust Region
by: Huang, Chao, et al.
Published: (2026)
by: Huang, Chao, et al.
Published: (2026)
The Amazing Agent Race: Strong Tool Users, Weak Navigators
by: Kim, Zae Myung, et al.
Published: (2026)
by: Kim, Zae Myung, et al.
Published: (2026)
On Giant's Shoulders: Effortless Weak to Strong by Dynamic Logits Fusion
by: Fan, Chenghao, et al.
Published: (2024)
by: Fan, Chenghao, et al.
Published: (2024)
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
by: Kang, Hao, et al.
Published: (2024)
by: Kang, Hao, et al.
Published: (2024)
Do pretrained Transformers Learn In-Context by Gradient Descent?
by: Shen, Lingfeng, et al.
Published: (2023)
by: Shen, Lingfeng, et al.
Published: (2023)
Leveraging Large Language Models for Structure Learning in Prompted Weak Supervision
by: Su, Jinyan, et al.
Published: (2024)
by: Su, Jinyan, et al.
Published: (2024)
Disentangling Latent Shifts of In-Context Learning with Weak Supervision
by: Jukić, Josip, et al.
Published: (2024)
by: Jukić, Josip, et al.
Published: (2024)
FESTA: Functionally Equivalent Sampling for Trust Assessment of Multimodal LLMs
by: Bhattacharya, Debarpan, et al.
Published: (2025)
by: Bhattacharya, Debarpan, et al.
Published: (2025)
When Should a Language Model Trust Itself? Same-Model Self-Verification as a Conditional Confidence Signal
by: Phalod, Aditya Ajay
Published: (2026)
by: Phalod, Aditya Ajay
Published: (2026)
Compressed Models are NOT Trust-equivalent to Their Large Counterparts
by: Rai, Rohit Raj, et al.
Published: (2025)
by: Rai, Rohit Raj, et al.
Published: (2025)
How Do Latent Reasoning Methods Perform Under Weak and Strong Supervision?
by: Cui, Yingqian, et al.
Published: (2026)
by: Cui, Yingqian, et al.
Published: (2026)
Trust The Typical
by: Ganguly, Debargha, et al.
Published: (2026)
by: Ganguly, Debargha, et al.
Published: (2026)
Alice: Proactive Learning with Teacher's Demonstrations for Weak-to-Strong Generalization
by: Wu, Shujin, et al.
Published: (2025)
by: Wu, Shujin, et al.
Published: (2025)
Benchmarking Procedural Language Understanding for Low-Resource Languages: A Case Study on Turkish
by: Uzunoglu, Arda, et al.
Published: (2023)
by: Uzunoglu, Arda, et al.
Published: (2023)
TRAM: Bridging Trust Regions and Sharpness Aware Minimization
by: Sherborne, Tom, et al.
Published: (2023)
by: Sherborne, Tom, et al.
Published: (2023)
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
by: Chen, Zixiang, et al.
Published: (2024)
by: Chen, Zixiang, et al.
Published: (2024)
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models
by: Yi, Hao, et al.
Published: (2024)
by: Yi, Hao, et al.
Published: (2024)
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
by: Liang, Xiao, et al.
Published: (2025)
by: Liang, Xiao, et al.
Published: (2025)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
by: Jung, Jaehun, et al.
Published: (2024)
by: Jung, Jaehun, et al.
Published: (2024)
IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
by: Mishra, Aayush, et al.
Published: (2025)
by: Mishra, Aayush, et al.
Published: (2025)
ToolExpander: Extending the Frontiers of Tool-Using Reinforcement Learning to Weak LLMs
by: Chen, Fu, et al.
Published: (2025)
by: Chen, Fu, et al.
Published: (2025)
Jointly Reinforcing Diversity and Quality in Language Model Generations
by: Li, Tianjian, et al.
Published: (2025)
by: Li, Tianjian, et al.
Published: (2025)
OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling
by: Lou, Yuxuan, et al.
Published: (2026)
by: Lou, Yuxuan, et al.
Published: (2026)
FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain
by: Hu, Tiansheng, et al.
Published: (2025)
by: Hu, Tiansheng, et al.
Published: (2025)
Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models
by: Zhou, Zhanhui, et al.
Published: (2024)
by: Zhou, Zhanhui, et al.
Published: (2024)
TrustLDM: Benchmarking Trustworthiness in Language Diffusion Models
by: Mo, Yichuan, et al.
Published: (2026)
by: Mo, Yichuan, et al.
Published: (2026)
CrEst: Credibility Estimation for Contexts in LLMs via Weak Supervision
by: Adila, Dyah, et al.
Published: (2025)
by: Adila, Dyah, et al.
Published: (2025)
STENCIL: Submodular Mutual Information Based Weak Supervision for Cold-Start Active Learning
by: Beck, Nathan, et al.
Published: (2024)
by: Beck, Nathan, et al.
Published: (2024)
Similar Items
-
The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
by: Uzunoglu, Arda, et al.
Published: (2025) -
Theoretical Analysis of Weak-to-Strong Generalization
by: Lang, Hunter, et al.
Published: (2024) -
WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment
by: Ou, Jiefu, et al.
Published: (2024) -
When to Trust the Cheap Check: Weak and Strong Verification for Reasoning
by: Kiyani, Shayan, et al.
Published: (2026) -
Bayesian WeakS-to-Strong from Text Classification to Generation
by: Cui, Ziyun, et al.
Published: (2024)