Salvato in:
| Autori principali: | Zheng, Chen, Sun, Ke, Zhou, Xun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2406.08657 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Balancing Enhancement, Harmlessness, and General Capabilities: Enhancing Conversational LLMs with Direct RLHF
di: Zheng, Chen, et al.
Pubblicazione: (2024)
di: Zheng, Chen, et al.
Pubblicazione: (2024)
Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
di: Zheng, Chen, et al.
Pubblicazione: (2025)
di: Zheng, Chen, et al.
Pubblicazione: (2025)
C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis
di: Luo, Miaosen, et al.
Pubblicazione: (2026)
di: Luo, Miaosen, et al.
Pubblicazione: (2026)
ICE-GRT: Instruction Context Enhancement by Generative Reinforcement based Transformers
di: Zheng, Chen, et al.
Pubblicazione: (2024)
di: Zheng, Chen, et al.
Pubblicazione: (2024)
Vuyko Mistral: Adapting LLMs for Low-Resource Dialectal Translation
di: Kyslyi, Roman, et al.
Pubblicazione: (2025)
di: Kyslyi, Roman, et al.
Pubblicazione: (2025)
Mistral-SPLADE: LLMs for better Learned Sparse Retrieval
di: Doshi, Meet, et al.
Pubblicazione: (2024)
di: Doshi, Meet, et al.
Pubblicazione: (2024)
Doc-V*:Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA
di: Zheng, Yuanlei, et al.
Pubblicazione: (2026)
di: Zheng, Yuanlei, et al.
Pubblicazione: (2026)
MAgICoRe: Multi-Agent, Iterative, Coarse-to-Fine Refinement for Reasoning
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2024)
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2024)
CFSP: An Efficient Structured Pruning Framework for LLMs with Coarse-to-Fine Activation Information
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
Taming Overconfidence in LLMs: Reward Calibration in RLHF
di: Leng, Jixuan, et al.
Pubblicazione: (2024)
di: Leng, Jixuan, et al.
Pubblicazione: (2024)
Reward-Robust RLHF in LLMs
di: Yan, Yuzi, et al.
Pubblicazione: (2024)
di: Yan, Yuzi, et al.
Pubblicazione: (2024)
Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging
di: Farn, Hua, et al.
Pubblicazione: (2024)
di: Farn, Hua, et al.
Pubblicazione: (2024)
CFMS: A Coarse-to-Fine Multimodal Synthesis Framework for Enhanced Tabular Reasoning
di: Huang, Qixian, et al.
Pubblicazione: (2026)
di: Huang, Qixian, et al.
Pubblicazione: (2026)
The Thinking Spectrum: An Empirical Study of Tunable Reasoning in LLMs through Model Merging
di: Lan, Xiaochong, et al.
Pubblicazione: (2025)
di: Lan, Xiaochong, et al.
Pubblicazione: (2025)
Joint Enhancement of Relational Reasoning for Long-Context LLMs
di: Chen, Zhirui, et al.
Pubblicazione: (2025)
di: Chen, Zhirui, et al.
Pubblicazione: (2025)
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
di: Ji, Jiaming, et al.
Pubblicazione: (2024)
di: Ji, Jiaming, et al.
Pubblicazione: (2024)
Continual SFT Matches Multimodal RLHF with Negative Supervision
di: Zhu, Ke, et al.
Pubblicazione: (2024)
di: Zhu, Ke, et al.
Pubblicazione: (2024)
ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model Merging
di: Yang, Junyao, et al.
Pubblicazione: (2026)
di: Yang, Junyao, et al.
Pubblicazione: (2026)
Linq-Embed-Mistral Technical Report
di: Choi, Chanyeol, et al.
Pubblicazione: (2024)
di: Choi, Chanyeol, et al.
Pubblicazione: (2024)
MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs
di: Zhao, Guojiang, et al.
Pubblicazione: (2025)
di: Zhao, Guojiang, et al.
Pubblicazione: (2025)
Advancing Translation Preference Modeling with RLHF: A Step Towards Cost-Effective Solution
di: Xu, Nuo, et al.
Pubblicazione: (2024)
di: Xu, Nuo, et al.
Pubblicazione: (2024)
From Bytes to Borsch: Fine-Tuning Gemma and Mistral for the Ukrainian Language Representation
di: Kiulian, Artur, et al.
Pubblicazione: (2024)
di: Kiulian, Artur, et al.
Pubblicazione: (2024)
TinyThinker: Distilling Reasoning through Coarse-to-Fine Knowledge Internalization with Self-Reflection
di: Piao, Shengmin, et al.
Pubblicazione: (2024)
di: Piao, Shengmin, et al.
Pubblicazione: (2024)
Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks
di: Wang, Zheng, et al.
Pubblicazione: (2024)
di: Wang, Zheng, et al.
Pubblicazione: (2024)
Removing RLHF Protections in GPT-4 via Fine-Tuning
di: Zhan, Qiusi, et al.
Pubblicazione: (2023)
di: Zhan, Qiusi, et al.
Pubblicazione: (2023)
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
di: Yu, Tianyu, et al.
Pubblicazione: (2023)
di: Yu, Tianyu, et al.
Pubblicazione: (2023)
Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
di: Chen, Xiaoshu, et al.
Pubblicazione: (2025)
di: Chen, Xiaoshu, et al.
Pubblicazione: (2025)
Harnessing RLHF for Robust Unanswerability Recognition and Trustworthy Response Generation in LLMs
di: Lin, Shuyuan, et al.
Pubblicazione: (2025)
di: Lin, Shuyuan, et al.
Pubblicazione: (2025)
RLHF Workflow: From Reward Modeling to Online RLHF
di: Dong, Hanze, et al.
Pubblicazione: (2024)
di: Dong, Hanze, et al.
Pubblicazione: (2024)
Towards Federated RLHF with Aggregated Client Preference for LLMs
di: Wu, Feijie, et al.
Pubblicazione: (2024)
di: Wu, Feijie, et al.
Pubblicazione: (2024)
Search-R2: Enhancing Search-Integrated Reasoning via Actor-Refiner Collaboration
di: He, Bowei, et al.
Pubblicazione: (2026)
di: He, Bowei, et al.
Pubblicazione: (2026)
Coarse-to-Fine Personalized LLM Impressions for Streamlined Radiology Reports
di: Sun, Chengbo, et al.
Pubblicazione: (2025)
di: Sun, Chengbo, et al.
Pubblicazione: (2025)
Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model
di: Yin, Yueqin, et al.
Pubblicazione: (2025)
di: Yin, Yueqin, et al.
Pubblicazione: (2025)
Actor Identification in Discourse: A Challenge for LLMs?
di: Barić, Ana, et al.
Pubblicazione: (2024)
di: Barić, Ana, et al.
Pubblicazione: (2024)
Multilevel Analysis of Cryptocurrency News using RAG Approach with Fine-Tuned Mistral Large Language Model
di: Pavlyshenko, Bohdan M.
Pubblicazione: (2025)
di: Pavlyshenko, Bohdan M.
Pubblicazione: (2025)
Reasoning Pattern Alignment Merging for Adaptive Reasoning
di: Zhong, Zhaofeng, et al.
Pubblicazione: (2026)
di: Zhong, Zhaofeng, et al.
Pubblicazione: (2026)
Effective Distillation of Table-based Reasoning Ability from LLMs
di: Yang, Bohao, et al.
Pubblicazione: (2023)
di: Yang, Bohao, et al.
Pubblicazione: (2023)
Fovea Transformer: Efficient Long-Context Modeling with Structured Fine-to-Coarse Attention
di: He, Ziwei, et al.
Pubblicazione: (2023)
di: He, Ziwei, et al.
Pubblicazione: (2023)
Document-Level Tabular Numerical Cross-Checking: A Coarse-to-Fine Approach
di: Pang, Chaoxu, et al.
Pubblicazione: (2025)
di: Pang, Chaoxu, et al.
Pubblicazione: (2025)
Adaptive Graph Refinement and Label Propagation with LLMs for Cost-Effective Entity Resolution
di: Wang, Hongtao, et al.
Pubblicazione: (2026)
di: Wang, Hongtao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Balancing Enhancement, Harmlessness, and General Capabilities: Enhancing Conversational LLMs with Direct RLHF
di: Zheng, Chen, et al.
Pubblicazione: (2024) -
Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
di: Zheng, Chen, et al.
Pubblicazione: (2025) -
C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis
di: Luo, Miaosen, et al.
Pubblicazione: (2026) -
ICE-GRT: Instruction Context Enhancement by Generative Reinforcement based Transformers
di: Zheng, Chen, et al.
Pubblicazione: (2024) -
Vuyko Mistral: Adapting LLMs for Low-Resource Dialectal Translation
di: Kyslyi, Roman, et al.
Pubblicazione: (2025)