Learning or Self-aligning? Rethinking Instruction Fine-tuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ren, Mengjie, Cao, Boxi, Lin, Hongyu, Liu, Cao, Han, Xianpei, Zeng, Ke, Wan, Guanglu, Cai, Xunliang, Sun, Le |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Life Cycle of Knowledge in Big Language Models: A Survey
von: Cao, Boxi, et al.
Veröffentlicht: (2023)
von: Cao, Boxi, et al.
Veröffentlicht: (2023)
StructEval: Deepen and Broaden Large Language Model Assessment via Structured Evaluation
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
Not All Contexts Are Equal: Teaching LLMs Credibility-aware Generation
von: Pan, Ruotong, et al.
Veröffentlicht: (2024)
von: Pan, Ruotong, et al.
Veröffentlicht: (2024)
Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards
von: Ren, Mengjie, et al.
Veröffentlicht: (2026)
von: Ren, Mengjie, et al.
Veröffentlicht: (2026)
Multi-Facet Counterfactual Learning for Content Quality Evaluation
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2024)
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2024)
Towards Universal Dense Blocking for Entity Resolution
von: Wang, Tianshu, et al.
Veröffentlicht: (2024)
von: Wang, Tianshu, et al.
Veröffentlicht: (2024)
Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2026)
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2026)
LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents
von: Peng, Xiaoxuan, et al.
Veröffentlicht: (2026)
von: Peng, Xiaoxuan, et al.
Veröffentlicht: (2026)
When Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning Models
von: Mao, Yingzhi, et al.
Veröffentlicht: (2025)
von: Mao, Yingzhi, et al.
Veröffentlicht: (2025)
Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2024)
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2024)
Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards
von: Ma, Zhengzhao, et al.
Veröffentlicht: (2026)
von: Ma, Zhengzhao, et al.
Veröffentlicht: (2026)
REInstruct: Building Instruction Data from Unlabeled Corpus
von: Chen, Shu, et al.
Veröffentlicht: (2024)
von: Chen, Shu, et al.
Veröffentlicht: (2024)
URL: Universal Referential Knowledge Linking via Task-instructed Representation Compression
von: Li, Zhuoqun, et al.
Veröffentlicht: (2024)
von: Li, Zhuoqun, et al.
Veröffentlicht: (2024)
Retentive or Forgetful? Diving into the Knowledge Memorizing Mechanism of Language Models
von: Cao, Boxi, et al.
Veröffentlicht: (2023)
von: Cao, Boxi, et al.
Veröffentlicht: (2023)
Towards Scalable Automated Alignment of LLMs: A Survey
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
von: Yuan, Qianhao, et al.
Veröffentlicht: (2026)
von: Yuan, Qianhao, et al.
Veröffentlicht: (2026)
AI for social science and social science of AI: A Survey
von: Xu, Ruoxi, et al.
Veröffentlicht: (2024)
von: Xu, Ruoxi, et al.
Veröffentlicht: (2024)
Critic-CoT: Boosting the reasoning abilities of large language model via Chain-of-thoughts Critic
von: Zheng, Xin, et al.
Veröffentlicht: (2024)
von: Zheng, Xin, et al.
Veröffentlicht: (2024)
Memorizing is Not Enough: Deep Knowledge Injection Through Reasoning
von: Xu, Ruoxi, et al.
Veröffentlicht: (2025)
von: Xu, Ruoxi, et al.
Veröffentlicht: (2025)
The Rise and Down of Babel Tower: Investigating the Evolution Process of Multilingual Code Large Language Model
von: Chen, Jiawei, et al.
Veröffentlicht: (2024)
von: Chen, Jiawei, et al.
Veröffentlicht: (2024)
Academically intelligent LLMs are not necessarily socially intelligent
von: Xu, Ruoxi, et al.
Veröffentlicht: (2024)
von: Xu, Ruoxi, et al.
Veröffentlicht: (2024)
Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering
von: Guan, Xinyan, et al.
Veröffentlicht: (2024)
von: Guan, Xinyan, et al.
Veröffentlicht: (2024)
ScaleBox: Enabling High-Fidelity and Scalable Code Verification for Large Language Models
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2026)
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2026)
On-Policy Self-Alignment with Fine-grained Knowledge Feedback for Hallucination Mitigation
von: Wen, Xueru, et al.
Veröffentlicht: (2024)
von: Wen, Xueru, et al.
Veröffentlicht: (2024)
All Languages Matter: Understanding and Mitigating Language Bias in Multilingual RAG
von: Wang, Dan, et al.
Veröffentlicht: (2026)
von: Wang, Dan, et al.
Veröffentlicht: (2026)
Tackling Length Inflation Without Trade-offs: Group Relative Reward Rescaling for Reinforcement Learning
von: Li, Zichao, et al.
Veröffentlicht: (2026)
von: Li, Zichao, et al.
Veröffentlicht: (2026)
Reflecting Twice before Speaking with Empathy: Self-Reflective Alternating Inference for Empathy-Aware End-to-End Spoken Dialogue
von: Jia, Yuhang, et al.
Veröffentlicht: (2026)
von: Jia, Yuhang, et al.
Veröffentlicht: (2026)
Rule or Story, Which is a Better Commonsense Expression for Talking with Large Language Models?
von: Bian, Ning, et al.
Veröffentlicht: (2024)
von: Bian, Ning, et al.
Veröffentlicht: (2024)
Open Grounded Planning: Challenges and Benchmark Construction
von: Guo, Shiguang, et al.
Veröffentlicht: (2024)
von: Guo, Shiguang, et al.
Veröffentlicht: (2024)
Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models
von: Yan, Xinru, et al.
Veröffentlicht: (2026)
von: Yan, Xinru, et al.
Veröffentlicht: (2026)
Fine-tuning Done Right in Model Editing
von: Yang, Wanli, et al.
Veröffentlicht: (2025)
von: Yang, Wanli, et al.
Veröffentlicht: (2025)
DBCopilot: Natural Language Querying over Massive Databases via Schema Routing
von: Wang, Tianshu, et al.
Veröffentlicht: (2023)
von: Wang, Tianshu, et al.
Veröffentlicht: (2023)
Executing Natural Language-Described Algorithms with Large Language Models: An Investigation
von: Zheng, Xin, et al.
Veröffentlicht: (2024)
von: Zheng, Xin, et al.
Veröffentlicht: (2024)
Meta-Cognitive Analysis: Evaluating Declarative and Procedural Knowledge in Datasets and Large Language Models
von: Li, Zhuoqun, et al.
Veröffentlicht: (2024)
von: Li, Zhuoqun, et al.
Veröffentlicht: (2024)
Large Language Models Often Say One Thing and Do Another
von: Xu, Ruoxi, et al.
Veröffentlicht: (2025)
von: Xu, Ruoxi, et al.
Veröffentlicht: (2025)
Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO
von: Yang, Xin, et al.
Veröffentlicht: (2026)
von: Yang, Xin, et al.
Veröffentlicht: (2026)
Rethinking Reward Model Evaluation: Are We Barking up the Wrong Tree?
von: Wen, Xueru, et al.
Veröffentlicht: (2024)
von: Wen, Xueru, et al.
Veröffentlicht: (2024)
Match, Compare, or Select? An Investigation of Large Language Models for Entity Matching
von: Wang, Tianshu, et al.
Veröffentlicht: (2024)
von: Wang, Tianshu, et al.
Veröffentlicht: (2024)
Self-Steering Optimization: Autonomous Preference Optimization for Large Language Models
von: Xiang, Hao, et al.
Veröffentlicht: (2024)
von: Xiang, Hao, et al.
Veröffentlicht: (2024)
Spiral of Silence: How is Large Language Model Killing Information Retrieval? -- A Case Study on Open Domain Question Answering
von: Chen, Xiaoyang, et al.
Veröffentlicht: (2024)
von: Chen, Xiaoyang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Life Cycle of Knowledge in Big Language Models: A Survey
von: Cao, Boxi, et al.
Veröffentlicht: (2023) -
StructEval: Deepen and Broaden Large Language Model Assessment via Structured Evaluation
von: Cao, Boxi, et al.
Veröffentlicht: (2024) -
Not All Contexts Are Equal: Teaching LLMs Credibility-aware Generation
von: Pan, Ruotong, et al.
Veröffentlicht: (2024) -
Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards
von: Ren, Mengjie, et al.
Veröffentlicht: (2026) -
Multi-Facet Counterfactual Learning for Content Quality Evaluation
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2024)