Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning
Fuente:
arXiv
Salvato in:
| Autori principali: | Patnaik, Sohan, Aggarwal, Milan, Bhatia, Sumit, Krishnamurthy, Balaji |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
It Helps to Take a Second Opinion: Teaching Smaller LLMs to Deliberate Mutually via Selective Rationale Optimisation
di: Patnaik, Sohan, et al.
Pubblicazione: (2025)
di: Patnaik, Sohan, et al.
Pubblicazione: (2025)
CABINET: Content Relevance based Noise Reduction for Table Question Answering
di: Patnaik, Sohan, et al.
Pubblicazione: (2024)
di: Patnaik, Sohan, et al.
Pubblicazione: (2024)
AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language Models
di: Patnaik, Sohan, et al.
Pubblicazione: (2025)
di: Patnaik, Sohan, et al.
Pubblicazione: (2025)
Dialogue Agents 101: A Beginner's Guide to Critical Ingredients for Designing Effective Conversational Systems
di: Kumar, Shivani, et al.
Pubblicazione: (2023)
di: Kumar, Shivani, et al.
Pubblicazione: (2023)
AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent
di: Anand, Neeraj, et al.
Pubblicazione: (2025)
di: Anand, Neeraj, et al.
Pubblicazione: (2025)
Towards Efficient CoT Distillation: Self-Guided Rationale Selector for Better Performance with Fewer Rationales
di: Yan, Jianzhi, et al.
Pubblicazione: (2025)
di: Yan, Jianzhi, et al.
Pubblicazione: (2025)
Architecture, Not Scale: Circuit Localization in Large Language Models
di: Venkatesh, Sohan
Pubblicazione: (2026)
di: Venkatesh, Sohan
Pubblicazione: (2026)
Teaching Human Behavior Improves Content Understanding Abilities Of LLMs
di: Singh, Somesh, et al.
Pubblicazione: (2024)
di: Singh, Somesh, et al.
Pubblicazione: (2024)
Are Rationales Necessary and Sufficient? Tuning LLMs for Explainable Misinformation Detection
di: Wang, Bing, et al.
Pubblicazione: (2026)
di: Wang, Bing, et al.
Pubblicazione: (2026)
TeaMs-RL: Teaching LLMs to Generate Better Instruction Datasets via Reinforcement Learning
di: Gu, Shangding, et al.
Pubblicazione: (2024)
di: Gu, Shangding, et al.
Pubblicazione: (2024)
Fine-Tuning and Prompt Optimization: Two Great Steps that Work Better Together
di: Soylu, Dilara, et al.
Pubblicazione: (2024)
di: Soylu, Dilara, et al.
Pubblicazione: (2024)
TriSum: Learning Summarization Ability from Large Language Models with Structured Rationale
di: Jiang, Pengcheng, et al.
Pubblicazione: (2024)
di: Jiang, Pengcheng, et al.
Pubblicazione: (2024)
Think Together and Work Better: Combining Humans' and LLMs' Think-Aloud Outcomes for Effective Text Evaluation
di: Chu, SeongYeub, et al.
Pubblicazione: (2024)
di: Chu, SeongYeub, et al.
Pubblicazione: (2024)
SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales
di: Xu, Tianyang, et al.
Pubblicazione: (2024)
di: Xu, Tianyang, et al.
Pubblicazione: (2024)
Better Together: Quantifying the Benefits of AI-Assisted Recruitment
di: Aka, Ada, et al.
Pubblicazione: (2025)
di: Aka, Ada, et al.
Pubblicazione: (2025)
Can LLMs Learn by Teaching for Better Reasoning? A Preliminary Study
di: Ning, Xuefei, et al.
Pubblicazione: (2024)
di: Ning, Xuefei, et al.
Pubblicazione: (2024)
SMART: Submodular Data Mixture Strategy for Instruction Tuning
di: Renduchintala, H S V N S Kowndinya, et al.
Pubblicazione: (2024)
di: Renduchintala, H S V N S Kowndinya, et al.
Pubblicazione: (2024)
Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability
di: Zhu, Chiwei, et al.
Pubblicazione: (2025)
di: Zhu, Chiwei, et al.
Pubblicazione: (2025)
Faster and Better LLMs via Latency-Aware Test-Time Scaling
di: Wang, Zili, et al.
Pubblicazione: (2025)
di: Wang, Zili, et al.
Pubblicazione: (2025)
LLMs Can Teach Themselves to Better Predict the Future
di: Turtel, Benjamin, et al.
Pubblicazione: (2025)
di: Turtel, Benjamin, et al.
Pubblicazione: (2025)
On the Effect of Instruction Tuning Loss on Generalization
di: Chatterjee, Anwoy, et al.
Pubblicazione: (2025)
di: Chatterjee, Anwoy, et al.
Pubblicazione: (2025)
Negative Before Positive: Asymmetric Valence Processing in Large Language Models
di: Venkatesh, Sohan
Pubblicazione: (2026)
di: Venkatesh, Sohan
Pubblicazione: (2026)
Towards Better Understanding of Cybercrime: The Role of Fine-Tuned LLMs in Translation
di: Valeros, Veronica, et al.
Pubblicazione: (2024)
di: Valeros, Veronica, et al.
Pubblicazione: (2024)
Teaching LLMs How to Learn with Contextual Fine-Tuning
di: Choi, Younwoo, et al.
Pubblicazione: (2025)
di: Choi, Younwoo, et al.
Pubblicazione: (2025)
Investigating the Impact of Rationales for LLMs on Natural Language Understanding
di: Shi, Wenhang, et al.
Pubblicazione: (2025)
di: Shi, Wenhang, et al.
Pubblicazione: (2025)
Fine-Tuning Small Embeddings for Elevated Performance
di: Silwal, Biraj
Pubblicazione: (2024)
di: Silwal, Biraj
Pubblicazione: (2024)
Teaching According to Talents! Instruction Tuning LLMs with Competence-Aware Curriculum Learning
di: Li, Yangning, et al.
Pubblicazione: (2025)
di: Li, Yangning, et al.
Pubblicazione: (2025)
Beyond GPT-5: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing
di: Zhang, Yiqun, et al.
Pubblicazione: (2025)
di: Zhang, Yiqun, et al.
Pubblicazione: (2025)
Rationale Behind Essay Scores: Enhancing S-LLM's Multi-Trait Essay Scoring with Rationale Generated by LLMs
di: Chu, SeongYeub, et al.
Pubblicazione: (2024)
di: Chu, SeongYeub, et al.
Pubblicazione: (2024)
Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs
di: Venkatesh, Sohan
Pubblicazione: (2026)
di: Venkatesh, Sohan
Pubblicazione: (2026)
SERPENT-VLM : Self-Refining Radiology Report Generation Using Vision Language Models
di: Kapadnis, Manav Nitin, et al.
Pubblicazione: (2024)
di: Kapadnis, Manav Nitin, et al.
Pubblicazione: (2024)
Step-On-Feet Tuning: Scaling Self-Alignment of LLMs via Bootstrapping
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
Measuring and Improving Persuasiveness of Large Language Models
di: Singh, Somesh, et al.
Pubblicazione: (2024)
di: Singh, Somesh, et al.
Pubblicazione: (2024)
Verbosity-Aware Rationale Reduction: Effective Reduction of Redundant Rationale via Principled Criteria
di: Jang, Joonwon, et al.
Pubblicazione: (2024)
di: Jang, Joonwon, et al.
Pubblicazione: (2024)
Scaling External Knowledge Input Beyond Context Windows of LLMs via Multi-Agent Collaboration
di: Liu, Zijun, et al.
Pubblicazione: (2025)
di: Liu, Zijun, et al.
Pubblicazione: (2025)
Self-Tuning: Instructing LLMs to Effectively Acquire New Knowledge through Self-Teaching
di: Zhang, Xiaoying, et al.
Pubblicazione: (2024)
di: Zhang, Xiaoying, et al.
Pubblicazione: (2024)
Teaching LLMs Human-Like Editing of Inappropriate Argumentation via Reinforcement Learning
di: Ziegenbein, Timon, et al.
Pubblicazione: (2026)
di: Ziegenbein, Timon, et al.
Pubblicazione: (2026)
Assessing and Mitigating Data Memorization Risks in Fine-Tuned Large Language Models
di: Ramakrishnan, Badrinath, et al.
Pubblicazione: (2025)
di: Ramakrishnan, Badrinath, et al.
Pubblicazione: (2025)
Calibrating LLMs with Preference Optimization on Thought Trees for Generating Rationale in Science Question Scoring
di: Li, Jiazheng, et al.
Pubblicazione: (2024)
di: Li, Jiazheng, et al.
Pubblicazione: (2024)
Small LLMs with Expert Blocks Are Good Enough for Hyperparamter Tuning
di: Naphade, Om, et al.
Pubblicazione: (2025)
di: Naphade, Om, et al.
Pubblicazione: (2025)
Documenti analoghi
-
It Helps to Take a Second Opinion: Teaching Smaller LLMs to Deliberate Mutually via Selective Rationale Optimisation
di: Patnaik, Sohan, et al.
Pubblicazione: (2025) -
CABINET: Content Relevance based Noise Reduction for Table Question Answering
di: Patnaik, Sohan, et al.
Pubblicazione: (2024) -
AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language Models
di: Patnaik, Sohan, et al.
Pubblicazione: (2025) -
Dialogue Agents 101: A Beginner's Guide to Critical Ingredients for Designing Effective Conversational Systems
di: Kumar, Shivani, et al.
Pubblicazione: (2023) -
AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent
di: Anand, Neeraj, et al.
Pubblicazione: (2025)