ThinkTuning: Instilling Cognitive Reflections without Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | RRV, Aswin, Dineen, Jacob, Handa, Divij, Uddin, Md Nayem, Parmar, Mihir, Baral, Chitta, Zhou, Ben |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
by: Handa, Divij, et al.
Published: (2025)
by: Handa, Divij, et al.
Published: (2025)
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models
by: RRV, Aswin, et al.
Published: (2026)
by: RRV, Aswin, et al.
Published: (2026)
When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers
by: Handa, Divij, et al.
Published: (2024)
by: Handa, Divij, et al.
Published: (2024)
Chaos with Keywords: Exposing Large Language Models Sycophantic Hallucination to Misleading Keywords and Evaluating Defense Strategies
by: RRV, Aswin, et al.
Published: (2024)
by: RRV, Aswin, et al.
Published: (2024)
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution
by: Dineen, Jacob, et al.
Published: (2026)
by: Dineen, Jacob, et al.
Published: (2026)
PHANTOM RECALL: When Familiar Puzzles Fool Smart Models
by: Mukhopadhyay, Souradeep, et al.
Published: (2025)
by: Mukhopadhyay, Souradeep, et al.
Published: (2025)
Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?
by: Tyagi, Nemika, et al.
Published: (2024)
by: Tyagi, Nemika, et al.
Published: (2024)
ToW: Thoughts of Words Improve Reasoning in Large Language Models
by: Xu, Zhikun, et al.
Published: (2024)
by: Xu, Zhikun, et al.
Published: (2024)
UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' Memorization
by: Uddin, Md Nayem, et al.
Published: (2024)
by: Uddin, Md Nayem, et al.
Published: (2024)
Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks
by: Saeidi, Amir, et al.
Published: (2024)
by: Saeidi, Amir, et al.
Published: (2024)
Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions
by: Parmar, Mihir, et al.
Published: (2022)
by: Parmar, Mihir, et al.
Published: (2022)
Triple Preference Optimization: Achieving Better Alignment using a Single Step Optimization
by: Saeidi, Amir, et al.
Published: (2024)
by: Saeidi, Amir, et al.
Published: (2024)
ActionReasoningBench: Reasoning about Actions with and without Ramification Constraints
by: Handa, Divij, et al.
Published: (2024)
by: Handa, Divij, et al.
Published: (2024)
QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
by: Dineen, Jacob, et al.
Published: (2025)
by: Dineen, Jacob, et al.
Published: (2025)
From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents
by: Uddin, Md Nayem, et al.
Published: (2026)
by: Uddin, Md Nayem, et al.
Published: (2026)
Hypothesis Generation for Materials Discovery and Design Using Goal-Driven and Constraint-Guided LLM Agents
by: Kumbhar, Shrinidhi, et al.
Published: (2025)
by: Kumbhar, Shrinidhi, et al.
Published: (2025)
Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark
by: Gupta, Himanshu, et al.
Published: (2024)
by: Gupta, Himanshu, et al.
Published: (2024)
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
by: Patel, Nisarg, et al.
Published: (2024)
by: Patel, Nisarg, et al.
Published: (2024)
Towards LogiGLUE: A Brief Survey and A Benchmark for Analyzing Logical Reasoning Capabilities of Language Models
by: Luo, Man, et al.
Published: (2023)
by: Luo, Man, et al.
Published: (2023)
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
by: Parmar, Mihir, et al.
Published: (2024)
by: Parmar, Mihir, et al.
Published: (2024)
PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem Solving
by: Parmar, Mihir, et al.
Published: (2025)
by: Parmar, Mihir, et al.
Published: (2025)
Investigating the Shortcomings of LLMs in Step-by-Step Legal Reasoning
by: Mishra, Venkatesh, et al.
Published: (2025)
by: Mishra, Venkatesh, et al.
Published: (2025)
Chronocept: Instilling a Sense of Time in Machines
by: Goel, Krish, et al.
Published: (2025)
by: Goel, Krish, et al.
Published: (2025)
PlanGEN: A Multi-Agent Framework for Generating Planning and Reasoning Trajectories for Complex Problem Solving
by: Parmar, Mihir, et al.
Published: (2025)
by: Parmar, Mihir, et al.
Published: (2025)
TarGEN: Targeted Data Generation with Large Language Models
by: Gupta, Himanshu, et al.
Published: (2023)
by: Gupta, Himanshu, et al.
Published: (2023)
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
by: Zhang, Zehua, et al.
Published: (2025)
by: Zhang, Zehua, et al.
Published: (2025)
SMART SLM: Structured Memory and Reasoning Transformer, A Small Language Model for Accurate Document Assistance
by: Dudeja, Divij, et al.
Published: (2025)
by: Dudeja, Divij, et al.
Published: (2025)
ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
by: Liu, Qin, et al.
Published: (2025)
by: Liu, Qin, et al.
Published: (2025)
Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes
by: Kenaston, Matthew W., et al.
Published: (2025)
by: Kenaston, Matthew W., et al.
Published: (2025)
Asking and Answering Questions to Extract Event-Argument Structures
by: Uddin, Md Nayem, et al.
Published: (2024)
by: Uddin, Md Nayem, et al.
Published: (2024)
Generating Uncontextualized and Contextualized Questions for Document-Level Event Argument Extraction
by: Uddin, Md Nayem, et al.
Published: (2024)
by: Uddin, Md Nayem, et al.
Published: (2024)
Indic-TunedLens: Interpreting Multilingual Models in Indian Languages
by: Panchal, Mihir, et al.
Published: (2026)
by: Panchal, Mihir, et al.
Published: (2026)
BOW: Reinforcement Learning for Bottlenecked Next Word Prediction
by: Shen, Ming, et al.
Published: (2025)
by: Shen, Ming, et al.
Published: (2025)
LIFT: Last-Mile Fine-Tuning for Table Explicitation
by: Khaitan, Divij, et al.
Published: (2026)
by: Khaitan, Divij, et al.
Published: (2026)
OptAgent: Optimizing Query Rewriting for E-commerce via Multi-Agent Simulation
by: Handa, Divij, et al.
Published: (2025)
by: Handa, Divij, et al.
Published: (2025)
Towards Enhancing Coherence in Extractive Summarization: Dataset and Experiments with LLMs
by: Parmar, Mihir, et al.
Published: (2024)
by: Parmar, Mihir, et al.
Published: (2024)
Thinking with DistilQwen: A Tale of Four Distilled Reasoning and Reward Model Series
by: Cai, Wenrui, et al.
Published: (2025)
by: Cai, Wenrui, et al.
Published: (2025)
Directional Optimization Asymmetry in Transformers: A Synthetic Stress Test
by: Sahasrabudhe, Mihir
Published: (2025)
by: Sahasrabudhe, Mihir
Published: (2025)
Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
Think Twice, Generate Once: Safeguarding by Progressive Self-Reflection
by: Phan, Hoang, et al.
Published: (2025)
by: Phan, Hoang, et al.
Published: (2025)
Similar Items
-
GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
by: Handa, Divij, et al.
Published: (2025) -
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models
by: RRV, Aswin, et al.
Published: (2026) -
When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers
by: Handa, Divij, et al.
Published: (2024) -
Chaos with Keywords: Exposing Large Language Models Sycophantic Hallucination to Misleading Keywords and Evaluating Defense Strategies
by: RRV, Aswin, et al.
Published: (2024) -
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution
by: Dineen, Jacob, et al.
Published: (2026)