Fine-Tuning LLMs to Analyze Multiple Dimensions of Code Review: A Maximum Entropy Regulated Long Chain-of-Thought Approach
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Yongda, Shi, Guohao, Wu, Xianwei, He, Haochuan, Gu, XueMing, Zhao, Qianqian, Liu, Kui, Wang, Qiushi, Tian, Zhao, Shen, Haifeng, Rong, Guoping |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Distilling Desired Comments for Enhanced Code Review with Large Language Models
by: Yu, Yongda, et al.
Published: (2024)
by: Yu, Yongda, et al.
Published: (2024)
SynthCoder: A Synthetical Strategy to Tune LLMs for Code Completion
by: Yu, Dongjun, et al.
Published: (2025)
by: Yu, Dongjun, et al.
Published: (2025)
Scattered Forest Search: Smarter Code Space Exploration with LLMs
by: Light, Jonathan, et al.
Published: (2024)
by: Light, Jonathan, et al.
Published: (2024)
Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers
by: Wang, Haochuan Kevin, et al.
Published: (2026)
by: Wang, Haochuan Kevin, et al.
Published: (2026)
TerraFormer: Automated Infrastructure-as-Code with LLMs Fine-Tuned via Policy-Guided Verifier Feedback
by: Jana, Prithwish, et al.
Published: (2026)
by: Jana, Prithwish, et al.
Published: (2026)
Diffusion-MPC in Discrete Domains: Feasibility Constraints, Horizon Effects, and Critic Alignment: Case study with Tetris
by: Wang, Haochuan Kevin
Published: (2026)
by: Wang, Haochuan Kevin
Published: (2026)
Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs
by: Pareja, Aldo, et al.
Published: (2024)
by: Pareja, Aldo, et al.
Published: (2024)
AutoCodeSherpa: Symbolic Explanations in AI Coding Agents
by: Kang, Sungmin, et al.
Published: (2025)
by: Kang, Sungmin, et al.
Published: (2025)
Exploring LLM-based Verilog Code Generation with Data-Efficient Fine-Tuning and Testbench Automation
by: Chen, Mu-Chi, et al.
Published: (2026)
by: Chen, Mu-Chi, et al.
Published: (2026)
Annotation Entropy Predicts Per-Example Learning Dynamics in LoRA Fine-Tuning
by: Steele, Brady
Published: (2026)
by: Steele, Brady
Published: (2026)
On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs
by: Zhao, Rosie, et al.
Published: (2026)
by: Zhao, Rosie, et al.
Published: (2026)
Fine-Tuning LLMs with Fine-Grained Human Feedback on Text Spans
by: CH-Wang, Sky, et al.
Published: (2025)
by: CH-Wang, Sky, et al.
Published: (2025)
Federated Fine-Tuning of LLMs on the Very Edge: The Good, the Bad, the Ugly
by: Woisetschläger, Herbert, et al.
Published: (2023)
by: Woisetschläger, Herbert, et al.
Published: (2023)
A Comparative Study of DSL Code Generation: Fine-Tuning vs. Optimized Retrieval Augmentation
by: Bassamzadeh, Nastaran, et al.
Published: (2024)
by: Bassamzadeh, Nastaran, et al.
Published: (2024)
FedMentalCare: Towards Privacy-Preserving Fine-Tuned LLMs to Analyze Mental Health Status Using Federated Learning Framework
by: Sarwar, Nobin
Published: (2025)
by: Sarwar, Nobin
Published: (2025)
Altered Thoughts, Altered Actions: Probing Chain-of-Thought Vulnerabilities in VLA Robotic Manipulation
by: Trinh, Tuan Duong, et al.
Published: (2026)
by: Trinh, Tuan Duong, et al.
Published: (2026)
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs
by: Rai, Daking, et al.
Published: (2024)
by: Rai, Daking, et al.
Published: (2024)
ME-IGM: Individual-Global-Max in Maximum Entropy Multi-Agent Reinforcement Learning
by: Chen, Wen-Tse, et al.
Published: (2024)
by: Chen, Wen-Tse, et al.
Published: (2024)
A Framework for Fine-Tuning LLMs using Heterogeneous Feedback
by: Aponte, Ryan, et al.
Published: (2024)
by: Aponte, Ryan, et al.
Published: (2024)
An Exploratory Study on Fine-Tuning Large Language Models for Secure Code Generation
by: Li, Junjie, et al.
Published: (2024)
by: Li, Junjie, et al.
Published: (2024)
Privately Fine-Tuned LLMs Preserve Temporal Dynamics in Tabular Data
by: Rosenblatt, Lucas, et al.
Published: (2026)
by: Rosenblatt, Lucas, et al.
Published: (2026)
Generalized Maximum Entropy: When and Why you need it
by: Ferro, Giuseppe M., et al.
Published: (2025)
by: Ferro, Giuseppe M., et al.
Published: (2025)
FED-FSTQ: Fisher-Guided Token Quantization for Communication-Efficient Federated Fine-Tuning of LLMs on Edge Devices
by: Li, Changyu, et al.
Published: (2026)
by: Li, Changyu, et al.
Published: (2026)
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
by: Ubukata, Shunsuke
Published: (2026)
by: Ubukata, Shunsuke
Published: (2026)
Augmented Relevance Datasets with Fine-Tuned Small LLMs
by: Fitte-Rey, Quentin, et al.
Published: (2025)
by: Fitte-Rey, Quentin, et al.
Published: (2025)
Chain of Unit-Physics: A Primitive-Centric Approach to Scientific Code Synthesis
by: Sharma, Vansh, et al.
Published: (2025)
by: Sharma, Vansh, et al.
Published: (2025)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
by: Pather, Kaviraj, et al.
Published: (2025)
by: Pather, Kaviraj, et al.
Published: (2025)
Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering
by: Fan, Lin, et al.
Published: (2026)
by: Fan, Lin, et al.
Published: (2026)
Sensorformer: Cross-patch attention with global-patch compression is effective for high-dimensional multivariate time series forecasting
by: Qin, Liyang, et al.
Published: (2025)
by: Qin, Liyang, et al.
Published: (2025)
A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
by: Balasubramanian, Sriram, et al.
Published: (2025)
by: Balasubramanian, Sriram, et al.
Published: (2025)
Directed Domain Fine-Tuning: Tailoring Separate Modalities for Specific Training Tasks
by: Wen, Daniel, et al.
Published: (2024)
by: Wen, Daniel, et al.
Published: (2024)
LLMDFA: Analyzing Dataflow in Code with Large Language Models
by: Wang, Chengpeng, et al.
Published: (2024)
by: Wang, Chengpeng, et al.
Published: (2024)
When Retrieval Hurts Code Completion: A Diagnostic Study of Stale Repository Context
by: Weng, Haojun, et al.
Published: (2026)
by: Weng, Haojun, et al.
Published: (2026)
Fine-Tuning LLMs on Small Medical Datasets: Text Classification and Normalization Effectiveness on Cardiology reports and Discharge records
by: Losch, Noah, et al.
Published: (2025)
by: Losch, Noah, et al.
Published: (2025)
STATe-of-Thoughts: Structured Action Templates for Tree-of-Thoughts
by: Bamberger, Zachary, et al.
Published: (2026)
by: Bamberger, Zachary, et al.
Published: (2026)
Who Sees What? Structured Thought-Action Sequences for Epistemic Reasoning in LLMs
by: Annese, Luca, et al.
Published: (2025)
by: Annese, Luca, et al.
Published: (2025)
The Last Word Often Wins: A Format Confound in Chain-of-Thought Corruption Studies
by: Garcia, Gabriel
Published: (2026)
by: Garcia, Gabriel
Published: (2026)
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
by: Oliveira, Daniel A. P., et al.
Published: (2025)
by: Oliveira, Daniel A. P., et al.
Published: (2025)
VeriPlan: Integrating Formal Verification and LLMs into End-User Planning
by: Lee, Christine, et al.
Published: (2025)
by: Lee, Christine, et al.
Published: (2025)
Agent-Facing Information Design in LLM Tool Registries
by: Wang, Haochuan Kevin
Published: (2026)
by: Wang, Haochuan Kevin
Published: (2026)
Similar Items
-
Distilling Desired Comments for Enhanced Code Review with Large Language Models
by: Yu, Yongda, et al.
Published: (2024) -
SynthCoder: A Synthetical Strategy to Tune LLMs for Code Completion
by: Yu, Dongjun, et al.
Published: (2025) -
Scattered Forest Search: Smarter Code Space Exploration with LLMs
by: Light, Jonathan, et al.
Published: (2024) -
Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers
by: Wang, Haochuan Kevin, et al.
Published: (2026) -
TerraFormer: Automated Infrastructure-as-Code with LLMs Fine-Tuned via Policy-Guided Verifier Feedback
by: Jana, Prithwish, et al.
Published: (2026)