Aligning Large Language Models to Follow Instructions and Hallucinate Less via Effective Data Filtering
Fuente:
arXiv
Saved in:
| Main Authors: | Si, Shuzheng, Zhao, Haozhe, Chen, Gang, Gao, Cheng, Bai, Yuzhuo, Wang, Zhitong, An, Kaikai, Luo, Kangyang, Qian, Chen, Qi, Fanchao, Chang, Baobao, Sun, Maosong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FaithLens: Detecting and Explaining Faithfulness Hallucination
by: Si, Shuzheng, et al.
Published: (2025)
by: Si, Shuzheng, et al.
Published: (2025)
A Goal Without a Plan Is Just a Wish: Efficient and Effective Global Planner Training for Long-Horizon Agent Tasks
by: Si, Shuzheng, et al.
Published: (2025)
by: Si, Shuzheng, et al.
Published: (2025)
GATEAU: Selecting Influential Samples for Long Context Alignment
by: Si, Shuzheng, et al.
Published: (2024)
by: Si, Shuzheng, et al.
Published: (2024)
Teaching Large Language Models to Maintain Contextual Faithfulness via Synthetic Tasks and Reinforcement Learning
by: Si, Shuzheng, et al.
Published: (2025)
by: Si, Shuzheng, et al.
Published: (2025)
InFi-Check: Interpretable and Fine-Grained Fact-Checking of LLMs
by: Bai, Yuzhuo, et al.
Published: (2026)
by: Bai, Yuzhuo, et al.
Published: (2026)
UltraIF: Advancing Instruction Following from the Wild
by: An, Kaikai, et al.
Published: (2025)
by: An, Kaikai, et al.
Published: (2025)
From Context to Skills: Can Language Models Learn from Context Skillfully?
by: Si, Shuzheng, et al.
Published: (2026)
by: Si, Shuzheng, et al.
Published: (2026)
Mitigating Language-Level Performance Disparity in mPLMs via Teacher Language Selection and Cross-lingual Self-Distillation
by: Zhao, Haozhe, et al.
Published: (2024)
by: Zhao, Haozhe, et al.
Published: (2024)
KARL: Mitigating Hallucinations in LLMs via Knowledge-Boundary-Aware Reinforcement Learning
by: Gao, Cheng, et al.
Published: (2026)
by: Gao, Cheng, et al.
Published: (2026)
Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
by: Zhao, Haozhe, et al.
Published: (2024)
by: Zhao, Haozhe, et al.
Published: (2024)
Rethinking Semantic Parsing for Large Language Models: Enhancing LLM Performance with Semantic Hints
by: An, Kaikai, et al.
Published: (2024)
by: An, Kaikai, et al.
Published: (2024)
Improving the Robustness of Distantly-Supervised Named Entity Recognition via Uncertainty-Aware Teacher Learning and Student-Student Collaborative Learning
by: Si, Shuzheng, et al.
Published: (2023)
by: Si, Shuzheng, et al.
Published: (2023)
MEIC-DT: Memory-Efficient Incremental Clustering for Long-Text Coreference Resolution with Dual-Threshold Constraints
by: Luo, Kangyang, et al.
Published: (2025)
by: Luo, Kangyang, et al.
Published: (2025)
UltraEdit: Instruction-based Fine-Grained Image Editing at Scale
by: Zhao, Haozhe, et al.
Published: (2024)
by: Zhao, Haozhe, et al.
Published: (2024)
RhinoInsight: Improving Deep Research through Control Mechanisms for Model Behavior and Context
by: Lei, Yu, et al.
Published: (2025)
by: Lei, Yu, et al.
Published: (2025)
ImCoref-CeS: An Improved Lightweight Pipeline for Coreference Resolution with LLM-based Checker-Splitter Refinement
by: Luo, Kangyang, et al.
Published: (2025)
by: Luo, Kangyang, et al.
Published: (2025)
MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
by: Zhao, Haozhe, et al.
Published: (2023)
by: Zhao, Haozhe, et al.
Published: (2023)
GLTW: Joint Improved Graph Transformer and LLM via Three-Word Language for Knowledge Graph Completion
by: Luo, Kangyang, et al.
Published: (2025)
by: Luo, Kangyang, et al.
Published: (2025)
From Context to EDUs: Faithful and Structured Context Compression via Elementary Discourse Unit Decomposition
by: Zhou, Yiqing, et al.
Published: (2025)
by: Zhou, Yiqing, et al.
Published: (2025)
WGRAMMAR: Leverage Prior Knowledge to Accelerate Structured Decoding
by: Wang, Ran, et al.
Published: (2025)
by: Wang, Ran, et al.
Published: (2025)
From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corpora
by: Shen, Yingli, et al.
Published: (2025)
by: Shen, Yingli, et al.
Published: (2025)
SANTA: Separate Strategies for Inaccurate and Incomplete Annotation Noise in Distantly-Supervised Named Entity Recognition
by: Si, Shuzheng, et al.
Published: (2023)
by: Si, Shuzheng, et al.
Published: (2023)
Improving MLLM Training Efficiency via Stage-Aware Sparsity
by: Shi, Kean, et al.
Published: (2025)
by: Shi, Kean, et al.
Published: (2025)
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs
by: Zhao, Haozhe, et al.
Published: (2026)
by: Zhao, Haozhe, et al.
Published: (2026)
An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
by: Chen, Liang, et al.
Published: (2024)
by: Chen, Liang, et al.
Published: (2024)
Thread: A Logic-Based Data Organization Paradigm for How-To Question Answering with Retrieval Augmented Generation
by: An, Kaikai, et al.
Published: (2024)
by: An, Kaikai, et al.
Published: (2024)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
by: Chen, Liang, et al.
Published: (2025)
by: Chen, Liang, et al.
Published: (2025)
H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs
by: Gao, Cheng, et al.
Published: (2025)
by: Gao, Cheng, et al.
Published: (2025)
Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning
by: Wu, Rujie, et al.
Published: (2026)
by: Wu, Rujie, et al.
Published: (2026)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
by: Li, Bin, et al.
Published: (2025)
by: Li, Bin, et al.
Published: (2025)
Finite-time stabilization of discontinuous fuzzy inertial Cohen–Grossberg neural networks with mixed time-varying delays.
by: Fanchao Kong
Published: (2021)
by: Fanchao Kong
Published: (2021)
MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
by: Zhao, Haozhe, et al.
Published: (2025)
by: Zhao, Haozhe, et al.
Published: (2025)
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
by: Ji, Jiaming, et al.
Published: (2024)
by: Ji, Jiaming, et al.
Published: (2024)
DCAD-2000: A Multilingual Dataset across 2000+ Languages with Data Cleaning as Anomaly Detection
by: Shen, Yingli, et al.
Published: (2025)
by: Shen, Yingli, et al.
Published: (2025)
GRAIT: Gradient-Driven Refusal-Aware Instruction Tuning for Effective Hallucination Mitigation
by: Zhu, Runchuan, et al.
Published: (2025)
by: Zhu, Runchuan, et al.
Published: (2025)
CROP: Expert-Aligned Image Cropping via Compositional Reasoning and Optimizing Preference
by: Dong, Zhitong, et al.
Published: (2026)
by: Dong, Zhitong, et al.
Published: (2026)
What Makes LVLMs Hallucinate Less? Unveiling the Architectural Factors Behind Hallucination Robustness
by: He, Yusheng, et al.
Published: (2026)
by: He, Yusheng, et al.
Published: (2026)
AI Organizations are More Effective but Less Aligned than Individual Agents
by: Shen, Judy Hanwen, et al.
Published: (2026)
by: Shen, Judy Hanwen, et al.
Published: (2026)
Optima: Optimizing Effectiveness and Efficiency for LLM-Based Multi-Agent System
by: Chen, Weize, et al.
Published: (2024)
by: Chen, Weize, et al.
Published: (2024)
Concavity property of minimal $L^{2}$ integrals with Lebesgue measurable gain II
by: Guan, Qi'an, et al.
Published: (2022)
by: Guan, Qi'an, et al.
Published: (2022)
Similar Items
-
FaithLens: Detecting and Explaining Faithfulness Hallucination
by: Si, Shuzheng, et al.
Published: (2025) -
A Goal Without a Plan Is Just a Wish: Efficient and Effective Global Planner Training for Long-Horizon Agent Tasks
by: Si, Shuzheng, et al.
Published: (2025) -
GATEAU: Selecting Influential Samples for Long Context Alignment
by: Si, Shuzheng, et al.
Published: (2024) -
Teaching Large Language Models to Maintain Contextual Faithfulness via Synthetic Tasks and Reinforcement Learning
by: Si, Shuzheng, et al.
Published: (2025) -
InFi-Check: Interpretable and Fine-Grained Fact-Checking of LLMs
by: Bai, Yuzhuo, et al.
Published: (2026)