DICE: Detecting In-distribution Contamination in LLM's Fine-tuning Phase for Math Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tu, Shangqing, Zhu, Kejian, Bai, Yushi, Yao, Zijun, Hou, Lei, Li, Juanzi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis
von: Zhu, Kejian, et al.
Veröffentlicht: (2025)
von: Zhu, Kejian, et al.
Veröffentlicht: (2025)
DeepPrune: Parallel Scaling without Inter-trace Redundancy
von: Tu, Shangqing, et al.
Veröffentlicht: (2025)
von: Tu, Shangqing, et al.
Veröffentlicht: (2025)
MM-MATH: Advancing Multimodal Math Evaluation with Process Evaluation and Fine-grained Classification
von: Sun, Kai, et al.
Veröffentlicht: (2024)
von: Sun, Kai, et al.
Veröffentlicht: (2024)
WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models
von: Tu, Shangqing, et al.
Veröffentlicht: (2023)
von: Tu, Shangqing, et al.
Veröffentlicht: (2023)
MM-THEBench: Do Reasoning MLLMs Think Reasonably?
von: Huang, Zhidian, et al.
Veröffentlicht: (2026)
von: Huang, Zhidian, et al.
Veröffentlicht: (2026)
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
von: Chen, Jianhui, et al.
Veröffentlicht: (2024)
von: Chen, Jianhui, et al.
Veröffentlicht: (2024)
Pre-training Distillation for Large Language Models: A Design Space Exploration
von: Peng, Hao, et al.
Veröffentlicht: (2024)
von: Peng, Hao, et al.
Veröffentlicht: (2024)
ChatLog: Carefully Evaluating the Evolution of ChatGPT Across Time
von: Tu, Shangqing, et al.
Veröffentlicht: (2023)
von: Tu, Shangqing, et al.
Veröffentlicht: (2023)
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
von: Bai, Yushi, et al.
Veröffentlicht: (2024)
von: Bai, Yushi, et al.
Veröffentlicht: (2024)
Auxiliary Metrics Help Decoding Skill Neurons in the Wild
von: Zhao, Yixiu, et al.
Veröffentlicht: (2025)
von: Zhao, Yixiu, et al.
Veröffentlicht: (2025)
Shifting Long-Context LLMs Research from Input to Output
von: Wu, Yuhao, et al.
Veröffentlicht: (2025)
von: Wu, Yuhao, et al.
Veröffentlicht: (2025)
AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning
von: Xin, Amy, et al.
Veröffentlicht: (2024)
von: Xin, Amy, et al.
Veröffentlicht: (2024)
MMGeoLM: Hard Negative Contrastive Learning for Fine-Grained Geometric Understanding in Large Multimodal Models
von: Sun, Kai, et al.
Veröffentlicht: (2025)
von: Sun, Kai, et al.
Veröffentlicht: (2025)
LongWriter-V: Enabling Ultra-Long and High-Fidelity Generation in Vision-Language Models
von: Tu, Shangqing, et al.
Veröffentlicht: (2025)
von: Tu, Shangqing, et al.
Veröffentlicht: (2025)
Untangle the KNOT: Interweaving Conflicting Knowledge and Reasoning Skills in Large Language Models
von: Liu, Yantao, et al.
Veröffentlicht: (2024)
von: Liu, Yantao, et al.
Veröffentlicht: (2024)
Aligning Teacher with Student Preferences for Tailored Training Data Generation
von: Liu, Yantao, et al.
Veröffentlicht: (2024)
von: Liu, Yantao, et al.
Veröffentlicht: (2024)
Evaluating Generative Language Models in Information Extraction as Subjective Question Correction
von: Fan, Yuchen, et al.
Veröffentlicht: (2024)
von: Fan, Yuchen, et al.
Veröffentlicht: (2024)
RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
von: Liu, Yantao, et al.
Veröffentlicht: (2024)
von: Liu, Yantao, et al.
Veröffentlicht: (2024)
PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament
von: Liu, Yantao, et al.
Veröffentlicht: (2025)
von: Liu, Yantao, et al.
Veröffentlicht: (2025)
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
von: Jing, Yi, et al.
Veröffentlicht: (2026)
von: Jing, Yi, et al.
Veröffentlicht: (2026)
Are Reasoning Models More Prone to Hallucination?
von: Yao, Zijun, et al.
Veröffentlicht: (2025)
von: Yao, Zijun, et al.
Veröffentlicht: (2025)
SIRI: Scaling Iterative Reinforcement Learning with Interleaved Compression
von: Wen, Haoming, et al.
Veröffentlicht: (2025)
von: Wen, Haoming, et al.
Veröffentlicht: (2025)
WildReward: Learning Reward Models from In-the-Wild Human Interactions
von: Peng, Hao, et al.
Veröffentlicht: (2026)
von: Peng, Hao, et al.
Veröffentlicht: (2026)
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
von: Zhu, Kejian, et al.
Veröffentlicht: (2025)
von: Zhu, Kejian, et al.
Veröffentlicht: (2025)
StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
von: Chen, Yanxu, et al.
Veröffentlicht: (2025)
von: Chen, Yanxu, et al.
Veröffentlicht: (2025)
LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA
von: Zhang, Jiajie, et al.
Veröffentlicht: (2024)
von: Zhang, Jiajie, et al.
Veröffentlicht: (2024)
LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-Encoder
von: Jing, Yi, et al.
Veröffentlicht: (2025)
von: Jing, Yi, et al.
Veröffentlicht: (2025)
LLMAEL: Large Language Models are Good Context Augmenters for Entity Linking
von: Xin, Amy, et al.
Veröffentlicht: (2024)
von: Xin, Amy, et al.
Veröffentlicht: (2024)
R-Eval: A Unified Toolkit for Evaluating Domain Knowledge of Retrieval Augmented Large Language Models
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
SoAy: A Solution-based LLM API-using Methodology for Academic Information Seeking
von: Wang, Yuanchun, et al.
Veröffentlicht: (2024)
von: Wang, Yuanchun, et al.
Veröffentlicht: (2024)
Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems
von: Peng, Hao, et al.
Veröffentlicht: (2025)
von: Peng, Hao, et al.
Veröffentlicht: (2025)
CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning
von: Qi, Ji, et al.
Veröffentlicht: (2024)
von: Qi, Ji, et al.
Veröffentlicht: (2024)
LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs
von: Bai, Yushi, et al.
Veröffentlicht: (2024)
von: Bai, Yushi, et al.
Veröffentlicht: (2024)
SeaKR: Self-aware Knowledge Retrieval for Adaptive Retrieval Augmented Generation
von: Yao, Zijun, et al.
Veröffentlicht: (2024)
von: Yao, Zijun, et al.
Veröffentlicht: (2024)
A Cause-Effect Look at Alleviating Hallucination of Knowledge-grounded Dialogue Generation
von: Yu, Jifan, et al.
Veröffentlicht: (2024)
von: Yu, Jifan, et al.
Veröffentlicht: (2024)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety
von: Zhang, Yuyou, et al.
Veröffentlicht: (2025)
von: Zhang, Yuyou, et al.
Veröffentlicht: (2025)
SuperWriter: Reflection-Driven Long-Form Generation with Large Language Models
von: Wu, Yuhao, et al.
Veröffentlicht: (2025)
von: Wu, Yuhao, et al.
Veröffentlicht: (2025)
Transferable and Efficient Non-Factual Content Detection via Probe Training with Offline Consistency Checking
von: Zhang, Xiaokang, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaokang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis
von: Zhu, Kejian, et al.
Veröffentlicht: (2025) -
DeepPrune: Parallel Scaling without Inter-trace Redundancy
von: Tu, Shangqing, et al.
Veröffentlicht: (2025) -
MM-MATH: Advancing Multimodal Math Evaluation with Process Evaluation and Fine-grained Classification
von: Sun, Kai, et al.
Veröffentlicht: (2024) -
WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models
von: Tu, Shangqing, et al.
Veröffentlicht: (2023) -
MM-THEBench: Do Reasoning MLLMs Think Reasonably?
von: Huang, Zhidian, et al.
Veröffentlicht: (2026)