Salvato in:
| Autori principali: | Li, Xiaoyuan, Li, Moxin, Men, Rui, Zhang, Yichang, Bao, Keqin, Wang, Wenjie, Feng, Fuli, Liu, Dayiheng, Lin, Junyang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2502.11393 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation
di: Li, Xiaoyuan, et al.
Pubblicazione: (2025)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2025)
What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks
di: Chizhov, Pavel, et al.
Pubblicazione: (2025)
di: Chizhov, Pavel, et al.
Pubblicazione: (2025)
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
di: Li, Xiaoyuan, et al.
Pubblicazione: (2025)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2025)
SAGE: Scalable Automated Robustness Augmentation for LLM Knowledge Evaluation
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
On Predicting the Post-training Potential of Pre-trained LLMs
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction
di: Li, Xiaoyuan, et al.
Pubblicazione: (2024)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2024)
Unified Data Selection for LLM Reasoning
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code
di: Bao, Keqin, et al.
Pubblicazione: (2025)
di: Bao, Keqin, et al.
Pubblicazione: (2025)
Chain of Execution Supervision Promotes General Reasoning in Large Language Models
di: Chen, Nuo, et al.
Pubblicazione: (2025)
di: Chen, Nuo, et al.
Pubblicazione: (2025)
Counterfactual Debating with Preset Stances for Hallucination Elimination of LLMs
di: Fang, Yi, et al.
Pubblicazione: (2024)
di: Fang, Yi, et al.
Pubblicazione: (2024)
Robust Prompt Optimization for Large Language Models Against Distribution Shifts
di: Li, Moxin, et al.
Pubblicazione: (2023)
di: Li, Moxin, et al.
Pubblicazione: (2023)
Teaching Language Models to Reason with Tools
di: Li, Chengpeng, et al.
Pubblicazione: (2025)
di: Li, Chengpeng, et al.
Pubblicazione: (2025)
Think Twice Before Trusting: Self-Detection for Large Language Models through Comprehensive Answer Reflection
di: Li, Moxin, et al.
Pubblicazione: (2024)
di: Li, Moxin, et al.
Pubblicazione: (2024)
CoRT: Code-integrated Reasoning within Thinking
di: Li, Chengpeng, et al.
Pubblicazione: (2025)
di: Li, Chengpeng, et al.
Pubblicazione: (2025)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
di: Fang, Yi, et al.
Pubblicazione: (2026)
di: Fang, Yi, et al.
Pubblicazione: (2026)
Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
di: Li, Yongxiang, et al.
Pubblicazione: (2026)
di: Li, Yongxiang, et al.
Pubblicazione: (2026)
Text-like Encoding of Collaborative Information in Large Language Models for Recommendation
di: Zhang, Yang, et al.
Pubblicazione: (2024)
di: Zhang, Yang, et al.
Pubblicazione: (2024)
TAT-LLM: A Specialized Language Model for Discrete Reasoning over Tabular and Textual Data
di: Zhu, Fengbin, et al.
Pubblicazione: (2024)
di: Zhu, Fengbin, et al.
Pubblicazione: (2024)
Assistant-Guided Mitigation of Teacher Preference Bias in LLM-as-a-Judge
di: Liu, Zhuo, et al.
Pubblicazione: (2025)
di: Liu, Zhuo, et al.
Pubblicazione: (2025)
Doc2SoarGraph: Discrete Reasoning over Visually-Rich Table-Text Documents via Semantic-Oriented Hierarchical Graphs
di: Zhu, Fengbin, et al.
Pubblicazione: (2023)
di: Zhu, Fengbin, et al.
Pubblicazione: (2023)
Byzantion Nea Hellás
Pubblicazione: (2011)
Pubblicazione: (2011)
Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts in Multi-Objective Alignment
di: Li, Moxin, et al.
Pubblicazione: (2025)
di: Li, Moxin, et al.
Pubblicazione: (2025)
Real-Time Personalization for LLM-based Recommendation with Customized In-Context Learning
di: Bao, Keqin, et al.
Pubblicazione: (2024)
di: Bao, Keqin, et al.
Pubblicazione: (2024)
Item-side Fairness of Large Language Model-based Recommendation System
di: Jiang, Meng, et al.
Pubblicazione: (2024)
di: Jiang, Meng, et al.
Pubblicazione: (2024)
Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations
di: Sun, Jiaxing, et al.
Pubblicazione: (2024)
di: Sun, Jiaxing, et al.
Pubblicazione: (2024)
Benchmarking Chinese Commonsense Reasoning with a Multi-hop Reasoning Perspective
di: You, Wangjie, et al.
Pubblicazione: (2025)
di: You, Wangjie, et al.
Pubblicazione: (2025)
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
di: Quan, Shanghaoran, et al.
Pubblicazione: (2025)
di: Quan, Shanghaoran, et al.
Pubblicazione: (2025)
SCoRE: Benchmarking Long-Chain Reasoning in Commonsense Scenarios
di: Zhan, Weidong, et al.
Pubblicazione: (2025)
di: Zhan, Weidong, et al.
Pubblicazione: (2025)
Prospect Personalized Recommendation on Large Language Model-based Agent Platform
di: Zhang, Jizhi, et al.
Pubblicazione: (2024)
di: Zhang, Jizhi, et al.
Pubblicazione: (2024)
Boosting Parameter Efficiency in LLM-Based Recommendation through Sophisticated Pruning
di: Zheng, Shanle, et al.
Pubblicazione: (2025)
di: Zheng, Shanle, et al.
Pubblicazione: (2025)
CoLLM: Integrating Collaborative Embeddings into Large Language Models for Recommendation
di: Zhang, Yang, et al.
Pubblicazione: (2023)
di: Zhang, Yang, et al.
Pubblicazione: (2023)
Decoding Matters: Addressing Amplification Bias and Homogeneity Issue for LLM-based Recommendation
di: Bao, Keqin, et al.
Pubblicazione: (2024)
di: Bao, Keqin, et al.
Pubblicazione: (2024)
A Bi-Step Grounding Paradigm for Large Language Models in Recommendation Systems
di: Bao, Keqin, et al.
Pubblicazione: (2023)
di: Bao, Keqin, et al.
Pubblicazione: (2023)
Dual-Phase Accelerated Prompt Optimization
di: Yang, Muchen, et al.
Pubblicazione: (2024)
di: Yang, Muchen, et al.
Pubblicazione: (2024)
Causal Debiasing for Visual Commonsense Reasoning
di: Zou, Jiayi, et al.
Pubblicazione: (2025)
di: Zou, Jiayi, et al.
Pubblicazione: (2025)
Leveraging LLMs for Influence Path Planning in Proactive Recommendation
di: Wang, Mingze, et al.
Pubblicazione: (2024)
di: Wang, Mingze, et al.
Pubblicazione: (2024)
Disentangling Reasoning Tokens and Boilerplate Tokens For Language Model Fine-tuning
di: Ye, Ziang, et al.
Pubblicazione: (2024)
di: Ye, Ziang, et al.
Pubblicazione: (2024)
GRASP: A Grid-Based Benchmark for Evaluating Commonsense Spatial Reasoning
di: Tang, Zhisheng, et al.
Pubblicazione: (2024)
di: Tang, Zhisheng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation
di: Li, Xiaoyuan, et al.
Pubblicazione: (2025) -
What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks
di: Chizhov, Pavel, et al.
Pubblicazione: (2025) -
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
di: Li, Xiaoyuan, et al.
Pubblicazione: (2025) -
SAGE: Scalable Automated Robustness Augmentation for LLM Knowledge Evaluation
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026) -
ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)