Evaluating LLM Understanding via Structured Tabular Decision Simulations
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Sichao, Xu, Xinyue, Li, Xiaomeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ELF-Gym: Evaluating Large Language Models Generated Features for Tabular Prediction
by: Zhang, Yanlin, et al.
Published: (2024)
by: Zhang, Yanlin, et al.
Published: (2024)
LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet Arena
by: Yang, Qingchuan, et al.
Published: (2025)
by: Yang, Qingchuan, et al.
Published: (2025)
RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Tabular Transfer Learning via Prompting LLMs
by: Nam, Jaehyun, et al.
Published: (2024)
by: Nam, Jaehyun, et al.
Published: (2024)
GPTree: Towards Explainable Decision-Making via LLM-powered Decision Trees
by: Xiong, Sichao, et al.
Published: (2024)
by: Xiong, Sichao, et al.
Published: (2024)
DiagnoLLM: A Hybrid Bayesian Neural Language Framework for Interpretable Disease Diagnosis
by: Xu, Bowen, et al.
Published: (2025)
by: Xu, Bowen, et al.
Published: (2025)
Anomaly Detection of Tabular Data Using LLMs
by: Li, Aodong, et al.
Published: (2024)
by: Li, Aodong, et al.
Published: (2024)
Understanding the planning of LLM agents: A survey
by: Huang, Xu, et al.
Published: (2024)
by: Huang, Xu, et al.
Published: (2024)
Large Scale Transfer Learning for Tabular Data via Language Modeling
by: Gardner, Josh, et al.
Published: (2024)
by: Gardner, Josh, et al.
Published: (2024)
Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
by: Polo, Felipe Maia, et al.
Published: (2025)
by: Polo, Felipe Maia, et al.
Published: (2025)
Understanding Generalization in Role-Playing Models via Information Theory
by: Li, Yongqi, et al.
Published: (2025)
by: Li, Yongqi, et al.
Published: (2025)
BPO: Staying Close to the Behavior LLM Creates Better Online LLM Alignment
by: Xu, Wenda, et al.
Published: (2024)
by: Xu, Wenda, et al.
Published: (2024)
Understanding LLM Embeddings for Regression
by: Tang, Eric, et al.
Published: (2024)
by: Tang, Eric, et al.
Published: (2024)
ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
by: Kim, Jeonghye, et al.
Published: (2025)
by: Kim, Jeonghye, et al.
Published: (2025)
Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking
by: Xu, Qinwu, et al.
Published: (2026)
by: Xu, Qinwu, et al.
Published: (2026)
Survey on Evaluation of LLM-based Agents
by: Yehudai, Asaf, et al.
Published: (2025)
by: Yehudai, Asaf, et al.
Published: (2025)
Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents
by: Wang, Zehong, et al.
Published: (2026)
by: Wang, Zehong, et al.
Published: (2026)
Teaching Your Models to Understand Code via Focal Preference Alignment
by: Wu, Jie, et al.
Published: (2025)
by: Wu, Jie, et al.
Published: (2025)
Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks
by: Li, Miaomiao, et al.
Published: (2025)
by: Li, Miaomiao, et al.
Published: (2025)
Predicting Decisions of AI Agents from Limited Interaction through Text-Tabular Modeling
by: Shapira, Eilam, et al.
Published: (2026)
by: Shapira, Eilam, et al.
Published: (2026)
Logical Structure as Knowledge: Enhancing LLM Reasoning via Structured Logical Knowledge Density Estimation
by: Bi, Zhen, et al.
Published: (2025)
by: Bi, Zhen, et al.
Published: (2025)
TabDLM: Free-Form Tabular Data Generation via Joint Numerical-Language Diffusion
by: Cai, Donghong, et al.
Published: (2026)
by: Cai, Donghong, et al.
Published: (2026)
Better LLM Reasoning via Dual-Play
by: Zhang, Zhengxin, et al.
Published: (2025)
by: Zhang, Zhengxin, et al.
Published: (2025)
Tuning LLM Judge Design Decisions for 1/1000 of the Cost
by: Salinas, David, et al.
Published: (2025)
by: Salinas, David, et al.
Published: (2025)
Is Escalation Worth It? A Decision-Theoretic Characterization of LLM Cascades
by: Bouchard, Dylan
Published: (2026)
by: Bouchard, Dylan
Published: (2026)
Understanding and Mitigating Dataset Corruption in LLM Steering
by: Anderson, Cullen, et al.
Published: (2026)
by: Anderson, Cullen, et al.
Published: (2026)
Understanding the Effects of RLHF on LLM Generalisation and Diversity
by: Kirk, Robert, et al.
Published: (2023)
by: Kirk, Robert, et al.
Published: (2023)
Pseudo-Siamese Network for Planning in Target-Oriented Proactive Dialogues
by: Kang, Xinyue, et al.
Published: (2026)
by: Kang, Xinyue, et al.
Published: (2026)
CreditAudit: 2$^\text{nd}$ Dimension for LLM Evaluation and Selection
by: Song, Yiliang, et al.
Published: (2026)
by: Song, Yiliang, et al.
Published: (2026)
Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
by: Hu, Xiaomeng, et al.
Published: (2025)
by: Hu, Xiaomeng, et al.
Published: (2025)
GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation
by: Li, Sijia, et al.
Published: (2026)
by: Li, Sijia, et al.
Published: (2026)
LLM-Guided Indoor Navigation with Multimodal Map Understanding
by: Coffrini, Alberto, et al.
Published: (2025)
by: Coffrini, Alberto, et al.
Published: (2025)
HindSight: Evaluating LLM-Generated Research Ideas via Future Impact
by: Jiang, Bo
Published: (2026)
by: Jiang, Bo
Published: (2026)
Understanding the Performance and Estimating the Cost of LLM Fine-Tuning
by: Xia, Yuchen, et al.
Published: (2024)
by: Xia, Yuchen, et al.
Published: (2024)
Multimodal Tabular Reasoning with Privileged Structured Information
by: Jiang, Jun-Peng, et al.
Published: (2025)
by: Jiang, Jun-Peng, et al.
Published: (2025)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
by: Liu, Yixin, et al.
Published: (2025)
by: Liu, Yixin, et al.
Published: (2025)
Towards Understanding Multi-Round Large Language Model Reasoning: Approximability, Learnability and Generalizability
by: Xu, Chenhui, et al.
Published: (2025)
by: Xu, Chenhui, et al.
Published: (2025)
AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
RelayLLM: Efficient Reasoning via Collaborative Decoding
by: Huang, Chengsong, et al.
Published: (2026)
by: Huang, Chengsong, et al.
Published: (2026)
Similar Items
-
ELF-Gym: Evaluating Large Language Models Generated Features for Tabular Prediction
by: Zhang, Yanlin, et al.
Published: (2024) -
LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet Arena
by: Yang, Qingchuan, et al.
Published: (2025) -
RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
by: Wang, Zihan, et al.
Published: (2025) -
Tabular Transfer Learning via Prompting LLMs
by: Nam, Jaehyun, et al.
Published: (2024) -
GPTree: Towards Explainable Decision-Making via LLM-powered Decision Trees
by: Xiong, Sichao, et al.
Published: (2024)