LLM-Cave: A benchmark and light environment for large language models reasoning and decision-making system
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Huanyu, Li, Zongyuan, Huang, Wei, Guo, Xian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Relational reasoning and inductive bias in transformers and large language models
by: Geerts, Jesse, et al.
Published: (2025)
by: Geerts, Jesse, et al.
Published: (2025)
Comprehensive benchmarking of large language models for RNA secondary structure prediction
by: Zablocki, L. I., et al.
Published: (2024)
by: Zablocki, L. I., et al.
Published: (2024)
A note on the impossibility of conditional PAC-efficient reasoning in large language models
by: Zeng, Hao
Published: (2025)
by: Zeng, Hao
Published: (2025)
A dataset and benchmark for hospital course summarization with adapted large language models
by: Aali, Asad, et al.
Published: (2024)
by: Aali, Asad, et al.
Published: (2024)
HARP: A challenging human-annotated math reasoning benchmark
by: Yue, Albert S., et al.
Published: (2024)
by: Yue, Albert S., et al.
Published: (2024)
Joint modeling for learning decision-making dynamics in behavioral experiments
by: Bian, Yuan, et al.
Published: (2025)
by: Bian, Yuan, et al.
Published: (2025)
Layer-wise dynamic rank for compressing large language models
by: Mi, Zhendong, et al.
Published: (2025)
by: Mi, Zhendong, et al.
Published: (2025)
Multi-step retrieval and reasoning improves radiology question answering with large language models
by: Wind, Sebastian, et al.
Published: (2025)
by: Wind, Sebastian, et al.
Published: (2025)
Post-training makes large language models less human-like
by: Binz, Marcel, et al.
Published: (2026)
by: Binz, Marcel, et al.
Published: (2026)
Beyond IID: data-driven decision-making in heterogeneous environments
by: Besbes, Omar, et al.
Published: (2022)
by: Besbes, Omar, et al.
Published: (2022)
How predictable is language model benchmark performance?
by: Owen, David
Published: (2024)
by: Owen, David
Published: (2024)
Anchor function: a type of benchmark functions for studying language models
by: Zhang, Zhongwang, et al.
Published: (2024)
by: Zhang, Zhongwang, et al.
Published: (2024)
Context information can be more important than reasoning for time series forecasting with a large language model
by: Yang, Janghoon
Published: (2025)
by: Yang, Janghoon
Published: (2025)
Generative models for decision-making under distributional shift
by: Cheng, Xiuyuan, et al.
Published: (2026)
by: Cheng, Xiuyuan, et al.
Published: (2026)
NIRVANA: Structured pruning reimagined for large language models compression
by: Ai, Mengting, et al.
Published: (2025)
by: Ai, Mengting, et al.
Published: (2025)
Making medical vision-language models think causally across modalities with retrieval-augmented cross-modal reasoning
by: Yang, Weiqin, et al.
Published: (2026)
by: Yang, Weiqin, et al.
Published: (2026)
Replacing thinking with tool usage enables reasoning in small language models
by: Rainone, Corrado, et al.
Published: (2025)
by: Rainone, Corrado, et al.
Published: (2025)
Long-form factuality in large language models
by: Wei, Jerry, et al.
Published: (2024)
by: Wei, Jerry, et al.
Published: (2024)
The SMeL Test: A simple benchmark for media literacy in language models
by: Ahdritz, Gustaf, et al.
Published: (2025)
by: Ahdritz, Gustaf, et al.
Published: (2025)
CausalDynamics: A large-scale benchmark for structural discovery of dynamical causal models
by: Herdeanu, Benjamin, et al.
Published: (2025)
by: Herdeanu, Benjamin, et al.
Published: (2025)
DevBench: A multimodal developmental benchmark for language learning
by: Tan, Alvin Wei Ming, et al.
Published: (2024)
by: Tan, Alvin Wei Ming, et al.
Published: (2024)
Look, Remember and Reason: Grounded reasoning in videos with language models
by: Bhattacharyya, Apratim, et al.
Published: (2023)
by: Bhattacharyya, Apratim, et al.
Published: (2023)
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents
by: Hariharan, Kaivalya, et al.
Published: (2025)
by: Hariharan, Kaivalya, et al.
Published: (2025)
Leveraging large language models for SQL behavior-based database intrusion detection
by: Shlezinger, Meital, et al.
Published: (2025)
by: Shlezinger, Meital, et al.
Published: (2025)
ECG-LLM -- training and evaluation of domain-specific large language models for electrocardiography
by: Ahrens, Lara, et al.
Published: (2025)
by: Ahrens, Lara, et al.
Published: (2025)
Visual cognition in multimodal large language models
by: Buschoff, Luca M. Schulze, et al.
Published: (2023)
by: Buschoff, Luca M. Schulze, et al.
Published: (2023)
Hypothesis generation and updating in large language models
by: Xiong, Hua-Dong
Published: (2026)
by: Xiong, Hua-Dong
Published: (2026)
Quantifying perturbation impacts for large language models
by: Rauba, Paulius, et al.
Published: (2024)
by: Rauba, Paulius, et al.
Published: (2024)
Alignment of large language models with constrained learning
by: Zhang, Botong, et al.
Published: (2025)
by: Zhang, Botong, et al.
Published: (2025)
Efficient selective attention LSTM for well log curve synthesis
by: Zhou, Yuankai, et al.
Published: (2023)
by: Zhou, Yuankai, et al.
Published: (2023)
BabySLM: language-acquisition-friendly benchmark of self-supervised spoken language models
by: Lavechin, Marvin, et al.
Published: (2023)
by: Lavechin, Marvin, et al.
Published: (2023)
Zero-shot generation of synthetic neurosurgical data with large language models
by: Barr, Austin A., et al.
Published: (2025)
by: Barr, Austin A., et al.
Published: (2025)
Question answering system of bridge design specification based on large language model
by: Zhang, Leye, et al.
Published: (2024)
by: Zhang, Leye, et al.
Published: (2024)
Representation in large language models
by: Yetman, Cameron
Published: (2025)
by: Yetman, Cameron
Published: (2025)
Make LLMs better zero-shot reasoners: Structure-orientated autonomous reasoning
by: He, Pengfei, et al.
Published: (2024)
by: He, Pengfei, et al.
Published: (2024)
Can a large language model be a gaslighter?
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
Multi-agent decision making: A Blackwell's informativeness approach
by: Zhang, Zheng, et al.
Published: (2026)
by: Zhang, Zheng, et al.
Published: (2026)
A large dataset curation and benchmark for drug target interaction
by: Golts, Alex, et al.
Published: (2024)
by: Golts, Alex, et al.
Published: (2024)
Leveraging large language models for nano synthesis mechanism explanation: solid foundations or mere conjectures?
by: Pu, Yingming, et al.
Published: (2024)
by: Pu, Yingming, et al.
Published: (2024)
Amortizing intractable inference in large language models
by: Hu, Edward J., et al.
Published: (2023)
by: Hu, Edward J., et al.
Published: (2023)
Similar Items
-
Relational reasoning and inductive bias in transformers and large language models
by: Geerts, Jesse, et al.
Published: (2025) -
Comprehensive benchmarking of large language models for RNA secondary structure prediction
by: Zablocki, L. I., et al.
Published: (2024) -
A note on the impossibility of conditional PAC-efficient reasoning in large language models
by: Zeng, Hao
Published: (2025) -
A dataset and benchmark for hospital course summarization with adapted large language models
by: Aali, Asad, et al.
Published: (2024) -
HARP: A challenging human-annotated math reasoning benchmark
by: Yue, Albert S., et al.
Published: (2024)