Gespeichert in:
| Hauptverfasser: | Panas, D., Seth, S., Belle, V. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2404.19432 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Probing for Arithmetic Errors in Language Models
von: Sun, Yucheng, et al.
Veröffentlicht: (2025)
von: Sun, Yucheng, et al.
Veröffentlicht: (2025)
T-REX: Table -- Refute or Entail eXplainer
von: Horstmann, Tim Luka, et al.
Veröffentlicht: (2025)
von: Horstmann, Tim Luka, et al.
Veröffentlicht: (2025)
Zero, Finite, and Infinite Belief History of Theory of Mind Reasoning in Large Language Models
von: Tang, Weizhi, et al.
Veröffentlicht: (2024)
von: Tang, Weizhi, et al.
Veröffentlicht: (2024)
ToM-LM: Delegating Theory of Mind Reasoning to External Symbolic Executors in Large Language Models
von: Tang, Weizhi, et al.
Veröffentlicht: (2024)
von: Tang, Weizhi, et al.
Veröffentlicht: (2024)
LTLBench: Towards Benchmarks for Evaluating Temporal Reasoning in Large Language Models
von: Tang, Weizhi, et al.
Veröffentlicht: (2024)
von: Tang, Weizhi, et al.
Veröffentlicht: (2024)
Talking the Talk Does Not Entail Walking the Walk: On the Limits of Large Language Models in Lexical Entailment Recognition
von: Greco, Candida M., et al.
Veröffentlicht: (2024)
von: Greco, Candida M., et al.
Veröffentlicht: (2024)
Can Large Language Models Infer Causal Relationships from Real-World Text?
von: Saklad, Ryan, et al.
Veröffentlicht: (2025)
von: Saklad, Ryan, et al.
Veröffentlicht: (2025)
Self-training Language Models for Arithmetic Reasoning
von: Kadlčík, Marek, et al.
Veröffentlicht: (2024)
von: Kadlčík, Marek, et al.
Veröffentlicht: (2024)
Arithmetic with Language Models: from Memorization to Computation
von: Maltoni, Davide, et al.
Veröffentlicht: (2023)
von: Maltoni, Davide, et al.
Veröffentlicht: (2023)
Probing Causality Manipulation of Large Language Models
von: Zhang, Chenyang, et al.
Veröffentlicht: (2024)
von: Zhang, Chenyang, et al.
Veröffentlicht: (2024)
Probing Neural Topology of Large Language Models
von: Zheng, Yu, et al.
Veröffentlicht: (2025)
von: Zheng, Yu, et al.
Veröffentlicht: (2025)
Probing the Robustness of Theory of Mind in Large Language Models
von: Nickel, Christian, et al.
Veröffentlicht: (2024)
von: Nickel, Christian, et al.
Veröffentlicht: (2024)
Probing the Difficulty Perception Mechanism of Large Language Models
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
Improving Arithmetic Reasoning Ability of Large Language Models through Relation Tuples, Verification and Dynamic Feedback
von: Miao, Zhongtao, et al.
Veröffentlicht: (2024)
von: Miao, Zhongtao, et al.
Veröffentlicht: (2024)
Can Large Language Models do Analytical Reasoning?
von: Hu, Yebowen, et al.
Veröffentlicht: (2024)
von: Hu, Yebowen, et al.
Veröffentlicht: (2024)
A Careful Examination of Large Language Model Performance on Grade School Arithmetic
von: Zhang, Hugh, et al.
Veröffentlicht: (2024)
von: Zhang, Hugh, et al.
Veröffentlicht: (2024)
Modular Arithmetic: Language Models Solve Math Digit by Digit
von: Baeumel, Tanja, et al.
Veröffentlicht: (2025)
von: Baeumel, Tanja, et al.
Veröffentlicht: (2025)
Stuck in the Matrix: Probing Spatial Reasoning in Large Language Models
von: Bai, Maggie, et al.
Veröffentlicht: (2025)
von: Bai, Maggie, et al.
Veröffentlicht: (2025)
Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models
von: Mohammadi, Seyedali, et al.
Veröffentlicht: (2026)
von: Mohammadi, Seyedali, et al.
Veröffentlicht: (2026)
Neural Probe-Based Hallucination Detection for Large Language Models
von: Liang, Shize, et al.
Veröffentlicht: (2025)
von: Liang, Shize, et al.
Veröffentlicht: (2025)
Counterfactual Probing for Hallucination Detection and Mitigation in Large Language Models
von: Feng, Yijun
Veröffentlicht: (2025)
von: Feng, Yijun
Veröffentlicht: (2025)
LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?
von: Wang, Jingyuan, et al.
Veröffentlicht: (2025)
von: Wang, Jingyuan, et al.
Veröffentlicht: (2025)
Break the Chain: Large Language Models Can be Shortcut Reasoners
von: Ding, Mengru, et al.
Veröffentlicht: (2024)
von: Ding, Mengru, et al.
Veröffentlicht: (2024)
Can Large Language Models Predict the Outcome of Judicial Decisions?
von: Kmainasi, Mohamed Bayan, et al.
Veröffentlicht: (2025)
von: Kmainasi, Mohamed Bayan, et al.
Veröffentlicht: (2025)
THiNK: Can Large Language Models Think-aloud?
von: Yu, Yongan, et al.
Veröffentlicht: (2025)
von: Yu, Yongan, et al.
Veröffentlicht: (2025)
Can Large Language Models Express Uncertainty Like Human?
von: Tao, Linwei, et al.
Veröffentlicht: (2025)
von: Tao, Linwei, et al.
Veröffentlicht: (2025)
Bias Vector: Mitigating Biases in Language Models with Task Arithmetic Approach
von: Shirafuji, Daiki, et al.
Veröffentlicht: (2024)
von: Shirafuji, Daiki, et al.
Veröffentlicht: (2024)
Probing Multimodal Large Language Models for Global and Local Semantic Representations
von: Tao, Mingxu, et al.
Veröffentlicht: (2024)
von: Tao, Mingxu, et al.
Veröffentlicht: (2024)
Chain-of-Description: What I can understand, I can put into words
von: Guo, Jiaxin, et al.
Veröffentlicht: (2025)
von: Guo, Jiaxin, et al.
Veröffentlicht: (2025)
Language Models are Few-Shot Graders
von: Zhao, Chenyan, et al.
Veröffentlicht: (2025)
von: Zhao, Chenyan, et al.
Veröffentlicht: (2025)
Task Arithmetic for Language Expansion in Speech Translation
von: Cheng, Yao-Fei, et al.
Veröffentlicht: (2024)
von: Cheng, Yao-Fei, et al.
Veröffentlicht: (2024)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
von: Gambardella, Andrew, et al.
Veröffentlicht: (2024)
von: Gambardella, Andrew, et al.
Veröffentlicht: (2024)
Large Language Models Fall Short: Understanding Complex Relationships in Detective Narratives
von: Zhao, Runcong, et al.
Veröffentlicht: (2024)
von: Zhao, Runcong, et al.
Veröffentlicht: (2024)
Large Language Models Can Self-Improve in Long-context Reasoning
von: Li, Siheng, et al.
Veröffentlicht: (2024)
von: Li, Siheng, et al.
Veröffentlicht: (2024)
Importance Weighting Can Help Large Language Models Self-Improve
von: Jiang, Chunyang, et al.
Veröffentlicht: (2024)
von: Jiang, Chunyang, et al.
Veröffentlicht: (2024)
Can Large Language Models Automatically Score Proficiency of Written Essays?
von: Mansour, Watheq, et al.
Veröffentlicht: (2024)
von: Mansour, Watheq, et al.
Veröffentlicht: (2024)
Cause and Effect: Can Large Language Models Truly Understand Causality?
von: Ashwani, Swagata, et al.
Veröffentlicht: (2024)
von: Ashwani, Swagata, et al.
Veröffentlicht: (2024)
Can Large Language Models Understand Real-World Complex Instructions?
von: He, Qianyu, et al.
Veröffentlicht: (2023)
von: He, Qianyu, et al.
Veröffentlicht: (2023)
Self-HarmLLM: Can Large Language Model Harm Itself?
von: Kim, Heehwan, et al.
Veröffentlicht: (2025)
von: Kim, Heehwan, et al.
Veröffentlicht: (2025)
Can Large Language Models Predict Associations Among Human Attitudes?
von: Ma, Ana, et al.
Veröffentlicht: (2025)
von: Ma, Ana, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Probing for Arithmetic Errors in Language Models
von: Sun, Yucheng, et al.
Veröffentlicht: (2025) -
T-REX: Table -- Refute or Entail eXplainer
von: Horstmann, Tim Luka, et al.
Veröffentlicht: (2025) -
Zero, Finite, and Infinite Belief History of Theory of Mind Reasoning in Large Language Models
von: Tang, Weizhi, et al.
Veröffentlicht: (2024) -
ToM-LM: Delegating Theory of Mind Reasoning to External Symbolic Executors in Large Language Models
von: Tang, Weizhi, et al.
Veröffentlicht: (2024) -
LTLBench: Towards Benchmarks for Evaluating Temporal Reasoning in Large Language Models
von: Tang, Weizhi, et al.
Veröffentlicht: (2024)