Gespeichert in:
| Hauptverfasser: | Bober-Irizar, Mikel, Banerjee, Soumya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2402.03507 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Skill Issues: An Analysis of CS:GO Skill Rating Systems
von: Bober-Irizar, Mikel, et al.
Veröffentlicht: (2024)
von: Bober-Irizar, Mikel, et al.
Veröffentlicht: (2024)
When can transformers reason with abstract symbols?
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2023)
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2023)
Artificial Expert Intelligence through PAC-reasoning
von: Shalev-Shwartz, Shai, et al.
Veröffentlicht: (2024)
von: Shalev-Shwartz, Shai, et al.
Veröffentlicht: (2024)
Is continuous CoT better suited for multi-lingual reasoning?
von: Bashir, Ali Hamza, et al.
Veröffentlicht: (2026)
von: Bashir, Ali Hamza, et al.
Veröffentlicht: (2026)
Sudoku-Bench: Evaluating creative reasoning with Sudoku variants
von: Seely, Jeffrey, et al.
Veröffentlicht: (2025)
von: Seely, Jeffrey, et al.
Veröffentlicht: (2025)
Are complicated loss functions necessary for teaching LLMs to reason?
von: Carrino, Gabriele, et al.
Veröffentlicht: (2026)
von: Carrino, Gabriele, et al.
Veröffentlicht: (2026)
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
von: Sprague, Zayne, et al.
Veröffentlicht: (2024)
von: Sprague, Zayne, et al.
Veröffentlicht: (2024)
Your thoughts tell who you are: Characterize the reasoning patterns of LRMs
von: Chen, Yida, et al.
Veröffentlicht: (2025)
von: Chen, Yida, et al.
Veröffentlicht: (2025)
Language models show human-like content effects on reasoning tasks
von: Dasgupta, Ishita, et al.
Veröffentlicht: (2022)
von: Dasgupta, Ishita, et al.
Veröffentlicht: (2022)
Neural machine translation of clinical procedure codes for medical diagnosis and uncertainty quantification
von: Chung, Pei-Hung, et al.
Veröffentlicht: (2024)
von: Chung, Pei-Hung, et al.
Veröffentlicht: (2024)
Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models
von: Sim, Shamus, et al.
Veröffentlicht: (2024)
von: Sim, Shamus, et al.
Veröffentlicht: (2024)
Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
von: Koishekenov, Yeskendir, et al.
Veröffentlicht: (2025)
von: Koishekenov, Yeskendir, et al.
Veröffentlicht: (2025)
A Statistical Framework for Data-dependent Retrieval-Augmented Models
von: Basu, Soumya, et al.
Veröffentlicht: (2024)
von: Basu, Soumya, et al.
Veröffentlicht: (2024)
Counterfactual reasoning: an analysis of in-context emergence
von: Miller, Moritz, et al.
Veröffentlicht: (2025)
von: Miller, Moritz, et al.
Veröffentlicht: (2025)
Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning
von: Sclar, Melanie, et al.
Veröffentlicht: (2024)
von: Sclar, Melanie, et al.
Veröffentlicht: (2024)
Multi-step retrieval and reasoning improves radiology question answering with large language models
von: Wind, Sebastian, et al.
Veröffentlicht: (2025)
von: Wind, Sebastian, et al.
Veröffentlicht: (2025)
LLMs cannot find reasoning errors, but can correct them given the error location
von: Tyen, Gladys, et al.
Veröffentlicht: (2023)
von: Tyen, Gladys, et al.
Veröffentlicht: (2023)
QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
Catching rationalization in the act: detecting motivated reasoning before and after CoT via activation probing
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2026)
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2026)
BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
von: Betley, Jan, et al.
Veröffentlicht: (2025)
von: Betley, Jan, et al.
Veröffentlicht: (2025)
Large Language Model Confidence Estimation via Black-Box Access
von: Pedapati, Tejaswini, et al.
Veröffentlicht: (2024)
von: Pedapati, Tejaswini, et al.
Veröffentlicht: (2024)
BertaQA: How Much Do Language Models Know About Local Culture?
von: Etxaniz, Julen, et al.
Veröffentlicht: (2024)
von: Etxaniz, Julen, et al.
Veröffentlicht: (2024)
Augmenting Lateral Thinking in Language Models with Humor and Riddle Data for the BRAINTEASER Task
von: Ghashami, Mina, et al.
Veröffentlicht: (2024)
von: Ghashami, Mina, et al.
Veröffentlicht: (2024)
Language hooks: a modular framework for augmenting LLM reasoning that decouples tool usage from the model and its prompt
von: de Mijolla, Damien, et al.
Veröffentlicht: (2024)
von: de Mijolla, Damien, et al.
Veröffentlicht: (2024)
Enabling robots to follow abstract instructions and complete complex dynamic tasks
von: Mon-Williams, Ruaridh, et al.
Veröffentlicht: (2024)
von: Mon-Williams, Ruaridh, et al.
Veröffentlicht: (2024)
Towards Efficient Neurally-Guided Program Induction for ARC-AGI
von: Ouellette, Simon
Veröffentlicht: (2024)
von: Ouellette, Simon
Veröffentlicht: (2024)
Deep learning and abstractive summarisation for radiological reports: an empirical study for adapting the PEGASUS models' family with scarce data
von: Benzoni, Claudio, et al.
Veröffentlicht: (2025)
von: Benzoni, Claudio, et al.
Veröffentlicht: (2025)
HiTZ at VarDial 2025 NorSID: Overcoming Data Scarcity with Language Transfer and Automatic Data Annotation
von: Bengoetxea, Jaione, et al.
Veröffentlicht: (2024)
von: Bengoetxea, Jaione, et al.
Veröffentlicht: (2024)
Towards Linguistic Neural Representation Learning and Sentence Retrieval from Electroencephalogram Recordings
von: Zhou, Jinzhao, et al.
Veröffentlicht: (2024)
von: Zhou, Jinzhao, et al.
Veröffentlicht: (2024)
ONNX-Net: Towards Universal Representations and Instant Performance Prediction for Neural Architectures
von: Qin, Shiwen, et al.
Veröffentlicht: (2025)
von: Qin, Shiwen, et al.
Veröffentlicht: (2025)
KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
Entertainment chatbot for the digital inclusion of elderly people without abstraction capabilities
von: García-Méndez, Silvia, et al.
Veröffentlicht: (2024)
von: García-Méndez, Silvia, et al.
Veröffentlicht: (2024)
SLOT: Structuring the Output of Large Language Models
von: Wang, Darren Yow-Bang, et al.
Veröffentlicht: (2025)
von: Wang, Darren Yow-Bang, et al.
Veröffentlicht: (2025)
Latxa: An Open Language Model and Evaluation Suite for Basque
von: Etxaniz, Julen, et al.
Veröffentlicht: (2024)
von: Etxaniz, Julen, et al.
Veröffentlicht: (2024)
Evil twins are not that evil: Qualitative insights into machine-generated prompts
von: Rakotonirina, Nathanaël Carraz, et al.
Veröffentlicht: (2024)
von: Rakotonirina, Nathanaël Carraz, et al.
Veröffentlicht: (2024)
PORTool: Importance-Aware Policy Optimization with Rewarded Tree for Multi-Tool-Integrated Reasoning
von: Wu, Feijie, et al.
Veröffentlicht: (2025)
von: Wu, Feijie, et al.
Veröffentlicht: (2025)
Explainable machine learning multi-label classification of Spanish legal judgements
von: de Arriba-Pérez, Francisco, et al.
Veröffentlicht: (2024)
von: de Arriba-Pérez, Francisco, et al.
Veröffentlicht: (2024)
Latent Reasoning in TRMs is Secretly a Policy Improvement Operator
von: Asadulaev, Arip, et al.
Veröffentlicht: (2025)
von: Asadulaev, Arip, et al.
Veröffentlicht: (2025)
Exposing propaganda: an analysis of stylistic cues comparing human annotations and machine classification
von: Faye, Géraud, et al.
Veröffentlicht: (2024)
von: Faye, Géraud, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Skill Issues: An Analysis of CS:GO Skill Rating Systems
von: Bober-Irizar, Mikel, et al.
Veröffentlicht: (2024) -
When can transformers reason with abstract symbols?
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2023) -
Artificial Expert Intelligence through PAC-reasoning
von: Shalev-Shwartz, Shai, et al.
Veröffentlicht: (2024) -
Is continuous CoT better suited for multi-lingual reasoning?
von: Bashir, Ali Hamza, et al.
Veröffentlicht: (2026) -
Sudoku-Bench: Evaluating creative reasoning with Sudoku variants
von: Seely, Jeffrey, et al.
Veröffentlicht: (2025)