Predicting generalization performance with correctness discriminators
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Yuekun, Koller, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Simple and effective data augmentation for compositional generalization
by: Yao, Yuekun, et al.
Published: (2024)
by: Yao, Yuekun, et al.
Published: (2024)
Language models can learn implicit multi-hop reasoning, but only if they have lots of training data
by: Yao, Yuekun, et al.
Published: (2025)
by: Yao, Yuekun, et al.
Published: (2025)
Barriers to Universal Reasoning With Transformers (And How to Overcome Them)
by: Kraus, Oliver, et al.
Published: (2026)
by: Kraus, Oliver, et al.
Published: (2026)
Fine-grained Controllable Text Generation through In-context Learning with Feedback
by: Thillainathan, Sarubi, et al.
Published: (2024)
by: Thillainathan, Sarubi, et al.
Published: (2024)
A Survey on Complex Tasks for Goal-Directed Interactive Agents
by: Hartmann, Mareike, et al.
Published: (2024)
by: Hartmann, Mareike, et al.
Published: (2024)
A Knapsack by Any Other Name: Presentation impacts LLM performance on NP-hard problems
by: Duchnowski, Alex, et al.
Published: (2025)
by: Duchnowski, Alex, et al.
Published: (2025)
Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMs
by: Yang, Xiulin, et al.
Published: (2025)
by: Yang, Xiulin, et al.
Published: (2025)
SIP: Injecting a Structural Inductive Bias into a Seq2Seq Model by Simulation
by: Lindemann, Matthias, et al.
Published: (2023)
by: Lindemann, Matthias, et al.
Published: (2023)
LLMs syntactically adapt their language use to their conversational partner
by: Kandra, Florian, et al.
Published: (2025)
by: Kandra, Florian, et al.
Published: (2025)
Collaborative Problem-Solving in an Optimization Game
by: Jeknic, Isidora, et al.
Published: (2025)
by: Jeknic, Isidora, et al.
Published: (2025)
A Dialogue Game for Eliciting Balanced Collaboration
by: Jeknić, Isidora, et al.
Published: (2024)
by: Jeknić, Isidora, et al.
Published: (2024)
Strengthening Structural Inductive Biases by Pre-training to Perform Syntactic Transformations
by: Lindemann, Matthias, et al.
Published: (2024)
by: Lindemann, Matthias, et al.
Published: (2024)
Positional Biases Shift as Inputs Approach Context Window Limits
by: Veseli, Blerta, et al.
Published: (2025)
by: Veseli, Blerta, et al.
Published: (2025)
Evaluating Spatiotemporal Consistency in Automatically Generated Sewing Instructions
by: Geiger, Luisa, et al.
Published: (2025)
by: Geiger, Luisa, et al.
Published: (2025)
Scope-enhanced Compositional Semantic Parsing for DRT
by: Yang, Xiulin, et al.
Published: (2024)
by: Yang, Xiulin, et al.
Published: (2024)
Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
by: Kohli, Harsh, et al.
Published: (2026)
by: Kohli, Harsh, et al.
Published: (2026)
Automating the Generation of Prompts for LLM-based Action Choice in PDDL Planning
by: Stein, Katharina, et al.
Published: (2023)
by: Stein, Katharina, et al.
Published: (2023)
LLMs Generate Kitsch
by: Klinge, Xenia, et al.
Published: (2026)
by: Klinge, Xenia, et al.
Published: (2026)
AuthorMix: Modular Authorship Style Transfer via Layer-wise Adapter Mixing
by: Thillainathan, Sarubi, et al.
Published: (2026)
by: Thillainathan, Sarubi, et al.
Published: (2026)
Reason to Rote: Rethinking Memorization in Reasoning
by: Du, Yupei, et al.
Published: (2025)
by: Du, Yupei, et al.
Published: (2025)
Improved Generalized Planning with LLMs through Strategy Refinement and Reflection
by: Stein, Katharina, et al.
Published: (2025)
by: Stein, Katharina, et al.
Published: (2025)
Greater accessibility can amplify discrimination in generative AI
by: Holtermann, Carolin, et al.
Published: (2026)
by: Holtermann, Carolin, et al.
Published: (2026)
Articulatory strategy in vowel production as a basis for speaker discrimination
by: Lo, Justin J. H., et al.
Published: (2025)
by: Lo, Justin J. H., et al.
Published: (2025)
Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests
by: Momentè, Filippo, et al.
Published: (2025)
by: Momentè, Filippo, et al.
Published: (2025)
On the Ability of Transformers to Verify Plans
by: Sarrof, Yash, et al.
Published: (2026)
by: Sarrof, Yash, et al.
Published: (2026)
ADaPT: As-Needed Decomposition and Planning with Language Models
by: Prasad, Archiki, et al.
Published: (2023)
by: Prasad, Archiki, et al.
Published: (2023)
Self-correction is Not An Innate Capability in Language Models
by: Liu, Guangliang, et al.
Published: (2024)
by: Liu, Guangliang, et al.
Published: (2024)
Tag and correct: high precision post-editing approach to correction of speech recognition errors
by: Ziętkiewicz, Tomasz
Published: (2024)
by: Ziętkiewicz, Tomasz
Published: (2024)
Predicting Text Preference Via Structured Comparative Reasoning
by: Yan, Jing Nathan, et al.
Published: (2023)
by: Yan, Jing Nathan, et al.
Published: (2023)
Cost Analysis of Human-corrected Transcription for Predominately Oral Languages
by: Diarra, Yacouba, et al.
Published: (2025)
by: Diarra, Yacouba, et al.
Published: (2025)
Self-Debias: Self-correcting for Debiasing Large Language Models
by: Feng, Xuan, et al.
Published: (2026)
by: Feng, Xuan, et al.
Published: (2026)
fastabx: A library for efficient computation of ABX discriminability
by: Poli, Maxime, et al.
Published: (2025)
by: Poli, Maxime, et al.
Published: (2025)
Efficient and Effective Internal Memory Retrieval for LLM-Based Healthcare Prediction
by: Li, Mingchen, et al.
Published: (2026)
by: Li, Mingchen, et al.
Published: (2026)
Exploring LLMs for Predicting Tutor Strategy and Student Outcomes in Dialogues
by: Ikram, Fareya, et al.
Published: (2025)
by: Ikram, Fareya, et al.
Published: (2025)
You Shall Know a Tool by the Traces it Leaves: The Predictability of Sentiment Analysis Tools
by: Baumartz, Daniel, et al.
Published: (2024)
by: Baumartz, Daniel, et al.
Published: (2024)
Chain-of-Though (CoT) prompting strategies for medical error detection and correction
by: Wu, Zhaolong, et al.
Published: (2024)
by: Wu, Zhaolong, et al.
Published: (2024)
CECOR: Correction-oriented synthetic data construction for factual error correction
by: Zhu, Lei, et al.
Published: (2026)
by: Zhu, Lei, et al.
Published: (2026)
Generalizing Fairness to Generative Language Models via Reformulation of Non-discrimination Criteria
by: Sterlie, Sara, et al.
Published: (2024)
by: Sterlie, Sara, et al.
Published: (2024)
Characterizing Language Use in a Collaborative Situated Game
by: Tomlin, Nicholas, et al.
Published: (2025)
by: Tomlin, Nicholas, et al.
Published: (2025)
Intrinsic Self-correction for Enhanced Morality: An Analysis of Internal Mechanisms and the Superficial Hypothesis
by: Liu, Guangliang, et al.
Published: (2024)
by: Liu, Guangliang, et al.
Published: (2024)
Similar Items
-
Simple and effective data augmentation for compositional generalization
by: Yao, Yuekun, et al.
Published: (2024) -
Language models can learn implicit multi-hop reasoning, but only if they have lots of training data
by: Yao, Yuekun, et al.
Published: (2025) -
Barriers to Universal Reasoning With Transformers (And How to Overcome Them)
by: Kraus, Oliver, et al.
Published: (2026) -
Fine-grained Controllable Text Generation through In-context Learning with Feedback
by: Thillainathan, Sarubi, et al.
Published: (2024) -
A Survey on Complex Tasks for Goal-Directed Interactive Agents
by: Hartmann, Mareike, et al.
Published: (2024)