Rule Extrapolation in Language Models: A Study of Compositional Generalization on OOD Prompts
Fuente:
arXiv
Salvato in:
| Autori principali: | Mészáros, Anna, Ujváry, Szilvia, Brendel, Wieland, Reizinger, Patrik, Huszár, Ferenc |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Position: Understanding LLMs Requires More Than Statistical Generalization
di: Reizinger, Patrik, et al.
Pubblicazione: (2024)
di: Reizinger, Patrik, et al.
Pubblicazione: (2024)
Out-of-distribution Tests Reveal Compositionality in Chess Transformers
di: Mészáros, Anna, et al.
Pubblicazione: (2025)
di: Mészáros, Anna, et al.
Pubblicazione: (2025)
Identifiable Exchangeable Mechanisms for Causal Structure and Representation Learning
di: Reizinger, Patrik, et al.
Pubblicazione: (2024)
di: Reizinger, Patrik, et al.
Pubblicazione: (2024)
An Interventional Perspective on Identifiability in Gaussian LTI Systems with Independent Component Analysis
di: Rajendran, Goutham, et al.
Pubblicazione: (2023)
di: Rajendran, Goutham, et al.
Pubblicazione: (2023)
Position: An Empirically Grounded Identifiability Theory Will Accelerate Self-Supervised Learning Research
di: Reizinger, Patrik, et al.
Pubblicazione: (2025)
di: Reizinger, Patrik, et al.
Pubblicazione: (2025)
Estimating Treatment Effects with Independent Component Analysis
di: Reizinger, Patrik, et al.
Pubblicazione: (2025)
di: Reizinger, Patrik, et al.
Pubblicazione: (2025)
InfoNCE: Identifying the Gap Between Theory and Practice
di: Rusak, Evgenia, et al.
Pubblicazione: (2024)
di: Rusak, Evgenia, et al.
Pubblicazione: (2024)
Who Guards the Guardians? The Challenges of Evaluating Identifiability of Learned Representations
di: Joshi, Shruti, et al.
Pubblicazione: (2026)
di: Joshi, Shruti, et al.
Pubblicazione: (2026)
Causality is Key for Interpretability Claims to Generalise
di: Joshi, Shruti, et al.
Pubblicazione: (2026)
di: Joshi, Shruti, et al.
Pubblicazione: (2026)
Skill Learning via Policy Diversity Yields Identifiable Representations for Reinforcement Learning
di: Reizinger, Patrik, et al.
Pubblicazione: (2025)
di: Reizinger, Patrik, et al.
Pubblicazione: (2025)
Cross-Entropy Is All You Need To Invert the Data Generating Process
di: Reizinger, Patrik, et al.
Pubblicazione: (2024)
di: Reizinger, Patrik, et al.
Pubblicazione: (2024)
LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025)
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025)
Learning Beyond Pattern Matching? Assaying Mathematical Understanding in LLMs
di: Guo, Siyuan, et al.
Pubblicazione: (2024)
di: Guo, Siyuan, et al.
Pubblicazione: (2024)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
Large Language Models as Interpolated and Extrapolated Event Predictors
di: Zhang, Libo, et al.
Pubblicazione: (2024)
di: Zhang, Libo, et al.
Pubblicazione: (2024)
Language Generation with Strictly Proper Scoring Rules
di: Shao, Chenze, et al.
Pubblicazione: (2024)
di: Shao, Chenze, et al.
Pubblicazione: (2024)
Prompt Perturbation Consistency Learning for Robust Language Models
di: Qiang, Yao, et al.
Pubblicazione: (2024)
di: Qiang, Yao, et al.
Pubblicazione: (2024)
Active Model Selection for Large Language Models
di: Durmazkeser, Yavuz, et al.
Pubblicazione: (2025)
di: Durmazkeser, Yavuz, et al.
Pubblicazione: (2025)
PolyPrompt: Automating Knowledge Extraction from Multilingual Language Models with Dynamic Prompt Generation
di: Roll, Nathan
Pubblicazione: (2025)
di: Roll, Nathan
Pubblicazione: (2025)
From Interpolation to Extrapolation: Complete Length Generalization for Arithmetic Transformers
di: Duan, Shaoxiong, et al.
Pubblicazione: (2023)
di: Duan, Shaoxiong, et al.
Pubblicazione: (2023)
Large Language Model Selection with Limited Annotations
di: Durmazkeser, Yavuz, et al.
Pubblicazione: (2026)
di: Durmazkeser, Yavuz, et al.
Pubblicazione: (2026)
Model Extrapolation Expedites Alignment
di: Zheng, Chujie, et al.
Pubblicazione: (2024)
di: Zheng, Chujie, et al.
Pubblicazione: (2024)
Enhancing Length Extrapolation in Sequential Models with Pointer-Augmented Neural Memory
di: Le, Hung, et al.
Pubblicazione: (2024)
di: Le, Hung, et al.
Pubblicazione: (2024)
Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models
di: Gao, Bo, et al.
Pubblicazione: (2025)
di: Gao, Bo, et al.
Pubblicazione: (2025)
JoPA:Explaining Large Language Model's Generation via Joint Prompt Attribution
di: Chang, Yurui, et al.
Pubblicazione: (2024)
di: Chang, Yurui, et al.
Pubblicazione: (2024)
Adaptive Prompt Structure Factorization: A Framework for Self-Discovering and Optimizing Compositional Prompt Programs
di: Liu, Haoyue, et al.
Pubblicazione: (2026)
di: Liu, Haoyue, et al.
Pubblicazione: (2026)
Automatic Prompt Selection for Large Language Models
di: Do, Viet-Tung, et al.
Pubblicazione: (2024)
di: Do, Viet-Tung, et al.
Pubblicazione: (2024)
RLPR: Extrapolating RLVR to General Domains without Verifiers
di: Yu, Tianyu, et al.
Pubblicazione: (2025)
di: Yu, Tianyu, et al.
Pubblicazione: (2025)
Beyond Rule-based Named Entity Recognition and Relation Extraction for Process Model Generation from Natural Language Text
di: Neuberger, Julian, et al.
Pubblicazione: (2023)
di: Neuberger, Julian, et al.
Pubblicazione: (2023)
CodeMixBench: Evaluating Large Language Models on Code Generation with Code-Mixed Prompts
di: Sheokand, Manik, et al.
Pubblicazione: (2025)
di: Sheokand, Manik, et al.
Pubblicazione: (2025)
Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks
di: Wiedemer, Thaddäus, et al.
Pubblicazione: (2025)
di: Wiedemer, Thaddäus, et al.
Pubblicazione: (2025)
PromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model Reasoning
di: Zhao, Xueliang, et al.
Pubblicazione: (2025)
di: Zhao, Xueliang, et al.
Pubblicazione: (2025)
Emergent Abilities in Reduced-Scale Generative Language Models
di: Muckatira, Sherin, et al.
Pubblicazione: (2024)
di: Muckatira, Sherin, et al.
Pubblicazione: (2024)
APPL: A Prompt Programming Language for Harmonious Integration of Programs and Large Language Model Prompts
di: Dong, Honghua, et al.
Pubblicazione: (2024)
di: Dong, Honghua, et al.
Pubblicazione: (2024)
Benchmarking Large Language Model Uncertainty for Prompt Optimization
di: Guo, Pei-Fu, et al.
Pubblicazione: (2024)
di: Guo, Pei-Fu, et al.
Pubblicazione: (2024)
Learning Extrapolative Sequence Transformations from Markov Chains
di: Hager, Sophia, et al.
Pubblicazione: (2025)
di: Hager, Sophia, et al.
Pubblicazione: (2025)
PromptKD: Distilling Student-Friendly Knowledge for Generative Language Models via Prompt Tuning
di: Kim, Gyeongman, et al.
Pubblicazione: (2024)
di: Kim, Gyeongman, et al.
Pubblicazione: (2024)
Make Prompts Adaptable: Bayesian Modeling for Vision-Language Prompt Learning with Data-Dependent Prior
di: Cho, Youngjae, et al.
Pubblicazione: (2024)
di: Cho, Youngjae, et al.
Pubblicazione: (2024)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
di: Yang, Wenkai, et al.
Pubblicazione: (2026)
di: Yang, Wenkai, et al.
Pubblicazione: (2026)
Learning Mathematical Rules with Large Language Models
di: Gorceix, Antoine, et al.
Pubblicazione: (2024)
di: Gorceix, Antoine, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Position: Understanding LLMs Requires More Than Statistical Generalization
di: Reizinger, Patrik, et al.
Pubblicazione: (2024) -
Out-of-distribution Tests Reveal Compositionality in Chess Transformers
di: Mészáros, Anna, et al.
Pubblicazione: (2025) -
Identifiable Exchangeable Mechanisms for Causal Structure and Representation Learning
di: Reizinger, Patrik, et al.
Pubblicazione: (2024) -
An Interventional Perspective on Identifiability in Gaussian LTI Systems with Independent Component Analysis
di: Rajendran, Goutham, et al.
Pubblicazione: (2023) -
Position: An Empirically Grounded Identifiability Theory Will Accelerate Self-Supervised Learning Research
di: Reizinger, Patrik, et al.
Pubblicazione: (2025)