Language models fail at extended rule following
Fuente:
arXiv
Saved in:
| Main Authors: | Dai, Tianxiang, Fan, Jonathan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Characterizing and Optimizing the Spatial Kernel of Multi Resolution Hash Encodings
by: Dai, Tianxiang, et al.
Published: (2026)
by: Dai, Tianxiang, et al.
Published: (2026)
FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMs
by: Fan, Zhiting, et al.
Published: (2024)
by: Fan, Zhiting, et al.
Published: (2024)
Negation Neglect: When models fail to learn negations in training
by: Mayne, Harry, et al.
Published: (2026)
by: Mayne, Harry, et al.
Published: (2026)
In-Memory Learning: A Declarative Learning Framework for Large Language Models
by: Wang, Bo, et al.
Published: (2024)
by: Wang, Bo, et al.
Published: (2024)
Chinese-Vicuna: A Chinese Instruction-following Llama-based Model
by: Fan, Chenghao, et al.
Published: (2025)
by: Fan, Chenghao, et al.
Published: (2025)
SQL-PaLM: Improved Large Language Model Adaptation for Text-to-SQL (extended)
by: Sun, Ruoxi, et al.
Published: (2023)
by: Sun, Ruoxi, et al.
Published: (2023)
Large Language Model for Multi-Domain Translation: Benchmarking and Domain CoT Fine-tuning
by: Hu, Tianxiang, et al.
Published: (2024)
by: Hu, Tianxiang, et al.
Published: (2024)
Dissecting the Ullman Variations with a SCALPEL: Why do LLMs fail at Trivial Alterations to the False Belief Task?
by: Pi, Zhiqiang, et al.
Published: (2024)
by: Pi, Zhiqiang, et al.
Published: (2024)
HardMTBench: Stress-Testing Chinese-English Translation on Knowledge-Intensive Domains
by: Li, Zheng, et al.
Published: (2026)
by: Li, Zheng, et al.
Published: (2026)
Emotion-Cause Pair Extraction in Conversations via Semantic Decoupling and Graph Alignment
by: Ma, Tianxiang, et al.
Published: (2026)
by: Ma, Tianxiang, et al.
Published: (2026)
AI as a deliberative partner fosters intercultural empathy for Americans but fails for Latin American participants
by: Villanueva, Isabel, et al.
Published: (2025)
by: Villanueva, Isabel, et al.
Published: (2025)
Enhance Graph Alignment for Large Language Models
by: Luo, Haitong, et al.
Published: (2024)
by: Luo, Haitong, et al.
Published: (2024)
Distilling Instruction-following Abilities of Large Language Models with Task-aware Curriculum Planning
by: Yue, Yuanhao, et al.
Published: (2024)
by: Yue, Yuanhao, et al.
Published: (2024)
Learnable Privacy Neurons Localization in Language Models
by: Chen, Ruizhe, et al.
Published: (2024)
by: Chen, Ruizhe, et al.
Published: (2024)
InstructPatentGPT: Training patent language models to follow instructions with human feedback
by: Lee, Jieh-Sheng
Published: (2024)
by: Lee, Jieh-Sheng
Published: (2024)
Evidence of conceptual mastery in the application of rules by Large Language Models
by: Nunes, José Luiz, et al.
Published: (2025)
by: Nunes, José Luiz, et al.
Published: (2025)
From Metaphor to Mechanism: How LLMs Decode Traditional Chinese Medicine Symbolic Language for Modern Clinical Relevance
by: Tang, Jiacheng, et al.
Published: (2025)
by: Tang, Jiacheng, et al.
Published: (2025)
Quantifying and extending the coverage of spatial categorization data sets
by: Li, Wanchun, et al.
Published: (2026)
by: Li, Wanchun, et al.
Published: (2026)
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
by: Berglund, Lukas, et al.
Published: (2023)
by: Berglund, Lukas, et al.
Published: (2023)
Scrambled text: training Language Models to correct OCR errors using synthetic data
by: Bourne, Jonathan
Published: (2024)
by: Bourne, Jonathan
Published: (2024)
CodeDelegator: Mitigating Context Pollution via Role Separation in Code-as-Action Agents
by: Fei, Tianxiang, et al.
Published: (2026)
by: Fei, Tianxiang, et al.
Published: (2026)
MoL-RL: Distilling Multi-Step Environmental Feedback into LLMs for Feedback-Independent Reasoning
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
Geographically-Informed Language Identification
by: Dunn, Jonathan, et al.
Published: (2024)
by: Dunn, Jonathan, et al.
Published: (2024)
Representational Analysis of Binding in Language Models
by: Dai, Qin, et al.
Published: (2024)
by: Dai, Qin, et al.
Published: (2024)
A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models
by: Ren, Weijieying, et al.
Published: (2025)
by: Ren, Weijieying, et al.
Published: (2025)
Breach in the Shield: Unveiling the Vulnerabilities of Large Language Models
by: Dai, Runpeng, et al.
Published: (2025)
by: Dai, Runpeng, et al.
Published: (2025)
Language Models Encode the Value of Numbers Linearly
by: Zhu, Fangwei, et al.
Published: (2024)
by: Zhu, Fangwei, et al.
Published: (2024)
Autoregressive Large Language Models are Computationally Universal
by: Schuurmans, Dale, et al.
Published: (2024)
by: Schuurmans, Dale, et al.
Published: (2024)
Cell-Based Representation of Relational Binding in Language Models
by: Dai, Qin, et al.
Published: (2026)
by: Dai, Qin, et al.
Published: (2026)
Retrieval-Pretrained Transformer: Long-range Language Modeling with Self-retrieval
by: Rubin, Ohad, et al.
Published: (2023)
by: Rubin, Ohad, et al.
Published: (2023)
We can still parse using syntactic rules
by: Hussein, Ghaly
Published: (2026)
by: Hussein, Ghaly
Published: (2026)
Agent Alignment in Evolving Social Norms
by: Li, Shimin, et al.
Published: (2024)
by: Li, Shimin, et al.
Published: (2024)
Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance
by: Ye, Jiasheng, et al.
Published: (2024)
by: Ye, Jiasheng, et al.
Published: (2024)
Reinforcement learning fine-tuning of language model for instruction following and math reasoning
by: Han, Yifu, et al.
Published: (2025)
by: Han, Yifu, et al.
Published: (2025)
Hybrid and Unitary PEFT for Resource-Efficient Large Language Models
by: Qi, Haomin, et al.
Published: (2025)
by: Qi, Haomin, et al.
Published: (2025)
Where Vision Becomes Text: Locating the OCR Routing Bottleneck in Vision-Language Models
by: Steinberg, Jonathan, et al.
Published: (2026)
by: Steinberg, Jonathan, et al.
Published: (2026)
Towards Efficient Vision-Language Tuning: More Information Density, More Generalizability
by: Hao, Tianxiang, et al.
Published: (2023)
by: Hao, Tianxiang, et al.
Published: (2023)
Generating Diverse Training Samples for Relation Extraction with Large Language Models
by: Li, Zexuan, et al.
Published: (2025)
by: Li, Zexuan, et al.
Published: (2025)
Large Language Model for Extracting Complex Contract Information in Industrial Scenes
by: Cao, Yunyang, et al.
Published: (2025)
by: Cao, Yunyang, et al.
Published: (2025)
UniICL: An Efficient Unified Framework Unifying Compression, Selection, and Generation
by: Gao, Jun, et al.
Published: (2024)
by: Gao, Jun, et al.
Published: (2024)
Similar Items
-
Characterizing and Optimizing the Spatial Kernel of Multi Resolution Hash Encodings
by: Dai, Tianxiang, et al.
Published: (2026) -
FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMs
by: Fan, Zhiting, et al.
Published: (2024) -
Negation Neglect: When models fail to learn negations in training
by: Mayne, Harry, et al.
Published: (2026) -
In-Memory Learning: A Declarative Learning Framework for Large Language Models
by: Wang, Bo, et al.
Published: (2024) -
Chinese-Vicuna: A Chinese Instruction-following Llama-based Model
by: Fan, Chenghao, et al.
Published: (2025)