It Ain't That Bad: Understanding the Mysterious Performance Drop in OOD Generalization for Generative Transformer Models
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Xingcheng, Pan, Zihao, Zhang, Haipeng, Yang, Yanqing |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Principled Understanding of Generalization for Generative Transformer Models in Arithmetic Reasoning Tasks
by: Xu, Xingcheng, et al.
Published: (2024)
by: Xu, Xingcheng, et al.
Published: (2024)
Capabilities Ain't All You Need: Measuring Propensities in AI
by: Romero-Alvarado, Daniel, et al.
Published: (2026)
by: Romero-Alvarado, Daniel, et al.
Published: (2026)
The Two-Stage Decision-Sampling Hypothesis: Understanding the Emergence of Self-Reflection in RL-Trained LLMs
by: Zhao, Zibo, et al.
Published: (2026)
by: Zhao, Zibo, et al.
Published: (2026)
Raising the Bar in Graph OOD Generalization: Invariant Learning Beyond Explicit Environment Modeling
by: Shen, Xu, et al.
Published: (2025)
by: Shen, Xu, et al.
Published: (2025)
Merging Smarter, Generalizing Better: Enhancing Model Merging on OOD Data
by: Zhang, Bingjie, et al.
Published: (2025)
by: Zhang, Bingjie, et al.
Published: (2025)
Toward a Unified Geometry Understanding: Riemannian Diffusion Framework for Graph Generation and Prediction
by: Gao, Yisen, et al.
Published: (2025)
by: Gao, Yisen, et al.
Published: (2025)
Epistemic Traps: Rational Misalignment Driven by Model Misspecification
by: Xu, Xingcheng, et al.
Published: (2026)
by: Xu, Xingcheng, et al.
Published: (2026)
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models
by: Xu, Xingcheng
Published: (2025)
by: Xu, Xingcheng
Published: (2025)
Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood
by: Yao, Qingmao, et al.
Published: (2025)
by: Yao, Qingmao, et al.
Published: (2025)
Tri-Level Navigator: LLM-Empowered Tri-Level Learning for Time Series OOD Generalization
by: Jian, Chengtao, et al.
Published: (2024)
by: Jian, Chengtao, et al.
Published: (2024)
Quantitative Analysis of Performance Drop in DeepSeek Model Quantization
by: Zhao, Enbo, et al.
Published: (2025)
by: Zhao, Enbo, et al.
Published: (2025)
Understanding the Training and Generalization of Pretrained Transformer for Sequential Decision Making
by: Wang, Hanzhao, et al.
Published: (2024)
by: Wang, Hanzhao, et al.
Published: (2024)
TAO-Net: Two-stage Adaptive OOD Classification Network for Fine-grained Encrypted Traffic Classification
by: Wang, Zihao, et al.
Published: (2025)
by: Wang, Zihao, et al.
Published: (2025)
RAG-GFM: Overcoming In-Memory Bottlenecks in Graph Foundation Models via Retrieval-Augmented Generation
by: Yuan, Haonan, et al.
Published: (2026)
by: Yuan, Haonan, et al.
Published: (2026)
Ain't Misbehavin' -- Using LLMs to Generate Expressive Robot Behavior in Conversations with the Tabletop Robot Haru
by: Wang, Zining, et al.
Published: (2024)
by: Wang, Zining, et al.
Published: (2024)
Three-dimensional attention Transformer for state evaluation in real-time strategy games
by: Ye, Yanqing, et al.
Published: (2025)
by: Ye, Yanqing, et al.
Published: (2025)
Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
by: Kachaev, Nikita, et al.
Published: (2025)
by: Kachaev, Nikita, et al.
Published: (2025)
Does Your Neural Network Extrapolate? Feature Engineering as Identifiability Bias for OOD Generalization
by: Aguilar, Leonel, et al.
Published: (2026)
by: Aguilar, Leonel, et al.
Published: (2026)
Offline Reinforcement Learning with OOD State Correction and OOD Action Suppression
by: Mao, Yixiu, et al.
Published: (2024)
by: Mao, Yixiu, et al.
Published: (2024)
TimeOmni-VL: Unified Models for Time Series Understanding and Generation
by: Guan, Tong, et al.
Published: (2026)
by: Guan, Tong, et al.
Published: (2026)
Population Aware Diffusion for Time Series Generation
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
ODICE: Revealing the Mystery of Distribution Correction Estimation via Orthogonal-gradient Update
by: Mao, Liyuan, et al.
Published: (2024)
by: Mao, Liyuan, et al.
Published: (2024)
Dynamic DropConnect: Enhancing Neural Network Robustness through Adaptive Edge Dropping Strategies
by: Yang, Yuan-Chih, et al.
Published: (2025)
by: Yang, Yuan-Chih, et al.
Published: (2025)
Scalable Ensemble Diversification for OOD Generalization and Detection
by: Rubinstein, Alexander, et al.
Published: (2024)
by: Rubinstein, Alexander, et al.
Published: (2024)
Traj-Transformer: Diffusion Models with Transformer for GPS Trajectory Generation
by: Zhang, Zhiyang, et al.
Published: (2025)
by: Zhang, Zhiyang, et al.
Published: (2025)
Offline Reinforcement Learning with Generative Trajectory Policies
by: Feng, Xinsong, et al.
Published: (2025)
by: Feng, Xinsong, et al.
Published: (2025)
FedSDWC: Federated Synergistic Dual-Representation Weak Causal Learning for OOD
by: Huang, Zhenyuan, et al.
Published: (2025)
by: Huang, Zhenyuan, et al.
Published: (2025)
GraphKeeper: Graph Domain-Incremental Learning via Knowledge Disentanglement and Preservation
by: Guo, Zihao, et al.
Published: (2025)
by: Guo, Zihao, et al.
Published: (2025)
Towards Theoretical Understandings of Self-Consuming Generative Models
by: Fu, Shi, et al.
Published: (2024)
by: Fu, Shi, et al.
Published: (2024)
From Small to Large: Generalization Bounds for Transformers on Variable-Size Inputs
by: Alokhina, Anastasiia, et al.
Published: (2025)
by: Alokhina, Anastasiia, et al.
Published: (2025)
VLBM: Variational Latent Basis Modeling for OOD Robust Multivariate Time Series Forecasting
by: Zhang, Xudong, et al.
Published: (2026)
by: Zhang, Xudong, et al.
Published: (2026)
ADEdgeDrop: Adversarial Edge Dropping for Robust Graph Neural Networks
by: Chen, Zhaoliang, et al.
Published: (2024)
by: Chen, Zhaoliang, et al.
Published: (2024)
GSTM-HMU: Generative Spatio-Temporal Modeling for Human Mobility Understanding
by: Luo, Wenying, et al.
Published: (2025)
by: Luo, Wenying, et al.
Published: (2025)
Large Language Models Badly Generalize across Option Length, Problem Types, and Irrelevant Noun Replacements
by: Zhao, Guangxiang, et al.
Published: (2025)
by: Zhao, Guangxiang, et al.
Published: (2025)
Understanding Generative AI Content with Embedding Models
by: Vargas, Max, et al.
Published: (2024)
by: Vargas, Max, et al.
Published: (2024)
Enhancing Parameter Efficiency and Generalization in Large-Scale Models: A Regularized and Masked Low-Rank Adaptation Approach
by: Mao, Yuzhu, et al.
Published: (2024)
by: Mao, Yuzhu, et al.
Published: (2024)
GraphMoRE: Mitigating Topological Heterogeneity via Mixture of Riemannian Experts
by: Guo, Zihao, et al.
Published: (2024)
by: Guo, Zihao, et al.
Published: (2024)
JTreeformer: Graph-Transformer via Latent-Diffusion Model for Molecular Generation
by: Shi, Ji, et al.
Published: (2025)
by: Shi, Ji, et al.
Published: (2025)
MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design
by: Duanmu, Haojie, et al.
Published: (2025)
by: Duanmu, Haojie, et al.
Published: (2025)
RL Fine-Tuning Heals OOD Forgetting in SFT
by: Jin, Hangzhan, et al.
Published: (2025)
by: Jin, Hangzhan, et al.
Published: (2025)
Similar Items
-
Principled Understanding of Generalization for Generative Transformer Models in Arithmetic Reasoning Tasks
by: Xu, Xingcheng, et al.
Published: (2024) -
Capabilities Ain't All You Need: Measuring Propensities in AI
by: Romero-Alvarado, Daniel, et al.
Published: (2026) -
The Two-Stage Decision-Sampling Hypothesis: Understanding the Emergence of Self-Reflection in RL-Trained LLMs
by: Zhao, Zibo, et al.
Published: (2026) -
Raising the Bar in Graph OOD Generalization: Invariant Learning Beyond Explicit Environment Modeling
by: Shen, Xu, et al.
Published: (2025) -
Merging Smarter, Generalizing Better: Enhancing Model Merging on OOD Data
by: Zhang, Bingjie, et al.
Published: (2025)