Teaching by Failure: Counter-Example-Driven Curricula for Transformer Self-Improvement
Fuente:
arXiv
Guardado en:
| Autor principal: | Vejendla, Harshil |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SliceMoE: Routing Embedding Slices Instead of Tokens for Fine-Grained and Balanced Transformer Scaling
por: Vejendla, Harshil
Publicado: (2025)
por: Vejendla, Harshil
Publicado: (2025)
RewriteNets: End-to-End Trainable String-Rewriting for Generative Sequence Modeling
por: Vejendla, Harshil
Publicado: (2026)
por: Vejendla, Harshil
Publicado: (2026)
Learning to Predict Chaos: Curriculum-Driven Training for Robust Forecasting of Chaotic Dynamics
por: Vejendla, Harshil
Publicado: (2025)
por: Vejendla, Harshil
Publicado: (2025)
LATTA: Langevin-Anchored Test-Time Adaptation for Enhanced Robustness and Stability
por: Vejendla, Harshil
Publicado: (2025)
por: Vejendla, Harshil
Publicado: (2025)
H1B-KV: Hybrid One-Bit Caches for Memory-Efficient Large Language Model Inference
por: Vejendla, Harshil
Publicado: (2025)
por: Vejendla, Harshil
Publicado: (2025)
Wave-PDE Nets: Trainable Wave-Equation Layers as an Alternative to Attention
por: Vejendla, Harshil
Publicado: (2025)
por: Vejendla, Harshil
Publicado: (2025)
Drift-Adapter: A Practical Approach to Near Zero-Downtime Embedding Model Upgrades in Vector Databases
por: Vejendla, Harshil
Publicado: (2025)
por: Vejendla, Harshil
Publicado: (2025)
Thermodynamics of Reinforcement Learning Curricula
por: Adamczyk, Jacob, et al.
Publicado: (2026)
por: Adamczyk, Jacob, et al.
Publicado: (2026)
Birdie: Advancing State Space Models with Reward-Driven Objectives and Curricula
por: Blouir, Sam, et al.
Publicado: (2024)
por: Blouir, Sam, et al.
Publicado: (2024)
Dist2ill: Distributional Distillation for One-Pass Uncertainty Estimation in Large Language Models
por: Zhao, Yicong, et al.
Publicado: (2025)
por: Zhao, Yicong, et al.
Publicado: (2025)
Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
por: Hübotter, Jonas, et al.
Publicado: (2025)
por: Hübotter, Jonas, et al.
Publicado: (2025)
Data Warmup: Complexity-Aware Curricula for Efficient Diffusion Training
por: Lin, Jinhong, et al.
Publicado: (2026)
por: Lin, Jinhong, et al.
Publicado: (2026)
Towards a Novel Perspective on Adversarial Examples Driven by Frequency
por: Zhang, Zhun, et al.
Publicado: (2024)
por: Zhang, Zhun, et al.
Publicado: (2024)
Toward In-Context Teaching: Adapting Examples to Students' Misconceptions
por: Ross, Alexis, et al.
Publicado: (2024)
por: Ross, Alexis, et al.
Publicado: (2024)
Curricula for Learning Robust Policies with Factored State Representations in Changing Environments
por: Panayiotou, Panayiotis, et al.
Publicado: (2024)
por: Panayiotou, Panayiotis, et al.
Publicado: (2024)
Alert-ME: An Explainability-Driven Defense Against Adversarial Examples in Transformer-Based Text Classification
por: Sabir, Bushra, et al.
Publicado: (2023)
por: Sabir, Bushra, et al.
Publicado: (2023)
DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning
por: Diaz-Bone, Leander, et al.
Publicado: (2025)
por: Diaz-Bone, Leander, et al.
Publicado: (2025)
A Deep Dive into Scaling RL for Code Generation with Synthetic Data and Curricula
por: Sancaktar, Cansu, et al.
Publicado: (2026)
por: Sancaktar, Cansu, et al.
Publicado: (2026)
Vision-Language Model Dialog Games for Self-Improvement
por: Konyushkova, Ksenia, et al.
Publicado: (2025)
por: Konyushkova, Ksenia, et al.
Publicado: (2025)
Provable and Practical In-Context Policy Optimization for Self-Improvement
por: Yu, Tianrun, et al.
Publicado: (2026)
por: Yu, Tianrun, et al.
Publicado: (2026)
Continual Driving Policy Optimization with Closed-Loop Individualized Curricula
por: Niu, Haoyi, et al.
Publicado: (2023)
por: Niu, Haoyi, et al.
Publicado: (2023)
Transformers Provably Implement In-Context Reinforcement Learning with Policy Improvement
por: Liang, Haodong, et al.
Publicado: (2026)
por: Liang, Haodong, et al.
Publicado: (2026)
Distinct Computations Emerge From Compositional Curricula in In-Context Learning
por: Lee, Jin Hwa, et al.
Publicado: (2025)
por: Lee, Jin Hwa, et al.
Publicado: (2025)
CF-VLM:CounterFactual Vision-Language Fine-tuning
por: Zhang, Jusheng, et al.
Publicado: (2025)
por: Zhang, Jusheng, et al.
Publicado: (2025)
Derivation of Closed Form of Expected Improvement for Gaussian Process Trained on Log-Transformed Objective
por: Watanabe, Shuhei
Publicado: (2024)
por: Watanabe, Shuhei
Publicado: (2024)
Self-Trained Verification for Training- and Test-Time Self-Improvement
por: Wu, Chen Henry, et al.
Publicado: (2026)
por: Wu, Chen Henry, et al.
Publicado: (2026)
Data-Efficient Contrastive Self-supervised Learning: Most Beneficial Examples for Supervised Learning Contribute the Least
por: Joshi, Siddharth, et al.
Publicado: (2023)
por: Joshi, Siddharth, et al.
Publicado: (2023)
Long-Term Prediction Accuracy Improvement of Data-Driven Medium-Range Global Weather Forecast
por: Hu, Yifan, et al.
Publicado: (2024)
por: Hu, Yifan, et al.
Publicado: (2024)
SIME: Enhancing Policy Self-Improvement with Modal-level Exploration
por: Jin, Yang, et al.
Publicado: (2025)
por: Jin, Yang, et al.
Publicado: (2025)
PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases
por: Vuddanti, Sri Vatsa, et al.
Publicado: (2025)
por: Vuddanti, Sri Vatsa, et al.
Publicado: (2025)
WIST: Web-Grounded Iterative Self-Play Tree for Domain-Targeted Reasoning Improvement
por: Li, Fangyuan, et al.
Publicado: (2026)
por: Li, Fangyuan, et al.
Publicado: (2026)
Self-Improvement in Language Models: The Sharpening Mechanism
por: Huang, Audrey, et al.
Publicado: (2024)
por: Huang, Audrey, et al.
Publicado: (2024)
CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples
por: Zhang, Jianrui, et al.
Publicado: (2024)
por: Zhang, Jianrui, et al.
Publicado: (2024)
Learning to Move Like Professional Counter-Strike Players
por: Durst, David, et al.
Publicado: (2024)
por: Durst, David, et al.
Publicado: (2024)
Debiasing Machine Unlearning with Counterfactual Examples
por: Chen, Ziheng, et al.
Publicado: (2024)
por: Chen, Ziheng, et al.
Publicado: (2024)
FairSTG: Countering performance heterogeneity via collaborative sample-level optimization
por: Lin, Gengyu, et al.
Publicado: (2024)
por: Lin, Gengyu, et al.
Publicado: (2024)
Latent Principle Discovery for Language Model Self-Improvement
por: Ramji, Keshav, et al.
Publicado: (2025)
por: Ramji, Keshav, et al.
Publicado: (2025)
Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains
por: Subramaniam, Vighnesh, et al.
Publicado: (2025)
por: Subramaniam, Vighnesh, et al.
Publicado: (2025)
Adaptive Machine Learning-Driven Multi-Fidelity Stratified Sampling for Failure Analysis of Nonlinear Stochastic Systems
por: Xu, Liuyun, et al.
Publicado: (2025)
por: Xu, Liuyun, et al.
Publicado: (2025)
Self-Improvement as Coherence Optimization: A Theoretical Account
por: Qiu, Tianyi, et al.
Publicado: (2026)
por: Qiu, Tianyi, et al.
Publicado: (2026)
Ejemplares similares
-
SliceMoE: Routing Embedding Slices Instead of Tokens for Fine-Grained and Balanced Transformer Scaling
por: Vejendla, Harshil
Publicado: (2025) -
RewriteNets: End-to-End Trainable String-Rewriting for Generative Sequence Modeling
por: Vejendla, Harshil
Publicado: (2026) -
Learning to Predict Chaos: Curriculum-Driven Training for Robust Forecasting of Chaotic Dynamics
por: Vejendla, Harshil
Publicado: (2025) -
LATTA: Langevin-Anchored Test-Time Adaptation for Enhanced Robustness and Stability
por: Vejendla, Harshil
Publicado: (2025) -
H1B-KV: Hybrid One-Bit Caches for Memory-Efficient Large Language Model Inference
por: Vejendla, Harshil
Publicado: (2025)