Learning to Answer from Correct Demonstrations
Fuente:
arXiv
Guardado en:
| Autores principales: | Joshi, Nirmit, Li, Gene, Bhandari, Siddharth, Kasiviswanathan, Shiva Prasad, Ma, Cong, Srebro, Nathan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Learning to Think from Multiple Thinkers
por: Joshi, Nirmit, et al.
Publicado: (2026)
por: Joshi, Nirmit, et al.
Publicado: (2026)
Debiasing Reward Models by Representation Learning with Guarantees
por: Ng, Ignavier, et al.
Publicado: (2025)
por: Ng, Ignavier, et al.
Publicado: (2025)
A Quantitative Characterization of Forgetting in Post-Training
por: Balasubramanian, Krishnakumar, et al.
Publicado: (2026)
por: Balasubramanian, Krishnakumar, et al.
Publicado: (2026)
A Theory of Learning with Autoregressive Chain of Thought
por: Joshi, Nirmit, et al.
Publicado: (2025)
por: Joshi, Nirmit, et al.
Publicado: (2025)
Noisy Interpolation Learning with Shallow Univariate ReLU Networks
por: Joshi, Nirmit, et al.
Publicado: (2023)
por: Joshi, Nirmit, et al.
Publicado: (2023)
A Classical View on Benign Overfitting: The Role of Sample Size
por: Park, Junhyung, et al.
Publicado: (2025)
por: Park, Junhyung, et al.
Publicado: (2025)
Benign Overfitting for Regression with Trained Two-Layer ReLU Networks
por: Park, Junhyung, et al.
Publicado: (2024)
por: Park, Junhyung, et al.
Publicado: (2024)
On the Complexity of Learning Sparse Functions with Statistical and Gradient Queries
por: Joshi, Nirmit, et al.
Publicado: (2024)
por: Joshi, Nirmit, et al.
Publicado: (2024)
Training Large Language Models To Reason In Parallel With Global Forking Tokens
por: Jia, Sheng, et al.
Publicado: (2025)
por: Jia, Sheng, et al.
Publicado: (2025)
Learning single-index models via harmonic decomposition
por: Joshi, Nirmit, et al.
Publicado: (2025)
por: Joshi, Nirmit, et al.
Publicado: (2025)
What Causes Postoperative Aspiration?
por: Nagesh, Supriya, et al.
Publicado: (2025)
por: Nagesh, Supriya, et al.
Publicado: (2025)
Research Program: Theory of Learning in Dynamical Systems
por: Hazan, Elad, et al.
Publicado: (2025)
por: Hazan, Elad, et al.
Publicado: (2025)
From Guess2Graph: When and How Can Unreliable Experts Safely Boost Causal Discovery in Finite Samples?
por: Hiremath, Sujai, et al.
Publicado: (2025)
por: Hiremath, Sujai, et al.
Publicado: (2025)
Data-Efficient Contrastive Self-supervised Learning: Most Beneficial Examples for Supervised Learning Contribute the Least
por: Joshi, Siddharth, et al.
Publicado: (2023)
por: Joshi, Siddharth, et al.
Publicado: (2023)
Agnostic Reinforcement Learning: Foundations and Algorithms
por: Li, Gene
Publicado: (2025)
por: Li, Gene
Publicado: (2025)
Limits of Approximating the Median Treatment Effect
por: Addanki, Raghavendra, et al.
Publicado: (2024)
por: Addanki, Raghavendra, et al.
Publicado: (2024)
When Correct Demonstrations Hurt: Rethinking the Role of Exemplars in In-Context Learning
por: Qiu, Chenghao, et al.
Publicado: (2026)
por: Qiu, Chenghao, et al.
Publicado: (2026)
The Role of Environment Access in Agnostic Reinforcement Learning
por: Krishnamurthy, Akshay, et al.
Publicado: (2025)
por: Krishnamurthy, Akshay, et al.
Publicado: (2025)
Optimistic Rates for Learning from Label Proportions
por: Li, Gene, et al.
Publicado: (2024)
por: Li, Gene, et al.
Publicado: (2024)
Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting
por: Lin, Hongxiang, et al.
Publicado: (2026)
por: Lin, Hongxiang, et al.
Publicado: (2026)
Dependence-Aware Label Aggregation for LLM-as-a-Judge via Ising Models
por: Balasubramanian, Krishnakumar, et al.
Publicado: (2026)
por: Balasubramanian, Krishnakumar, et al.
Publicado: (2026)
Learning Quadruped Walking from Seconds of Demonstration
por: Zhang, Ruipeng, et al.
Publicado: (2026)
por: Zhang, Ruipeng, et al.
Publicado: (2026)
Are Human-generated Demonstrations Necessary for In-context Learning?
por: Li, Rui, et al.
Publicado: (2023)
por: Li, Rui, et al.
Publicado: (2023)
Proving that Cryptic Crossword Clue Answers are Correct
por: Andrews, Martin, et al.
Publicado: (2024)
por: Andrews, Martin, et al.
Publicado: (2024)
Modeling Feature Maps for Quantum Machine Learning
por: Singh, Navneet, et al.
Publicado: (2025)
por: Singh, Navneet, et al.
Publicado: (2025)
Differentially Private Conditional Independence Testing
por: Kalemaj, Iden, et al.
Publicado: (2023)
por: Kalemaj, Iden, et al.
Publicado: (2023)
An Independent Implementation of Quantum Machine Learning Algorithms in Qiskit for Genomic Data
por: Singh, Navneet, et al.
Publicado: (2024)
por: Singh, Navneet, et al.
Publicado: (2024)
Can We Predict the Unpredictable? Leveraging DisasterNet-LLM for Multimodal Disaster Classification
por: Kulahara, Manaswi, et al.
Publicado: (2025)
por: Kulahara, Manaswi, et al.
Publicado: (2025)
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
por: Guo, Yipin, et al.
Publicado: (2026)
por: Guo, Yipin, et al.
Publicado: (2026)
Optimal Bayesian Stopping for Efficient Inference of Consistent LLM Answers
por: Huang, Jingkai, et al.
Publicado: (2026)
por: Huang, Jingkai, et al.
Publicado: (2026)
New Insights on Unfolding and Fine-tuning Quantum Federated Learning
por: Nanayakkara, Shanika Iroshi, et al.
Publicado: (2025)
por: Nanayakkara, Shanika Iroshi, et al.
Publicado: (2025)
Learning Safety Constraints from Demonstrations with Unknown Rewards
por: Lindner, David, et al.
Publicado: (2023)
por: Lindner, David, et al.
Publicado: (2023)
Learning Parameterized Skills from Demonstrations
por: Gupta, Vedant, et al.
Publicado: (2025)
por: Gupta, Vedant, et al.
Publicado: (2025)
Learning to Select In-Context Demonstration Preferred by Large Language Model
por: Zhang, Zheng, et al.
Publicado: (2025)
por: Zhang, Zheng, et al.
Publicado: (2025)
A Comprehensive Machine Learning Framework for Heart Disease Prediction: Performance Evaluation and Future Perspectives
por: Lamir, Ali Azimi, et al.
Publicado: (2025)
por: Lamir, Ali Azimi, et al.
Publicado: (2025)
Hierarchical Apprenticeship Learning from Imperfect Demonstrations with Evolving Rewards
por: Islam, Md Mirajul, et al.
Publicado: (2026)
por: Islam, Md Mirajul, et al.
Publicado: (2026)
Skill-Enhanced Reinforcement Learning Acceleration from Heterogeneous Demonstrations
por: Zhang, Hanping, et al.
Publicado: (2024)
por: Zhang, Hanping, et al.
Publicado: (2024)
Learning Causally Invariant Reward Functions from Diverse Demonstrations
por: Ovinnikov, Ivan, et al.
Publicado: (2024)
por: Ovinnikov, Ivan, et al.
Publicado: (2024)
From Answers to Arguments: Toward Trustworthy Clinical Diagnostic Reasoning with Toulmin-Guided Curriculum Goal-Conditioned Learning
por: Zhan, Chen, et al.
Publicado: (2026)
por: Zhan, Chen, et al.
Publicado: (2026)
Imitation Learning from Suboptimal Demonstrations via Meta-Learning An Action Ranker
por: Fan, Jiangdong, et al.
Publicado: (2024)
por: Fan, Jiangdong, et al.
Publicado: (2024)
Ejemplares similares
-
Learning to Think from Multiple Thinkers
por: Joshi, Nirmit, et al.
Publicado: (2026) -
Debiasing Reward Models by Representation Learning with Guarantees
por: Ng, Ignavier, et al.
Publicado: (2025) -
A Quantitative Characterization of Forgetting in Post-Training
por: Balasubramanian, Krishnakumar, et al.
Publicado: (2026) -
A Theory of Learning with Autoregressive Chain of Thought
por: Joshi, Nirmit, et al.
Publicado: (2025) -
Noisy Interpolation Learning with Shallow Univariate ReLU Networks
por: Joshi, Nirmit, et al.
Publicado: (2023)