Learning to Discover at Test Time
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yuksekgonul, Mert, Koceja, Daniel, Li, Xinhao, Bianchi, Federico, McCaleb, Jed, Wang, Xiaolong, Kautz, Jan, Choi, Yejin, Zou, James, Guestrin, Carlos, Sun, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
End-to-End Test-Time Training for Long Context
von: Tandon, Arnuv, et al.
Veröffentlicht: (2025)
von: Tandon, Arnuv, et al.
Veröffentlicht: (2025)
metaTextGrad: Automatically optimizing language model optimizers
von: Xu, Guowei, et al.
Veröffentlicht: (2025)
von: Xu, Guowei, et al.
Veröffentlicht: (2025)
Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory
von: Suzgun, Mirac, et al.
Veröffentlicht: (2025)
von: Suzgun, Mirac, et al.
Veröffentlicht: (2025)
TextGrad: Automatic "Differentiation" via Text
von: Yuksekgonul, Mert, et al.
Veröffentlicht: (2024)
von: Yuksekgonul, Mert, et al.
Veröffentlicht: (2024)
One-Minute Video Generation with Test-Time Training
von: Dalal, Karan, et al.
Veröffentlicht: (2025)
von: Dalal, Karan, et al.
Veröffentlicht: (2025)
Thinking agents for zero-shot generalization to qualitatively novel tasks
von: Miconi, Thomas, et al.
Veröffentlicht: (2025)
von: Miconi, Thomas, et al.
Veröffentlicht: (2025)
Inefficiencies of Meta Agents for Agent Design
von: El, Batu, et al.
Veröffentlicht: (2025)
von: El, Batu, et al.
Veröffentlicht: (2025)
Sparse Reward Subsystem in Large Language Models
von: Xu, Guowei, et al.
Veröffentlicht: (2026)
von: Xu, Guowei, et al.
Veröffentlicht: (2026)
Learning to (Learn at Test Time)
von: Sun, Yu, et al.
Veröffentlicht: (2023)
von: Sun, Yu, et al.
Veröffentlicht: (2023)
How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis
von: Bianchi, Federico, et al.
Veröffentlicht: (2024)
von: Bianchi, Federico, et al.
Veröffentlicht: (2024)
SiriuS: Self-improving Multi-agent Systems via Bootstrapped Reasoning
von: Zhao, Wanjia, et al.
Veröffentlicht: (2025)
von: Zhao, Wanjia, et al.
Veröffentlicht: (2025)
Goal-Directed Search Outperforms Goal-Agnostic Memory Compression in Long-Context Memory Tasks
von: Zheng, Yicong, et al.
Veröffentlicht: (2025)
von: Zheng, Yicong, et al.
Veröffentlicht: (2025)
Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2026)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2026)
Learning to (Learn at Test Time): RNNs with Expressive Hidden States
von: Sun, Yu, et al.
Veröffentlicht: (2024)
von: Sun, Yu, et al.
Veröffentlicht: (2024)
Discovering Implicit Large Language Model Alignment Objectives
von: Chen, Edward, et al.
Veröffentlicht: (2026)
von: Chen, Edward, et al.
Veröffentlicht: (2026)
Cost-of-Pass: An Economic Framework for Evaluating Language Models
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2025)
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2025)
Model Equality Testing: Which Model Is This API Serving?
von: Gao, Irena, et al.
Veröffentlicht: (2024)
von: Gao, Irena, et al.
Veröffentlicht: (2024)
Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content
von: Bianchi, Federico, et al.
Veröffentlicht: (2024)
von: Bianchi, Federico, et al.
Veröffentlicht: (2024)
Diversity of Thought Improves Reasoning Abilities of LLMs
von: Naik, Ranjita, et al.
Veröffentlicht: (2023)
von: Naik, Ranjita, et al.
Veröffentlicht: (2023)
Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
von: Thakkar, Nitya, et al.
Veröffentlicht: (2025)
von: Thakkar, Nitya, et al.
Veröffentlicht: (2025)
Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the Wild
von: Mireshghallah, Niloofar, et al.
Veröffentlicht: (2024)
von: Mireshghallah, Niloofar, et al.
Veröffentlicht: (2024)
Societal Impacts Research Requires Benchmarks for Creative Composition Tasks
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2025)
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2025)
Benchmarking Distributional Alignment of Large Language Models
von: Meister, Nicole, et al.
Veröffentlicht: (2024)
von: Meister, Nicole, et al.
Veröffentlicht: (2024)
The Extractive-Abstractive Spectrum: Uncovering Verifiability Trade-offs in LLM Generations
von: Worledge, Theodora, et al.
Veröffentlicht: (2024)
von: Worledge, Theodora, et al.
Veröffentlicht: (2024)
Exploring the use of AI authors and reviewers at Agents4Science
von: Bianchi, Federico, et al.
Veröffentlicht: (2025)
von: Bianchi, Federico, et al.
Veröffentlicht: (2025)
LoRA-TTT: Low-Rank Test-Time Training for Vision-Language Models
von: Kojima, Yuto, et al.
Veröffentlicht: (2025)
von: Kojima, Yuto, et al.
Veröffentlicht: (2025)
iGRPO: Self-Feedback-Driven LLM Reasoning
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2026)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2026)
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
von: Liu, Mingjie, et al.
Veröffentlicht: (2025)
von: Liu, Mingjie, et al.
Veröffentlicht: (2025)
Thinking While Listening: Simple Test Time Scaling For Audio Classification
von: Verma, Prateek, et al.
Veröffentlicht: (2025)
von: Verma, Prateek, et al.
Veröffentlicht: (2025)
"Sorry, I Didn't Catch That": How Speech Models Miss What Matters Most
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2026)
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2026)
Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions
von: Acuna, David, et al.
Veröffentlicht: (2025)
von: Acuna, David, et al.
Veröffentlicht: (2025)
Peer Review in Nursing: A Guide for Early Career Scholars
von: Hye Ri Choi, et al.
Veröffentlicht: (2025)
von: Hye Ri Choi, et al.
Veröffentlicht: (2025)
A ghost mechanism: An analytical model of abrupt learning in recurrent networks
von: Dinc, Fatih, et al.
Veröffentlicht: (2025)
von: Dinc, Fatih, et al.
Veröffentlicht: (2025)
Can LLMs Reason with Rules? Logic Scaffolding for Stress-Testing and Improving LLMs
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
Test-Time Training on Video Streams
von: Wang, Renhao, et al.
Veröffentlicht: (2023)
von: Wang, Renhao, et al.
Veröffentlicht: (2023)
RLP: Reinforcement as a Pretraining Objective
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2025)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2025)
Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs
von: Hong, Yining, et al.
Veröffentlicht: (2026)
von: Hong, Yining, et al.
Veröffentlicht: (2026)
Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning
von: Yu, Qinan, et al.
Veröffentlicht: (2026)
von: Yu, Qinan, et al.
Veröffentlicht: (2026)
ACORN: Performant and Predicate-Agnostic Search Over Vector Embeddings and Structured Data
von: Patel, Liana, et al.
Veröffentlicht: (2024)
von: Patel, Liana, et al.
Veröffentlicht: (2024)
Opt-ICL at LeWiDi-2025: Maximizing In-Context Signal from Rater Examples via Meta-Learning
von: Sorensen, Taylor, et al.
Veröffentlicht: (2025)
von: Sorensen, Taylor, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
End-to-End Test-Time Training for Long Context
von: Tandon, Arnuv, et al.
Veröffentlicht: (2025) -
metaTextGrad: Automatically optimizing language model optimizers
von: Xu, Guowei, et al.
Veröffentlicht: (2025) -
Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory
von: Suzgun, Mirac, et al.
Veröffentlicht: (2025) -
TextGrad: Automatic "Differentiation" via Text
von: Yuksekgonul, Mert, et al.
Veröffentlicht: (2024) -
One-Minute Video Generation with Test-Time Training
von: Dalal, Karan, et al.
Veröffentlicht: (2025)