TR-ICRL: Test-Time Rethinking for In-Context Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Wenxuan, Zuo, Yuxin, Zhang, Zijian, Wu, Xuecheng, Fan, Zining, Liu, Wenxuan, Chen, Li, Li, Xiaoyu, Cao, Xuezhi, Jin, Xiaolong, Liu, Ninghao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ICRL: Learning to Internalize Self-Critique with Reinforcement Learning
by: Lin, Jianbo, et al.
Published: (2026)
by: Lin, Jianbo, et al.
Published: (2026)
KnowCoder-X: Boosting Multilingual Information Extraction via Code
by: Zuo, Yuxin, et al.
Published: (2024)
by: Zuo, Yuxin, et al.
Published: (2024)
Testing APS conjecture on regular graphs
by: Tao, Wenxuan, et al.
Published: (2025)
by: Tao, Wenxuan, et al.
Published: (2025)
TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs
by: Zhang, Zijian, et al.
Published: (2025)
by: Zhang, Zijian, et al.
Published: (2025)
Towards Event Extraction with Massive Types: LLM-based Collaborative Annotation and Partitioning Extraction
by: Liu, Wenxuan, et al.
Published: (2025)
by: Liu, Wenxuan, et al.
Published: (2025)
HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs
by: Zhang, Zijian, et al.
Published: (2025)
by: Zhang, Zijian, et al.
Published: (2025)
Dos escritores chinos hablan de literatura
by: Cao Wenxuan
Published: (2003)
by: Cao Wenxuan
Published: (2003)
SWIFT: Mapping Sub-series with Wavelet Decomposition Improves Time Series Forecasting
by: Xie, Wenxuan, et al.
Published: (2025)
by: Xie, Wenxuan, et al.
Published: (2025)
Testing and Evaluation of Large Language Models: Correctness, Non-Toxicity, and Fairness
by: Wang, Wenxuan
Published: (2024)
by: Wang, Wenxuan
Published: (2024)
A Refined Algorithm For the EPR model
by: Tao, Wenxuan, et al.
Published: (2025)
by: Tao, Wenxuan, et al.
Published: (2025)
Reinforcement Learning in a Safety-Embedded MDP with Trajectory Optimization
by: Yang, Fan, et al.
Published: (2023)
by: Yang, Fan, et al.
Published: (2023)
Do passive investors influence corporate social responsibility? A risk‐management perspective
by: Wenxuan Hou, et al.
Published: (2024)
by: Wenxuan Hou, et al.
Published: (2024)
QuIVer: Rethinking ANN Graph Topology via Training-Free Binary Quantization
by: Xiao, Wenxuan, et al.
Published: (2026)
by: Xiao, Wenxuan, et al.
Published: (2026)
Mitigating Data Scarcity in Time Series Analysis: A Foundation Model with Series-Symbol Data Generation
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Synthetic Series-Symbol Data Generation for Time Series Foundation Models
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Matcha: Mitigating Graph Structure Shifts with Test-Time Adaptation
by: Bao, Wenxuan, et al.
Published: (2024)
by: Bao, Wenxuan, et al.
Published: (2024)
TTRL: Test-Time Reinforcement Learning
by: Zuo, Yuxin, et al.
Published: (2025)
by: Zuo, Yuxin, et al.
Published: (2025)
RESC: A Reinforcement Learning Based Search-to-Control Framework for Quadrotor Local Planning in Dense Environments
by: Liu, Zhaohong, et al.
Published: (2024)
by: Liu, Zhaohong, et al.
Published: (2024)
Mint: A Simple Test-Time Adaptation of Vision-Language Models against Common Corruptions
by: Bao, Wenxuan, et al.
Published: (2025)
by: Bao, Wenxuan, et al.
Published: (2025)
Panda: Test-Time Adaptation with Negative Data Augmentation
by: Deng, Ruxi, et al.
Published: (2025)
by: Deng, Ruxi, et al.
Published: (2025)
A Mobile Magnetic Manipulation Platform for Gastrointestinal Navigation with Deep Reinforcement Learning Control
by: Yan, Zhifan, et al.
Published: (2026)
by: Yan, Zhifan, et al.
Published: (2026)
Sim2Real Manipulation on Unknown Objects with Tactile-based Reinforcement Learning
by: Su, Entong, et al.
Published: (2024)
by: Su, Entong, et al.
Published: (2024)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
by: Li, Wenxuan, et al.
Published: (2025)
by: Li, Wenxuan, et al.
Published: (2025)
A Survey of Deep Learning for Geometry Problem Solving
by: Ma, Jianzhe, et al.
Published: (2025)
by: Ma, Jianzhe, et al.
Published: (2025)
Common Inpainted Objects In-N-Out of Context
by: Yang, Tianze, et al.
Published: (2025)
by: Yang, Tianze, et al.
Published: (2025)
Learning to Model Diverse Driving Behaviors in Highly Interactive Autonomous Driving Scenarios with Multi-Agent Reinforcement Learning
by: Weiwei, Liu, et al.
Published: (2024)
by: Weiwei, Liu, et al.
Published: (2024)
Ramen: Robust Test-Time Adaptation of Vision-Language Models with Active Sample Selection
by: Bao, Wenxuan, et al.
Published: (2026)
by: Bao, Wenxuan, et al.
Published: (2026)
HTseaat/AD-MPC: AD-MPC
by: Wenxuan Yu
Published: (2025)
by: Wenxuan Yu
Published: (2025)
Covariance Structure and Coordinate Heterogeneity Govern Binary Quantization of Contrastive Embeddings
by: Xiao, Wenxuan
Published: (2026)
by: Xiao, Wenxuan
Published: (2026)
QGHNN: A quantum graph Hamiltonian neural network
by: Wang, Wenxuan
Published: (2025)
by: Wang, Wenxuan
Published: (2025)
Noise-resistant adaptive Hamiltonian learning
by: Wang, Wenxuan
Published: (2025)
by: Wang, Wenxuan
Published: (2025)
On weak convergence of stochastic wave equation with colored noise on $\mathbb{R}$
by: Tao, Wenxuan
Published: (2024)
by: Tao, Wenxuan
Published: (2024)
Latte: Collaborative Test-Time Adaptation of Vision-Language Models in Federated Learning
by: Bao, Wenxuan, et al.
Published: (2025)
by: Bao, Wenxuan, et al.
Published: (2025)
FreeChunker: A Cross-Granularity Chunking Framework
by: Zhang, Wenxuan, et al.
Published: (2025)
by: Zhang, Wenxuan, et al.
Published: (2025)
Rethinking the Unsolvable: When In-Context Search Meets Test-Time Scaling
by: Xia, Fanzeng, et al.
Published: (2025)
by: Xia, Fanzeng, et al.
Published: (2025)
Time-Series Learning for Proactive Fault Prediction in Distributed Systems with Deep Neural Structures
by: Wang, Yang, et al.
Published: (2025)
by: Wang, Yang, et al.
Published: (2025)
Flexible Full‐Inorganic Ultrathin Films with Stable Circularly Polarized Luminescence Covering the Visible to Near‐Infrared Region
by: Wenxuan Wu, et al.
Published: (2024)
by: Wenxuan Wu, et al.
Published: (2024)
BOBA: Byzantine-Robust Federated Learning with Label Skewness
by: Bao, Wenxuan, et al.
Published: (2022)
by: Bao, Wenxuan, et al.
Published: (2022)
Test and Prediction of ESET Fracture Toughness of Fiber‐Reinforced Composite Under Different Orientations
by: Xuecheng Liu, et al.
Published: (2026)
by: Xuecheng Liu, et al.
Published: (2026)
HQ-DiT: Efficient Diffusion Transformer with FP4 Hybrid Quantization
by: Liu, Wenxuan, et al.
Published: (2024)
by: Liu, Wenxuan, et al.
Published: (2024)
Similar Items
-
ICRL: Learning to Internalize Self-Critique with Reinforcement Learning
by: Lin, Jianbo, et al.
Published: (2026) -
KnowCoder-X: Boosting Multilingual Information Extraction via Code
by: Zuo, Yuxin, et al.
Published: (2024) -
Testing APS conjecture on regular graphs
by: Tao, Wenxuan, et al.
Published: (2025) -
TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs
by: Zhang, Zijian, et al.
Published: (2025) -
Towards Event Extraction with Massive Types: LLM-based Collaborative Annotation and Partitioning Extraction
by: Liu, Wenxuan, et al.
Published: (2025)