TR-ICRL: Test-Time Rethinking for In-Context Reinforcement Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Jiang, Wenxuan, Zuo, Yuxin, Zhang, Zijian, Wu, Xuecheng, Fan, Zining, Liu, Wenxuan, Chen, Li, Li, Xiaoyu, Cao, Xuezhi, Jin, Xiaolong, Liu, Ninghao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ICRL: Learning to Internalize Self-Critique with Reinforcement Learning
di: Lin, Jianbo, et al.
Pubblicazione: (2026)
di: Lin, Jianbo, et al.
Pubblicazione: (2026)
KnowCoder-X: Boosting Multilingual Information Extraction via Code
di: Zuo, Yuxin, et al.
Pubblicazione: (2024)
di: Zuo, Yuxin, et al.
Pubblicazione: (2024)
Testing APS conjecture on regular graphs
di: Tao, Wenxuan, et al.
Pubblicazione: (2025)
di: Tao, Wenxuan, et al.
Pubblicazione: (2025)
TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs
di: Zhang, Zijian, et al.
Pubblicazione: (2025)
di: Zhang, Zijian, et al.
Pubblicazione: (2025)
Towards Event Extraction with Massive Types: LLM-based Collaborative Annotation and Partitioning Extraction
di: Liu, Wenxuan, et al.
Pubblicazione: (2025)
di: Liu, Wenxuan, et al.
Pubblicazione: (2025)
HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs
di: Zhang, Zijian, et al.
Pubblicazione: (2025)
di: Zhang, Zijian, et al.
Pubblicazione: (2025)
Dos escritores chinos hablan de literatura
di: Cao Wenxuan
Pubblicazione: (2003)
di: Cao Wenxuan
Pubblicazione: (2003)
SWIFT: Mapping Sub-series with Wavelet Decomposition Improves Time Series Forecasting
di: Xie, Wenxuan, et al.
Pubblicazione: (2025)
di: Xie, Wenxuan, et al.
Pubblicazione: (2025)
Testing and Evaluation of Large Language Models: Correctness, Non-Toxicity, and Fairness
di: Wang, Wenxuan
Pubblicazione: (2024)
di: Wang, Wenxuan
Pubblicazione: (2024)
A Refined Algorithm For the EPR model
di: Tao, Wenxuan, et al.
Pubblicazione: (2025)
di: Tao, Wenxuan, et al.
Pubblicazione: (2025)
Do passive investors influence corporate social responsibility? A risk‐management perspective
di: Wenxuan Hou, et al.
Pubblicazione: (2024)
di: Wenxuan Hou, et al.
Pubblicazione: (2024)
Reinforcement Learning in a Safety-Embedded MDP with Trajectory Optimization
di: Yang, Fan, et al.
Pubblicazione: (2023)
di: Yang, Fan, et al.
Pubblicazione: (2023)
QuIVer: Rethinking ANN Graph Topology via Training-Free Binary Quantization
di: Xiao, Wenxuan, et al.
Pubblicazione: (2026)
di: Xiao, Wenxuan, et al.
Pubblicazione: (2026)
Mitigating Data Scarcity in Time Series Analysis: A Foundation Model with Series-Symbol Data Generation
di: Wang, Wenxuan, et al.
Pubblicazione: (2025)
di: Wang, Wenxuan, et al.
Pubblicazione: (2025)
Synthetic Series-Symbol Data Generation for Time Series Foundation Models
di: Wang, Wenxuan, et al.
Pubblicazione: (2025)
di: Wang, Wenxuan, et al.
Pubblicazione: (2025)
Matcha: Mitigating Graph Structure Shifts with Test-Time Adaptation
di: Bao, Wenxuan, et al.
Pubblicazione: (2024)
di: Bao, Wenxuan, et al.
Pubblicazione: (2024)
TTRL: Test-Time Reinforcement Learning
di: Zuo, Yuxin, et al.
Pubblicazione: (2025)
di: Zuo, Yuxin, et al.
Pubblicazione: (2025)
RESC: A Reinforcement Learning Based Search-to-Control Framework for Quadrotor Local Planning in Dense Environments
di: Liu, Zhaohong, et al.
Pubblicazione: (2024)
di: Liu, Zhaohong, et al.
Pubblicazione: (2024)
Mint: A Simple Test-Time Adaptation of Vision-Language Models against Common Corruptions
di: Bao, Wenxuan, et al.
Pubblicazione: (2025)
di: Bao, Wenxuan, et al.
Pubblicazione: (2025)
Panda: Test-Time Adaptation with Negative Data Augmentation
di: Deng, Ruxi, et al.
Pubblicazione: (2025)
di: Deng, Ruxi, et al.
Pubblicazione: (2025)
A Mobile Magnetic Manipulation Platform for Gastrointestinal Navigation with Deep Reinforcement Learning Control
di: Yan, Zhifan, et al.
Pubblicazione: (2026)
di: Yan, Zhifan, et al.
Pubblicazione: (2026)
Sim2Real Manipulation on Unknown Objects with Tactile-based Reinforcement Learning
di: Su, Entong, et al.
Pubblicazione: (2024)
di: Su, Entong, et al.
Pubblicazione: (2024)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
di: Li, Wenxuan, et al.
Pubblicazione: (2025)
di: Li, Wenxuan, et al.
Pubblicazione: (2025)
A Survey of Deep Learning for Geometry Problem Solving
di: Ma, Jianzhe, et al.
Pubblicazione: (2025)
di: Ma, Jianzhe, et al.
Pubblicazione: (2025)
Common Inpainted Objects In-N-Out of Context
di: Yang, Tianze, et al.
Pubblicazione: (2025)
di: Yang, Tianze, et al.
Pubblicazione: (2025)
Learning to Model Diverse Driving Behaviors in Highly Interactive Autonomous Driving Scenarios with Multi-Agent Reinforcement Learning
di: Weiwei, Liu, et al.
Pubblicazione: (2024)
di: Weiwei, Liu, et al.
Pubblicazione: (2024)
Ramen: Robust Test-Time Adaptation of Vision-Language Models with Active Sample Selection
di: Bao, Wenxuan, et al.
Pubblicazione: (2026)
di: Bao, Wenxuan, et al.
Pubblicazione: (2026)
HTseaat/AD-MPC: AD-MPC
di: Wenxuan Yu
Pubblicazione: (2025)
di: Wenxuan Yu
Pubblicazione: (2025)
Covariance Structure and Coordinate Heterogeneity Govern Binary Quantization of Contrastive Embeddings
di: Xiao, Wenxuan
Pubblicazione: (2026)
di: Xiao, Wenxuan
Pubblicazione: (2026)
QGHNN: A quantum graph Hamiltonian neural network
di: Wang, Wenxuan
Pubblicazione: (2025)
di: Wang, Wenxuan
Pubblicazione: (2025)
Noise-resistant adaptive Hamiltonian learning
di: Wang, Wenxuan
Pubblicazione: (2025)
di: Wang, Wenxuan
Pubblicazione: (2025)
On weak convergence of stochastic wave equation with colored noise on $\mathbb{R}$
di: Tao, Wenxuan
Pubblicazione: (2024)
di: Tao, Wenxuan
Pubblicazione: (2024)
FreeChunker: A Cross-Granularity Chunking Framework
di: Zhang, Wenxuan, et al.
Pubblicazione: (2025)
di: Zhang, Wenxuan, et al.
Pubblicazione: (2025)
Latte: Collaborative Test-Time Adaptation of Vision-Language Models in Federated Learning
di: Bao, Wenxuan, et al.
Pubblicazione: (2025)
di: Bao, Wenxuan, et al.
Pubblicazione: (2025)
Rethinking the Unsolvable: When In-Context Search Meets Test-Time Scaling
di: Xia, Fanzeng, et al.
Pubblicazione: (2025)
di: Xia, Fanzeng, et al.
Pubblicazione: (2025)
Time-Series Learning for Proactive Fault Prediction in Distributed Systems with Deep Neural Structures
di: Wang, Yang, et al.
Pubblicazione: (2025)
di: Wang, Yang, et al.
Pubblicazione: (2025)
Flexible Full‐Inorganic Ultrathin Films with Stable Circularly Polarized Luminescence Covering the Visible to Near‐Infrared Region
di: Wenxuan Wu, et al.
Pubblicazione: (2024)
di: Wenxuan Wu, et al.
Pubblicazione: (2024)
BOBA: Byzantine-Robust Federated Learning with Label Skewness
di: Bao, Wenxuan, et al.
Pubblicazione: (2022)
di: Bao, Wenxuan, et al.
Pubblicazione: (2022)
Test and Prediction of ESET Fracture Toughness of Fiber‐Reinforced Composite Under Different Orientations
di: Xuecheng Liu, et al.
Pubblicazione: (2026)
di: Xuecheng Liu, et al.
Pubblicazione: (2026)
HQ-DiT: Efficient Diffusion Transformer with FP4 Hybrid Quantization
di: Liu, Wenxuan, et al.
Pubblicazione: (2024)
di: Liu, Wenxuan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ICRL: Learning to Internalize Self-Critique with Reinforcement Learning
di: Lin, Jianbo, et al.
Pubblicazione: (2026) -
KnowCoder-X: Boosting Multilingual Information Extraction via Code
di: Zuo, Yuxin, et al.
Pubblicazione: (2024) -
Testing APS conjecture on regular graphs
di: Tao, Wenxuan, et al.
Pubblicazione: (2025) -
TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs
di: Zhang, Zijian, et al.
Pubblicazione: (2025) -
Towards Event Extraction with Massive Types: LLM-based Collaborative Annotation and Partitioning Extraction
di: Liu, Wenxuan, et al.
Pubblicazione: (2025)