Meta-TTRL: A Metacognitive Framework for Self-Improving Test-Time Reinforcement Learning in Unified Multimodal Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Tan, Lit Sin, Chen, Junzhe, Fu, Xiaolong, Ma, Lichen, Huang, Junshi, Shi, Jianzhong, Li, Yan, Wen, Lijie |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
TTRL: Test-Time Reinforcement Learning
par: Zuo, Yuxin, et autres
Publié: (2025)
par: Zuo, Yuxin, et autres
Publié: (2025)
UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing
par: Ma, Lichen, et autres
Publié: (2026)
par: Ma, Lichen, et autres
Publié: (2026)
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning
par: Zhang, Haoyu, et autres
Publié: (2025)
par: Zhang, Haoyu, et autres
Publié: (2025)
Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling
par: Fu, Xiaolong, et autres
Publié: (2025)
par: Fu, Xiaolong, et autres
Publié: (2025)
CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models
par: Hosseini, Peyman, et autres
Publié: (2025)
par: Hosseini, Peyman, et autres
Publié: (2025)
FrequencyBooster: Full-Frequency Modeling for High-Fidelity Pixel Diffusion
par: Ma, Lichen, et autres
Publié: (2026)
par: Ma, Lichen, et autres
Publié: (2026)
Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition
par: He, Yu, et autres
Publié: (2026)
par: He, Yu, et autres
Publié: (2026)
CoT2-Meta: Budgeted Metacognitive Control for Test-Time Reasoning
par: Ma, Siyuan, et autres
Publié: (2026)
par: Ma, Siyuan, et autres
Publié: (2026)
Beyond Meta-Reasoning: Metacognitive Consolidation for Self-Improving LLM Reasoning
par: Zhuang, Ziqing, et autres
Publié: (2026)
par: Zhuang, Ziqing, et autres
Publié: (2026)
HyperDiT: Hyper-Connected Transformers for High-Fidelity Pixel-Space Diffusion
par: He, Yu, et autres
Publié: (2026)
par: He, Yu, et autres
Publié: (2026)
RePainter: Empowering E-commerce Object Removal via Spatial-matting Reinforcement Learning
par: Guo, Zipeng, et autres
Publié: (2025)
par: Guo, Zipeng, et autres
Publié: (2025)
SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards
par: Hong, Jixiang, et autres
Publié: (2025)
par: Hong, Jixiang, et autres
Publié: (2025)
Irec: A Metacognitive Scaffolding for Self-Regulated Learning through Just-in-Time Insight Recall: A Conceptual Framework and System Prototype
par: Hou, Xuefei, et autres
Publié: (2025)
par: Hou, Xuefei, et autres
Publié: (2025)
Do We Really Need External Tools to Mitigate Hallucinations? SIRA: Shared-Prefix Internal Reconstruction of Attribution
par: Qin, Tian, et autres
Publié: (2026)
par: Qin, Tian, et autres
Publié: (2026)
UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
par: Mao, Weijia, et autres
Publié: (2025)
par: Mao, Weijia, et autres
Publié: (2025)
Metacognitive Sensitivity for Test-Time Dynamic Model Selection
par: Trinh, Le Tuan Minh, et autres
Publié: (2025)
par: Trinh, Le Tuan Minh, et autres
Publié: (2025)
UniT: Unified Multimodal Chain-of-Thought Test-time Scaling
par: Chen, Leon Liangyu, et autres
Publié: (2026)
par: Chen, Leon Liangyu, et autres
Publié: (2026)
MetaCogAgent: A Metacognitive Multi-Agent LLM Framework with Self-Aware Task Delegation
par: Wang, Chenyu, et autres
Publié: (2026)
par: Wang, Chenyu, et autres
Publié: (2026)
LiWi: Layering in the Wild
par: He, Yu, et autres
Publié: (2026)
par: He, Yu, et autres
Publié: (2026)
MetaCLASS: Metacognitive Coaching for Learning with Adaptive Self-regulation Support
par: Liu, Naiming, et autres
Publié: (2026)
par: Liu, Naiming, et autres
Publié: (2026)
Truly Self-Improving Agents Require Intrinsic Metacognitive Learning
par: Liu, Tennison, et autres
Publié: (2025)
par: Liu, Tennison, et autres
Publié: (2025)
MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs
par: Shan, Liang, et autres
Publié: (2025)
par: Shan, Liang, et autres
Publié: (2025)
Exploring the Output of Software Testing Tools through a Visual Comparative Analysis
par: Lit, Brandon, et autres
Publié: (2026)
par: Lit, Brandon, et autres
Publié: (2026)
The Debate on the Dietary Guidelines for Americans (2025–2030) and Implications for China's Nutritional Policy
par: Junshi Chen
Publié: (2026)
par: Junshi Chen
Publié: (2026)
Test-Time Meta-Adaptation with Self-Synthesis
par: Kaya, Zeyneb N., et autres
Publié: (2026)
par: Kaya, Zeyneb N., et autres
Publié: (2026)
Test-Time Regret Minimization in Meta Reinforcement Learning
par: Mutti, Mirco, et autres
Publié: (2024)
par: Mutti, Mirco, et autres
Publié: (2024)
Curiosity and Metacognition: Towards a Unified Framework for Learning and Education in the Age of AI
par: Desvaux, Chloé, et autres
Publié: (2026)
par: Desvaux, Chloé, et autres
Publié: (2026)
MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
par: Xiao, Zilin, et autres
Publié: (2025)
par: Xiao, Zilin, et autres
Publié: (2025)
The PCP-like Theorem for Sub-linear Time Inapproximability
par: Ma, Hengzhao, et autres
Publié: (2021)
par: Ma, Hengzhao, et autres
Publié: (2021)
ImAgent: A Unified Multimodal Agent Framework for Test-Time Scalable Image Generation
par: Wang, Kaishen, et autres
Publié: (2025)
par: Wang, Kaishen, et autres
Publié: (2025)
UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision
par: Han, Ruiyan, et autres
Publié: (2026)
par: Han, Ruiyan, et autres
Publié: (2026)
A Unified Information-Theoretic Framework for Meta-Learning Generalization
par: Wen, Wen, et autres
Publié: (2025)
par: Wen, Wen, et autres
Publié: (2025)
Meta Fusion: A Unified Framework For Multimodality Fusion with Mutual Learning
par: Liang, Ziyi, et autres
Publié: (2025)
par: Liang, Ziyi, et autres
Publié: (2025)
PER-DPP Sampling Framework and Its Application in Path Planning
par: Wang, Junzhe
Publié: (2025)
par: Wang, Junzhe
Publié: (2025)
Constrained Meta Reinforcement Learning with Provable Test-Time Safety
par: Ni, Tingting, et autres
Publié: (2026)
par: Ni, Tingting, et autres
Publié: (2026)
Grouped Competition Test with Unified False Discovery Rate Control
par: Deng, Mingzhou, et autres
Publié: (2025)
par: Deng, Mingzhou, et autres
Publié: (2025)
MUSE: Manipulating Unified Framework for Synthesizing Emotions in Images via Test-Time Optimization
par: Xia, Yingjie, et autres
Publié: (2025)
par: Xia, Yingjie, et autres
Publié: (2025)
Self-Improving LLM Agents at Test-Time
par: Acikgoz, Emre Can, et autres
Publié: (2025)
par: Acikgoz, Emre Can, et autres
Publié: (2025)
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
par: Jiang, Jingjing, et autres
Publié: (2025)
par: Jiang, Jingjing, et autres
Publié: (2025)
UR$^2$: Unify RAG and Reasoning through Reinforcement Learning
par: Li, Weitao, et autres
Publié: (2025)
par: Li, Weitao, et autres
Publié: (2025)
Documents similaires
-
TTRL: Test-Time Reinforcement Learning
par: Zuo, Yuxin, et autres
Publié: (2025) -
UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing
par: Ma, Lichen, et autres
Publié: (2026) -
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning
par: Zhang, Haoyu, et autres
Publié: (2025) -
Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling
par: Fu, Xiaolong, et autres
Publié: (2025) -
CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models
par: Hosseini, Peyman, et autres
Publié: (2025)