Long Live The Balance: Information Bottleneck Driven Tree-based Policy Optimization
Fuente:
arXiv
Guardado en:
| Autores principales: | Jiang, Hao, Li, Shurui, Bu, Tianpeng, Xu, Bowen, Liu, Xin, Chen, Qihua, Duan, Hongtao, Hu, Lulu, Yang, Bin, Zhang, Minying |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents
por: Bu, Tianpeng, et al.
Publicado: (2026)
por: Bu, Tianpeng, et al.
Publicado: (2026)
Mixture of Balanced Information Bottlenecks for Long-Tailed Visual Recognition
por: Lan, Yifan, et al.
Publicado: (2025)
por: Lan, Yifan, et al.
Publicado: (2025)
D-CORE: Incentivizing Task Decomposition in Large Reasoning Models for Complex Tool Use
por: Xu, Bowen, et al.
Publicado: (2026)
por: Xu, Bowen, et al.
Publicado: (2026)
Research on Personalized Medical Intervention Strategy Generation System based on Group Relative Policy Optimization and Time-Series Data Fusion
por: Lu, Dingxin, et al.
Publicado: (2025)
por: Lu, Dingxin, et al.
Publicado: (2025)
Transformer Enhanced Relation Classification: A Comparative Analysis of Contextuality, Data Efficiency and Sequence Complexity
por: Jing, Bowen, et al.
Publicado: (2025)
por: Jing, Bowen, et al.
Publicado: (2025)
IIB-LPO: Latent Policy Optimization via Iterative Information Bottleneck
por: Deng, Huilin, et al.
Publicado: (2026)
por: Deng, Huilin, et al.
Publicado: (2026)
Expression Syntax Information Bottleneck for Math Word Problems
por: Xiong, Jing, et al.
Publicado: (2023)
por: Xiong, Jing, et al.
Publicado: (2023)
SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention
por: Xu, Hongtao, et al.
Publicado: (2026)
por: Xu, Hongtao, et al.
Publicado: (2026)
Information-Bottleneck Driven Binary Neural Network for Change Detection
por: Yin, Kaijie, et al.
Publicado: (2025)
por: Yin, Kaijie, et al.
Publicado: (2025)
Abandoning the Prize: Local Policy Termination of Science and Technology Awards in China
por: Hu Xi, et al.
Publicado: (2026)
por: Hu Xi, et al.
Publicado: (2026)
Debiasing Graph Representation Learning based on Information Bottleneck
por: Zhang, Ziyi, et al.
Publicado: (2024)
por: Zhang, Ziyi, et al.
Publicado: (2024)
La RoSA: Enhancing LLM Efficiency via Layerwise Rotated Sparse Activation
por: Liu, Kai, et al.
Publicado: (2025)
por: Liu, Kai, et al.
Publicado: (2025)
Mixture-of-Instructions: Aligning Large Language Models via Mixture Prompting
por: Xu, Bowen, et al.
Publicado: (2024)
por: Xu, Bowen, et al.
Publicado: (2024)
To be Green Is to Live Forever: The Impact of Environmental Information Types on Green Consumption Behavior
por: Huiying Zhang, et al.
Publicado: (2024)
por: Huiying Zhang, et al.
Publicado: (2024)
Enhancing Adversarial Transferability via Information Bottleneck Constraints
por: Qi, Biqing, et al.
Publicado: (2024)
por: Qi, Biqing, et al.
Publicado: (2024)
Repeated Modifications in Policy Diffusion: Evidence From River Management Regulations in China's Cities
por: Hu Xi, et al.
Publicado: (2026)
por: Hu Xi, et al.
Publicado: (2026)
Learning Unsupervised Gaze Representation via Eye Mask Driven Information Bottleneck
por: Jiang, Yangzhou, et al.
Publicado: (2024)
por: Jiang, Yangzhou, et al.
Publicado: (2024)
Revisiting Counterfactual Regression through the Lens of Gromov-Wasserstein Information Bottleneck
por: Yang, Hao, et al.
Publicado: (2024)
por: Yang, Hao, et al.
Publicado: (2024)
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale
por: Bu, Tianci, et al.
Publicado: (2026)
por: Bu, Tianci, et al.
Publicado: (2026)
Leveraging LLM Agents for Automated Optimization Modeling for SASP Problems: A Graph-RAG based Approach
por: Pan, Tianpeng, et al.
Publicado: (2025)
por: Pan, Tianpeng, et al.
Publicado: (2025)
Dynamic path planning for mobile robots based on artificial potential field enhanced improved multiobjective snake optimization (APF‐IMOSO)
por: Qilin Li, et al.
Publicado: (2024)
por: Qilin Li, et al.
Publicado: (2024)
Information Must Flow: Recursive Bootstrapping for Information Bottleneck in Optimal Transport
por: Li, Xin
Publicado: (2025)
por: Li, Xin
Publicado: (2025)
Beautimeter: Harnessing GPT for Assessing Architectural and Urban Beauty based on the 15 Properties of Living Structure
por: Jiang, Bin
Publicado: (2024)
por: Jiang, Bin
Publicado: (2024)
Design of Uniform Polymer Nanoparticles via Living Crystallization Driven Self‐Assembly
por: Bowen Zheng, et al.
Publicado: (2025)
por: Bowen Zheng, et al.
Publicado: (2025)
Protecting Your LLMs with Information Bottleneck
por: Liu, Zichuan, et al.
Publicado: (2024)
por: Liu, Zichuan, et al.
Publicado: (2024)
Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck
por: Hu, Zhetao, et al.
Publicado: (2026)
por: Hu, Zhetao, et al.
Publicado: (2026)
Amplify Graph Learning for Recommendation via Sparsity Completion
por: Yuan, Peng, et al.
Publicado: (2024)
por: Yuan, Peng, et al.
Publicado: (2024)
Human-assisted Robotic Policy Refinement via Action Preference Optimization
por: Xia, Wenke, et al.
Publicado: (2025)
por: Xia, Wenke, et al.
Publicado: (2025)
Prototypical Information Bottlenecking and Disentangling for Multimodal Cancer Survival Prediction
por: Zhang, Yilan, et al.
Publicado: (2024)
por: Zhang, Yilan, et al.
Publicado: (2024)
TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
por: Tang, Canhui, et al.
Publicado: (2025)
por: Tang, Canhui, et al.
Publicado: (2025)
MASQuant: Modality-Aware Smoothing Quantization for Multimodal Large Language Models
por: Hu, Lulu, et al.
Publicado: (2026)
por: Hu, Lulu, et al.
Publicado: (2026)
DDTime: Dataset Distillation with Spectral Alignment and Information Bottleneck for Time-Series Forecasting
por: Li, Yuqi, et al.
Publicado: (2025)
por: Li, Yuqi, et al.
Publicado: (2025)
Agentic Entropy-Balanced Policy Optimization
por: Dong, Guanting, et al.
Publicado: (2025)
por: Dong, Guanting, et al.
Publicado: (2025)
Projection Head is Secretly an Information Bottleneck
por: Ouyang, Zhuo, et al.
Publicado: (2025)
por: Ouyang, Zhuo, et al.
Publicado: (2025)
Perfect AI Mimicry and the Epistemology of Consciousness: A Solipsistic Dilemma
por: Li, Shurui
Publicado: (2025)
por: Li, Shurui
Publicado: (2025)
Interpretable Prototype-based Graph Information Bottleneck
por: Seo, Sangwoo, et al.
Publicado: (2023)
por: Seo, Sangwoo, et al.
Publicado: (2023)
The Devil is in the Sources! Knowledge Enhanced Cross-Domain Recommendation in an Information Bottleneck Perspective
por: Hu, Binbin, et al.
Publicado: (2024)
por: Hu, Binbin, et al.
Publicado: (2024)
Adaptive Redundancy Regulation for Balanced Multimodal Information Refinement
por: Yang, Zhe, et al.
Publicado: (2025)
por: Yang, Zhe, et al.
Publicado: (2025)
Dual-branch Distilled Transformer for Efficient Asymmetric UAV Tracking
por: Yang, Hongtao, et al.
Publicado: (2026)
por: Yang, Hongtao, et al.
Publicado: (2026)
Enhancing Transferability of Targeted Adversarial Examples: A Self-Universal Perspective
por: Peng, Bowen, et al.
Publicado: (2024)
por: Peng, Bowen, et al.
Publicado: (2024)
Ejemplares similares
-
Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents
por: Bu, Tianpeng, et al.
Publicado: (2026) -
Mixture of Balanced Information Bottlenecks for Long-Tailed Visual Recognition
por: Lan, Yifan, et al.
Publicado: (2025) -
D-CORE: Incentivizing Task Decomposition in Large Reasoning Models for Complex Tool Use
por: Xu, Bowen, et al.
Publicado: (2026) -
Research on Personalized Medical Intervention Strategy Generation System based on Group Relative Policy Optimization and Time-Series Data Fusion
por: Lu, Dingxin, et al.
Publicado: (2025) -
Transformer Enhanced Relation Classification: A Comparative Analysis of Contextuality, Data Efficiency and Sequence Complexity
por: Jing, Bowen, et al.
Publicado: (2025)