Extending Differential Temporal Difference Methods for Episodic Problems
Fuente:
arXiv
Guardado en:
| Autores principales: | De Asis, Kris, Elsayed, Mohamed, He, Jiamin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
An Idiosyncrasy of Time-discretization in Reinforcement Learning
por: De Asis, Kris, et al.
Publicado: (2024)
por: De Asis, Kris, et al.
Publicado: (2024)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
por: Hedar, Abdel-Rahman, et al.
Publicado: (2024)
por: Hedar, Abdel-Rahman, et al.
Publicado: (2024)
ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing
por: Kadu, Ankush, et al.
Publicado: (2025)
por: Kadu, Ankush, et al.
Publicado: (2025)
Machine Learning vs Deep Learning: The Generalization Problem
por: Bay, Yong Yi, et al.
Publicado: (2024)
por: Bay, Yong Yi, et al.
Publicado: (2024)
Extending NGU to Multi-Agent RL: A Preliminary Study
por: Hernandez, Juan, et al.
Publicado: (2025)
por: Hernandez, Juan, et al.
Publicado: (2025)
A Geometric Perspective for High-Dimensional Multiplex Graphs
por: Abdous, Kamel, et al.
Publicado: (2025)
por: Abdous, Kamel, et al.
Publicado: (2025)
How Reliable and Stable are Explanations of XAI Methods?
por: Ribeiro, José, et al.
Publicado: (2024)
por: Ribeiro, José, et al.
Publicado: (2024)
One-vs.-One Mitigation of Intersectional Bias: A General Method to Extend Fairness-Aware Binary Classification
por: Kobayashi, Kenji, et al.
Publicado: (2020)
por: Kobayashi, Kenji, et al.
Publicado: (2020)
Embedded Safety-Aligned Intelligence via Differentiable Internal Alignment Embeddings
por: Rathva, Harsh, et al.
Publicado: (2025)
por: Rathva, Harsh, et al.
Publicado: (2025)
Differentiable Symbolic Planning: A Neural Architecture for Constraint Reasoning with Learned Feasibility
por: Oruganti, Venkatakrishna Reddy
Publicado: (2026)
por: Oruganti, Venkatakrishna Reddy
Publicado: (2026)
Reasoning Large Language Model Errors Arise from Hallucinating Critical Problem Features
por: Heyman, Alex, et al.
Publicado: (2025)
por: Heyman, Alex, et al.
Publicado: (2025)
VN Network: Embedding Newly Emerging Entities with Virtual Neighbors
por: He, Yongquan, et al.
Publicado: (2024)
por: He, Yongquan, et al.
Publicado: (2024)
Fusion-Based Neural Generalization for Predicting Temperature Fields in Industrial PET Preform Heating
por: Alsheikh, Ahmad, et al.
Publicado: (2025)
por: Alsheikh, Ahmad, et al.
Publicado: (2025)
The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary
por: Guo, Dongxin, et al.
Publicado: (2026)
por: Guo, Dongxin, et al.
Publicado: (2026)
ParalESN: Enabling parallel information processing in Reservoir Computing
por: Pinna, Matteo, et al.
Publicado: (2026)
por: Pinna, Matteo, et al.
Publicado: (2026)
Understanding Goal Generalisation in Sequential Reinforcement Learning
por: Brown, Jason Ross, et al.
Publicado: (2026)
por: Brown, Jason Ross, et al.
Publicado: (2026)
What changes after deployment? A survey on On-device Learning in TinyML
por: Pavan, Massimo, et al.
Publicado: (2026)
por: Pavan, Massimo, et al.
Publicado: (2026)
Bounded Ratio Reinforcement Learning
por: Ao, Yunke, et al.
Publicado: (2026)
por: Ao, Yunke, et al.
Publicado: (2026)
Dynamics Reveals Structure: Challenging the Linear Propagation Assumption
por: Chang, Hoyeon, et al.
Publicado: (2026)
por: Chang, Hoyeon, et al.
Publicado: (2026)
Low-Dimensional Execution Manifolds in Transformer Learning Dynamics: Evidence from Modular Arithmetic Tasks
por: Xu, Yongzhong
Publicado: (2026)
por: Xu, Yongzhong
Publicado: (2026)
Democratic Preference Alignment via Sortition-Weighted RLHF
por: Sana, Suvadip, et al.
Publicado: (2026)
por: Sana, Suvadip, et al.
Publicado: (2026)
Architectural Proprioception in State Space Models: Thermodynamic Training Induces Anticipatory Halt Detection
por: Noon, Jay
Publicado: (2026)
por: Noon, Jay
Publicado: (2026)
Interestingness as an Inductive Heuristic for Future Compression Progress
por: Herrmann, Vincent, et al.
Publicado: (2026)
por: Herrmann, Vincent, et al.
Publicado: (2026)
Behavior Learning (BL): Learning Hierarchical Optimization Structures from Data
por: Ma, Zhenyao, et al.
Publicado: (2026)
por: Ma, Zhenyao, et al.
Publicado: (2026)
Spectral Compact Training: Pre-Training Large Language Models via Permanent Truncated SVD and Stiefel QR Retraction
por: Kohlberger, Björn Roman
Publicado: (2026)
por: Kohlberger, Björn Roman
Publicado: (2026)
Market-Alignment Risk in Pricing Agents: Trace Diagnostics and Trace-Prior RL under Hidden Competitor State
por: Zhu, Peiying, et al.
Publicado: (2026)
por: Zhu, Peiying, et al.
Publicado: (2026)
Superposition Is Not Necessary: A Mechanistic Interpretability Analysis of Transformer Representations for Time Series Forecasting
por: Yıldırım, Alper
Publicado: (2026)
por: Yıldırım, Alper
Publicado: (2026)
Path-Coupled Bellman Flows for Distributional Reinforcement Learning
por: Xu, Boyang, et al.
Publicado: (2026)
por: Xu, Boyang, et al.
Publicado: (2026)
TACO: Tackling Over-correction in Federated Learning with Tailored Adaptive Correction
por: Liu, Weijie, et al.
Publicado: (2025)
por: Liu, Weijie, et al.
Publicado: (2025)
DataRater: Meta-Learned Dataset Curation
por: Calian, Dan A., et al.
Publicado: (2025)
por: Calian, Dan A., et al.
Publicado: (2025)
Extrinsicaly Rewarded Soft Q Imitation Learning with Discriminator
por: Furuyama, Ryoma, et al.
Publicado: (2024)
por: Furuyama, Ryoma, et al.
Publicado: (2024)
Manipulating Predictions over Discrete Inputs in Machine Teaching
por: Wu, Xiaodong, et al.
Publicado: (2024)
por: Wu, Xiaodong, et al.
Publicado: (2024)
Axiomatic Characterisations of Sample-based Explainers
por: Amgoud, Leila, et al.
Publicado: (2024)
por: Amgoud, Leila, et al.
Publicado: (2024)
The Lattice Geometry of Neural Network Quantization -- A Short Equivalence Proof of GPTQ and Babai's Algorithm
por: Birnick, Johann
Publicado: (2025)
por: Birnick, Johann
Publicado: (2025)
Evaluation of post-hoc interpretability methods in time-series classification
por: Turbé, Hugues, et al.
Publicado: (2022)
por: Turbé, Hugues, et al.
Publicado: (2022)
Resilience to the Flowing Unknown: an Open Set Recognition Framework for Data Streams
por: Barcina-Blanco, Marcos, et al.
Publicado: (2024)
por: Barcina-Blanco, Marcos, et al.
Publicado: (2024)
Scaling Value Iteration Networks to 5000 Layers for Extreme Long-Term Planning
por: Wang, Yuhui, et al.
Publicado: (2024)
por: Wang, Yuhui, et al.
Publicado: (2024)
Unlearning's Blind Spots: Over-Unlearning and Prototypical Relearning Attack
por: Ha, SeungBum, et al.
Publicado: (2025)
por: Ha, SeungBum, et al.
Publicado: (2025)
Residual Reservoir Memory Networks
por: Pinna, Matteo, et al.
Publicado: (2025)
por: Pinna, Matteo, et al.
Publicado: (2025)
Individual Fairness Through Reweighting and Tuning
por: Mahamadou, Abdoul Jalil Djiberou, et al.
Publicado: (2024)
por: Mahamadou, Abdoul Jalil Djiberou, et al.
Publicado: (2024)
Ejemplares similares
-
An Idiosyncrasy of Time-discretization in Reinforcement Learning
por: De Asis, Kris, et al.
Publicado: (2024) -
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
por: Hedar, Abdel-Rahman, et al.
Publicado: (2024) -
ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing
por: Kadu, Ankush, et al.
Publicado: (2025) -
Machine Learning vs Deep Learning: The Generalization Problem
por: Bay, Yong Yi, et al.
Publicado: (2024) -
Extending NGU to Multi-Agent RL: A Preliminary Study
por: Hernandez, Juan, et al.
Publicado: (2025)