Recursive Reward Aggregation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Yuting, Zhang, Yivan, Ackermann, Johannes, Zhang, Yu-Jie, Nishimori, Soichiro, Sugiyama, Masashi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enriching Disentanglement: From Logical Definitions to Quantitative Metrics
von: Zhang, Yivan, et al.
Veröffentlicht: (2023)
von: Zhang, Yivan, et al.
Veröffentlicht: (2023)
A Category-theoretical Meta-analysis of Definitions of Disentanglement
von: Zhang, Yivan, et al.
Veröffentlicht: (2023)
von: Zhang, Yivan, et al.
Veröffentlicht: (2023)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
Offline Reinforcement Learning with Domain-Unlabeled Data
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2024)
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2024)
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2025)
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2025)
Learning Is a Kan Extension
von: Pugh, Matthew, et al.
Veröffentlicht: (2025)
von: Pugh, Matthew, et al.
Veröffentlicht: (2025)
Reinforcement Learning in Categorical Cybernetics
von: Hedges, Jules, et al.
Veröffentlicht: (2024)
von: Hedges, Jules, et al.
Veröffentlicht: (2024)
Stochastic Neural Network Symmetrisation in Markov Categories
von: Cornish, Rob
Veröffentlicht: (2024)
von: Cornish, Rob
Veröffentlicht: (2024)
Weaves, Wires, and Morphisms: Formalizing and Implementing the Algebra of Deep Learning
von: Abbott, Vincent, et al.
Veröffentlicht: (2026)
von: Abbott, Vincent, et al.
Veröffentlicht: (2026)
Learners' Languages
von: Spivak, David I.
Veröffentlicht: (2021)
von: Spivak, David I.
Veröffentlicht: (2021)
The Theory behind UMAP?
von: Wegmann, David
Veröffentlicht: (2026)
von: Wegmann, David
Veröffentlicht: (2026)
Generalized Gradient Descent is a Hypergraph Functor
von: Hanks, Tyler, et al.
Veröffentlicht: (2024)
von: Hanks, Tyler, et al.
Veröffentlicht: (2024)
Using Enriched Category Theory to Construct the Nearest Neighbour Classification Algorithm
von: Pugh, Matthew, et al.
Veröffentlicht: (2023)
von: Pugh, Matthew, et al.
Veröffentlicht: (2023)
The Topos of Transformer Networks
von: Villani, Mattia Jacopo, et al.
Veröffentlicht: (2024)
von: Villani, Mattia Jacopo, et al.
Veröffentlicht: (2024)
Identifiable Equivariant Networks are Layerwise Equivariant
von: Shahverdi, Vahid, et al.
Veröffentlicht: (2026)
von: Shahverdi, Vahid, et al.
Veröffentlicht: (2026)
Accelerating Machine Learning Systems via Category Theory: Applications to Spherical Attention for Gene Regulatory Networks
von: Abbott, Vincent, et al.
Veröffentlicht: (2025)
von: Abbott, Vincent, et al.
Veröffentlicht: (2025)
Compact Matrix Quantum Group Equivariant Neural Networks
von: Pearce-Crump, Edward
Veröffentlicht: (2023)
von: Pearce-Crump, Edward
Veröffentlicht: (2023)
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
von: Ackermann, Johannes, et al.
Veröffentlicht: (2025)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2025)
Aggregating time-series and image data: functors and double functors
von: Diehl, Joscha
Veröffentlicht: (2025)
von: Diehl, Joscha
Veröffentlicht: (2025)
Gaussian Sheaf Neural Networks
von: Ribeiro, André, et al.
Veröffentlicht: (2026)
von: Ribeiro, André, et al.
Veröffentlicht: (2026)
Fundamental Components of Deep Learning: A category-theoretic approach
von: Gavranović, Bruno
Veröffentlicht: (2024)
von: Gavranović, Bruno
Veröffentlicht: (2024)
Towards structure-preserving quantum encodings
von: Parzygnat, Arthur J., et al.
Veröffentlicht: (2024)
von: Parzygnat, Arthur J., et al.
Veröffentlicht: (2024)
Typing Tensor Calculus in 2-Categories (I)
von: Ahmadi, Fatimah Rita
Veröffentlicht: (2019)
von: Ahmadi, Fatimah Rita
Veröffentlicht: (2019)
Towards a Categorical Foundation of Deep Learning: A Survey
von: Crescenzi, Francesco Riccardo
Veröffentlicht: (2024)
von: Crescenzi, Francesco Riccardo
Veröffentlicht: (2024)
Categorical and geometric methods in statistical, manifold, and machine learning
von: Lê, Hông Vân, et al.
Veröffentlicht: (2025)
von: Lê, Hông Vân, et al.
Veröffentlicht: (2025)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
von: Ackermann, Johannes, et al.
Veröffentlicht: (2026)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2026)
The Relativity of Causal Knowledge
von: D'Acunto, Gabriele, et al.
Veröffentlicht: (2025)
von: D'Acunto, Gabriele, et al.
Veröffentlicht: (2025)
The Gauss-Markov Adjunction Provides Categorical Semantics of Residuals in Supervised Learning
von: Kamiura, Moto
Veröffentlicht: (2025)
von: Kamiura, Moto
Veröffentlicht: (2025)
DiagrammaticLearning: A Graphical Language for Compositional Training Regimes
von: Lary, Mason, et al.
Veröffentlicht: (2025)
von: Lary, Mason, et al.
Veröffentlicht: (2025)
Position: Categorical Deep Learning is an Algebraic Theory of All Architectures
von: Gavranović, Bruno, et al.
Veröffentlicht: (2024)
von: Gavranović, Bruno, et al.
Veröffentlicht: (2024)
Language Modeling with Reduced Densities
von: Bradley, Tai-Danae, et al.
Veröffentlicht: (2020)
von: Bradley, Tai-Danae, et al.
Veröffentlicht: (2020)
Reduce, Reuse, Recycle: Categories for Compositional Reinforcement Learning
von: Bakirtzis, Georgios, et al.
Veröffentlicht: (2024)
von: Bakirtzis, Georgios, et al.
Veröffentlicht: (2024)
On Meta-Prompting
von: de Wynter, Adrian, et al.
Veröffentlicht: (2023)
von: de Wynter, Adrian, et al.
Veröffentlicht: (2023)
Functorial Neural Architectures from Higher Inductive Types
von: Sargsyan, Karen
Veröffentlicht: (2026)
von: Sargsyan, Karen
Veröffentlicht: (2026)
Completeness of quantaloid-enriched categories up to Morita equivalence
von: Tang, Xiaoye
Veröffentlicht: (2025)
von: Tang, Xiaoye
Veröffentlicht: (2025)
Monoidal Ringel duality and monoidal highest weight envelopes
von: Flake, Johannes, et al.
Veröffentlicht: (2025)
von: Flake, Johannes, et al.
Veröffentlicht: (2025)
Monads and Distributive Laws in Substructural Contexts (Extended Version)
von: Fujii, Soichiro, et al.
Veröffentlicht: (2026)
von: Fujii, Soichiro, et al.
Veröffentlicht: (2026)
Offline Reinforcement Learning from Datasets with Structured Non-Stationarity
von: Ackermann, Johannes, et al.
Veröffentlicht: (2024)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2024)
Weisfeiler and Lehman Go Categorical
von: Choi, Seongjin, et al.
Veröffentlicht: (2026)
von: Choi, Seongjin, et al.
Veröffentlicht: (2026)
Towards Compositional Interpretability for XAI
von: Tull, Sean, et al.
Veröffentlicht: (2024)
von: Tull, Sean, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Enriching Disentanglement: From Logical Definitions to Quantitative Metrics
von: Zhang, Yivan, et al.
Veröffentlicht: (2023) -
A Category-theoretical Meta-analysis of Definitions of Disentanglement
von: Zhang, Yivan, et al.
Veröffentlicht: (2023) -
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026) -
Offline Reinforcement Learning with Domain-Unlabeled Data
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2024) -
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2025)