Ethics2vec: aligning automatic agents and human preferences
Fuente:
arXiv
Saved in:
| Main Author: | Bontempi, Gianluca |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FRAUD-RLA: A new reinforcement learning adversarial attack against credit card fraud detection
by: Lunghi, Daniele, et al.
Published: (2025)
by: Lunghi, Daniele, et al.
Published: (2025)
mini-vec2vec: Scaling Universal Geometry Alignment with Linear Transformations
by: Dar, Guy
Published: (2025)
by: Dar, Guy
Published: (2025)
Task2vec Readiness: Diagnostics for Federated Learning from Pre-Training Embeddings
by: Mafuz, Cristiano, et al.
Published: (2026)
by: Mafuz, Cristiano, et al.
Published: (2026)
A Direct Classification Approach for Reliable Wind Ramp Event Forecasting under Severe Class Imbalance
by: Morales-Hernández, Alejandro, et al.
Published: (2026)
by: Morales-Hernández, Alejandro, et al.
Published: (2026)
Aligning LLM agents with human learning and adjustment behavior: a dual agent approach
by: Liu, Tianming, et al.
Published: (2025)
by: Liu, Tianming, et al.
Published: (2025)
player2vec: A Language Modeling Approach to Understand Player Behavior in Games
by: Wang, Tianze, et al.
Published: (2024)
by: Wang, Tianze, et al.
Published: (2024)
Tell me why: Training preferences-based RL with human preferences and step-level explanations
by: Karalus, Jakob
Published: (2024)
by: Karalus, Jakob
Published: (2024)
A density estimation perspective on learning from pairwise human preferences
by: Dumoulin, Vincent, et al.
Published: (2023)
by: Dumoulin, Vincent, et al.
Published: (2023)
ffstruc2vec: Flat, Flexible and Scalable Learning of Node Representations from Structural Identities
by: Heidrich, Mario, et al.
Published: (2025)
by: Heidrich, Mario, et al.
Published: (2025)
mRNA2vec: mRNA Embedding with Language Model in the 5'UTR-CDS for mRNA Design
by: Zhang, Honggen, et al.
Published: (2024)
by: Zhang, Honggen, et al.
Published: (2024)
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
by: Wijk, Hjalmar, et al.
Published: (2024)
by: Wijk, Hjalmar, et al.
Published: (2024)
Constrained multi-fidelity Bayesian optimization with automatic stop condition
by: Foumani, Zahra Zanjani, et al.
Published: (2025)
by: Foumani, Zahra Zanjani, et al.
Published: (2025)
Human-aligned Chess with a Bit of Search
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
Bayesian preference elicitation for decision support in multiobjective optimization
by: Huber, Felix, et al.
Published: (2025)
by: Huber, Felix, et al.
Published: (2025)
Knowledge distillation through geometry-aware representational alignment
by: Bhattarai, Prajjwal, et al.
Published: (2025)
by: Bhattarai, Prajjwal, et al.
Published: (2025)
Towards Stable Preferences for Stakeholder-aligned Machine Learning
by: Sheraz, Haleema, et al.
Published: (2024)
by: Sheraz, Haleema, et al.
Published: (2024)
Vision Transformer attention alignment with human visual perception in aesthetic object evaluation
by: Carrasco, Miguel, et al.
Published: (2025)
by: Carrasco, Miguel, et al.
Published: (2025)
PyAWD: A Library for Generating Large Synthetic Datasets of Acoustic Wave Propagation
by: Tribel, Pascal, et al.
Published: (2024)
by: Tribel, Pascal, et al.
Published: (2024)
Easydiagnos: a framework for accurate feature selection for automatic diagnosis in smart healthcare
by: Maji, Prasenjit, et al.
Published: (2024)
by: Maji, Prasenjit, et al.
Published: (2024)
ProtAlign: Contrastive learning paradigm for Sequence and structure alignment
by: Ranganath, Aditya, et al.
Published: (2026)
by: Ranganath, Aditya, et al.
Published: (2026)
Gram: Assessing sabotage propensities via automated alignment auditing
by: Lindner, David, et al.
Published: (2026)
by: Lindner, David, et al.
Published: (2026)
Understanding the performance gap between online and offline alignment algorithms
by: Tang, Yunhao, et al.
Published: (2024)
by: Tang, Yunhao, et al.
Published: (2024)
When should we prefer Decision Transformers for Offline Reinforcement Learning?
by: Bhargava, Prajjwal, et al.
Published: (2023)
by: Bhargava, Prajjwal, et al.
Published: (2023)
Self-supervised video pretraining yields robust and more human-aligned visual representations
by: Parthasarathy, Nikhil, et al.
Published: (2022)
by: Parthasarathy, Nikhil, et al.
Published: (2022)
Evaluating alignment between humans and neural network representations in image-based learning tasks
by: Demircan, Can, et al.
Published: (2023)
by: Demircan, Can, et al.
Published: (2023)
CLUE: Neural Networks Calibration via Learning Uncertainty-Error alignment
by: Mendes, Pedro, et al.
Published: (2025)
by: Mendes, Pedro, et al.
Published: (2025)
Diffusion Tree Sampling: Scalable inference-time alignment of diffusion models
by: Jain, Vineet, et al.
Published: (2025)
by: Jain, Vineet, et al.
Published: (2025)
A learning-driven automatic planning framework for proton PBS treatments of H&N cancers
by: Wang, Qingqing, et al.
Published: (2025)
by: Wang, Qingqing, et al.
Published: (2025)
Dimensions underlying the representational alignment of deep neural networks with humans
by: Mahner, Florian P., et al.
Published: (2024)
by: Mahner, Florian P., et al.
Published: (2024)
Fitness aligned structural modeling enables scalable virtual screening with AuroBind
by: Zhang, Zhongyue, et al.
Published: (2025)
by: Zhang, Zhongyue, et al.
Published: (2025)
Decomposed Trust: Privacy, Adversarial Robustness, Ethics, and Fairness in Low-Rank LLMs
by: Asante, Daniel Agyei, et al.
Published: (2025)
by: Asante, Daniel Agyei, et al.
Published: (2025)
Cross-model Fairness: Empirical Study of Fairness and Ethics Under Model Multiplicity
by: Sokol, Kacper, et al.
Published: (2022)
by: Sokol, Kacper, et al.
Published: (2022)
Are aligned neural networks adversarially aligned?
by: Carlini, Nicholas, et al.
Published: (2023)
by: Carlini, Nicholas, et al.
Published: (2023)
Capturing waste collection planning expert knowledge in a fitness function through preference learning
by: Díaz, Laura Fernández, et al.
Published: (2024)
by: Díaz, Laura Fernández, et al.
Published: (2024)
OSIL: Learning Offline Safe Imitation Policies with Safety Inferred from Non-preferred Trajectories
by: Burnwal, Returaj, et al.
Published: (2026)
by: Burnwal, Returaj, et al.
Published: (2026)
Dynamic LLM Routing and Selection based on User Preferences: Balancing Performance, Cost, and Ethics
by: Piskala, Deepak Babu, et al.
Published: (2025)
by: Piskala, Deepak Babu, et al.
Published: (2025)
On Goodhart's law, with an application to value alignment
by: El-Mhamdi, El-Mahdi, et al.
Published: (2024)
by: El-Mhamdi, El-Mahdi, et al.
Published: (2024)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
by: Beaglehole, Daniel, et al.
Published: (2024)
by: Beaglehole, Daniel, et al.
Published: (2024)
Introduction to AI Safety, Ethics, and Society
by: Hendrycks, Dan
Published: (2024)
by: Hendrycks, Dan
Published: (2024)
Getting aligned on representational alignment
by: Sucholutsky, Ilia, et al.
Published: (2023)
by: Sucholutsky, Ilia, et al.
Published: (2023)
Similar Items
-
FRAUD-RLA: A new reinforcement learning adversarial attack against credit card fraud detection
by: Lunghi, Daniele, et al.
Published: (2025) -
mini-vec2vec: Scaling Universal Geometry Alignment with Linear Transformations
by: Dar, Guy
Published: (2025) -
Task2vec Readiness: Diagnostics for Federated Learning from Pre-Training Embeddings
by: Mafuz, Cristiano, et al.
Published: (2026) -
A Direct Classification Approach for Reliable Wind Ramp Event Forecasting under Severe Class Imbalance
by: Morales-Hernández, Alejandro, et al.
Published: (2026) -
Aligning LLM agents with human learning and adjustment behavior: a dual agent approach
by: Liu, Tianming, et al.
Published: (2025)