Saved in:
| Main Authors: | Lee, Hyunin, Park, Chanwoo, Abel, David, Jin, Ming |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2407.18422 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Position: AI Safety Must Embrace an Antifragile Perspective
by: Jin, Ming, et al.
Published: (2025)
by: Jin, Ming, et al.
Published: (2025)
Preparing for Black Swans: The Antifragility Imperative for Machine Learning
by: Jin, Ming
Published: (2024)
by: Jin, Ming
Published: (2024)
RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation
by: Park, Chanwoo, et al.
Published: (2024)
by: Park, Chanwoo, et al.
Published: (2024)
TraM : Enhancing User Sleep Prediction with Transformer-based Multivariate Time Series Modeling and Machine Learning Ensembles
by: Kim, Jinjae, et al.
Published: (2024)
by: Kim, Jinjae, et al.
Published: (2024)
Hypothesis Testing the Circuit Hypothesis in LLMs
by: Shi, Claudia, et al.
Published: (2024)
by: Shi, Claudia, et al.
Published: (2024)
Do LLM Agents Have Regret? A Case Study in Online Learning and Games
by: Park, Chanwoo, et al.
Published: (2024)
by: Park, Chanwoo, et al.
Published: (2024)
More Than Irrational: Modeling Belief-Biased Agents
by: Zhu, Yifan, et al.
Published: (2025)
by: Zhu, Yifan, et al.
Published: (2025)
NeuroAI for AI Safety
by: Mineault, Patrick, et al.
Published: (2024)
by: Mineault, Patrick, et al.
Published: (2024)
Superficial Safety Alignment Hypothesis
by: Li, Jianwei, et al.
Published: (2024)
by: Li, Jianwei, et al.
Published: (2024)
Synthetic Mixed Training: Scaling Parametric Knowledge Acquisition Beyond RAG
by: Han, Seungju, et al.
Published: (2026)
by: Han, Seungju, et al.
Published: (2026)
CellCLIP -- Learning Perturbation Effects in Cell Painting via Text-Guided Contrastive Learning
by: Lu, Mingyu, et al.
Published: (2025)
by: Lu, Mingyu, et al.
Published: (2025)
Pausing Policy Learning in Non-stationary Reinforcement Learning
by: Lee, Hyunin, et al.
Published: (2024)
by: Lee, Hyunin, et al.
Published: (2024)
The Linear Representation Hypothesis and the Geometry of Large Language Models
by: Park, Kiho, et al.
Published: (2023)
by: Park, Kiho, et al.
Published: (2023)
PnPXAI: A Universal XAI Framework Providing Automatic Explanations Across Diverse Modalities and Models
by: Kim, Seongun, et al.
Published: (2025)
by: Kim, Seongun, et al.
Published: (2025)
Memory Allocation in Resource-Constrained Reinforcement Learning
by: Tamborski, Massimiliano, et al.
Published: (2025)
by: Tamborski, Massimiliano, et al.
Published: (2025)
SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
by: Kim, Geon-Hyeong, et al.
Published: (2025)
by: Kim, Geon-Hyeong, et al.
Published: (2025)
Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation
by: Gu, Shangding, et al.
Published: (2024)
by: Gu, Shangding, et al.
Published: (2024)
Physics Informed Distillation for Diffusion Models
by: Tee, Joshua Tian Jin, et al.
Published: (2024)
by: Tee, Joshua Tian Jin, et al.
Published: (2024)
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
Hypothesis Hunting with Evolving Networks of Autonomous Scientific Agents
by: Liu, Tennison, et al.
Published: (2025)
by: Liu, Tennison, et al.
Published: (2025)
Expert-Guided LLM Reasoning for Battery Discovery: From AI-Driven Hypothesis to Synthesis and Characterization
by: Liu, Shengchao, et al.
Published: (2025)
by: Liu, Shengchao, et al.
Published: (2025)
MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-Making
by: Kim, Yubin, et al.
Published: (2024)
by: Kim, Yubin, et al.
Published: (2024)
Shutdown Safety Valves for Advanced AI
by: Conitzer, Vincent
Published: (2026)
by: Conitzer, Vincent
Published: (2026)
The Neural Pruning Law Hypothesis
by: Barbulescu, Eugen, et al.
Published: (2025)
by: Barbulescu, Eugen, et al.
Published: (2025)
The Safety-Aware Denoiser for Text Diffusion Models
by: Yusuf, Amman, et al.
Published: (2026)
by: Yusuf, Amman, et al.
Published: (2026)
HyperFlow: Gradient-Free Emulation of Few-Shot Fine-Tuning
by: Kim, Donggyun, et al.
Published: (2025)
by: Kim, Donggyun, et al.
Published: (2025)
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
by: Tan, Zhewen, et al.
Published: (2026)
by: Tan, Zhewen, et al.
Published: (2026)
AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing
by: Bright-Thonney, Samuel, et al.
Published: (2025)
by: Bright-Thonney, Samuel, et al.
Published: (2025)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
by: Siu, Vincent, et al.
Published: (2025)
by: Siu, Vincent, et al.
Published: (2025)
Computational Safety for Generative AI: A Signal Processing Perspective
by: Chen, Pin-Yu
Published: (2025)
by: Chen, Pin-Yu
Published: (2025)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
by: Griffin, Charlie, et al.
Published: (2024)
by: Griffin, Charlie, et al.
Published: (2024)
On the Sparsity of the Strong Lottery Ticket Hypothesis
by: Natale, Emanuele, et al.
Published: (2024)
by: Natale, Emanuele, et al.
Published: (2024)
Scaling Law Hypothesis for Multimodal Model
by: Sun, Qingyun, et al.
Published: (2024)
by: Sun, Qingyun, et al.
Published: (2024)
MaD-Scientist: AI-based Scientist solving Convection-Diffusion-Reaction Equations Using Massive PINN-Based Prior Data
by: Kang, Mingu, et al.
Published: (2024)
by: Kang, Mingu, et al.
Published: (2024)
International AI Safety Report
by: Bengio, Yoshua, et al.
Published: (2025)
by: Bengio, Yoshua, et al.
Published: (2025)
Hypothesis-Conditioned Query Rewriting for Decision-Useful Retrieval
by: Chang, Hangeol, et al.
Published: (2026)
by: Chang, Hangeol, et al.
Published: (2026)
Holistic Safety and Responsibility Evaluations of Advanced AI Models
by: Weidinger, Laura, et al.
Published: (2024)
by: Weidinger, Laura, et al.
Published: (2024)
A Classical View on Benign Overfitting: The Role of Sample Size
by: Park, Junhyung, et al.
Published: (2025)
by: Park, Junhyung, et al.
Published: (2025)
Three Dogmas of Reinforcement Learning
by: Abel, David, et al.
Published: (2024)
by: Abel, David, et al.
Published: (2024)
Learning to Transfer Human Hand Skills for Robot Manipulations
by: Park, Sungjae, et al.
Published: (2025)
by: Park, Sungjae, et al.
Published: (2025)
Similar Items
-
Position: AI Safety Must Embrace an Antifragile Perspective
by: Jin, Ming, et al.
Published: (2025) -
Preparing for Black Swans: The Antifragility Imperative for Machine Learning
by: Jin, Ming
Published: (2024) -
RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation
by: Park, Chanwoo, et al.
Published: (2024) -
TraM : Enhancing User Sleep Prediction with Transformer-based Multivariate Time Series Modeling and Machine Learning Ensembles
by: Kim, Jinjae, et al.
Published: (2024) -
Hypothesis Testing the Circuit Hypothesis in LLMs
by: Shi, Claudia, et al.
Published: (2024)