Research Superalignment Should Advance Now with Alternating Competence and Conformity Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, HyunJin, Yi, Xiaoyuan, Yao, Jing, Huang, Muhua, Bak, JinYeong, Evans, James, Xie, Xing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment
von: Kim, HyunJin, et al.
Veröffentlicht: (2024)
von: Kim, HyunJin, et al.
Veröffentlicht: (2024)
PEMA: An Offsite-Tunable Plug-in External Memory Adaptation for Language Models
von: Kim, HyunJin, et al.
Veröffentlicht: (2023)
von: Kim, HyunJin, et al.
Veröffentlicht: (2023)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
von: Choi, Sooyung, et al.
Veröffentlicht: (2025)
von: Choi, Sooyung, et al.
Veröffentlicht: (2025)
Memoria: Resolving Fateful Forgetting Problem through Human-Inspired Memory Architecture
von: Park, Sangjun, et al.
Veröffentlicht: (2023)
von: Park, Sangjun, et al.
Veröffentlicht: (2023)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2026)
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2026)
KpopMT: Translation Dataset with Terminology for Kpop Fandom
von: Kim, JiWoo, et al.
Veröffentlicht: (2024)
von: Kim, JiWoo, et al.
Veröffentlicht: (2024)
Beyond Turn-taking: Introducing Text-based Overlap into Human-LLM Interactions
von: Kim, JiWoo, et al.
Veröffentlicht: (2025)
von: Kim, JiWoo, et al.
Veröffentlicht: (2025)
CASHG: Context-Aware Stylized Online Handwriting Generation
von: Shin, Jinsu, et al.
Veröffentlicht: (2026)
von: Shin, Jinsu, et al.
Veröffentlicht: (2026)
Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2024)
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2024)
MentalAgora: A Gateway to Advanced Personalized Care in Mental Health through Multi-Agent Debating and Attribute Control
von: Lee, Yeonji, et al.
Veröffentlicht: (2024)
von: Lee, Yeonji, et al.
Veröffentlicht: (2024)
Perturb-and-Compare Approach for Detecting Out-of-Distribution Samples in Constrained Access Environments
von: Lee, Heeyoung, et al.
Veröffentlicht: (2024)
von: Lee, Heeyoung, et al.
Veröffentlicht: (2024)
Translating Hanja Historical Documents to Contemporary Korean and English
von: Son, Juhee, et al.
Veröffentlicht: (2022)
von: Son, Juhee, et al.
Veröffentlicht: (2022)
A Temporally Correlated Latent Exploration for Reinforcement Learning
von: Oh, SuMin, et al.
Veröffentlicht: (2024)
von: Oh, SuMin, et al.
Veröffentlicht: (2024)
Talking with Tables for Better LLM Factual Data Interactions
von: Oh, Jio, et al.
Veröffentlicht: (2024)
von: Oh, Jio, et al.
Veröffentlicht: (2024)
On the Dynamics of Multi-Agent LLM Communities Driven by Value Diversity
von: Huang, Muhua, et al.
Veröffentlicht: (2025)
von: Huang, Muhua, et al.
Veröffentlicht: (2025)
CLAVE: An Adaptive Framework for Evaluating Values of LLM Generated Responses
von: Yao, Jing, et al.
Veröffentlicht: (2024)
von: Yao, Jing, et al.
Veröffentlicht: (2024)
IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
von: Bai, Yuzhuo, et al.
Veröffentlicht: (2025)
von: Bai, Yuzhuo, et al.
Veröffentlicht: (2025)
CAReDiO: Cultural Alignment via Representativeness and Distinctiveness Guided Data Optimization
von: Yao, Jing, et al.
Veröffentlicht: (2025)
von: Yao, Jing, et al.
Veröffentlicht: (2025)
Superalignment with Dynamic Human Values
von: Mai, Florian, et al.
Veröffentlicht: (2025)
von: Mai, Florian, et al.
Veröffentlicht: (2025)
Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
von: Guo, Hanze, et al.
Veröffentlicht: (2025)
von: Guo, Hanze, et al.
Veröffentlicht: (2025)
The Superalignment of Superhuman Intelligence with Large Language Models
von: Huang, Minlie, et al.
Veröffentlicht: (2024)
von: Huang, Minlie, et al.
Veröffentlicht: (2024)
Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values
von: Yao, Jing, et al.
Veröffentlicht: (2025)
von: Yao, Jing, et al.
Veröffentlicht: (2025)
Embedding an Ethical Mind: Aligning Text-to-Image Synthesis via Lightweight Value Optimization
von: Wang, Xingqi, et al.
Veröffentlicht: (2024)
von: Wang, Xingqi, et al.
Veröffentlicht: (2024)
Designing AI-Agents with Personalities: A Psychometric Approach
von: Huang, Muhua, et al.
Veröffentlicht: (2024)
von: Huang, Muhua, et al.
Veröffentlicht: (2024)
Optimized Conformal Selection: Powerful Selective Inference After Conformity Score Optimization
von: Bai, Tian, et al.
Veröffentlicht: (2024)
von: Bai, Tian, et al.
Veröffentlicht: (2024)
A Superalignment Framework in Autonomous Driving with Large Language Models
von: Kong, Xiangrui, et al.
Veröffentlicht: (2024)
von: Kong, Xiangrui, et al.
Veröffentlicht: (2024)
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
von: Guo, Jianyuan, et al.
Veröffentlicht: (2024)
von: Guo, Jianyuan, et al.
Veröffentlicht: (2024)
AI Evaluation Should Require Standardized Item-Level Data Releases
von: Jiang, Han, et al.
Veröffentlicht: (2026)
von: Jiang, Han, et al.
Veröffentlicht: (2026)
"Should I Give Up Now?" Investigating LLM Pitfalls in Software Engineering
von: Tie, Jiessie, et al.
Veröffentlicht: (2024)
von: Tie, Jiessie, et al.
Veröffentlicht: (2024)
Topological Structure Learning Should Be A Research Priority for LLM-Based Multi-Agent Systems
von: Yang, Jiaxi, et al.
Veröffentlicht: (2025)
von: Yang, Jiaxi, et al.
Veröffentlicht: (2025)
The Incomplete Bridge: How AI Research (Mis)Engages with Psychology
von: Jiang, Han, et al.
Veröffentlicht: (2025)
von: Jiang, Han, et al.
Veröffentlicht: (2025)
A Moral Imperative: The Need for Continual Superalignment of Large Language Models
von: Puthumanaillam, Gokul, et al.
Veröffentlicht: (2024)
von: Puthumanaillam, Gokul, et al.
Veröffentlicht: (2024)
Reinforcing Stromal Cell Spheroid Through Red‐Light Preconditioning for Advanced Vascularization
von: Yu‐Jin Kim, et al.
Veröffentlicht: (2025)
von: Yu‐Jin Kim, et al.
Veröffentlicht: (2025)
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
von: Jiang, Han, et al.
Veröffentlicht: (2025)
von: Jiang, Han, et al.
Veröffentlicht: (2025)
Should I Stay or Should I Go Now? An Investigation into Gender Differences in the Impact of Switching Jobs on Earnings
von: Winskill, Emily
Veröffentlicht: (2025)
von: Winskill, Emily
Veröffentlicht: (2025)
Correspondence for “Developing a Virtual Endoscopic Surgery Planning System to Optimize Surgical Outcomes”
von: Hyun Jin Min
Veröffentlicht: (2025)
von: Hyun Jin Min
Veröffentlicht: (2025)
Reinforcing Stromal Cell Spheroid Through Red‐Light Preconditioning for Advanced Vascularization (Adv. Sci. 29/2025)
von: Yu‐Jin Kim, et al.
Veröffentlicht: (2025)
von: Yu‐Jin Kim, et al.
Veröffentlicht: (2025)
Advancing Equity Planning Now
Veröffentlicht: (2023)
Veröffentlicht: (2023)
Recent advances on MOF ‐based colorimetric sensors
von: Solmin Lee, et al.
Veröffentlicht: (2025)
von: Solmin Lee, et al.
Veröffentlicht: (2025)
Convex Optimization Algorithm for Maximizing Directivity of Conformal Array With DRR Constraint
von: Chao Liu, et al.
Veröffentlicht: (2024)
von: Chao Liu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment
von: Kim, HyunJin, et al.
Veröffentlicht: (2024) -
PEMA: An Offsite-Tunable Plug-in External Memory Adaptation for Language Models
von: Kim, HyunJin, et al.
Veröffentlicht: (2023) -
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
von: Choi, Sooyung, et al.
Veröffentlicht: (2025) -
Memoria: Resolving Fateful Forgetting Problem through Human-Inspired Memory Architecture
von: Park, Sangjun, et al.
Veröffentlicht: (2023) -
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2026)