The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, HyunJin, Yi, Xiaoyuan, Yao, Jing, Lian, Jianxun, Huang, Muhua, Duan, Shitong, Bak, JinYeong, Xie, Xing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Research Superalignment Should Advance Now with Alternating Competence and Conformity Optimization
by: Kim, HyunJin, et al.
Published: (2025)
by: Kim, HyunJin, et al.
Published: (2025)
PEMA: An Offsite-Tunable Plug-in External Memory Adaptation for Language Models
by: Kim, HyunJin, et al.
Published: (2023)
by: Kim, HyunJin, et al.
Published: (2023)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
by: Choi, Sooyung, et al.
Published: (2025)
by: Choi, Sooyung, et al.
Published: (2025)
Memoria: Resolving Fateful Forgetting Problem through Human-Inspired Memory Architecture
by: Park, Sangjun, et al.
Published: (2023)
by: Park, Sangjun, et al.
Published: (2023)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
by: Lee, Jaehyeok, et al.
Published: (2026)
by: Lee, Jaehyeok, et al.
Published: (2026)
CASHG: Context-Aware Stylized Online Handwriting Generation
by: Shin, Jinsu, et al.
Published: (2026)
by: Shin, Jinsu, et al.
Published: (2026)
Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
by: Lee, Jaehyeok, et al.
Published: (2024)
by: Lee, Jaehyeok, et al.
Published: (2024)
KpopMT: Translation Dataset with Terminology for Kpop Fandom
by: Kim, JiWoo, et al.
Published: (2024)
by: Kim, JiWoo, et al.
Published: (2024)
Beyond Turn-taking: Introducing Text-based Overlap into Human-LLM Interactions
by: Kim, JiWoo, et al.
Published: (2025)
by: Kim, JiWoo, et al.
Published: (2025)
MentalAgora: A Gateway to Advanced Personalized Care in Mental Health through Multi-Agent Debating and Attribute Control
by: Lee, Yeonji, et al.
Published: (2024)
by: Lee, Yeonji, et al.
Published: (2024)
Towards Cybersecurity SuperIntelligence (CSI): What's the best harness for cybersecurity?
by: Mayoral-Vilches, Víctor, et al.
Published: (2026)
by: Mayoral-Vilches, Víctor, et al.
Published: (2026)
Perturb-and-Compare Approach for Detecting Out-of-Distribution Samples in Constrained Access Environments
by: Lee, Heeyoung, et al.
Published: (2024)
by: Lee, Heeyoung, et al.
Published: (2024)
Translating Hanja Historical Documents to Contemporary Korean and English
by: Son, Juhee, et al.
Published: (2022)
by: Son, Juhee, et al.
Published: (2022)
IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
by: Bai, Yuzhuo, et al.
Published: (2025)
by: Bai, Yuzhuo, et al.
Published: (2025)
On the Dynamics of Multi-Agent LLM Communities Driven by Value Diversity
by: Huang, Muhua, et al.
Published: (2025)
by: Huang, Muhua, et al.
Published: (2025)
MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?
by: Yong, Xixian, et al.
Published: (2025)
by: Yong, Xixian, et al.
Published: (2025)
Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values
by: Yao, Jing, et al.
Published: (2025)
by: Yao, Jing, et al.
Published: (2025)
Talking with Tables for Better LLM Factual Data Interactions
by: Oh, Jio, et al.
Published: (2024)
by: Oh, Jio, et al.
Published: (2024)
A Temporally Correlated Latent Exploration for Reinforcement Learning
by: Oh, SuMin, et al.
Published: (2024)
by: Oh, SuMin, et al.
Published: (2024)
Contextualized Privacy Defense for LLM Agents
by: Wen, Yule, et al.
Published: (2026)
by: Wen, Yule, et al.
Published: (2026)
CLAVE: An Adaptive Framework for Evaluating Values of LLM Generated Responses
by: Yao, Jing, et al.
Published: (2024)
by: Yao, Jing, et al.
Published: (2024)
The Superalignment of Superhuman Intelligence with Large Language Models
by: Huang, Minlie, et al.
Published: (2024)
by: Huang, Minlie, et al.
Published: (2024)
Denevil: Towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning
by: Duan, Shitong, et al.
Published: (2023)
by: Duan, Shitong, et al.
Published: (2023)
RecExplainer: Aligning Large Language Models for Explaining Recommendation Models
by: Lei, Yuxuan, et al.
Published: (2023)
by: Lei, Yuxuan, et al.
Published: (2023)
Recommender AI Agent: Integrating Large Language Models for Interactive Recommendations
by: Huang, Xu, et al.
Published: (2023)
by: Huang, Xu, et al.
Published: (2023)
Aligning Language Models for Versatile Text-based Item Retrieval
by: Lei, Yuxuan, et al.
Published: (2024)
by: Lei, Yuxuan, et al.
Published: (2024)
AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
by: Yao, Jing, et al.
Published: (2025)
by: Yao, Jing, et al.
Published: (2025)
Superalignment with Dynamic Human Values
by: Mai, Florian, et al.
Published: (2025)
by: Mai, Florian, et al.
Published: (2025)
RecAI: Leveraging Large Language Models for Next-Generation Recommender Systems
by: Lian, Jianxun, et al.
Published: (2024)
by: Lian, Jianxun, et al.
Published: (2024)
MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions
by: Zhu, Yanxu, et al.
Published: (2025)
by: Zhu, Yanxu, et al.
Published: (2025)
On the Essence and Prospect: An Investigation of Alignment Approaches for Big Models
by: Wang, Xinpeng, et al.
Published: (2024)
by: Wang, Xinpeng, et al.
Published: (2024)
LLM-powered Multi-agent Framework for Goal-oriented Learning in Intelligent Tutoring System
by: Wang, Tianfu, et al.
Published: (2025)
by: Wang, Tianfu, et al.
Published: (2025)
Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization
by: Duan, Shitong, et al.
Published: (2024)
by: Duan, Shitong, et al.
Published: (2024)
Ada-Retrieval: An Adaptive Multi-Round Retrieval Paradigm for Sequential Recommendations
by: Li, Lei, et al.
Published: (2024)
by: Li, Lei, et al.
Published: (2024)
CAReDiO: Cultural Alignment via Representativeness and Distinctiveness Guided Data Optimization
by: Yao, Jing, et al.
Published: (2025)
by: Yao, Jing, et al.
Published: (2025)
Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
by: Guo, Hanze, et al.
Published: (2025)
by: Guo, Hanze, et al.
Published: (2025)
Agentic Artificial Intelligence in Finance: A Comprehensive Survey
by: Aldridge, Irene, et al.
Published: (2026)
by: Aldridge, Irene, et al.
Published: (2026)
Generative Artificial Intelligence for Navigating Synthesizable Chemical Space
by: Gao, Wenhao, et al.
Published: (2024)
by: Gao, Wenhao, et al.
Published: (2024)
HumanLLM: Towards Personalized Understanding and Simulation of Human Nature
by: Lei, Yuxuan, et al.
Published: (2026)
by: Lei, Yuxuan, et al.
Published: (2026)
Artificial Intelligence in Landscape Architecture: A Survey
by: Xing, Yue, et al.
Published: (2024)
by: Xing, Yue, et al.
Published: (2024)
Similar Items
-
Research Superalignment Should Advance Now with Alternating Competence and Conformity Optimization
by: Kim, HyunJin, et al.
Published: (2025) -
PEMA: An Offsite-Tunable Plug-in External Memory Adaptation for Language Models
by: Kim, HyunJin, et al.
Published: (2023) -
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
by: Choi, Sooyung, et al.
Published: (2025) -
Memoria: Resolving Fateful Forgetting Problem through Human-Inspired Memory Architecture
by: Park, Sangjun, et al.
Published: (2023) -
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
by: Lee, Jaehyeok, et al.
Published: (2026)