On the Essence and Prospect: An Investigation of Alignment Approaches for Big Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Xinpeng, Duan, Shitong, Yi, Xiaoyuan, Yao, Jing, Zhou, Shanlin, Wei, Zhihua, Zhang, Peng, Xu, Dongkuan, Sun, Maosong, Xie, Xing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
di: Yao, Jing, et al.
Pubblicazione: (2025)
di: Yao, Jing, et al.
Pubblicazione: (2025)
IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
di: Bai, Yuzhuo, et al.
Pubblicazione: (2025)
di: Bai, Yuzhuo, et al.
Pubblicazione: (2025)
Leveraging Implicit Sentiments: Enhancing Reliability and Validity in Psychological Trait Evaluation of LLMs
di: Ma, Huanhuan, et al.
Pubblicazione: (2025)
di: Ma, Huanhuan, et al.
Pubblicazione: (2025)
Beyond Human Norms: Unveiling Unique Values of Large Language Models through Interdisciplinary Approaches
di: Biedma, Pablo, et al.
Pubblicazione: (2024)
di: Biedma, Pablo, et al.
Pubblicazione: (2024)
Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization
di: Duan, Shitong, et al.
Pubblicazione: (2024)
di: Duan, Shitong, et al.
Pubblicazione: (2024)
Denevil: Towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning
di: Duan, Shitong, et al.
Pubblicazione: (2023)
di: Duan, Shitong, et al.
Pubblicazione: (2023)
ToolNet: Connecting Large Language Models with Massive Tools via Tool Graph
di: Liu, Xukun, et al.
Pubblicazione: (2024)
di: Liu, Xukun, et al.
Pubblicazione: (2024)
CAReDiO: Cultural Alignment via Representativeness and Distinctiveness Guided Data Optimization
di: Yao, Jing, et al.
Pubblicazione: (2025)
di: Yao, Jing, et al.
Pubblicazione: (2025)
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
di: Jiang, Han, et al.
Pubblicazione: (2025)
di: Jiang, Han, et al.
Pubblicazione: (2025)
CLAVE: An Adaptive Framework for Evaluating Values of LLM Generated Responses
di: Yao, Jing, et al.
Pubblicazione: (2024)
di: Yao, Jing, et al.
Pubblicazione: (2024)
Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing
di: Jiang, Han, et al.
Pubblicazione: (2024)
di: Jiang, Han, et al.
Pubblicazione: (2024)
Chinese Short-Form Creative Content Generation via Explanation-Oriented Multi-Objective Optimization
di: Zhou, Shanlin, et al.
Pubblicazione: (2025)
di: Zhou, Shanlin, et al.
Pubblicazione: (2025)
Elephant in the Room: Unveiling the Impact of Reward Model Quality in Alignment
di: Liu, Yan, et al.
Pubblicazione: (2024)
di: Liu, Yan, et al.
Pubblicazione: (2024)
Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities
di: Zhang, Xiangxu, et al.
Pubblicazione: (2026)
di: Zhang, Xiangxu, et al.
Pubblicazione: (2026)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
di: Lee, Jaehyeok, et al.
Pubblicazione: (2026)
di: Lee, Jaehyeok, et al.
Pubblicazione: (2026)
Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values
di: Yao, Jing, et al.
Pubblicazione: (2025)
di: Yao, Jing, et al.
Pubblicazione: (2025)
Towards Robust Pruning: An Adaptive Knowledge-Retention Pruning Strategy for Language Models
di: Li, Jianwei, et al.
Pubblicazione: (2023)
di: Li, Jianwei, et al.
Pubblicazione: (2023)
Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
di: Guo, Hanze, et al.
Pubblicazione: (2025)
di: Guo, Hanze, et al.
Pubblicazione: (2025)
DAMRO: Dive into the Attention Mechanism of LVLM to Reduce Object Hallucination
di: Gong, Xuan, et al.
Pubblicazione: (2024)
di: Gong, Xuan, et al.
Pubblicazione: (2024)
MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?
di: Yong, Xixian, et al.
Pubblicazione: (2025)
di: Yong, Xixian, et al.
Pubblicazione: (2025)
MoVa: Towards Generalizable Classification of Human Morals and Values
di: Chen, Ziyu, et al.
Pubblicazione: (2025)
di: Chen, Ziyu, et al.
Pubblicazione: (2025)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
di: Choi, Sooyung, et al.
Pubblicazione: (2025)
di: Choi, Sooyung, et al.
Pubblicazione: (2025)
Incremental Computation: What Is the Essence?
di: Liu, Yanhong A.
Pubblicazione: (2023)
di: Liu, Yanhong A.
Pubblicazione: (2023)
Fantastic Semantics and Where to Find Them: Investigating Which Layers of Generative LLMs Reflect Lexical Semantics
di: Liu, Zhu, et al.
Pubblicazione: (2024)
di: Liu, Zhu, et al.
Pubblicazione: (2024)
"Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?
di: Chen, Beiduo, et al.
Pubblicazione: (2024)
di: Chen, Beiduo, et al.
Pubblicazione: (2024)
Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization
di: Duan, Shaohua, et al.
Pubblicazione: (2025)
di: Duan, Shaohua, et al.
Pubblicazione: (2025)
FastFiD: Improve Inference Efficiency of Open Domain Question Answering via Sentence Selection
di: Huang, Yufei, et al.
Pubblicazione: (2024)
di: Huang, Yufei, et al.
Pubblicazione: (2024)
Adaptive Draft-Verification for Efficient Large Language Model Decoding
di: Liu, Xukun, et al.
Pubblicazione: (2024)
di: Liu, Xukun, et al.
Pubblicazione: (2024)
MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions
di: Zhu, Yanxu, et al.
Pubblicazione: (2025)
di: Zhu, Yanxu, et al.
Pubblicazione: (2025)
Breaking through Deterministic Barriers: Randomized Pruning Mask Generation and Selection
di: Li, Jianwei, et al.
Pubblicazione: (2023)
di: Li, Jianwei, et al.
Pubblicazione: (2023)
The Incomplete Bridge: How AI Research (Mis)Engages with Psychology
di: Jiang, Han, et al.
Pubblicazione: (2025)
di: Jiang, Han, et al.
Pubblicazione: (2025)
Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
di: Guo, Yiju, et al.
Pubblicazione: (2024)
di: Guo, Yiju, et al.
Pubblicazione: (2024)
Agent AI with LangGraph: A Modular Framework for Enhancing Machine Translation Using Large Language Models
di: Wang, Jialin, et al.
Pubblicazione: (2024)
di: Wang, Jialin, et al.
Pubblicazione: (2024)
Implicit Sentiment Analysis Based on Chain of Thought Prompting
di: Duan, Zhihua, et al.
Pubblicazione: (2024)
di: Duan, Zhihua, et al.
Pubblicazione: (2024)
Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation
di: Chen, Wei-Rui, et al.
Pubblicazione: (2025)
di: Chen, Wei-Rui, et al.
Pubblicazione: (2025)
CDEval: A Benchmark for Measuring the Cultural Dimensions of Large Language Models
di: Wang, Yuhang, et al.
Pubblicazione: (2023)
di: Wang, Yuhang, et al.
Pubblicazione: (2023)
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
di: Cao, Maosong, et al.
Pubblicazione: (2025)
di: Cao, Maosong, et al.
Pubblicazione: (2025)
The Essence of Contextual Understanding in Theory of Mind: A Study on Question Answering with Story Characters
di: Zhou, Chulun, et al.
Pubblicazione: (2025)
di: Zhou, Chulun, et al.
Pubblicazione: (2025)
The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment
di: Kim, HyunJin, et al.
Pubblicazione: (2024)
di: Kim, HyunJin, et al.
Pubblicazione: (2024)
Enhancing Multi-Agent Consensus through Third-Party LLM Integration: Analyzing Uncertainty and Mitigating Hallucinations in Large Language Models
di: Duan, Zhihua, et al.
Pubblicazione: (2024)
di: Duan, Zhihua, et al.
Pubblicazione: (2024)
Documenti analoghi
-
AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
di: Yao, Jing, et al.
Pubblicazione: (2025) -
IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
di: Bai, Yuzhuo, et al.
Pubblicazione: (2025) -
Leveraging Implicit Sentiments: Enhancing Reliability and Validity in Psychological Trait Evaluation of LLMs
di: Ma, Huanhuan, et al.
Pubblicazione: (2025) -
Beyond Human Norms: Unveiling Unique Values of Large Language Models through Interdisciplinary Approaches
di: Biedma, Pablo, et al.
Pubblicazione: (2024) -
Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization
di: Duan, Shitong, et al.
Pubblicazione: (2024)