Capability-Oriented Training Induced Alignment Risk
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Yujun, Huang, Yue, Bao, Han, Guo, Kehan, Liang, Zhenwen, Chen, Pin-Yu, Gao, Tian, Geyer, Werner, Moniz, Nuno, Chawla, Nitesh V, Zhang, Xiangliang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
by: Zhou, Yujun, et al.
Published: (2024)
by: Zhou, Yujun, et al.
Published: (2024)
Defending Jailbreak Prompts via In-Context Adversarial Game
by: Zhou, Yujun, et al.
Published: (2024)
by: Zhou, Yujun, et al.
Published: (2024)
AnyLoss: Transforming Classification Metrics into Loss Functions
by: Han, Doheon, et al.
Published: (2024)
by: Han, Doheon, et al.
Published: (2024)
Fast Explanations via Policy Gradient-Optimized Explainer
by: Pan, Deng, et al.
Published: (2024)
by: Pan, Deng, et al.
Published: (2024)
Conformalized Selective Regression
by: Sokol, Anna, et al.
Published: (2024)
by: Sokol, Anna, et al.
Published: (2024)
Differentially-Private Data Synthetisation for Efficient Re-Identification Risk Control
by: Carvalho, Tânia, et al.
Published: (2022)
by: Carvalho, Tânia, et al.
Published: (2022)
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
by: Ye, Jiayi, et al.
Published: (2024)
by: Ye, Jiayi, et al.
Published: (2024)
Intersectional Divergence: Measuring Fairness in Regression
by: Germino, Joe, et al.
Published: (2025)
by: Germino, Joe, et al.
Published: (2025)
BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks
by: Sokol, Anna, et al.
Published: (2024)
by: Sokol, Anna, et al.
Published: (2024)
Class-Aware Contrastive Optimization for Imbalanced Text Classification
by: Khvatskii, Grigorii, et al.
Published: (2024)
by: Khvatskii, Grigorii, et al.
Published: (2024)
From Verification Burden to Trusted Collaboration: Design Goals for LLM-Assisted Literature Reviews
by: Nogueira, Brenda, et al.
Published: (2025)
by: Nogueira, Brenda, et al.
Published: (2025)
Are we making much progress? Revisiting chemical reaction yield prediction from an imbalanced regression perspective
by: Ma, Yihong, et al.
Published: (2024)
by: Ma, Yihong, et al.
Published: (2024)
AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration - Learning from Cheap, Optimizing Expensive
by: Guo, Taicheng, et al.
Published: (2026)
by: Guo, Taicheng, et al.
Published: (2026)
Context Attribution with Multi-Armed Bandit Optimization
by: Pan, Deng, et al.
Published: (2025)
by: Pan, Deng, et al.
Published: (2025)
PolicyLLM: Towards Excellent Comprehension of Public Policy for Large Language Models
by: Bao, Han, et al.
Published: (2026)
by: Bao, Han, et al.
Published: (2026)
Dissecting Logical Reasoning in LLMs: A Fine-Grained Evaluation and Supervision Study
by: Zhou, Yujun, et al.
Published: (2025)
by: Zhou, Yujun, et al.
Published: (2025)
Spectral Manifold Harmonization for Graph Imbalanced Regression
by: Nogueira, Brenda, et al.
Published: (2025)
by: Nogueira, Brenda, et al.
Published: (2025)
Genotype-Conditioned Molecular Generation via Evidence-Grounded Multi-Objective Latent Perturbation in Diffusion Models
by: Nogueira, Brenda, et al.
Published: (2026)
by: Nogueira, Brenda, et al.
Published: (2026)
Causally-Enhanced Reinforcement Policy Optimization
by: Wang, Xiangqi, et al.
Published: (2025)
by: Wang, Xiangqi, et al.
Published: (2025)
SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark
by: Liang, Zhenwen, et al.
Published: (2024)
by: Liang, Zhenwen, et al.
Published: (2024)
Artificial Intelligence in Spectroscopy: Advancing Chemistry from Prediction to Generation and Beyond
by: Guo, Kehan, et al.
Published: (2025)
by: Guo, Kehan, et al.
Published: (2025)
Emergent Social Intelligence Risks in Generative Multi-Agent Systems
by: Huang, Yue, et al.
Published: (2026)
by: Huang, Yue, et al.
Published: (2026)
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data
by: Liang, Zhenwen, et al.
Published: (2026)
by: Liang, Zhenwen, et al.
Published: (2026)
AI Alignment Breaks at the Edge
by: Bao, Han, et al.
Published: (2026)
by: Bao, Han, et al.
Published: (2026)
SPECTRA: Spectral Domain-Aware Graph Generation for Imbalanced Molecular Property Regression
by: Nogueira, Brenda, et al.
Published: (2025)
by: Nogueira, Brenda, et al.
Published: (2025)
AdaReasoner: Adaptive Reasoning Enables More Flexible Thinking in Large Language Models
by: Wang, Xiangqi, et al.
Published: (2025)
by: Wang, Xiangqi, et al.
Published: (2025)
Prioritization First, Principles Second: An Adaptive Interpretation of Helpful, Honest, and Harmless Principles
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
ReactionTeam: Teaming Experts for Divergent Thinking Beyond Typical Reaction Patterns
by: Guo, Taicheng, et al.
Published: (2023)
by: Guo, Taicheng, et al.
Published: (2023)
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
by: Tian, Yijun, et al.
Published: (2024)
by: Tian, Yijun, et al.
Published: (2024)
Beyond Single-Value Metrics: Evaluating and Enhancing LLM Unlearning with Cognitive Diagnosis
by: Lang, Yicheng, et al.
Published: (2025)
by: Lang, Yicheng, et al.
Published: (2025)
UGMAE: A Unified Framework for Graph Masked Autoencoders
by: Tian, Yijun, et al.
Published: (2024)
by: Tian, Yijun, et al.
Published: (2024)
NGQA: A Nutritional Graph Question Answering Benchmark for Personalized Health-aware Nutritional Reasoning
by: Zhang, Zheyuan, et al.
Published: (2024)
by: Zhang, Zheyuan, et al.
Published: (2024)
ChemOrch: Empowering LLMs with Chemical Intelligence via Synthetic Instructions
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
ProbeLLM: Automating Principled Diagnosis of LLM Failures
by: Huang, Yue, et al.
Published: (2026)
by: Huang, Yue, et al.
Published: (2026)
Dual Optimal: Make Your LLM Peer-like with Dignity
by: Wang, Xiangqi, et al.
Published: (2026)
by: Wang, Xiangqi, et al.
Published: (2026)
SaSR-Net: Source-Aware Semantic Representation Network for Enhancing Audio-Visual Question Answering
by: Yang, Tianyu, et al.
Published: (2024)
by: Yang, Tianyu, et al.
Published: (2024)
Manipulating Predictions over Discrete Inputs in Machine Teaching
by: Wu, Xiaodong, et al.
Published: (2024)
by: Wu, Xiaodong, et al.
Published: (2024)
Improving LLM Reasoning through Scaling Inference Computation with Collaborative Verification
by: Liang, Zhenwen, et al.
Published: (2024)
by: Liang, Zhenwen, et al.
Published: (2024)
Relevance-aware Algorithmic Recourse
by: Kim, Dongwhi, et al.
Published: (2024)
by: Kim, Dongwhi, et al.
Published: (2024)
A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning
by: Yang, Tianyu, et al.
Published: (2026)
by: Yang, Tianyu, et al.
Published: (2026)
Similar Items
-
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
by: Zhou, Yujun, et al.
Published: (2024) -
Defending Jailbreak Prompts via In-Context Adversarial Game
by: Zhou, Yujun, et al.
Published: (2024) -
AnyLoss: Transforming Classification Metrics into Loss Functions
by: Han, Doheon, et al.
Published: (2024) -
Fast Explanations via Policy Gradient-Optimized Explainer
by: Pan, Deng, et al.
Published: (2024) -
Conformalized Selective Regression
by: Sokol, Anna, et al.
Published: (2024)