Capability-Oriented Training Induced Alignment Risk
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhou, Yujun, Huang, Yue, Bao, Han, Guo, Kehan, Liang, Zhenwen, Chen, Pin-Yu, Gao, Tian, Geyer, Werner, Moniz, Nuno, Chawla, Nitesh V, Zhang, Xiangliang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
por: Zhou, Yujun, et al.
Publicado: (2024)
por: Zhou, Yujun, et al.
Publicado: (2024)
Defending Jailbreak Prompts via In-Context Adversarial Game
por: Zhou, Yujun, et al.
Publicado: (2024)
por: Zhou, Yujun, et al.
Publicado: (2024)
AnyLoss: Transforming Classification Metrics into Loss Functions
por: Han, Doheon, et al.
Publicado: (2024)
por: Han, Doheon, et al.
Publicado: (2024)
Fast Explanations via Policy Gradient-Optimized Explainer
por: Pan, Deng, et al.
Publicado: (2024)
por: Pan, Deng, et al.
Publicado: (2024)
Conformalized Selective Regression
por: Sokol, Anna, et al.
Publicado: (2024)
por: Sokol, Anna, et al.
Publicado: (2024)
Differentially-Private Data Synthetisation for Efficient Re-Identification Risk Control
por: Carvalho, Tânia, et al.
Publicado: (2022)
por: Carvalho, Tânia, et al.
Publicado: (2022)
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
por: Ye, Jiayi, et al.
Publicado: (2024)
por: Ye, Jiayi, et al.
Publicado: (2024)
Intersectional Divergence: Measuring Fairness in Regression
por: Germino, Joe, et al.
Publicado: (2025)
por: Germino, Joe, et al.
Publicado: (2025)
BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks
por: Sokol, Anna, et al.
Publicado: (2024)
por: Sokol, Anna, et al.
Publicado: (2024)
Class-Aware Contrastive Optimization for Imbalanced Text Classification
por: Khvatskii, Grigorii, et al.
Publicado: (2024)
por: Khvatskii, Grigorii, et al.
Publicado: (2024)
From Verification Burden to Trusted Collaboration: Design Goals for LLM-Assisted Literature Reviews
por: Nogueira, Brenda, et al.
Publicado: (2025)
por: Nogueira, Brenda, et al.
Publicado: (2025)
Are we making much progress? Revisiting chemical reaction yield prediction from an imbalanced regression perspective
por: Ma, Yihong, et al.
Publicado: (2024)
por: Ma, Yihong, et al.
Publicado: (2024)
AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration - Learning from Cheap, Optimizing Expensive
por: Guo, Taicheng, et al.
Publicado: (2026)
por: Guo, Taicheng, et al.
Publicado: (2026)
Context Attribution with Multi-Armed Bandit Optimization
por: Pan, Deng, et al.
Publicado: (2025)
por: Pan, Deng, et al.
Publicado: (2025)
PolicyLLM: Towards Excellent Comprehension of Public Policy for Large Language Models
por: Bao, Han, et al.
Publicado: (2026)
por: Bao, Han, et al.
Publicado: (2026)
Dissecting Logical Reasoning in LLMs: A Fine-Grained Evaluation and Supervision Study
por: Zhou, Yujun, et al.
Publicado: (2025)
por: Zhou, Yujun, et al.
Publicado: (2025)
Spectral Manifold Harmonization for Graph Imbalanced Regression
por: Nogueira, Brenda, et al.
Publicado: (2025)
por: Nogueira, Brenda, et al.
Publicado: (2025)
Genotype-Conditioned Molecular Generation via Evidence-Grounded Multi-Objective Latent Perturbation in Diffusion Models
por: Nogueira, Brenda, et al.
Publicado: (2026)
por: Nogueira, Brenda, et al.
Publicado: (2026)
Causally-Enhanced Reinforcement Policy Optimization
por: Wang, Xiangqi, et al.
Publicado: (2025)
por: Wang, Xiangqi, et al.
Publicado: (2025)
SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark
por: Liang, Zhenwen, et al.
Publicado: (2024)
por: Liang, Zhenwen, et al.
Publicado: (2024)
Artificial Intelligence in Spectroscopy: Advancing Chemistry from Prediction to Generation and Beyond
por: Guo, Kehan, et al.
Publicado: (2025)
por: Guo, Kehan, et al.
Publicado: (2025)
Emergent Social Intelligence Risks in Generative Multi-Agent Systems
por: Huang, Yue, et al.
Publicado: (2026)
por: Huang, Yue, et al.
Publicado: (2026)
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data
por: Liang, Zhenwen, et al.
Publicado: (2026)
por: Liang, Zhenwen, et al.
Publicado: (2026)
AI Alignment Breaks at the Edge
por: Bao, Han, et al.
Publicado: (2026)
por: Bao, Han, et al.
Publicado: (2026)
SPECTRA: Spectral Domain-Aware Graph Generation for Imbalanced Molecular Property Regression
por: Nogueira, Brenda, et al.
Publicado: (2025)
por: Nogueira, Brenda, et al.
Publicado: (2025)
AdaReasoner: Adaptive Reasoning Enables More Flexible Thinking in Large Language Models
por: Wang, Xiangqi, et al.
Publicado: (2025)
por: Wang, Xiangqi, et al.
Publicado: (2025)
Prioritization First, Principles Second: An Adaptive Interpretation of Helpful, Honest, and Harmless Principles
por: Huang, Yue, et al.
Publicado: (2025)
por: Huang, Yue, et al.
Publicado: (2025)
ReactionTeam: Teaming Experts for Divergent Thinking Beyond Typical Reaction Patterns
por: Guo, Taicheng, et al.
Publicado: (2023)
por: Guo, Taicheng, et al.
Publicado: (2023)
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
por: Tian, Yijun, et al.
Publicado: (2024)
por: Tian, Yijun, et al.
Publicado: (2024)
Beyond Single-Value Metrics: Evaluating and Enhancing LLM Unlearning with Cognitive Diagnosis
por: Lang, Yicheng, et al.
Publicado: (2025)
por: Lang, Yicheng, et al.
Publicado: (2025)
UGMAE: A Unified Framework for Graph Masked Autoencoders
por: Tian, Yijun, et al.
Publicado: (2024)
por: Tian, Yijun, et al.
Publicado: (2024)
NGQA: A Nutritional Graph Question Answering Benchmark for Personalized Health-aware Nutritional Reasoning
por: Zhang, Zheyuan, et al.
Publicado: (2024)
por: Zhang, Zheyuan, et al.
Publicado: (2024)
ChemOrch: Empowering LLMs with Chemical Intelligence via Synthetic Instructions
por: Huang, Yue, et al.
Publicado: (2025)
por: Huang, Yue, et al.
Publicado: (2025)
ProbeLLM: Automating Principled Diagnosis of LLM Failures
por: Huang, Yue, et al.
Publicado: (2026)
por: Huang, Yue, et al.
Publicado: (2026)
Dual Optimal: Make Your LLM Peer-like with Dignity
por: Wang, Xiangqi, et al.
Publicado: (2026)
por: Wang, Xiangqi, et al.
Publicado: (2026)
SaSR-Net: Source-Aware Semantic Representation Network for Enhancing Audio-Visual Question Answering
por: Yang, Tianyu, et al.
Publicado: (2024)
por: Yang, Tianyu, et al.
Publicado: (2024)
Manipulating Predictions over Discrete Inputs in Machine Teaching
por: Wu, Xiaodong, et al.
Publicado: (2024)
por: Wu, Xiaodong, et al.
Publicado: (2024)
Improving LLM Reasoning through Scaling Inference Computation with Collaborative Verification
por: Liang, Zhenwen, et al.
Publicado: (2024)
por: Liang, Zhenwen, et al.
Publicado: (2024)
Relevance-aware Algorithmic Recourse
por: Kim, Dongwhi, et al.
Publicado: (2024)
por: Kim, Dongwhi, et al.
Publicado: (2024)
A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning
por: Yang, Tianyu, et al.
Publicado: (2026)
por: Yang, Tianyu, et al.
Publicado: (2026)
Ejemplares similares
-
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
por: Zhou, Yujun, et al.
Publicado: (2024) -
Defending Jailbreak Prompts via In-Context Adversarial Game
por: Zhou, Yujun, et al.
Publicado: (2024) -
AnyLoss: Transforming Classification Metrics into Loss Functions
por: Han, Doheon, et al.
Publicado: (2024) -
Fast Explanations via Policy Gradient-Optimized Explainer
por: Pan, Deng, et al.
Publicado: (2024) -
Conformalized Selective Regression
por: Sokol, Anna, et al.
Publicado: (2024)