InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Pengyu, Zhang, Dong, Li, Linyang, Tan, Chenkun, Wang, Xinghao, Ren, Ke, Jiang, Botian, Qiu, Xipeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Alignment in MLLM
von: Tan, Chenkun, et al.
Veröffentlicht: (2025)
von: Tan, Chenkun, et al.
Veröffentlicht: (2025)
MetaAlign: Align Large Language Models with Diverse Preferences during Inference Time
von: Zhang, Mozhi, et al.
Veröffentlicht: (2024)
von: Zhang, Mozhi, et al.
Veröffentlicht: (2024)
UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets
von: Wang, Pengyu, et al.
Veröffentlicht: (2025)
von: Wang, Pengyu, et al.
Veröffentlicht: (2025)
Sparser Block-Sparse Attention via Token Permutation
von: Wang, Xinghao, et al.
Veröffentlicht: (2025)
von: Wang, Xinghao, et al.
Veröffentlicht: (2025)
BitStack: Any-Size Compression of Large Language Models in Variable Memory Environments
von: Wang, Xinghao, et al.
Veröffentlicht: (2024)
von: Wang, Xinghao, et al.
Veröffentlicht: (2024)
DenoSent: A Denoising Objective for Self-Supervised Sentence Representation Learning
von: Wang, Xinghao, et al.
Veröffentlicht: (2024)
von: Wang, Xinghao, et al.
Veröffentlicht: (2024)
Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment
von: Roy, Abhinaba, et al.
Veröffentlicht: (2025)
von: Roy, Abhinaba, et al.
Veröffentlicht: (2025)
Dishonesty in Helpful and Harmless Alignment
von: Huang, Youcheng, et al.
Veröffentlicht: (2024)
von: Huang, Youcheng, et al.
Veröffentlicht: (2024)
Prism: Spectral-Aware Block-Sparse Attention
von: Wang, Xinghao, et al.
Veröffentlicht: (2026)
von: Wang, Xinghao, et al.
Veröffentlicht: (2026)
LongSafety: Enhance Safety for Long-Context LLMs
von: Huang, Mianqiu, et al.
Veröffentlicht: (2024)
von: Huang, Mianqiu, et al.
Veröffentlicht: (2024)
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance
von: Huang, Caishuang, et al.
Veröffentlicht: (2024)
von: Huang, Caishuang, et al.
Veröffentlicht: (2024)
UnitCoder: Scalable Iterative Code Synthesis with Unit Test Guidance
von: Ma, Yichuan, et al.
Veröffentlicht: (2025)
von: Ma, Yichuan, et al.
Veröffentlicht: (2025)
Inference-Time Language Model Alignment via Integrated Value Guidance
von: Liu, Zhixuan, et al.
Veröffentlicht: (2024)
von: Liu, Zhixuan, et al.
Veröffentlicht: (2024)
Aligner: Efficient Alignment by Learning to Correct
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
The Open-World Lottery Ticket Hypothesis for OOD Intent Classification
von: Zhou, Yunhua, et al.
Veröffentlicht: (2022)
von: Zhou, Yunhua, et al.
Veröffentlicht: (2022)
SpeechAgents: Human-Communication Simulation with Multi-Modal Multi-Agent Systems
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment
von: Shen, Gerald, et al.
Veröffentlicht: (2024)
von: Shen, Gerald, et al.
Veröffentlicht: (2024)
Balanced Data Sampling for Language Model Training with Clustering
von: Shao, Yunfan, et al.
Veröffentlicht: (2024)
von: Shao, Yunfan, et al.
Veröffentlicht: (2024)
Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance
von: Jin, Luozhijie, et al.
Veröffentlicht: (2025)
von: Jin, Luozhijie, et al.
Veröffentlicht: (2025)
Aligners: Decoupling LLMs and Alignment
von: Ngweta, Lilian, et al.
Veröffentlicht: (2024)
von: Ngweta, Lilian, et al.
Veröffentlicht: (2024)
ArcAligner: Adaptive Recursive Aligner for Compressed Context Embeddings in RAG
von: Li, Jianbo, et al.
Veröffentlicht: (2026)
von: Li, Jianbo, et al.
Veröffentlicht: (2026)
P-Aligner: Enabling Pre-Alignment of Language Models via Principled Instruction Synthesis
von: Song, Feifan, et al.
Veröffentlicht: (2025)
von: Song, Feifan, et al.
Veröffentlicht: (2025)
Model Utility Law: Evaluating LLMs beyond Performance through Mechanism Interpretable Metric
von: Cao, Yixin, et al.
Veröffentlicht: (2025)
von: Cao, Yixin, et al.
Veröffentlicht: (2025)
Safe Inputs but Unsafe Output: Benchmarking Cross-modality Safety Alignment of Large Vision-Language Model
von: Wang, Siyin, et al.
Veröffentlicht: (2024)
von: Wang, Siyin, et al.
Veröffentlicht: (2024)
MetaAligner: Towards Generalizable Multi-Objective Alignment of Language Models
von: Yang, Kailai, et al.
Veröffentlicht: (2024)
von: Yang, Kailai, et al.
Veröffentlicht: (2024)
Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction
von: Lou, Hantao, et al.
Veröffentlicht: (2025)
von: Lou, Hantao, et al.
Veröffentlicht: (2025)
UnifiedMLLM: Enabling Unified Representation for Multi-modal Multi-tasks With Large Language Model
von: Li, Zhaowei, et al.
Veröffentlicht: (2024)
von: Li, Zhaowei, et al.
Veröffentlicht: (2024)
Emergent Structured Representations Support Flexible In-Context Inference in Large Language Models
von: Xu, Ningyu, et al.
Veröffentlicht: (2026)
von: Xu, Ningyu, et al.
Veröffentlicht: (2026)
LLatrieval: LLM-Verified Retrieval for Verifiable Generation
von: Li, Xiaonan, et al.
Veröffentlicht: (2023)
von: Li, Xiaonan, et al.
Veröffentlicht: (2023)
Inference-Time Decontamination: Reusing Leaked Benchmarks for Large Language Model Evaluation
von: Zhu, Qin, et al.
Veröffentlicht: (2024)
von: Zhu, Qin, et al.
Veröffentlicht: (2024)
Agent Alignment in Evolving Social Norms
von: Li, Shimin, et al.
Veröffentlicht: (2024)
von: Li, Shimin, et al.
Veröffentlicht: (2024)
Multi-hop Reasoning via Early Knowledge Alignment
von: Wang, Yuxin, et al.
Veröffentlicht: (2025)
von: Wang, Yuxin, et al.
Veröffentlicht: (2025)
SpeechAlign: Aligning Speech Generation to Human Preferences
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
Alignment for Honesty
von: Yang, Yuqing, et al.
Veröffentlicht: (2023)
von: Yang, Yuqing, et al.
Veröffentlicht: (2023)
Revisiting Backdoor Attacks on LLMs: A Stealthy and Practical Poisoning Framework via Harmless Inputs
von: Kong, Jiawei, et al.
Veröffentlicht: (2025)
von: Kong, Jiawei, et al.
Veröffentlicht: (2025)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
von: Jiang, Botian, et al.
Veröffentlicht: (2024)
von: Jiang, Botian, et al.
Veröffentlicht: (2024)
Speculative Thinking: Enhancing Small-Model Reasoning with Large Model Guidance at Inference Time
von: Yang, Wang, et al.
Veröffentlicht: (2025)
von: Yang, Wang, et al.
Veröffentlicht: (2025)
SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
von: Wang, Qiaolin, et al.
Veröffentlicht: (2025)
von: Wang, Qiaolin, et al.
Veröffentlicht: (2025)
Balancing Enhancement, Harmlessness, and General Capabilities: Enhancing Conversational LLMs with Direct RLHF
von: Zheng, Chen, et al.
Veröffentlicht: (2024)
von: Zheng, Chen, et al.
Veröffentlicht: (2024)
MentorCollab: Selective Large-to-Small Inference-Time Guidance for Efficient Reasoning
von: Wang, Haojin, et al.
Veröffentlicht: (2026)
von: Wang, Haojin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Alignment in MLLM
von: Tan, Chenkun, et al.
Veröffentlicht: (2025) -
MetaAlign: Align Large Language Models with Diverse Preferences during Inference Time
von: Zhang, Mozhi, et al.
Veröffentlicht: (2024) -
UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets
von: Wang, Pengyu, et al.
Veröffentlicht: (2025) -
Sparser Block-Sparse Attention via Token Permutation
von: Wang, Xinghao, et al.
Veröffentlicht: (2025) -
BitStack: Any-Size Compression of Large Language Models in Variable Memory Environments
von: Wang, Xinghao, et al.
Veröffentlicht: (2024)