QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
Fuente:
arXiv
Saved in:
| Main Authors: | Dineen, Jacob, RRV, Aswin, Liu, Qin, Xu, Zhikun, Ye, Xiao, Shen, Ming, Li, Zhaonan, Lu, Shijie, Baral, Chitta, Chen, Muhao, Zhou, Ben |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ToW: Thoughts of Words Improve Reasoning in Large Language Models
by: Xu, Zhikun, et al.
Published: (2024)
by: Xu, Zhikun, et al.
Published: (2024)
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution
by: Dineen, Jacob, et al.
Published: (2026)
by: Dineen, Jacob, et al.
Published: (2026)
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models
by: RRV, Aswin, et al.
Published: (2026)
by: RRV, Aswin, et al.
Published: (2026)
ThinkTuning: Instilling Cognitive Reflections without Distillation
by: RRV, Aswin, et al.
Published: (2025)
by: RRV, Aswin, et al.
Published: (2025)
CC-LEARN: Cohort-based Consistency Learning
by: Ye, Xiao, et al.
Published: (2025)
by: Ye, Xiao, et al.
Published: (2025)
BOW: Reinforcement Learning for Bottlenecked Next Word Prediction
by: Shen, Ming, et al.
Published: (2025)
by: Shen, Ming, et al.
Published: (2025)
Triple Preference Optimization: Achieving Better Alignment using a Single Step Optimization
by: Saeidi, Amir, et al.
Published: (2024)
by: Saeidi, Amir, et al.
Published: (2024)
GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
by: Handa, Divij, et al.
Published: (2025)
by: Handa, Divij, et al.
Published: (2025)
Chaos with Keywords: Exposing Large Language Models Sycophantic Hallucination to Misleading Keywords and Evaluating Defense Strategies
by: RRV, Aswin, et al.
Published: (2024)
by: RRV, Aswin, et al.
Published: (2024)
When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers
by: Handa, Divij, et al.
Published: (2024)
by: Handa, Divij, et al.
Published: (2024)
Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?
by: Tyagi, Nemika, et al.
Published: (2024)
by: Tyagi, Nemika, et al.
Published: (2024)
PHANTOM RECALL: When Familiar Puzzles Fool Smart Models
by: Mukhopadhyay, Souradeep, et al.
Published: (2025)
by: Mukhopadhyay, Souradeep, et al.
Published: (2025)
Unbiased Visual Reasoning with Controlled Visual Inputs
by: Li, Zhaonan, et al.
Published: (2025)
by: Li, Zhaonan, et al.
Published: (2025)
Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
by: Ye, Xiao, et al.
Published: (2025)
by: Ye, Xiao, et al.
Published: (2025)
UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' Memorization
by: Uddin, Md Nayem, et al.
Published: (2024)
by: Uddin, Md Nayem, et al.
Published: (2024)
Answering Questions in Stages: Prompt Chaining for Contract QA
by: Roegiest, Adam, et al.
Published: (2024)
by: Roegiest, Adam, et al.
Published: (2024)
VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images
by: Li, Zhaonan, et al.
Published: (2026)
by: Li, Zhaonan, et al.
Published: (2026)
ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
by: Liu, Qin, et al.
Published: (2025)
by: Liu, Qin, et al.
Published: (2025)
NativQA: Multilingual Culturally-Aligned Natural Query for LLMs
by: Hasan, Md. Arid, et al.
Published: (2024)
by: Hasan, Md. Arid, et al.
Published: (2024)
Skill Reuse as Compression in Agentic RL
by: Xu, Zhikun, et al.
Published: (2026)
by: Xu, Zhikun, et al.
Published: (2026)
VideoQA in the Era of LLMs: An Empirical Study
by: Xiao, Junbin, et al.
Published: (2024)
by: Xiao, Junbin, et al.
Published: (2024)
MedConceptsQA: Open Source Medical Concepts QA Benchmark
by: Shoham, Ofir Ben, et al.
Published: (2024)
by: Shoham, Ofir Ben, et al.
Published: (2024)
LatentQA: Teaching LLMs to Decode Activations Into Natural Language
by: Pan, Alexander, et al.
Published: (2024)
by: Pan, Alexander, et al.
Published: (2024)
RJUA-QA: A Comprehensive QA Dataset for Urology
by: Lyu, Shiwei, et al.
Published: (2023)
by: Lyu, Shiwei, et al.
Published: (2023)
CON-QA: Privacy-Preserving QA using cloud LLMs in Contract Domain
by: Singh, Ajeet Kumar, et al.
Published: (2025)
by: Singh, Ajeet Kumar, et al.
Published: (2025)
MapQA: Open-domain Geospatial Question Answering on Map Data
by: Li, Zekun, et al.
Published: (2025)
by: Li, Zekun, et al.
Published: (2025)
SpatialMem: Metric-Aligned Long-Horizon Video Memory for Language Grounding and QA
by: Zheng, Xinyi, et al.
Published: (2026)
by: Zheng, Xinyi, et al.
Published: (2026)
MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA
by: Ye, Hanrong, et al.
Published: (2024)
by: Ye, Hanrong, et al.
Published: (2024)
SEC-QA: A Systematic Evaluation Corpus for Financial QA
by: Lai, Viet Dac, et al.
Published: (2024)
by: Lai, Viet Dac, et al.
Published: (2024)
ChatQA: Surpassing GPT-4 on Conversational QA and RAG
by: Liu, Zihan, et al.
Published: (2024)
by: Liu, Zihan, et al.
Published: (2024)
Beyond MedQA: Towards Real-world Clinical Decision Making in the Era of LLMs
by: Xiao, Yunpeng, et al.
Published: (2025)
by: Xiao, Yunpeng, et al.
Published: (2025)
GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
by: Siingh, Shikhhar, et al.
Published: (2025)
by: Siingh, Shikhhar, et al.
Published: (2025)
CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning
by: Lu, Zhiyuan, et al.
Published: (2026)
by: Lu, Zhiyuan, et al.
Published: (2026)
Learning to Correct for QA Reasoning with Black-box LLMs
by: Kim, Jaehyung, et al.
Published: (2024)
by: Kim, Jaehyung, et al.
Published: (2024)
Agentic Adversarial QA for Improving Domain-Specific LLMs
by: Grari, Vincent, et al.
Published: (2026)
by: Grari, Vincent, et al.
Published: (2026)
LocalRQA: From Generating Data to Locally Training, Testing, and Deploying Retrieval-Augmented QA Systems
by: Yu, Xiao, et al.
Published: (2024)
by: Yu, Xiao, et al.
Published: (2024)
AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation
by: Huang, Tiancheng, et al.
Published: (2025)
by: Huang, Tiancheng, et al.
Published: (2025)
When Facts Change: Probing LLMs on Evolving Knowledge with evolveQA
by: Nakshatri, Nishanth Sridhar, et al.
Published: (2025)
by: Nakshatri, Nishanth Sridhar, et al.
Published: (2025)
CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
by: Zhou, Xingcheng, et al.
Published: (2026)
by: Zhou, Xingcheng, et al.
Published: (2026)
SciTaRC: Benchmarking QA on Scientific Tabular Data that Requires Language Reasoning and Complex Computation
by: Wang, Hexuan, et al.
Published: (2026)
by: Wang, Hexuan, et al.
Published: (2026)
Similar Items
-
ToW: Thoughts of Words Improve Reasoning in Large Language Models
by: Xu, Zhikun, et al.
Published: (2024) -
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution
by: Dineen, Jacob, et al.
Published: (2026) -
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models
by: RRV, Aswin, et al.
Published: (2026) -
ThinkTuning: Instilling Cognitive Reflections without Distillation
by: RRV, Aswin, et al.
Published: (2025) -
CC-LEARN: Cohort-based Consistency Learning
by: Ye, Xiao, et al.
Published: (2025)