Model Spec Midtraining: Improving How Alignment Training Generalizes
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Chloe, Wichers, Nevan, Price, Sara, Marks, Samuel, Kutasov, Jon |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Visualizing Neural Network Imagination
por: Wichers, Nevan, et al.
Publicado: (2024)
por: Wichers, Nevan, et al.
Publicado: (2024)
Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation
por: Cao, Meng, et al.
Publicado: (2024)
por: Cao, Meng, et al.
Publicado: (2024)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
por: Shu, Lei, et al.
Publicado: (2023)
por: Shu, Lei, et al.
Publicado: (2023)
Optimizing AI Agent Attacks With Synthetic Data
por: Loughridge, Chloe, et al.
Publicado: (2025)
por: Loughridge, Chloe, et al.
Publicado: (2025)
Evaluating Control Protocols for Untrusted AI Agents
por: Kutasov, Jon, et al.
Publicado: (2025)
por: Kutasov, Jon, et al.
Publicado: (2025)
Gradient-Based Language Model Red Teaming
por: Wichers, Nevan, et al.
Publicado: (2024)
por: Wichers, Nevan, et al.
Publicado: (2024)
Recontextualization Mitigates Specification Gaming without Modifying the Specification
por: Azarbal, Ariana, et al.
Publicado: (2025)
por: Azarbal, Ariana, et al.
Publicado: (2025)
SpecAlign: A Semantic Alignment Framework for SystemVerilog Assertion Generation
por: Imperial, Jaime Rafael, et al.
Publicado: (2026)
por: Imperial, Jaime Rafael, et al.
Publicado: (2026)
MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining
por: Wen, Bingbing, et al.
Publicado: (2026)
por: Wen, Bingbing, et al.
Publicado: (2026)
The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
por: Marks, Samuel, et al.
Publicado: (2023)
por: Marks, Samuel, et al.
Publicado: (2023)
Robustly Improving LLM Fairness in Realistic Settings via Interpretability
por: Karvonen, Adam, et al.
Publicado: (2025)
por: Karvonen, Adam, et al.
Publicado: (2025)
EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training
por: Du, Yiyang, et al.
Publicado: (2026)
por: Du, Yiyang, et al.
Publicado: (2026)
Unsupervised decoding of encoded reasoning using language model interpretability
por: Fang, Ching, et al.
Publicado: (2025)
por: Fang, Ching, et al.
Publicado: (2025)
SpecMoE: Spectral Mixture-of-Experts Foundation Model for Cross-Species EEG Decoding
por: Darankoum, Davy, et al.
Publicado: (2026)
por: Darankoum, Davy, et al.
Publicado: (2026)
Improving Code Generation by Training with Natural Language Feedback
por: Chen, Angelica, et al.
Publicado: (2023)
por: Chen, Angelica, et al.
Publicado: (2023)
VeriMoA: A Mixture-of-Agents Framework for Spec-to-HDL Generation
por: Ping, Heng, et al.
Publicado: (2025)
por: Ping, Heng, et al.
Publicado: (2025)
SpecDB: LLM-Generated Customized Databases via Feature-Oriented Decomposition
por: Lou, Yunkai, et al.
Publicado: (2026)
por: Lou, Yunkai, et al.
Publicado: (2026)
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
por: Zhou, Yongchao, et al.
Publicado: (2023)
por: Zhou, Yongchao, et al.
Publicado: (2023)
DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
por: Li, Guanghao, et al.
Publicado: (2025)
por: Li, Guanghao, et al.
Publicado: (2025)
Liars' Bench: Evaluating Lie Detectors for Language Models
por: Kretschmar, Kieron, et al.
Publicado: (2025)
por: Kretschmar, Kieron, et al.
Publicado: (2025)
SpecPylot: Python Specification Generation using Large Language Models
por: Ayon, Ragib Shahariar, et al.
Publicado: (2026)
por: Ayon, Ragib Shahariar, et al.
Publicado: (2026)
Score-of-Mixture Training: Training One-Step Generative Models Made Simple via Score Estimation of Mixture Distributions
por: Jayashankar, Tejas, et al.
Publicado: (2025)
por: Jayashankar, Tejas, et al.
Publicado: (2025)
Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
por: Campbell, David, et al.
Publicado: (2026)
por: Campbell, David, et al.
Publicado: (2026)
SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding
por: Li, Shenggui, et al.
Publicado: (2026)
por: Li, Shenggui, et al.
Publicado: (2026)
Believe It or Not: How Deeply do LLMs Believe Implanted Facts?
por: Slocum, Stewart, et al.
Publicado: (2025)
por: Slocum, Stewart, et al.
Publicado: (2025)
Steering Evaluation-Aware Language Models to Act Like They Are Deployed
por: Hua, Tim Tian, et al.
Publicado: (2025)
por: Hua, Tim Tian, et al.
Publicado: (2025)
LLM-Assisted Repository-Level Generation with Structured Spec-Driven Engineering
por: Feng, Shuzhao, et al.
Publicado: (2026)
por: Feng, Shuzhao, et al.
Publicado: (2026)
DualSpec: Text-to-spatial-audio Generation via Dual-Spectrogram Guided Diffusion Model
por: Zhao, Lei, et al.
Publicado: (2025)
por: Zhao, Lei, et al.
Publicado: (2025)
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models
por: RRV, Aswin, et al.
Publicado: (2026)
por: RRV, Aswin, et al.
Publicado: (2026)
Improving Human-AI Coordination through Online Adversarial Training and Generative Models
por: Chaudhary, Paresh, et al.
Publicado: (2025)
por: Chaudhary, Paresh, et al.
Publicado: (2025)
SpecTM: Spectral Targeted Masking for Trustworthy Foundation Models
por: Imtiaz, Syed Usama, et al.
Publicado: (2026)
por: Imtiaz, Syed Usama, et al.
Publicado: (2026)
Introspection Adapters: Training LLMs to Report Their Learned Behaviors
por: Shenoy, Keshav, et al.
Publicado: (2026)
por: Shenoy, Keshav, et al.
Publicado: (2026)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
por: Huang, Haiduo, et al.
Publicado: (2025)
por: Huang, Haiduo, et al.
Publicado: (2025)
Stress-Testing Model Specs Reveals Character Differences among Language Models
por: Zhang, Jifan, et al.
Publicado: (2025)
por: Zhang, Jifan, et al.
Publicado: (2025)
LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
por: Yang, Chenghao, et al.
Publicado: (2025)
por: Yang, Chenghao, et al.
Publicado: (2025)
How Useful Is Cross-Domain Generalization for Training LLM Monitors?
por: Martin, Sam, et al.
Publicado: (2026)
por: Martin, Sam, et al.
Publicado: (2026)
Maia-2: A Unified Model for Human-AI Alignment in Chess
por: Tang, Zhenwei, et al.
Publicado: (2024)
por: Tang, Zhenwei, et al.
Publicado: (2024)
General Information Metrics for Improving AI Model Training Efficiency
por: Xu, Jianfeng, et al.
Publicado: (2025)
por: Xu, Jianfeng, et al.
Publicado: (2025)
Progressive Prompt Detailing for Improved Alignment in Text-to-Image Generative Models
por: Saichandran, Ketan Suhaas, et al.
Publicado: (2025)
por: Saichandran, Ketan Suhaas, et al.
Publicado: (2025)
AutoReSpec: A Framework for Generating Specification using Large Language Models
por: Ayon, Ragib Shahariar, et al.
Publicado: (2026)
por: Ayon, Ragib Shahariar, et al.
Publicado: (2026)
Ejemplares similares
-
Visualizing Neural Network Imagination
por: Wichers, Nevan, et al.
Publicado: (2024) -
Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation
por: Cao, Meng, et al.
Publicado: (2024) -
Fusion-Eval: Integrating Assistant Evaluators with LLMs
por: Shu, Lei, et al.
Publicado: (2023) -
Optimizing AI Agent Attacks With Synthetic Data
por: Loughridge, Chloe, et al.
Publicado: (2025) -
Evaluating Control Protocols for Untrusted AI Agents
por: Kutasov, Jon, et al.
Publicado: (2025)