A Statistical Framework for Alignment with Biased AI Feedback
Fuente:
arXiv
Guardado en:
| Autores principales: | Xia, Xintao, Xia, Zhiqiu, Zhang, Linjun, Cai, Zhanrui |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Distribution-Free Uncertainty Quantification for Continuous AI Agent Evaluation
por: Gao, Yuxuan, et al.
Publicado: (2026)
por: Gao, Yuxuan, et al.
Publicado: (2026)
Augur: Modeling Covariate Causal Associations in Time Series via Large Language Models
por: Cui, Zhiqing, et al.
Publicado: (2025)
por: Cui, Zhiqing, et al.
Publicado: (2025)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
por: Liu, Fangxin, et al.
Publicado: (2025)
por: Liu, Fangxin, et al.
Publicado: (2025)
Enhancing Automated Essay Scoring with Three Techniques: Two-Stage Fine-Tuning, Score Alignment, and Self-Training
por: Choi, Hongseok, et al.
Publicado: (2026)
por: Choi, Hongseok, et al.
Publicado: (2026)
Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
por: Geuter, Jonathan, et al.
Publicado: (2025)
por: Geuter, Jonathan, et al.
Publicado: (2025)
Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
por: Ahmadian, Arash, et al.
Publicado: (2024)
por: Ahmadian, Arash, et al.
Publicado: (2024)
A Framework for Fine-Tuning LLMs using Heterogeneous Feedback
por: Aponte, Ryan, et al.
Publicado: (2024)
por: Aponte, Ryan, et al.
Publicado: (2024)
Constraint-Driven Small Language Models Based on Agent and OpenAlex Knowledge Graph: Mining Conceptual Pathways and Discovering Innovation Points in Academic Papers
por: Xia, Ziye, et al.
Publicado: (2025)
por: Xia, Ziye, et al.
Publicado: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
A Hybrid Framework for Real-Time Data Drift and Anomaly Identification Using Hierarchical Temporal Memory and Statistical Tests
por: Bandyopadhyay, Subhadip, et al.
Publicado: (2025)
por: Bandyopadhyay, Subhadip, et al.
Publicado: (2025)
Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs
por: Easley, Eric, et al.
Publicado: (2026)
por: Easley, Eric, et al.
Publicado: (2026)
ICE: Intervention-Consistent Explanation Evaluation with Statistical Grounding for LLMs
por: Basu, Abhinaba, et al.
Publicado: (2026)
por: Basu, Abhinaba, et al.
Publicado: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
por: Fadli, Samih
Publicado: (2025)
por: Fadli, Samih
Publicado: (2025)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
por: Schesch, Benedikt, et al.
Publicado: (2026)
por: Schesch, Benedikt, et al.
Publicado: (2026)
Multidimensional Analysis of Specific Language Impairment Using Unsupervised Learning Through PCA and Clustering
por: Selvanayagam, Niruthiha
Publicado: (2025)
por: Selvanayagam, Niruthiha
Publicado: (2025)
Attention Drift: What Autoregressive Speculative Decoding Models Learn
por: Eldenk, Doğaç, et al.
Publicado: (2026)
por: Eldenk, Doğaç, et al.
Publicado: (2026)
Why LoRA Resists Label Noise: A Theoretical Framework for Noise-Robust Parameter-Efficient Fine-Tuning
por: Steele, Brady
Publicado: (2026)
por: Steele, Brady
Publicado: (2026)
Lightweight Quantum-Enhanced ResNet for Coronary Angiography Classification: A Hybrid Quantum-Classical Feature Enhancement Framework
por: Xia, Jingsong
Publicado: (2026)
por: Xia, Jingsong
Publicado: (2026)
A methodological analysis of prompt perturbations and their effect on attack success rates
por: Machado, Tiago, et al.
Publicado: (2025)
por: Machado, Tiago, et al.
Publicado: (2025)
LLM Unlearning on Noisy Forget Sets: A Study of Incomplete, Rewritten, and Watermarked Data
por: Wang, Changsheng, et al.
Publicado: (2025)
por: Wang, Changsheng, et al.
Publicado: (2025)
RAPTOR-AI for Disaster OODA Loop: Hierarchical Multimodal RAG with Experience-Driven Agentic Decision-Making
por: Yasuno, Takato
Publicado: (2026)
por: Yasuno, Takato
Publicado: (2026)
An experimental study of KV cache reuse strategies in chunk-level caching systems
por: Cestola, Samuel, et al.
Publicado: (2026)
por: Cestola, Samuel, et al.
Publicado: (2026)
Behavioural Analysis of Alignment Faking
por: Hadida, Nathaniel Mitrani, et al.
Publicado: (2026)
por: Hadida, Nathaniel Mitrani, et al.
Publicado: (2026)
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
por: Zhang, Yizhuo, et al.
Publicado: (2025)
por: Zhang, Yizhuo, et al.
Publicado: (2025)
ProactBench: Beyond What The User Asked For
por: Harfi, Sepehr, et al.
Publicado: (2026)
por: Harfi, Sepehr, et al.
Publicado: (2026)
Latent-Autoregressive GP-VAE Language Model
por: Ruffenach, Yves
Publicado: (2025)
por: Ruffenach, Yves
Publicado: (2025)
Beyond the Black Box: A Statistical Model for LLM Reasoning and Inference
por: Dalal, Siddhartha, et al.
Publicado: (2024)
por: Dalal, Siddhartha, et al.
Publicado: (2024)
Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
por: Ye, Hua, et al.
Publicado: (2025)
por: Ye, Hua, et al.
Publicado: (2025)
Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs
por: Zhang, Zhaowei, et al.
Publicado: (2025)
por: Zhang, Zhaowei, et al.
Publicado: (2025)
EvilGenie: A Reward Hacking Benchmark
por: Gabor, Jonathan, et al.
Publicado: (2025)
por: Gabor, Jonathan, et al.
Publicado: (2025)
Statistical Measures for Explainable Aspect-Based Sentiment Analysis: A Case Study on Environmental Discourse in Reddit
por: Stracqualursi, Luisa, et al.
Publicado: (2026)
por: Stracqualursi, Luisa, et al.
Publicado: (2026)
Memory-Efficient Differentially Private Training with Gradient Random Projection
por: Mulrooney, Alex, et al.
Publicado: (2025)
por: Mulrooney, Alex, et al.
Publicado: (2025)
Random Heterogeneous Neurochaos Learning Architecture for Data Classification
por: S, Remya Ajai A, et al.
Publicado: (2024)
por: S, Remya Ajai A, et al.
Publicado: (2024)
The Hidden Attention of Mamba Models
por: Ali, Ameen, et al.
Publicado: (2024)
por: Ali, Ameen, et al.
Publicado: (2024)
The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback
por: Adapala, Sai Teja Reddy
Publicado: (2025)
por: Adapala, Sai Teja Reddy
Publicado: (2025)
FETILDA: An Effective Framework For Fin-tuned Embeddings For Long Financial Text Documents
por: Xia, Bolun "Namir", et al.
Publicado: (2022)
por: Xia, Bolun "Namir", et al.
Publicado: (2022)
Hierarchical Shift Mixing -- Beyond Dense Attention in Transformers
por: Forchheimer, Robert
Publicado: (2026)
por: Forchheimer, Robert
Publicado: (2026)
On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs
por: Zhao, Rosie, et al.
Publicado: (2026)
por: Zhao, Rosie, et al.
Publicado: (2026)
Procedural Environment Generation for Tool-Use Agents
por: Sullivan, Michael, et al.
Publicado: (2025)
por: Sullivan, Michael, et al.
Publicado: (2025)
On the Challenges of Creating Datasets for Analyzing Commercial Sex Advertisements to Assess Human Trafficking Risk and Organized Activity
por: Rivas, Pablo, et al.
Publicado: (2024)
por: Rivas, Pablo, et al.
Publicado: (2024)
Ejemplares similares
-
Distribution-Free Uncertainty Quantification for Continuous AI Agent Evaluation
por: Gao, Yuxuan, et al.
Publicado: (2026) -
Augur: Modeling Covariate Causal Associations in Time Series via Large Language Models
por: Cui, Zhiqing, et al.
Publicado: (2025) -
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
por: Liu, Fangxin, et al.
Publicado: (2025) -
Enhancing Automated Essay Scoring with Three Techniques: Two-Stage Fine-Tuning, Score Alignment, and Self-Training
por: Choi, Hongseok, et al.
Publicado: (2026) -
Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
por: Geuter, Jonathan, et al.
Publicado: (2025)