Intermediate direct preference optimization
Fuente:
arXiv
Salvato in:
| Autore principale: | Kojima, Atsushi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Self-Supervised Learning for Multi-Channel Neural Transducer
di: Kojima, Atsushi
Pubblicazione: (2024)
di: Kojima, Atsushi
Pubblicazione: (2024)
Mitigating LLM biases toward spurious social contexts using direct preference optimization
di: Nam, Hyunji, et al.
Pubblicazione: (2026)
di: Nam, Hyunji, et al.
Pubblicazione: (2026)
REFA: Reference Free Alignment for multi-preference optimization
di: Gupta, Taneesh, et al.
Pubblicazione: (2024)
di: Gupta, Taneesh, et al.
Pubblicazione: (2024)
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
di: Sudo, Yui, et al.
Pubblicazione: (2025)
di: Sudo, Yui, et al.
Pubblicazione: (2025)
AC/DC: LLM-based Audio Comprehension via Dialogue Continuation
di: Fujita, Yusuke, et al.
Pubblicazione: (2025)
di: Fujita, Yusuke, et al.
Pubblicazione: (2025)
Automatic design optimization of preference-based subjective evaluation with online learning in crowdsourcing environment
di: Yasuda, Yusuke, et al.
Pubblicazione: (2024)
di: Yasuda, Yusuke, et al.
Pubblicazione: (2024)
REFINER: Reasoning Feedback on Intermediate Representations
di: Paul, Debjit, et al.
Pubblicazione: (2023)
di: Paul, Debjit, et al.
Pubblicazione: (2023)
Enhancing BERTopic with Intermediate Layer Representations
di: Koterwa, Dominik, et al.
Pubblicazione: (2025)
di: Koterwa, Dominik, et al.
Pubblicazione: (2025)
Aligning language models with human preferences
di: Korbak, Tomasz
Pubblicazione: (2024)
di: Korbak, Tomasz
Pubblicazione: (2024)
Unimodal Intermediate Training for Multimodal Meme Sentiment Classification
di: Hazman, Muzhaffar, et al.
Pubblicazione: (2023)
di: Hazman, Muzhaffar, et al.
Pubblicazione: (2023)
Uncovering Intermediate Variables in Transformers using Circuit Probing
di: Lepori, Michael A., et al.
Pubblicazione: (2023)
di: Lepori, Michael A., et al.
Pubblicazione: (2023)
Compositional preference models for aligning LMs
di: Go, Dongyoung, et al.
Pubblicazione: (2023)
di: Go, Dongyoung, et al.
Pubblicazione: (2023)
Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition
di: Shi, Hao, et al.
Pubblicazione: (2025)
di: Shi, Hao, et al.
Pubblicazione: (2025)
UniSparse: An Intermediate Language for General Sparse Format Customization
di: Liu, Jie, et al.
Pubblicazione: (2024)
di: Liu, Jie, et al.
Pubblicazione: (2024)
LIRE: listwise reward enhancement for preference alignment
di: Zhu, Mingye, et al.
Pubblicazione: (2024)
di: Zhu, Mingye, et al.
Pubblicazione: (2024)
THOUGHTSCULPT: Reasoning with Intermediate Revision and Search
di: Chi, Yizhou, et al.
Pubblicazione: (2024)
di: Chi, Yizhou, et al.
Pubblicazione: (2024)
Dafny as Verification-Aware Intermediate Language for Code Generation
di: Li, Yue Chen, et al.
Pubblicazione: (2025)
di: Li, Yue Chen, et al.
Pubblicazione: (2025)
Exploring the Effectiveness and Consistency of Task Selection in Intermediate-Task Transfer Learning
di: Lin, Pin-Jie, et al.
Pubblicazione: (2024)
di: Lin, Pin-Jie, et al.
Pubblicazione: (2024)
The Uneven Impact of Post-Training Quantization in Machine Translation
di: Marie, Benjamin, et al.
Pubblicazione: (2025)
di: Marie, Benjamin, et al.
Pubblicazione: (2025)
Reservoir Computing as a Language Model
di: Köster, Felix, et al.
Pubblicazione: (2025)
di: Köster, Felix, et al.
Pubblicazione: (2025)
Transcribe, Translate, or Transliterate: An Investigation of Intermediate Representations in Spoken Language Models
di: Ògúnrèmí, Tolúlopé, et al.
Pubblicazione: (2025)
di: Ògúnrèmí, Tolúlopé, et al.
Pubblicazione: (2025)
NL-Debugging: Exploiting Natural Language as an Intermediate Representation for Code Debugging
di: Zhang, Weiming, et al.
Pubblicazione: (2025)
di: Zhang, Weiming, et al.
Pubblicazione: (2025)
LiteToken: Removing Intermediate Merge Residues From BPE Tokenizers
di: Sun, Yike, et al.
Pubblicazione: (2026)
di: Sun, Yike, et al.
Pubblicazione: (2026)
Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics
di: Al-Khalili, Zena, et al.
Pubblicazione: (2025)
di: Al-Khalili, Zena, et al.
Pubblicazione: (2025)
Deduplicating and Ranking Solution Programs for Suggesting Reference Solutions
di: Shirafuji, Atsushi, et al.
Pubblicazione: (2023)
di: Shirafuji, Atsushi, et al.
Pubblicazione: (2023)
Everyone prefers human writers, including AI
di: Haverals, Wouter, et al.
Pubblicazione: (2025)
di: Haverals, Wouter, et al.
Pubblicazione: (2025)
The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious
di: Chua, James, et al.
Pubblicazione: (2026)
di: Chua, James, et al.
Pubblicazione: (2026)
A Joint Study of Phrase Grounding and Task Performance in Vision and Language Models
di: Kojima, Noriyuki, et al.
Pubblicazione: (2023)
di: Kojima, Noriyuki, et al.
Pubblicazione: (2023)
MiCoTA: Bridging the Learnability Gap with Intermediate CoT and Teacher Assistants
di: Ding, Dongyi, et al.
Pubblicazione: (2025)
di: Ding, Dongyi, et al.
Pubblicazione: (2025)
IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators
di: Paul, Indraneil, et al.
Pubblicazione: (2024)
di: Paul, Indraneil, et al.
Pubblicazione: (2024)
On the Multilingual Ability of Decoder-based Pre-trained Language Models: Finding and Controlling Language-Specific Neurons
di: Kojima, Takeshi, et al.
Pubblicazione: (2024)
di: Kojima, Takeshi, et al.
Pubblicazione: (2024)
Dynamic Injection of Entity Knowledge into Dense Retrievers
di: Yamada, Ikuya, et al.
Pubblicazione: (2025)
di: Yamada, Ikuya, et al.
Pubblicazione: (2025)
Automated Refinement of Essay Scoring Rubrics for Language Models via Reflect-and-Revise
di: Harada, Keno, et al.
Pubblicazione: (2025)
di: Harada, Keno, et al.
Pubblicazione: (2025)
Inference time LLM alignment in single and multidomain preference spectrum
di: Shahriar, Sadat, et al.
Pubblicazione: (2024)
di: Shahriar, Sadat, et al.
Pubblicazione: (2024)
MOSLIM:Align with diverse preferences in prompts through reward classification
di: Zhang, Yu, et al.
Pubblicazione: (2025)
di: Zhang, Yu, et al.
Pubblicazione: (2025)
Step Potential Advantage Estimation: Harnessing Intermediate Confidence and Correctness for Efficient Mathematical Reasoning
di: Wu, Fei, et al.
Pubblicazione: (2026)
di: Wu, Fei, et al.
Pubblicazione: (2026)
Beyond the Final Layer: Intermediate Representations for Better Multilingual Calibration in Large Language Models
di: Zhou, Ej, et al.
Pubblicazione: (2025)
di: Zhou, Ej, et al.
Pubblicazione: (2025)
Task-Specific Knowledge Distillation via Intermediate Probes
di: Brown, Ryan, et al.
Pubblicazione: (2026)
di: Brown, Ryan, et al.
Pubblicazione: (2026)
PizzaCommonSense: Learning to Model Commonsense Reasoning about Intermediate Steps in Cooking Recipes
di: Diallo, Aissatou, et al.
Pubblicazione: (2024)
di: Diallo, Aissatou, et al.
Pubblicazione: (2024)
Explainable Multi-hop Question Generation: An End-to-End Approach without Intermediate Question Labeling
di: Hwang, Seonjeong, et al.
Pubblicazione: (2024)
di: Hwang, Seonjeong, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Self-Supervised Learning for Multi-Channel Neural Transducer
di: Kojima, Atsushi
Pubblicazione: (2024) -
Mitigating LLM biases toward spurious social contexts using direct preference optimization
di: Nam, Hyunji, et al.
Pubblicazione: (2026) -
REFA: Reference Free Alignment for multi-preference optimization
di: Gupta, Taneesh, et al.
Pubblicazione: (2024) -
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
di: Sudo, Yui, et al.
Pubblicazione: (2025) -
AC/DC: LLM-based Audio Comprehension via Dialogue Continuation
di: Fujita, Yusuke, et al.
Pubblicazione: (2025)