Is Crowdsourcing Breaking Your Bank? Cost-Effective Fine-Tuning of Pre-trained Language Models with Proximal Policy Optimization
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Shuo, Kasneci, Gjergji |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
por: Kocak, Aysenur, et al.
Publicado: (2025)
por: Kocak, Aysenur, et al.
Publicado: (2025)
Probabilistic Aggregation and Targeted Embedding Optimization for Collective Moral Reasoning in Large Language Models
por: Yuan, Chenchen, et al.
Publicado: (2025)
por: Yuan, Chenchen, et al.
Publicado: (2025)
Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents
por: Kirchhof, Michael, et al.
Publicado: (2025)
por: Kirchhof, Michael, et al.
Publicado: (2025)
Emergent Abilities in Large Language Models: A Survey
por: Berti, Leonardo, et al.
Publicado: (2025)
por: Berti, Leonardo, et al.
Publicado: (2025)
From Confidence to Collapse in LLM Factual Robustness
por: Fastowski, Alina, et al.
Publicado: (2025)
por: Fastowski, Alina, et al.
Publicado: (2025)
Adoption of Explainable Natural Language Processing: Perspectives from Industry and Academia on Practices and Challenges
por: Dhaini, Mahdi, et al.
Publicado: (2025)
por: Dhaini, Mahdi, et al.
Publicado: (2025)
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
por: Leemann, Tobias, et al.
Publicado: (2024)
por: Leemann, Tobias, et al.
Publicado: (2024)
Can LLM-Generated Textual Explanations Enhance Model Classification Performance? An Empirical Study
por: Dhaini, Mahdi, et al.
Publicado: (2025)
por: Dhaini, Mahdi, et al.
Publicado: (2025)
Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
por: Kasneci, Enkelejda, et al.
Publicado: (2026)
por: Kasneci, Enkelejda, et al.
Publicado: (2026)
EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models
por: Dhaini, Mahdi, et al.
Publicado: (2025)
por: Dhaini, Mahdi, et al.
Publicado: (2025)
Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
por: Fastowski, Alina, et al.
Publicado: (2025)
por: Fastowski, Alina, et al.
Publicado: (2025)
Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods
por: Dhaini, Mahdi, et al.
Publicado: (2025)
por: Dhaini, Mahdi, et al.
Publicado: (2025)
Dynamic Adaptive Optimization for Effective Sentiment Analysis Fine-Tuning on Large Language Models
por: Ding, Hongcheng, et al.
Publicado: (2024)
por: Ding, Hongcheng, et al.
Publicado: (2024)
Proximal Supervised Fine-Tuning
por: Zhu, Wenhong, et al.
Publicado: (2025)
por: Zhu, Wenhong, et al.
Publicado: (2025)
CamemBERT-bio: Leveraging Continual Pre-training for Cost-Effective Models on French Biomedical Data
por: Touchent, Rian, et al.
Publicado: (2023)
por: Touchent, Rian, et al.
Publicado: (2023)
Legal Documents Drafting with Fine-Tuned Pre-Trained Large Language Model
por: Lin, Chun-Hsien, et al.
Publicado: (2024)
por: Lin, Chun-Hsien, et al.
Publicado: (2024)
Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs
por: Yuan, Chenchen, et al.
Publicado: (2026)
por: Yuan, Chenchen, et al.
Publicado: (2026)
Doubling Your Data in Minutes: Ultra-fast Tabular Data Generation via LLM-Induced Dependency Graphs
por: Yang, Shuo, et al.
Publicado: (2025)
por: Yang, Shuo, et al.
Publicado: (2025)
SPAFIT: Stratified Progressive Adaptation Fine-tuning for Pre-trained Large Language Models
por: Arora, Samir, et al.
Publicado: (2024)
por: Arora, Samir, et al.
Publicado: (2024)
Model Fusion through Bayesian Optimization in Language Model Fine-Tuning
por: Jang, Chaeyun, et al.
Publicado: (2024)
por: Jang, Chaeyun, et al.
Publicado: (2024)
Fine-Tuning Language Models with Reward Learning on Policy
por: Lang, Hao, et al.
Publicado: (2024)
por: Lang, Hao, et al.
Publicado: (2024)
Topic Over Source: The Key to Effective Data Mixing for Language Models Pre-training
por: Peng, Jiahui, et al.
Publicado: (2025)
por: Peng, Jiahui, et al.
Publicado: (2025)
Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
por: Lin, Tzu-Quan, et al.
Publicado: (2025)
por: Lin, Tzu-Quan, et al.
Publicado: (2025)
Active Tabular Augmentation via Policy-Guided Diffusion Inpainting
por: Zhang, Zheyu, et al.
Publicado: (2026)
por: Zhang, Zheyu, et al.
Publicado: (2026)
Sparse is Enough in Fine-tuning Pre-trained Large Language Models
por: Song, Weixi, et al.
Publicado: (2023)
por: Song, Weixi, et al.
Publicado: (2023)
Amuro and Char: Analyzing the Relationship between Pre-Training and Fine-Tuning of Large Language Models
por: Sun, Kaiser, et al.
Publicado: (2024)
por: Sun, Kaiser, et al.
Publicado: (2024)
Online Causal Kalman Filtering for Stable and Effective Policy Optimization
por: He, Shuo, et al.
Publicado: (2026)
por: He, Shuo, et al.
Publicado: (2026)
Beyond Fine-Tuning: Effective Strategies for Mitigating Hallucinations in Large Language Models for Data Analytics
por: Rumiantsau, Mikhail, et al.
Publicado: (2024)
por: Rumiantsau, Mikhail, et al.
Publicado: (2024)
Enriching Tabular Data with Contextual LLM Embeddings: A Comprehensive Ablation Study for Ensemble Classifiers
por: Kasneci, Gjergji, et al.
Publicado: (2024)
por: Kasneci, Gjergji, et al.
Publicado: (2024)
Cost-efficient Crowdsourcing for Span-based Sequence Labeling: Worker Selection and Data Augmentation
por: Wang, Yujie, et al.
Publicado: (2023)
por: Wang, Yujie, et al.
Publicado: (2023)
Not All Features Deserve Attention: Graph-Guided Dependency Learning for Tabular Data Generation with Language Models
por: Zhang, Zheyu, et al.
Publicado: (2025)
por: Zhang, Zheyu, et al.
Publicado: (2025)
DEFT: Data Efficient Fine-Tuning for Pre-Trained Language Models via Unsupervised Core-Set Selection
por: Das, Devleena, et al.
Publicado: (2023)
por: Das, Devleena, et al.
Publicado: (2023)
DataMan: Data Manager for Pre-training Large Language Models
por: Peng, Ru, et al.
Publicado: (2025)
por: Peng, Ru, et al.
Publicado: (2025)
Fine Tuning Large Language Models for Medicine: The Role and Importance of Direct Preference Optimization
por: Savage, Thomas, et al.
Publicado: (2024)
por: Savage, Thomas, et al.
Publicado: (2024)
Optimizing Language Models for Grammatical Acceptability: A Comparative Study of Fine-Tuning Techniques
por: Ratan, Shobhit, et al.
Publicado: (2025)
por: Ratan, Shobhit, et al.
Publicado: (2025)
Overtrained Language Models Are Harder to Fine-Tune
por: Springer, Jacob Mitchell, et al.
Publicado: (2025)
por: Springer, Jacob Mitchell, et al.
Publicado: (2025)
Memorization in Fine-Tuned Large Language Models
por: Savine, Danil
Publicado: (2025)
por: Savine, Danil
Publicado: (2025)
Probing Language Models for Pre-training Data Detection
por: Liu, Zhenhua, et al.
Publicado: (2024)
por: Liu, Zhenhua, et al.
Publicado: (2024)
RAZOR: Sharpening Knowledge by Cutting Bias with Unsupervised Text Rewriting
por: Yang, Shuo, et al.
Publicado: (2024)
por: Yang, Shuo, et al.
Publicado: (2024)
Consolidating Rewarded Perturbations for LLM Post-Training
por: Zhang, Zheyu, et al.
Publicado: (2026)
por: Zhang, Zheyu, et al.
Publicado: (2026)
Ejemplares similares
-
CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
por: Kocak, Aysenur, et al.
Publicado: (2025) -
Probabilistic Aggregation and Targeted Embedding Optimization for Collective Moral Reasoning in Large Language Models
por: Yuan, Chenchen, et al.
Publicado: (2025) -
Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents
por: Kirchhof, Michael, et al.
Publicado: (2025) -
Emergent Abilities in Large Language Models: A Survey
por: Berti, Leonardo, et al.
Publicado: (2025) -
From Confidence to Collapse in LLM Factual Robustness
por: Fastowski, Alina, et al.
Publicado: (2025)