Exploring Design Choices for Building Language-Specific LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Tejaswi, Atula, Gupta, Nilesh, Choi, Eunsol |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RARe: Retrieval Augmented Retrieval with In-Context Examples
di: Tejaswi, Atula, et al.
Pubblicazione: (2024)
di: Tejaswi, Atula, et al.
Pubblicazione: (2024)
Entropy Aware Reward Guidance for Diffusion Language Model Alignment
di: Tejaswi, Atula, et al.
Pubblicazione: (2026)
di: Tejaswi, Atula, et al.
Pubblicazione: (2026)
SVFT: Parameter-Efficient Fine-Tuning with Singular Vectors
di: Lingam, Vijay, et al.
Pubblicazione: (2024)
di: Lingam, Vijay, et al.
Pubblicazione: (2024)
PropMEND: Hypernetworks for Knowledge Propagation in LLMs
di: Liu, Zeyu Leo, et al.
Pubblicazione: (2025)
di: Liu, Zeyu Leo, et al.
Pubblicazione: (2025)
CaLMQA: Exploring culturally specific long-form question answering across 23 languages
di: Arora, Shane, et al.
Pubblicazione: (2024)
di: Arora, Shane, et al.
Pubblicazione: (2024)
Contrastive Learning to Improve Retrieval for Real-world Fact Checking
di: Sriram, Aniruddh, et al.
Pubblicazione: (2024)
di: Sriram, Aniruddh, et al.
Pubblicazione: (2024)
M2Lingual: Enhancing Multilingual, Multi-Turn Instruction Alignment in Large Language Models
di: Maheshwary, Rishabh, et al.
Pubblicazione: (2024)
di: Maheshwary, Rishabh, et al.
Pubblicazione: (2024)
Focus On This, Not That! Steering LLMs with Adaptive Feature Specification
di: Lamb, Tom A., et al.
Pubblicazione: (2024)
di: Lamb, Tom A., et al.
Pubblicazione: (2024)
Solo Connection: A Parameter Efficient Fine-Tuning Technique for Transformers
di: Pathak, Harsh Nilesh, et al.
Pubblicazione: (2025)
di: Pathak, Harsh Nilesh, et al.
Pubblicazione: (2025)
Scaling Rich Style-Prompted Text-to-Speech Datasets
di: Diwan, Anuj, et al.
Pubblicazione: (2025)
di: Diwan, Anuj, et al.
Pubblicazione: (2025)
Specialised or Generic? Tokenization Choices for Radiology Language Models
di: Warr, Hermione, et al.
Pubblicazione: (2025)
di: Warr, Hermione, et al.
Pubblicazione: (2025)
Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A
di: Plaut, Benjamin, et al.
Pubblicazione: (2024)
di: Plaut, Benjamin, et al.
Pubblicazione: (2024)
Agentic Adversarial QA for Improving Domain-Specific LLMs
di: Grari, Vincent, et al.
Pubblicazione: (2026)
di: Grari, Vincent, et al.
Pubblicazione: (2026)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
di: Chandak, Nikhil, et al.
Pubblicazione: (2025)
di: Chandak, Nikhil, et al.
Pubblicazione: (2025)
Multiple Choice Learning of Low-Rank Adapters for Language Modeling
di: Letzelter, Victor, et al.
Pubblicazione: (2025)
di: Letzelter, Victor, et al.
Pubblicazione: (2025)
TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs
di: Hu, Lanxiang, et al.
Pubblicazione: (2024)
di: Hu, Lanxiang, et al.
Pubblicazione: (2024)
UnibucLLM: Harnessing LLMs for Automated Prediction of Item Difficulty and Response Time for Multiple-Choice Questions
di: Rogoz, Ana-Cristina, et al.
Pubblicazione: (2024)
di: Rogoz, Ana-Cristina, et al.
Pubblicazione: (2024)
Configurable Foundation Models: Building LLMs from a Modular Perspective
di: Xiao, Chaojun, et al.
Pubblicazione: (2024)
di: Xiao, Chaojun, et al.
Pubblicazione: (2024)
Benchmarking for Domain-Specific LLMs: A Case Study on Academia and Beyond
di: Chen, Rubing, et al.
Pubblicazione: (2025)
di: Chen, Rubing, et al.
Pubblicazione: (2025)
Shears: Unstructured Sparsity with Neural Low-rank Adapter Search
di: Muñoz, J. Pablo, et al.
Pubblicazione: (2024)
di: Muñoz, J. Pablo, et al.
Pubblicazione: (2024)
SQFT: Low-cost Model Adaptation in Low-precision Sparse Foundation Models
di: Muñoz, Juan Pablo, et al.
Pubblicazione: (2024)
di: Muñoz, Juan Pablo, et al.
Pubblicazione: (2024)
Low-Rank Adapters Meet Neural Architecture Search for LLM Compression
di: Muñoz, J. Pablo, et al.
Pubblicazione: (2025)
di: Muñoz, J. Pablo, et al.
Pubblicazione: (2025)
Auto-Cypher: Improving LLMs on Cypher generation via LLM-supervised generation-verification framework
di: Tiwari, Aman, et al.
Pubblicazione: (2024)
di: Tiwari, Aman, et al.
Pubblicazione: (2024)
Oh! We Freeze: Improving Quantized Knowledge Distillation via Signal Propagation Analysis for Large Language Models
di: Bhardwaj, Kartikeya, et al.
Pubblicazione: (2024)
di: Bhardwaj, Kartikeya, et al.
Pubblicazione: (2024)
A Study on Large Language Models' Limitations in Multiple-Choice Question Answering
di: Khatun, Aisha, et al.
Pubblicazione: (2024)
di: Khatun, Aisha, et al.
Pubblicazione: (2024)
Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences
di: Pattnaik, Pulkit, et al.
Pubblicazione: (2024)
di: Pattnaik, Pulkit, et al.
Pubblicazione: (2024)
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
di: Lee, Yooseop, et al.
Pubblicazione: (2025)
di: Lee, Yooseop, et al.
Pubblicazione: (2025)
Generation and De-Identification of Indian Clinical Discharge Summaries using LLMs
di: Singh, Sanjeet, et al.
Pubblicazione: (2024)
di: Singh, Sanjeet, et al.
Pubblicazione: (2024)
Exploring the Hidden Capacity of LLMs for One-Step Text Generation
di: Mezentsev, Gleb, et al.
Pubblicazione: (2025)
di: Mezentsev, Gleb, et al.
Pubblicazione: (2025)
Expanding Foundational Language Capabilities in Open-Source LLMs through a Korean Case Study
di: Lim, Junghwan, et al.
Pubblicazione: (2025)
di: Lim, Junghwan, et al.
Pubblicazione: (2025)
Small or Large? Zero-Shot or Finetuned? Guiding Language Model Choice for Specialized Applications in Healthcare
di: Gondara, Lovedeep, et al.
Pubblicazione: (2025)
di: Gondara, Lovedeep, et al.
Pubblicazione: (2025)
IITK at SemEval-2024 Task 2: Exploring the Capabilities of LLMs for Safe Biomedical Natural Language Inference for Clinical Trials
di: Mandal, Shreyasi, et al.
Pubblicazione: (2024)
di: Mandal, Shreyasi, et al.
Pubblicazione: (2024)
Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena
di: Luo, Haipeng, et al.
Pubblicazione: (2024)
di: Luo, Haipeng, et al.
Pubblicazione: (2024)
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
di: Sclar, Melanie, et al.
Pubblicazione: (2023)
di: Sclar, Melanie, et al.
Pubblicazione: (2023)
KVCrush: Key value cache size-reduction using similarity in head-behaviour
di: Jha, Gopi Krishna, et al.
Pubblicazione: (2025)
di: Jha, Gopi Krishna, et al.
Pubblicazione: (2025)
When Should LLMs Be Less Specific? Selective Abstraction for Reliable Long-Form Text Generation
di: Goren, Shani, et al.
Pubblicazione: (2026)
di: Goren, Shani, et al.
Pubblicazione: (2026)
COPAL: Continual Pruning in Large Language Generative Models
di: Malla, Srikanth, et al.
Pubblicazione: (2024)
di: Malla, Srikanth, et al.
Pubblicazione: (2024)
DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024)
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024)
Learn To be Efficient: Build Structured Sparsity in Large Language Models
di: Zheng, Haizhong, et al.
Pubblicazione: (2024)
di: Zheng, Haizhong, et al.
Pubblicazione: (2024)
Building Decision Making Models Through Language Model Regime
di: Zhang, Yu, et al.
Pubblicazione: (2024)
di: Zhang, Yu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
RARe: Retrieval Augmented Retrieval with In-Context Examples
di: Tejaswi, Atula, et al.
Pubblicazione: (2024) -
Entropy Aware Reward Guidance for Diffusion Language Model Alignment
di: Tejaswi, Atula, et al.
Pubblicazione: (2026) -
SVFT: Parameter-Efficient Fine-Tuning with Singular Vectors
di: Lingam, Vijay, et al.
Pubblicazione: (2024) -
PropMEND: Hypernetworks for Knowledge Propagation in LLMs
di: Liu, Zeyu Leo, et al.
Pubblicazione: (2025) -
CaLMQA: Exploring culturally specific long-form question answering across 23 languages
di: Arora, Shane, et al.
Pubblicazione: (2024)