SOLAR 10.7B: Scaling Large Language Models with Simple yet Effective Depth Up-Scaling
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Dahyun, Park, Chanjun, Kim, Sanghoon, Lee, Wonsung, Song, Wonho, Kim, Yunsu, Kim, Hyeonwoo, Kim, Yungi, Lee, Hyeonju, Kim, Jihoo, Ahn, Changbae, Yang, Seonghoon, Lee, Sukyung, Park, Hyunbyung, Gim, Gyoungjin, Cha, Mikyoung, Lee, Hwalsuk, Kim, Sunghun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Dataverse: Open-Source ETL (Extract, Transform, Load) Pipeline for Large Language Models
di: Park, Hyunbyung, et al.
Pubblicazione: (2024)
di: Park, Hyunbyung, et al.
Pubblicazione: (2024)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
di: Park, Chanjun, et al.
Pubblicazione: (2024)
di: Park, Chanjun, et al.
Pubblicazione: (2024)
SAAS: Solving Ability Amplification Strategy for Enhanced Mathematical Reasoning in Large Language Models
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
sDPO: Don't Use Your Data All at Once
di: Kim, Dahyun, et al.
Pubblicazione: (2024)
di: Kim, Dahyun, et al.
Pubblicazione: (2024)
Evalverse: Unified and Accessible Library for Large Language Model Evaluation
di: Kim, Jihoo, et al.
Pubblicazione: (2024)
di: Kim, Jihoo, et al.
Pubblicazione: (2024)
1 Trillion Token (1TT) Platform: A Novel Framework for Efficient Data Sharing and Compensation in Large Language Models
di: Park, Chanjun, et al.
Pubblicazione: (2024)
di: Park, Chanjun, et al.
Pubblicazione: (2024)
LP Data Pipeline: Lightweight, Purpose-driven Data Pipeline for Large Language Models
di: Kim, Yungi, et al.
Pubblicazione: (2024)
di: Kim, Yungi, et al.
Pubblicazione: (2024)
Rethinking KenLM: Good and Bad Model Ensembles for Efficient Text Quality Filtering in Large Web Corpora
di: Kim, Yungi, et al.
Pubblicazione: (2024)
di: Kim, Yungi, et al.
Pubblicazione: (2024)
Representing the Under-Represented: Cultural and Core Capability Benchmarks for Developing Thai Large Language Models
di: Kim, Dahyun, et al.
Pubblicazione: (2024)
di: Kim, Dahyun, et al.
Pubblicazione: (2024)
InstaTrans: An Instruction-Aware Translation Framework for Non-English Instruction Datasets
di: Kim, Yungi, et al.
Pubblicazione: (2024)
di: Kim, Yungi, et al.
Pubblicazione: (2024)
Understanding LLM Development Through Longitudinal Study: Insights from the Open Ko-LLM Leaderboard
di: Park, Chanjun, et al.
Pubblicazione: (2024)
di: Park, Chanjun, et al.
Pubblicazione: (2024)
Solar Open Technical Report
di: Park, Sungrae, et al.
Pubblicazione: (2026)
di: Park, Sungrae, et al.
Pubblicazione: (2026)
Model-Based Data-Centric AI: Bridging the Divide Between Academic Ideals and Industrial Pragmatism
di: Park, Chanjun, et al.
Pubblicazione: (2024)
di: Park, Chanjun, et al.
Pubblicazione: (2024)
Evaluating Route Choice Models: A Comparative Analysis
di: Hyeonwoo Lee, et al.
Pubblicazione: (2025)
di: Hyeonwoo Lee, et al.
Pubblicazione: (2025)
Can Code-Switched Texts Activate a Knowledge Switch in LLMs? A Case Study on English-Korean Code-Switching
di: Kim, Seoyeon, et al.
Pubblicazione: (2024)
di: Kim, Seoyeon, et al.
Pubblicazione: (2024)
When Is Enough Not Enough? Illusory Completion in Search Agents
di: Ko, Dayoon, et al.
Pubblicazione: (2026)
di: Ko, Dayoon, et al.
Pubblicazione: (2026)
Effective Test-Time Scaling of Discrete Diffusion through Iterative Refinement
di: Lee, Sanghyun, et al.
Pubblicazione: (2025)
di: Lee, Sanghyun, et al.
Pubblicazione: (2025)
Type-I Blowup Solutions for Yang-Mills Flow
di: Kim, Jaehwan, et al.
Pubblicazione: (2024)
di: Kim, Jaehwan, et al.
Pubblicazione: (2024)
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
di: Kim, Haechan, et al.
Pubblicazione: (2026)
di: Kim, Haechan, et al.
Pubblicazione: (2026)
AVIN-Chat: An Audio-Visual Interactive Chatbot System with Emotional State Tuning
di: Park, Chanhyuk, et al.
Pubblicazione: (2024)
di: Park, Chanhyuk, et al.
Pubblicazione: (2024)
Hybrid Deep Searcher: Scalable Parallel and Sequential Search Reasoning
di: Ko, Dayoon, et al.
Pubblicazione: (2025)
di: Ko, Dayoon, et al.
Pubblicazione: (2025)
Laboratory Characterization of Cryogenic Frost‐Point Hygrometers Using an Upper Air Simulator
di: Sang‐Wook Lee, et al.
Pubblicazione: (2026)
di: Sang‐Wook Lee, et al.
Pubblicazione: (2026)
Mobility-edge-embedded Hofstadter butterfly from a tilt-induced quasiperiodic potential
di: Lee, Sanghoon, et al.
Pubblicazione: (2026)
di: Lee, Sanghoon, et al.
Pubblicazione: (2026)
Revealing the Impact of Aggregations in the Graph‐Based Molecular Machine Learning: Electrostatic Interaction Versus Pooling Methods
di: Sanghoon Lee, et al.
Pubblicazione: (2025)
di: Sanghoon Lee, et al.
Pubblicazione: (2025)
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
di: Park, Gunho, et al.
Pubblicazione: (2022)
di: Park, Gunho, et al.
Pubblicazione: (2022)
SNeRV: Spectra-preserving Neural Representation for Video
di: Kim, Jina, et al.
Pubblicazione: (2025)
di: Kim, Jina, et al.
Pubblicazione: (2025)
Optimizing Korean-Centric LLMs via Token Pruning
di: Kim, Hoyeol, et al.
Pubblicazione: (2026)
di: Kim, Hoyeol, et al.
Pubblicazione: (2026)
Separating Novel Features for Logical Anomaly Detection: A Straightforward yet Effective Approach
di: Lee, Kangil, et al.
Pubblicazione: (2024)
di: Lee, Kangil, et al.
Pubblicazione: (2024)
LASTIST: LArge-Scale Target-Independent STance dataset
di: Kim, DongJae, et al.
Pubblicazione: (2025)
di: Kim, DongJae, et al.
Pubblicazione: (2025)
Towards Next-Level Post-Training Quantization of Hyper-Scale Transformers
di: Kim, Junhan, et al.
Pubblicazione: (2024)
di: Kim, Junhan, et al.
Pubblicazione: (2024)
Scale‐Up Strategies for Redox‐Mediated Electrodialysis for Desalination: The Role of Electrode and Channel Stacks
di: Gamin Kim, et al.
Pubblicazione: (2025)
di: Gamin Kim, et al.
Pubblicazione: (2025)
Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs
di: Lee, Jongseo, et al.
Pubblicazione: (2026)
di: Lee, Jongseo, et al.
Pubblicazione: (2026)
Cross-lingual Transfer for Automatic Question Generation by Learning Interrogative Structures in Target Languages
di: Hwang, Seonjeong, et al.
Pubblicazione: (2024)
di: Hwang, Seonjeong, et al.
Pubblicazione: (2024)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
di: Jeon, Yejin, et al.
Pubblicazione: (2024)
di: Jeon, Yejin, et al.
Pubblicazione: (2024)
Autoregressive Score Generation for Multi-trait Essay Scoring
di: Do, Heejin, et al.
Pubblicazione: (2024)
di: Do, Heejin, et al.
Pubblicazione: (2024)
Leveraging the Interplay Between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause Formation
di: Jeon, Yejin, et al.
Pubblicazione: (2024)
di: Jeon, Yejin, et al.
Pubblicazione: (2024)
Explainable Multi-hop Question Generation: An End-to-End Approach without Intermediate Question Labeling
di: Hwang, Seonjeong, et al.
Pubblicazione: (2024)
di: Hwang, Seonjeong, et al.
Pubblicazione: (2024)
BemaGANv2: Discriminator Combination Strategies for GAN-based Vocoders in Long-Term Audio Generation
di: Park, Taesoo, et al.
Pubblicazione: (2025)
di: Park, Taesoo, et al.
Pubblicazione: (2025)
Large‐Area Floating Display with Wafer‐Scale Manufactured Metalens Arrays
di: Joohoon Kim, et al.
Pubblicazione: (2024)
di: Joohoon Kim, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Dataverse: Open-Source ETL (Extract, Transform, Load) Pipeline for Large Language Models
di: Park, Hyunbyung, et al.
Pubblicazione: (2024) -
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024) -
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
di: Park, Chanjun, et al.
Pubblicazione: (2024) -
SAAS: Solving Ability Amplification Strategy for Enhanced Mathematical Reasoning in Large Language Models
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024) -
sDPO: Don't Use Your Data All at Once
di: Kim, Dahyun, et al.
Pubblicazione: (2024)