FLM-101B: An Open LLM and How to Train It with $100K Budget
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Xiang, Yao, Yiqun, Jiang, Xin, Fang, Xuezhi, Meng, Xuying, Fan, Siqi, Han, Peng, Li, Jing, Du, Li, Qin, Bowen, Zhang, Zheng, Sun, Aixin, Wang, Yequan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training
di: Yao, Yiqun, et al.
Pubblicazione: (2025)
di: Yao, Yiqun, et al.
Pubblicazione: (2025)
Open-domain Implicit Format Control for Large Language Model Generation
di: Yao, Yiqun, et al.
Pubblicazione: (2024)
di: Yao, Yiqun, et al.
Pubblicazione: (2024)
Sketch: A Toolkit for Streamlining LLM Operations
di: Jiang, Xin, et al.
Pubblicazione: (2024)
di: Jiang, Xin, et al.
Pubblicazione: (2024)
Tele-FLM Technical Report
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
EgoMem: Lifelong Memory Agent for Full-duplex Omnimodal Models
di: Yao, Yiqun, et al.
Pubblicazione: (2025)
di: Yao, Yiqun, et al.
Pubblicazione: (2025)
nanoLM: an Affordable LLM Pre-training Benchmark via Accurate Loss Prediction across Scales
di: Yao, Yiqun, et al.
Pubblicazione: (2023)
di: Yao, Yiqun, et al.
Pubblicazione: (2023)
52B to 1T: Lessons Learned via Tele-FLM Series
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
RoboEgo System Card: An Omnimodal Model with Native Full Duplexity
di: Yao, Yiqun, et al.
Pubblicazione: (2025)
di: Yao, Yiqun, et al.
Pubblicazione: (2025)
If an LLM Were a Character, Would It Know Its Own Story? Evaluating Lifelong Learning in LLMs
di: Fan, Siqi, et al.
Pubblicazione: (2025)
di: Fan, Siqi, et al.
Pubblicazione: (2025)
GCRE-GPT: A Generative Model for Comparative Relation Extraction
di: Wang, Yequan, et al.
Pubblicazione: (2023)
di: Wang, Yequan, et al.
Pubblicazione: (2023)
Not All Layers of LLMs Are Necessary During Inference
di: Fan, Siqi, et al.
Pubblicazione: (2024)
di: Fan, Siqi, et al.
Pubblicazione: (2024)
The Price of a Second Thought: On the Evaluation of Reasoning Efficiency in Large Language Models
di: Fan, Siqi, et al.
Pubblicazione: (2025)
di: Fan, Siqi, et al.
Pubblicazione: (2025)
Masked Structural Growth for 2x Faster Language Model Pre-training
di: Yao, Yiqun, et al.
Pubblicazione: (2023)
di: Yao, Yiqun, et al.
Pubblicazione: (2023)
Toward Embodied AGI: A Review of Embodied AI and the Road Ahead
di: Wang, Yequan, et al.
Pubblicazione: (2025)
di: Wang, Yequan, et al.
Pubblicazione: (2025)
Evaluating LLM Adaptation to Sociodemographic Factors: User Profile vs. Dialogue History
di: Zhong, Qishuai, et al.
Pubblicazione: (2025)
di: Zhong, Qishuai, et al.
Pubblicazione: (2025)
Position-Aware Depth Decay Decoding ($D^3$): Boosting Large Language Model Inference Efficiency
di: Fan, Siqi, et al.
Pubblicazione: (2025)
di: Fan, Siqi, et al.
Pubblicazione: (2025)
NetGPT: Generative Pretrained Transformer for Network Traffic
di: Meng, Xuying, et al.
Pubblicazione: (2023)
di: Meng, Xuying, et al.
Pubblicazione: (2023)
OmniDFA: A Unified Framework for Open Set Synthesis Image Detection and Few-Shot Attribution
di: Wu, Shiyu, et al.
Pubblicazione: (2025)
di: Wu, Shiyu, et al.
Pubblicazione: (2025)
SARDet-100K: Towards Open-Source Benchmark and ToolKit for Large-Scale SAR Object Detection
di: Li, Yuxuan, et al.
Pubblicazione: (2024)
di: Li, Yuxuan, et al.
Pubblicazione: (2024)
Few-Shot Learner Generalizes Across AI-Generated Image Detection
di: Wu, Shiyu, et al.
Pubblicazione: (2025)
di: Wu, Shiyu, et al.
Pubblicazione: (2025)
OpenLLM-RTL: Open Dataset and Benchmark for LLM-Aided Design RTL Generation
di: Liu, Shang, et al.
Pubblicazione: (2025)
di: Liu, Shang, et al.
Pubblicazione: (2025)
Spectral-Based Graph Neural Networks for Complementary Item Recommendation
di: Luo, Haitong, et al.
Pubblicazione: (2024)
di: Luo, Haitong, et al.
Pubblicazione: (2024)
Ein simulationsbasierter Ansatz zur Auslegung additiv gefertigter FLM-Faserverbundstrukturen
di: Völkl, Harald
Pubblicazione: (2025)
di: Völkl, Harald
Pubblicazione: (2025)
LocateEdit-Bench: A Benchmark for Instruction-Based Editing Localization
di: Wu, Shiyu, et al.
Pubblicazione: (2026)
di: Wu, Shiyu, et al.
Pubblicazione: (2026)
Can LLM Safety Be Ensured by Constraining Parameter Regions?
di: Li, Zongmin, et al.
Pubblicazione: (2026)
di: Li, Zongmin, et al.
Pubblicazione: (2026)
Skill-as-Pseudocode: Refactoring Skill Libraries to Pseudocode for LLM Agents
di: Li, Xinze, et al.
Pubblicazione: (2026)
di: Li, Xinze, et al.
Pubblicazione: (2026)
Poor Man's Training on MCUs: A Memory-Efficient Quantized Back-Propagation-Free Approach
di: Zhao, Yequan, et al.
Pubblicazione: (2024)
di: Zhao, Yequan, et al.
Pubblicazione: (2024)
Efficient Budget Allocation for Large-Scale LLM-Enabled Virtual Screening
di: Li, Zaile, et al.
Pubblicazione: (2024)
di: Li, Zaile, et al.
Pubblicazione: (2024)
Multimodal Reasoning with Multimodal Knowledge Graph
di: Lee, Junlin, et al.
Pubblicazione: (2024)
di: Lee, Junlin, et al.
Pubblicazione: (2024)
BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens
di: Wen, Hao, et al.
Pubblicazione: (2025)
di: Wen, Hao, et al.
Pubblicazione: (2025)
HDINO: A Concise and Efficient Open-Vocabulary Detector
di: Zhang, Hao, et al.
Pubblicazione: (2026)
di: Zhang, Hao, et al.
Pubblicazione: (2026)
Tensor-Compressed and Fully-Quantized Training of Neural PDE Solvers
di: Lu, Jinming, et al.
Pubblicazione: (2025)
di: Lu, Jinming, et al.
Pubblicazione: (2025)
LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
di: Zhao, Yi, et al.
Pubblicazione: (2025)
di: Zhao, Yi, et al.
Pubblicazione: (2025)
LatentGuard: Controllable Latent Steering for Robust Refusal of Attacks and Reliable Response Generation
di: Shu, Huizhen, et al.
Pubblicazione: (2025)
di: Shu, Huizhen, et al.
Pubblicazione: (2025)
Targeting the Core: A Simple and Effective Method to Attack RAG-based Agents via Direct LLM Manipulation
di: Li, Xuying, et al.
Pubblicazione: (2024)
di: Li, Xuying, et al.
Pubblicazione: (2024)
A Minimalist Prompt for Zero-Shot Policy Learning
di: Song, Meng, et al.
Pubblicazione: (2024)
di: Song, Meng, et al.
Pubblicazione: (2024)
How to Rent GPUs on a Budget
di: Li, Zhouzi, et al.
Pubblicazione: (2024)
di: Li, Zhouzi, et al.
Pubblicazione: (2024)
How Far Can Unsupervised RLVR Scale LLM Training?
di: He, Bingxiang, et al.
Pubblicazione: (2026)
di: He, Bingxiang, et al.
Pubblicazione: (2026)
The Hidden Cost of Readability: How Code Formatting Silently Consumes Your LLM Budget
di: Pan, Dangfeng, et al.
Pubblicazione: (2025)
di: Pan, Dangfeng, et al.
Pubblicazione: (2025)
Active Hypothesis Testing under Computational Budgets with Applications to GWAS and LLM
di: Kuang, Qi, et al.
Pubblicazione: (2025)
di: Kuang, Qi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training
di: Yao, Yiqun, et al.
Pubblicazione: (2025) -
Open-domain Implicit Format Control for Large Language Model Generation
di: Yao, Yiqun, et al.
Pubblicazione: (2024) -
Sketch: A Toolkit for Streamlining LLM Operations
di: Jiang, Xin, et al.
Pubblicazione: (2024) -
Tele-FLM Technical Report
di: Li, Xiang, et al.
Pubblicazione: (2024) -
EgoMem: Lifelong Memory Agent for Full-duplex Omnimodal Models
di: Yao, Yiqun, et al.
Pubblicazione: (2025)