Mini Minds: Exploring Bebeshka and Zlata Baby Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Proskurina, Irina, Metzler, Guillaume, Velcin, Julien |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Quantization Affects Confidence of Large Language Models?
von: Proskurina, Irina, et al.
Veröffentlicht: (2024)
von: Proskurina, Irina, et al.
Veröffentlicht: (2024)
Fair-GPTQ: Bias-Aware Quantization for Large Language Models
von: Proskurina, Irina, et al.
Veröffentlicht: (2025)
von: Proskurina, Irina, et al.
Veröffentlicht: (2025)
HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
von: Proskurina, Irina, et al.
Veröffentlicht: (2025)
von: Proskurina, Irina, et al.
Veröffentlicht: (2025)
Histoires Morales: A French Dataset for Assessing Moral Alignment
von: Leteno, Thibaud, et al.
Veröffentlicht: (2025)
von: Leteno, Thibaud, et al.
Veröffentlicht: (2025)
Baby Scale: Investigating Models Trained on Individual Children's Language Input
von: Feng, Steven Y., et al.
Veröffentlicht: (2026)
von: Feng, Steven Y., et al.
Veröffentlicht: (2026)
Capturing Style in Author and Document Representation
von: Terreau, Enzo, et al.
Veröffentlicht: (2024)
von: Terreau, Enzo, et al.
Veröffentlicht: (2024)
Bringing Up a Bilingual BabyLM: Investigating Multilingual Language Acquisition Using Small-Scale Models
von: Zeng, Linda, et al.
Veröffentlicht: (2026)
von: Zeng, Linda, et al.
Veröffentlicht: (2026)
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
von: MiniMax, et al.
Veröffentlicht: (2026)
von: MiniMax, et al.
Veröffentlicht: (2026)
Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning
von: Sclar, Melanie, et al.
Veröffentlicht: (2024)
von: Sclar, Melanie, et al.
Veröffentlicht: (2024)
Mini-Giants: "Small" Language Models and Open Source Win-Win
von: Zhou, Zhengping, et al.
Veröffentlicht: (2023)
von: Zhou, Zhengping, et al.
Veröffentlicht: (2023)
Mini-batch Coresets for Memory-efficient Language Model Training on Data Mixtures
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
von: Liu, Akide, et al.
Veröffentlicht: (2024)
von: Liu, Akide, et al.
Veröffentlicht: (2024)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
MOM: Memory-Efficient Offloaded Mini-Sequence Inference for Long Context Language Models
von: Zhang, Junyang, et al.
Veröffentlicht: (2025)
von: Zhang, Junyang, et al.
Veröffentlicht: (2025)
LoRA-Mini : Adaptation Matrices Decomposition and Selective Training
von: Singh, Ayush, et al.
Veröffentlicht: (2024)
von: Singh, Ayush, et al.
Veröffentlicht: (2024)
A Notion of Complexity for Theory of Mind via Discrete World Models
von: Huang, X. Angelo, et al.
Veröffentlicht: (2024)
von: Huang, X. Angelo, et al.
Veröffentlicht: (2024)
When Should Models Change Their Minds? Contextual Belief Management in Large Language Models
von: Xu, Haoming, et al.
Veröffentlicht: (2026)
von: Xu, Haoming, et al.
Veröffentlicht: (2026)
Exposing propaganda: an analysis of stylistic cues comparing human annotations and machine classification
von: Faye, Géraud, et al.
Veröffentlicht: (2024)
von: Faye, Géraud, et al.
Veröffentlicht: (2024)
PrivacyMind: Large Language Models Can Be Contextual Privacy Protection Learners
von: Xiao, Yijia, et al.
Veröffentlicht: (2023)
von: Xiao, Yijia, et al.
Veröffentlicht: (2023)
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2025)
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
von: Microsoft, et al.
Veröffentlicht: (2025)
von: Microsoft, et al.
Veröffentlicht: (2025)
MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models
von: Wen, Yilin, et al.
Veröffentlicht: (2023)
von: Wen, Yilin, et al.
Veröffentlicht: (2023)
Selective Deficits in LLM Mental Self-Modeling in a Behavior-Based Test of Theory of Mind
von: Ackerman, Christopher
Veröffentlicht: (2026)
von: Ackerman, Christopher
Veröffentlicht: (2026)
Improving VTE Identification through Language Models from Radiology Reports: A Comparative Study of Mamba, Phi-3 Mini, and BERT
von: Deng, Jamie, et al.
Veröffentlicht: (2024)
von: Deng, Jamie, et al.
Veröffentlicht: (2024)
MediaMind: Revolutionizing Media Monitoring using Agentification
von: Gunduz, Ahmet, et al.
Veröffentlicht: (2025)
von: Gunduz, Ahmet, et al.
Veröffentlicht: (2025)
Automated Meta Prompt Engineering for Alignment with the Theory of Mind
von: Baughman, Aaron, et al.
Veröffentlicht: (2025)
von: Baughman, Aaron, et al.
Veröffentlicht: (2025)
GPT-4o Lacks Core Features of Theory of Mind
von: Muchovej, John, et al.
Veröffentlicht: (2026)
von: Muchovej, John, et al.
Veröffentlicht: (2026)
Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2024)
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2024)
SEMQA: Semi-Extractive Multi-Source Question Answering
von: Schuster, Tal, et al.
Veröffentlicht: (2023)
von: Schuster, Tal, et al.
Veröffentlicht: (2023)
Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models
von: Roger, Alexis, et al.
Veröffentlicht: (2025)
von: Roger, Alexis, et al.
Veröffentlicht: (2025)
Mind the Gap: A Review of Arabic Post-Training Datasets and Their Limitations
von: Alkhowaiter, Mohammed, et al.
Veröffentlicht: (2025)
von: Alkhowaiter, Mohammed, et al.
Veröffentlicht: (2025)
Exploring Memorization in Fine-tuned Language Models
von: Zeng, Shenglai, et al.
Veröffentlicht: (2023)
von: Zeng, Shenglai, et al.
Veröffentlicht: (2023)
Exploring Scaling Laws for EHR Foundation Models
von: Zhang, Sheng, et al.
Veröffentlicht: (2025)
von: Zhang, Sheng, et al.
Veröffentlicht: (2025)
Large Language Models Explore by Latent Distilling
von: Zeng, Yuanhao, et al.
Veröffentlicht: (2026)
von: Zeng, Yuanhao, et al.
Veröffentlicht: (2026)
Speculative Streaming: Fast LLM Inference without Auxiliary Models
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2024)
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2024)
Exploring and Benchmarking the Planning Capabilities of Large Language Models
von: Bohnet, Bernd, et al.
Veröffentlicht: (2024)
von: Bohnet, Bernd, et al.
Veröffentlicht: (2024)
Beyond Arrow's Impossibility: Fairness as an Emergent Property of Multi-Agent Collaboration
von: Chaki, Sayan Kumar, et al.
Veröffentlicht: (2026)
von: Chaki, Sayan Kumar, et al.
Veröffentlicht: (2026)
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning
von: Sarangi, Sneheel, et al.
Veröffentlicht: (2025)
von: Sarangi, Sneheel, et al.
Veröffentlicht: (2025)
Labrador: Exploring the Limits of Masked Language Modeling for Laboratory Data
von: Bellamy, David R., et al.
Veröffentlicht: (2023)
von: Bellamy, David R., et al.
Veröffentlicht: (2023)
To Think or Not to Think: Exploring the Unthinking Vulnerability in Large Reasoning Models
von: Zhu, Zihao, et al.
Veröffentlicht: (2025)
von: Zhu, Zihao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
When Quantization Affects Confidence of Large Language Models?
von: Proskurina, Irina, et al.
Veröffentlicht: (2024) -
Fair-GPTQ: Bias-Aware Quantization for Large Language Models
von: Proskurina, Irina, et al.
Veröffentlicht: (2025) -
HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
von: Proskurina, Irina, et al.
Veröffentlicht: (2025) -
Histoires Morales: A French Dataset for Assessing Moral Alignment
von: Leteno, Thibaud, et al.
Veröffentlicht: (2025) -
Baby Scale: Investigating Models Trained on Individual Children's Language Input
von: Feng, Steven Y., et al.
Veröffentlicht: (2026)