Many Minds from One Model: Bayesian-Inspired Transformers for Population Diversity
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yang, Diji, Zhang, Yi |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Beyond Introspection: Reinforcing Thinking via Externalist Behavioral Feedback
par: Yang, Diji, et autres
Publié: (2024)
par: Yang, Diji, et autres
Publié: (2024)
Flextron: Many-in-One Flexible Large Language Model
par: Cai, Ruisi, et autres
Publié: (2024)
par: Cai, Ruisi, et autres
Publié: (2024)
Extended Mind Transformers
par: Klett, Phoebe, et autres
Publié: (2024)
par: Klett, Phoebe, et autres
Publié: (2024)
Out of One, Many: Using Language Models to Simulate Human Samples
par: Argyle, Lisa P., et autres
Publié: (2022)
par: Argyle, Lisa P., et autres
Publié: (2022)
When 2D Tasks Meet 1D Serialization: On Serialization Friction in Structured Tasks
par: Lo, Chung-Hsiang, et autres
Publié: (2026)
par: Lo, Chung-Hsiang, et autres
Publié: (2026)
Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
par: Kamigaito, Hidetaka, et autres
Publié: (2025)
par: Kamigaito, Hidetaka, et autres
Publié: (2025)
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
par: Zhang, Jiazheng, et autres
Publié: (2025)
par: Zhang, Jiazheng, et autres
Publié: (2025)
Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information
par: Tutnov, Rasul, et autres
Publié: (2025)
par: Tutnov, Rasul, et autres
Publié: (2025)
Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries
par: Yan, Tianyi Lorena, et autres
Publié: (2025)
par: Yan, Tianyi Lorena, et autres
Publié: (2025)
Projectable Models: One-Shot Generation of Small Specialized Transformers from Large Ones
par: Zhmoginov, Andrey, et autres
Publié: (2025)
par: Zhmoginov, Andrey, et autres
Publié: (2025)
More Than One Teacher: Adaptive Multi-Guidance Policy Optimization for Diverse Exploration
par: Yuan, Xiaoyang, et autres
Publié: (2025)
par: Yuan, Xiaoyang, et autres
Publié: (2025)
One Goal, Many Challenges: Robust Preference Optimization Amid Content-Aware and Multi-Source Noise
par: Afzali, Amirabbas, et autres
Publié: (2025)
par: Afzali, Amirabbas, et autres
Publié: (2025)
Momentum Streams for Optimizer-Inspired Transformers
par: Gai, Jingchu, et autres
Publié: (2026)
par: Gai, Jingchu, et autres
Publié: (2026)
One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMs
par: Li, Yinghui, et autres
Publié: (2025)
par: Li, Yinghui, et autres
Publié: (2025)
Train Once, Answer All: Many Pretraining Experiments for the Cost of One
par: Bordt, Sebastian, et autres
Publié: (2025)
par: Bordt, Sebastian, et autres
Publié: (2025)
Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
par: Song, Yuda, et autres
Publié: (2024)
par: Song, Yuda, et autres
Publié: (2024)
DLM-One: Diffusion Language Models for One-Step Sequence Generation
par: Chen, Tianqi, et autres
Publié: (2025)
par: Chen, Tianqi, et autres
Publié: (2025)
Provable Knowledge Acquisition and Extraction in One-Layer Transformers
par: Xu, Ruichen, et autres
Publié: (2025)
par: Xu, Ruichen, et autres
Publié: (2025)
In-Context Learning of a Linear Transformer Block: Benefits of the MLP Component and One-Step GD Initialization
par: Zhang, Ruiqi, et autres
Publié: (2024)
par: Zhang, Ruiqi, et autres
Publié: (2024)
MULTIVERSE: Exposing Large Language Model Alignment Problems in Diverse Worlds
par: Jin, Xiaolong, et autres
Publié: (2024)
par: Jin, Xiaolong, et autres
Publié: (2024)
Bayesian Mixture of Experts For Large Language Models
par: Dialameh, Maryam, et autres
Publié: (2025)
par: Dialameh, Maryam, et autres
Publié: (2025)
The Few Govern the Many:Unveiling Few-Layer Dominance for Time Series Models
par: Qiu, Xin, et autres
Publié: (2025)
par: Qiu, Xin, et autres
Publié: (2025)
Comparing Human and AI Rater Effects Using the Many-Facet Rasch Model
par: Jiao, Hong, et autres
Publié: (2025)
par: Jiao, Hong, et autres
Publié: (2025)
CoreInfer: Accelerating Large Language Model Inference with Semantics-Inspired Adaptive Sparse Activation
par: Wang, Qinsi, et autres
Publié: (2024)
par: Wang, Qinsi, et autres
Publié: (2024)
Many-to-English Machine Translation Tools, Data, and Pretrained Models
par: Gowda, Thamme, et autres
Publié: (2021)
par: Gowda, Thamme, et autres
Publié: (2021)
Can Large Language Models Transform Computational Social Science?
par: Ziems, Caleb, et autres
Publié: (2023)
par: Ziems, Caleb, et autres
Publié: (2023)
Jointly Reinforcing Diversity and Quality in Language Model Generations
par: Li, Tianjian, et autres
Publié: (2025)
par: Li, Tianjian, et autres
Publié: (2025)
Spectral Generative Flow Models: A Physics-Inspired Replacement for Vectorized Large Language Models
par: Kiruluta, Andrew
Publié: (2026)
par: Kiruluta, Andrew
Publié: (2026)
An Innovative CGL-MHA Model for Sarcasm Sentiment Recognition Using the MindSpore Framework
par: Qin, Zhenkai, et autres
Publié: (2024)
par: Qin, Zhenkai, et autres
Publié: (2024)
GOFA: A Generative One-For-All Model for Joint Graph Language Modeling
par: Kong, Lecheng, et autres
Publié: (2024)
par: Kong, Lecheng, et autres
Publié: (2024)
Learning Multiplex Representations on Text-Attributed Graphs with One Language Model Encoder
par: Jin, Bowen, et autres
Publié: (2023)
par: Jin, Bowen, et autres
Publié: (2023)
Classifying and Addressing the Diversity of Errors in Retrieval-Augmented Generation Systems
par: Leung, Kin Kwan, et autres
Publié: (2025)
par: Leung, Kin Kwan, et autres
Publié: (2025)
Adaptive Feature-based Low-Rank Compression of Large Language Models via Bayesian Optimization
par: Ji, Yixin, et autres
Publié: (2024)
par: Ji, Yixin, et autres
Publié: (2024)
Test-Time Adaptation via Many-Shot Prompting: Benefits, Limits, and Pitfalls
par: Upasani, Shubhangi, et autres
Publié: (2026)
par: Upasani, Shubhangi, et autres
Publié: (2026)
Mixture-of-Personas Language Models for Population Simulation
par: Bui, Ngoc, et autres
Publié: (2025)
par: Bui, Ngoc, et autres
Publié: (2025)
MillStone: How Open-Minded Are LLMs?
par: Triedman, Harold, et autres
Publié: (2025)
par: Triedman, Harold, et autres
Publié: (2025)
MetaGreen: Meta-Learning Inspired Transformer Selection for Green Semantic Communication
par: Mukherjee, Shubhabrata, et autres
Publié: (2024)
par: Mukherjee, Shubhabrata, et autres
Publié: (2024)
Orchestrating Dual-Boundaries: An Arithmetic Intensity Inspired Acceleration Framework for Diffusion Language Models
par: Wei, Linye, et autres
Publié: (2025)
par: Wei, Linye, et autres
Publié: (2025)
Many-Shot In-Context Learning
par: Agarwal, Rishabh, et autres
Publié: (2024)
par: Agarwal, Rishabh, et autres
Publié: (2024)
Sentence-level Reward Model can Generalize Better for Aligning LLM from Human Preference
par: Qiu, Wenjie, et autres
Publié: (2025)
par: Qiu, Wenjie, et autres
Publié: (2025)
Documents similaires
-
Beyond Introspection: Reinforcing Thinking via Externalist Behavioral Feedback
par: Yang, Diji, et autres
Publié: (2024) -
Flextron: Many-in-One Flexible Large Language Model
par: Cai, Ruisi, et autres
Publié: (2024) -
Extended Mind Transformers
par: Klett, Phoebe, et autres
Publié: (2024) -
Out of One, Many: Using Language Models to Simulate Human Samples
par: Argyle, Lisa P., et autres
Publié: (2022) -
When 2D Tasks Meet 1D Serialization: On Serialization Friction in Structured Tasks
par: Lo, Chung-Hsiang, et autres
Publié: (2026)