Small Models, Smarter Learning: The Power of Joint Task Training
Fuente:
arXiv
Saved in:
| Main Authors: | Both, Csaba, Hoover, Benjamin, Strobelt, Hendrik, Krotov, Dmitry, Weidele, Daniel Karl I., Martino, Mauro, Dehmamy, Nima |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Memory in Plain Sight: Surveying the Uncanny Resemblances of Associative Memories and Diffusion Models
by: Hoover, Benjamin, et al.
Published: (2023)
by: Hoover, Benjamin, et al.
Published: (2023)
Dense Associative Memory Through the Lens of Random Features
by: Hoover, Benjamin, et al.
Published: (2024)
by: Hoover, Benjamin, et al.
Published: (2024)
NRGPT: An Energy-based Alternative for GPT
by: Dehmamy, Nima, et al.
Published: (2025)
by: Dehmamy, Nima, et al.
Published: (2025)
Tensor Cache: Eviction-conditioned Associative Memory for Transformers
by: Swain, Kabir, et al.
Published: (2026)
by: Swain, Kabir, et al.
Published: (2026)
Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers
by: Swain, Kabir, et al.
Published: (2026)
by: Swain, Kabir, et al.
Published: (2026)
Hurwitz Quaternion Multiplicative Quantization for KV Cache Compression
by: Swain, Kabir, et al.
Published: (2026)
by: Swain, Kabir, et al.
Published: (2026)
Modern Methods in Associative Memory
by: Krotov, Dmitry, et al.
Published: (2025)
by: Krotov, Dmitry, et al.
Published: (2025)
Abstraction Alignment: Comparing Model-Learned and Human-Encoded Conceptual Relationships
by: Boggust, Angie, et al.
Published: (2024)
by: Boggust, Angie, et al.
Published: (2024)
Interactive Visual Learning for Stable Diffusion
by: Lee, Seongmin, et al.
Published: (2024)
by: Lee, Seongmin, et al.
Published: (2024)
DEX-AR: A Dynamic Explainability Method for Autoregressive Vision-Language Models
by: Bousselham, Walid, et al.
Published: (2026)
by: Bousselham, Walid, et al.
Published: (2026)
GPT-2 Through the Lens of Vector Symbolic Architectures
by: Knittel, Johannes, et al.
Published: (2024)
by: Knittel, Johannes, et al.
Published: (2024)
Operator Learning for Reconstructing Flow Fields from Sparse Measurements: an Energy Transformer Approach
by: Zhang, Qian, et al.
Published: (2025)
by: Zhang, Qian, et al.
Published: (2025)
Dense Associative Memory with Epanechnikov Energy
by: Hoover, Benjamin, et al.
Published: (2025)
by: Hoover, Benjamin, et al.
Published: (2025)
Enhancing Explainability with Multimodal Context Representations for Smarter Robots
by: Viswanath, Anargh, et al.
Published: (2025)
by: Viswanath, Anargh, et al.
Published: (2025)
Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion
by: Lee, Seongmin, et al.
Published: (2023)
by: Lee, Seongmin, et al.
Published: (2023)
Smarter Together: Combining Large Language Models and Small Models for Physiological Signals Visual Inspection
by: Li, Huayu, et al.
Published: (2025)
by: Li, Huayu, et al.
Published: (2025)
Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data
by: Pham, Bao, et al.
Published: (2026)
by: Pham, Bao, et al.
Published: (2026)
Why context matters in VQA and Reasoning: Semantic interventions for VLM input modalities
by: Amara, Kenza, et al.
Published: (2024)
by: Amara, Kenza, et al.
Published: (2024)
HotelMatch-LLM: Joint Multi-Task Training of Small and Large Language Models for Efficient Multimodal Hotel Retrieval
by: Askari, Arian, et al.
Published: (2025)
by: Askari, Arian, et al.
Published: (2025)
Towards Harnessing the Collaborative Power of Large and Small Models for Domain Tasks
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Understanding Large Language Model Behaviors through Interactive Counterfactual Generation and Analysis
by: Cheng, Furui, et al.
Published: (2024)
by: Cheng, Furui, et al.
Published: (2024)
Understanding Mode Connectivity via Parameter Space Symmetry
by: Zhao, Bo, et al.
Published: (2025)
by: Zhao, Bo, et al.
Published: (2025)
Latent Space Symmetry Discovery
by: Yang, Jianke, et al.
Published: (2023)
by: Yang, Jianke, et al.
Published: (2023)
Are Large Language Models Truly Smarter Than Humans?
by: M, Eshwar Reddy, et al.
Published: (2026)
by: M, Eshwar Reddy, et al.
Published: (2026)
Boosting the Power of Small Multimodal Reasoning Models to Match Larger Models with Self-Consistency Training
by: Tan, Cheng, et al.
Published: (2023)
by: Tan, Cheng, et al.
Published: (2023)
Machines Learn to Infer Stellar Parameters Just by Looking at a Large Number of Spectra
by: Sedaghat, Nima, et al.
Published: (2020)
by: Sedaghat, Nima, et al.
Published: (2020)
Diffusion Models for Smarter UAVs: Decision-Making and Modeling
by: Emami, Yousef, et al.
Published: (2025)
by: Emami, Yousef, et al.
Published: (2025)
SPARQL Query Generation with LLMs: Measuring the Impact of Training Data Memorization and Knowledge Injection
by: Gashkov, Aleksandr, et al.
Published: (2025)
by: Gashkov, Aleksandr, et al.
Published: (2025)
Benchmarks Saturate When The Model Gets Smarter Than The Judge
by: Ballon, Marthe, et al.
Published: (2026)
by: Ballon, Marthe, et al.
Published: (2026)
Merging Smarter, Generalizing Better: Enhancing Model Merging on OOD Data
by: Zhang, Bingjie, et al.
Published: (2025)
by: Zhang, Bingjie, et al.
Published: (2025)
Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks
by: Wang, Yang, et al.
Published: (2025)
by: Wang, Yang, et al.
Published: (2025)
Post-hoc LLM-Supported Debugging of Distributed Processes
by: Schiese, Dennis, et al.
Published: (2025)
by: Schiese, Dennis, et al.
Published: (2025)
SUGAR: Leveraging Contextual Confidence for Smarter Retrieval
by: Zubkova, Hanna, et al.
Published: (2025)
by: Zubkova, Hanna, et al.
Published: (2025)
Smarter Together: Creating Agentic Communities of Practice through Shared Experiential Learning
by: Tablan, Valentin, et al.
Published: (2025)
by: Tablan, Valentin, et al.
Published: (2025)
RUMI: Rummaging Using Mutual Information
by: Zhong, Sheng, et al.
Published: (2024)
by: Zhong, Sheng, et al.
Published: (2024)
Knowledge-Guided Manipulation Using Multi-Task Reinforcement Learning
by: Narendra, Aditya, et al.
Published: (2026)
by: Narendra, Aditya, et al.
Published: (2026)
Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
by: Zhang, Zhaoyang, et al.
Published: (2023)
by: Zhang, Zhaoyang, et al.
Published: (2023)
Mapping Smarter, Not Harder: A Test-Time Reinforcement Learning Agent That Improves Without Labels or Model Updates
by: Tsao, Wen-Kwang, et al.
Published: (2025)
by: Tsao, Wen-Kwang, et al.
Published: (2025)
Reasoning with Sampling: Your Base Model is Smarter Than You Think
by: Karan, Aayush, et al.
Published: (2025)
by: Karan, Aayush, et al.
Published: (2025)
Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization
by: Yu, Zishun, et al.
Published: (2025)
by: Yu, Zishun, et al.
Published: (2025)
Similar Items
-
Memory in Plain Sight: Surveying the Uncanny Resemblances of Associative Memories and Diffusion Models
by: Hoover, Benjamin, et al.
Published: (2023) -
Dense Associative Memory Through the Lens of Random Features
by: Hoover, Benjamin, et al.
Published: (2024) -
NRGPT: An Energy-based Alternative for GPT
by: Dehmamy, Nima, et al.
Published: (2025) -
Tensor Cache: Eviction-conditioned Associative Memory for Transformers
by: Swain, Kabir, et al.
Published: (2026) -
Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers
by: Swain, Kabir, et al.
Published: (2026)