Saved in:
| Main Authors: | Donaldson, Artur, Balaji, Bharathan, Oriekezie, Cajetan, Kumar, Manish, Patouillard, Laure |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.19886 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpiderGen: Towards Procedure Generation For Carbon Life Cycle Assessments with Generative AI
by: Sitaraman, Anupama, et al.
Published: (2025)
by: Sitaraman, Anupama, et al.
Published: (2025)
CF-RAG: A Dataset and Method for Carbon Footprint QA Using Retrieval-Augmented Generation
by: Zhao, Kaiwen, et al.
Published: (2025)
by: Zhao, Kaiwen, et al.
Published: (2025)
Structured Outputs Enable General-Purpose LLMs to be Medical Experts
by: Guo, Guangfu, et al.
Published: (2025)
by: Guo, Guangfu, et al.
Published: (2025)
FAME-MT Dataset: Formality Awareness Made Easy for Machine Translation Purposes
by: Wiśniewski, Dawid, et al.
Published: (2024)
by: Wiśniewski, Dawid, et al.
Published: (2024)
Skill-LLM: Repurposing General-Purpose LLMs for Skill Extraction
by: Herandi, Amirhossein, et al.
Published: (2024)
by: Herandi, Amirhossein, et al.
Published: (2024)
s2n-bignum-bench: A practical benchmark for evaluating low-level code reasoning of LLMs
by: Rao, Balaji, et al.
Published: (2026)
by: Rao, Balaji, et al.
Published: (2026)
A General-Purpose Device for Interaction with LLMs
by: Xu, Jiajun, et al.
Published: (2024)
by: Xu, Jiajun, et al.
Published: (2024)
Text2Model: Generating dynamic chemical reactor models using large language models (LLMs)
by: Rupprecht, Sophia, et al.
Published: (2025)
by: Rupprecht, Sophia, et al.
Published: (2025)
AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs
by: Fathullah, Yassir, et al.
Published: (2023)
by: Fathullah, Yassir, et al.
Published: (2023)
Tag-LLM: Repurposing General-Purpose LLMs for Specialized Domains
by: Shen, Junhong, et al.
Published: (2024)
by: Shen, Junhong, et al.
Published: (2024)
TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts
by: Wang, Ruida, et al.
Published: (2024)
by: Wang, Ruida, et al.
Published: (2024)
A Mixture-of-Experts Model for Multimodal Emotion Recognition in Conversations
by: Dutta, Soumya, et al.
Published: (2026)
by: Dutta, Soumya, et al.
Published: (2026)
Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs
by: Wu, Wei, et al.
Published: (2024)
by: Wu, Wei, et al.
Published: (2024)
LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training
by: Wang, Yiming, et al.
Published: (2025)
by: Wang, Yiming, et al.
Published: (2025)
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
From Beginner to Expert: Modeling Medical Knowledge into General LLMs
by: Li, Qiang, et al.
Published: (2023)
by: Li, Qiang, et al.
Published: (2023)
ExpertSteer: Intervening in LLMs through Expert Knowledge
by: Wang, Weixuan, et al.
Published: (2025)
by: Wang, Weixuan, et al.
Published: (2025)
CangjieBench: Benchmarking LLMs on a Low-Resource General-Purpose Programming Language
by: Cheng, Junhang, et al.
Published: (2026)
by: Cheng, Junhang, et al.
Published: (2026)
WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics
by: Maurya, Sneha, et al.
Published: (2026)
by: Maurya, Sneha, et al.
Published: (2026)
Suvach -- Generated Hindi QA benchmark
by: Narayanan, Vaishak, et al.
Published: (2024)
by: Narayanan, Vaishak, et al.
Published: (2024)
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning
by: Patnaik, Sohan, et al.
Published: (2025)
by: Patnaik, Sohan, et al.
Published: (2025)
Learning High-Quality and General-Purpose Phrase Representations
by: Chen, Lihu, et al.
Published: (2024)
by: Chen, Lihu, et al.
Published: (2024)
AdaptBPE: From General Purpose to Specialized Tokenizers
by: Liyanage, Vijini, et al.
Published: (2026)
by: Liyanage, Vijini, et al.
Published: (2026)
Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs
by: Bai, Jun, et al.
Published: (2025)
by: Bai, Jun, et al.
Published: (2025)
It Helps to Take a Second Opinion: Teaching Smaller LLMs to Deliberate Mutually via Selective Rationale Optimisation
by: Patnaik, Sohan, et al.
Published: (2025)
by: Patnaik, Sohan, et al.
Published: (2025)
Mixture of Experts for Low-Resource LLMs
by: Joseph, Ori Bar, et al.
Published: (2026)
by: Joseph, Ori Bar, et al.
Published: (2026)
Implicit Causality-biases in humans and LLMs as a tool for benchmarking LLM discourse capabilities
by: Kankowski, Florian, et al.
Published: (2025)
by: Kankowski, Florian, et al.
Published: (2025)
Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs
by: Abdaljalil, Samir, et al.
Published: (2026)
by: Abdaljalil, Samir, et al.
Published: (2026)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
by: Wang, Chonghua, et al.
Published: (2024)
by: Wang, Chonghua, et al.
Published: (2024)
RedunCut: Measurement-Driven Sampling and Accuracy Performance Modeling for Low-Cost Live Video Analytics
by: Sela, Gur-Eyal, et al.
Published: (2025)
by: Sela, Gur-Eyal, et al.
Published: (2025)
General Purpose Verification for Chain of Thought Prompting
by: Vacareanu, Robert, et al.
Published: (2024)
by: Vacareanu, Robert, et al.
Published: (2024)
Do LLMs Understand Why We Write Diaries? A Method for Purpose Extraction and Clustering
by: Goloviznina, Valeriya, et al.
Published: (2025)
by: Goloviznina, Valeriya, et al.
Published: (2025)
A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
by: Wang, Andrew Z., et al.
Published: (2025)
by: Wang, Andrew Z., et al.
Published: (2025)
LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models
by: Nguyen, Nam V., et al.
Published: (2024)
by: Nguyen, Nam V., et al.
Published: (2024)
MultiLoKo: a multilingual local knowledge benchmark for LLMs spanning 31 languages
by: Hupkes, Dieuwke, et al.
Published: (2025)
by: Hupkes, Dieuwke, et al.
Published: (2025)
Clinical named entity recognition in the Portuguese language: a benchmark of modern BERT models and LLMs
by: de Almeida, Vinicius Anjos, et al.
Published: (2026)
by: de Almeida, Vinicius Anjos, et al.
Published: (2026)
Infusing Knowledge into Large Language Models with Contextual Prompts
by: Vasisht, Kinshuk, et al.
Published: (2024)
by: Vasisht, Kinshuk, et al.
Published: (2024)
Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
by: Zhou, Yixiao, et al.
Published: (2025)
by: Zhou, Yixiao, et al.
Published: (2025)
Understanding Multilingualism in Mixture-of-Experts LLMs: Routing Mechanism, Expert Specialization, and Layerwise Steering
by: Chen, Yuxin, et al.
Published: (2026)
by: Chen, Yuxin, et al.
Published: (2026)
LCA and energy efficiency in buildings: mapping more than twenty years of research
by: Asdrubali, F., et al.
Published: (2024)
by: Asdrubali, F., et al.
Published: (2024)
Similar Items
-
SpiderGen: Towards Procedure Generation For Carbon Life Cycle Assessments with Generative AI
by: Sitaraman, Anupama, et al.
Published: (2025) -
CF-RAG: A Dataset and Method for Carbon Footprint QA Using Retrieval-Augmented Generation
by: Zhao, Kaiwen, et al.
Published: (2025) -
Structured Outputs Enable General-Purpose LLMs to be Medical Experts
by: Guo, Guangfu, et al.
Published: (2025) -
FAME-MT Dataset: Formality Awareness Made Easy for Machine Translation Purposes
by: Wiśniewski, Dawid, et al.
Published: (2024) -
Skill-LLM: Repurposing General-Purpose LLMs for Skill Extraction
by: Herandi, Amirhossein, et al.
Published: (2024)