Diverse LLMs or Diverse Question Interpretations? That is the Ensembling Question
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rosales, Rafael, Miret, Santiago |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
von: Badshah, Sher, et al.
Veröffentlicht: (2024)
von: Badshah, Sher, et al.
Veröffentlicht: (2024)
Chatbots put to the test in math and logic problems: A preliminary comparison and assessment of ChatGPT-3.5, ChatGPT-4, and Google Bard
von: Plevris, Vagelis, et al.
Veröffentlicht: (2023)
von: Plevris, Vagelis, et al.
Veröffentlicht: (2023)
Prompt Tuned Embedding Classification for Multi-Label Industry Sector Allocation
von: Buchner, Valentin Leonhard, et al.
Veröffentlicht: (2023)
von: Buchner, Valentin Leonhard, et al.
Veröffentlicht: (2023)
Large Language Models Report Subjective Experience Under Self-Referential Processing
von: Berg, Cameron, et al.
Veröffentlicht: (2025)
von: Berg, Cameron, et al.
Veröffentlicht: (2025)
CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
von: Badshah, Sher, et al.
Veröffentlicht: (2025)
von: Badshah, Sher, et al.
Veröffentlicht: (2025)
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
von: Sarkar, Nilesh, et al.
Veröffentlicht: (2026)
von: Sarkar, Nilesh, et al.
Veröffentlicht: (2026)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
Data and AI governance: Promoting equity, ethics, and fairness in large language models
von: Abhishek, Alok, et al.
Veröffentlicht: (2025)
von: Abhishek, Alok, et al.
Veröffentlicht: (2025)
BEATS: Bias Evaluation and Assessment Test Suite for Large Language Models
von: Abhishek, Alok, et al.
Veröffentlicht: (2025)
von: Abhishek, Alok, et al.
Veröffentlicht: (2025)
SHARP: Social Harm Analysis via Risk Profiles for Measuring Inequities in Large Language Models
von: Abhishek, Alok, et al.
Veröffentlicht: (2026)
von: Abhishek, Alok, et al.
Veröffentlicht: (2026)
Reasoning Promotes Robustness in Theory of Mind Tasks
von: de Haan, Ian B., et al.
Veröffentlicht: (2026)
von: de Haan, Ian B., et al.
Veröffentlicht: (2026)
GIM: Evaluating models via tasks that integrate multiple cognitive domains
von: Patel, Rohit, et al.
Veröffentlicht: (2026)
von: Patel, Rohit, et al.
Veröffentlicht: (2026)
Generative AI for Enhancing Active Learning in Education: A Comparative Study of GPT-3.5 and GPT-4 in Crafting Customized Test Questions
von: Rouzegar, Hamdireza, et al.
Veröffentlicht: (2024)
von: Rouzegar, Hamdireza, et al.
Veröffentlicht: (2024)
JAM: Controllable and Responsible Text Generation via Causal Reasoning and Latent Vector Manipulation
von: Huang, Yingbing, et al.
Veröffentlicht: (2025)
von: Huang, Yingbing, et al.
Veröffentlicht: (2025)
Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
von: Mahale, Ajay Pravin
Veröffentlicht: (2026)
von: Mahale, Ajay Pravin
Veröffentlicht: (2026)
CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2025)
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2025)
A Graph-based Approach for Multi-Modal Question Answering from Flowcharts in Telecom Documents
von: Soman, Sumit, et al.
Veröffentlicht: (2025)
von: Soman, Sumit, et al.
Veröffentlicht: (2025)
Multi-Model Synthetic Training for Mission-Critical Small Language Models
von: Platt, Nolan, et al.
Veröffentlicht: (2025)
von: Platt, Nolan, et al.
Veröffentlicht: (2025)
Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
A Confidence-Diversity Framework for Calibrating AI Judgement in Accessible Qualitative Coding Tasks
von: Zhao, Zhilong, et al.
Veröffentlicht: (2025)
von: Zhao, Zhilong, et al.
Veröffentlicht: (2025)
Evaluation of RAG Metrics for Question Answering in the Telecom Domain
von: Roychowdhury, Sujoy, et al.
Veröffentlicht: (2024)
von: Roychowdhury, Sujoy, et al.
Veröffentlicht: (2024)
Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs
von: Anshumann, et al.
Veröffentlicht: (2025)
von: Anshumann, et al.
Veröffentlicht: (2025)
Triad: A Framework Leveraging a Multi-Role LLM-based Agent to Solve Knowledge Base Question Answering
von: Zong, Chang, et al.
Veröffentlicht: (2024)
von: Zong, Chang, et al.
Veröffentlicht: (2024)
Survey and Evaluation of Converging Architecture in LLMs based on Footsteps of Operations
von: Kim, Seongho, et al.
Veröffentlicht: (2024)
von: Kim, Seongho, et al.
Veröffentlicht: (2024)
Vibe-Creation: The Epistemology of Human-AI Emergent Cognition
von: Levin, Ilya
Veröffentlicht: (2026)
von: Levin, Ilya
Veröffentlicht: (2026)
Hybrid Gated Flow (HGF): Stabilizing 1.58-bit LLMs via Selective Low-Rank Correction
von: Pizzo, David Alejandro Trejo
Veröffentlicht: (2026)
von: Pizzo, David Alejandro Trejo
Veröffentlicht: (2026)
Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data?
von: Kovač, Grgur, et al.
Veröffentlicht: (2025)
von: Kovač, Grgur, et al.
Veröffentlicht: (2025)
On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance
von: Casanova, Etienne, et al.
Veröffentlicht: (2026)
von: Casanova, Etienne, et al.
Veröffentlicht: (2026)
Evaluating Prompting Strategies for Chart Question Answering with Large Language Models
von: Naikar, Ruthuparna, et al.
Veröffentlicht: (2026)
von: Naikar, Ruthuparna, et al.
Veröffentlicht: (2026)
Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation
von: Edin, Joakim, et al.
Veröffentlicht: (2025)
von: Edin, Joakim, et al.
Veröffentlicht: (2025)
CAPE: Corrective Actions from Precondition Errors using Large Language Models
von: Raman, Shreyas Sundara, et al.
Veröffentlicht: (2022)
von: Raman, Shreyas Sundara, et al.
Veröffentlicht: (2022)
Knowledge Distillation of Domain-adapted LLMs for Question-Answering in Telecom
von: Sen, Rishika, et al.
Veröffentlicht: (2025)
von: Sen, Rishika, et al.
Veröffentlicht: (2025)
Mitigating Position-Shift Failures in Text-Based Modular Arithmetic via Position Curriculum and Template Diversity
von: Yudin, Nikolay
Veröffentlicht: (2026)
von: Yudin, Nikolay
Veröffentlicht: (2026)
Measuring Intent Comprehension in LLMs
von: Kunievsky, Nadav, et al.
Veröffentlicht: (2025)
von: Kunievsky, Nadav, et al.
Veröffentlicht: (2025)
Integrating External Tools with Large Language Models to Improve Accuracy
von: Niketan, Nripesh, et al.
Veröffentlicht: (2025)
von: Niketan, Nripesh, et al.
Veröffentlicht: (2025)
ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2025)
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2025)
NRR-Core: Non-Resolution Reasoning as a Computational Framework for Contextual Identity and Ambiguity Preservation
von: Saito, Kei
Veröffentlicht: (2025)
von: Saito, Kei
Veröffentlicht: (2025)
RWKV-7 "Goose" with Expressive Dynamic State Evolution
von: Peng, Bo, et al.
Veröffentlicht: (2025)
von: Peng, Bo, et al.
Veröffentlicht: (2025)
Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory
von: Goldstein, Daniel, et al.
Veröffentlicht: (2026)
von: Goldstein, Daniel, et al.
Veröffentlicht: (2026)
Change Is the Only Constant: Dynamic LLM Slicing based on Layer Redundancy
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
von: Badshah, Sher, et al.
Veröffentlicht: (2024) -
Chatbots put to the test in math and logic problems: A preliminary comparison and assessment of ChatGPT-3.5, ChatGPT-4, and Google Bard
von: Plevris, Vagelis, et al.
Veröffentlicht: (2023) -
Prompt Tuned Embedding Classification for Multi-Label Industry Sector Allocation
von: Buchner, Valentin Leonhard, et al.
Veröffentlicht: (2023) -
Large Language Models Report Subjective Experience Under Self-Referential Processing
von: Berg, Cameron, et al.
Veröffentlicht: (2025) -
CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
von: Badshah, Sher, et al.
Veröffentlicht: (2025)