When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Galeone, Cosimo, Park, Minsu, Ettorre, Giuseppe, Ligorio, Daniele |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
por: Xu, Xiaoyu, et al.
Publicado: (2025)
por: Xu, Xiaoyu, et al.
Publicado: (2025)
When Privacy Isn't Synthetic: Hidden Data Leakage in Generative AI Models
por: Mustaqim, S. M., et al.
Publicado: (2025)
por: Mustaqim, S. M., et al.
Publicado: (2025)
Inverse Scaling: When Bigger Isn't Better
por: McKenzie, Ian R., et al.
Publicado: (2023)
por: McKenzie, Ian R., et al.
Publicado: (2023)
Why Isn't Relational Learning Taking Over the World?
por: Poole, David
Publicado: (2025)
por: Poole, David
Publicado: (2025)
SLOT: Structuring the Output of Large Language Models
por: Wang, Darren Yow-Bang, et al.
Publicado: (2025)
por: Wang, Darren Yow-Bang, et al.
Publicado: (2025)
Domain-Adapted Small Language Models for Reliable Clinical Triage
por: Aljohani, Manar, et al.
Publicado: (2026)
por: Aljohani, Manar, et al.
Publicado: (2026)
When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning
por: Barale, Claire, et al.
Publicado: (2025)
por: Barale, Claire, et al.
Publicado: (2025)
Understanding Dataset Difficulty with $\mathcal{V}$-Usable Information
por: Ethayarajh, Kawin, et al.
Publicado: (2021)
por: Ethayarajh, Kawin, et al.
Publicado: (2021)
Clarify: Improving Model Robustness With Natural Language Corrections
por: Lee, Yoonho, et al.
Publicado: (2024)
por: Lee, Yoonho, et al.
Publicado: (2024)
Structured Prompts Improve Evaluation of Language Models
por: Aali, Asad, et al.
Publicado: (2025)
por: Aali, Asad, et al.
Publicado: (2025)
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
por: Barkett, Emilio, et al.
Publicado: (2025)
por: Barkett, Emilio, et al.
Publicado: (2025)
The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
por: Cho, Seonglae, et al.
Publicado: (2026)
por: Cho, Seonglae, et al.
Publicado: (2026)
An Evaluation on Large Language Model Outputs: Discourse and Memorization
por: de Wynter, Adrian, et al.
Publicado: (2023)
por: de Wynter, Adrian, et al.
Publicado: (2023)
Leviathan: Decoupling Input and Output Representations in Language Models
por: Batley, Reza T., et al.
Publicado: (2026)
por: Batley, Reza T., et al.
Publicado: (2026)
Adapt-Pruner: Adaptive Structural Pruning for Efficient Small Language Model Training
por: Pan, Rui, et al.
Publicado: (2025)
por: Pan, Rui, et al.
Publicado: (2025)
Improving Language Models with Intentional Analysis
por: Yin, Yuwei, et al.
Publicado: (2025)
por: Yin, Yuwei, et al.
Publicado: (2025)
Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic Inference
por: Zhao, Stephen, et al.
Publicado: (2025)
por: Zhao, Stephen, et al.
Publicado: (2025)
Graph-based Uncertainty Metrics for Long-form Language Model Outputs
por: Jiang, Mingjian, et al.
Publicado: (2024)
por: Jiang, Mingjian, et al.
Publicado: (2024)
Group-Aware Reinforcement Learning for Output Diversity in Large Language Models
por: Anschel, Oron, et al.
Publicado: (2025)
por: Anschel, Oron, et al.
Publicado: (2025)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
por: Zhang, Yue, et al.
Publicado: (2026)
por: Zhang, Yue, et al.
Publicado: (2026)
Integrating Locality-Aware Attention with Transformers for General Geometry PDEs
por: Koh, Minsu, et al.
Publicado: (2025)
por: Koh, Minsu, et al.
Publicado: (2025)
Reliable, Adaptable, and Attributable Language Models with Retrieval
por: Asai, Akari, et al.
Publicado: (2024)
por: Asai, Akari, et al.
Publicado: (2024)
Context Dependence and Reliability in Autoregressive Language Models
por: Sengupta, Poushali, et al.
Publicado: (2026)
por: Sengupta, Poushali, et al.
Publicado: (2026)
Stacking Small Language Models for Generalizability
por: Liang, Laurence
Publicado: (2024)
por: Liang, Laurence
Publicado: (2024)
All Language Models Large and Small
por: Chen, Zhixun, et al.
Publicado: (2024)
por: Chen, Zhixun, et al.
Publicado: (2024)
Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification
por: Fadeeva, Ekaterina, et al.
Publicado: (2024)
por: Fadeeva, Ekaterina, et al.
Publicado: (2024)
Kakugo: Distillation of Low-Resource Languages into Small Language Models
por: Devine, Peter, et al.
Publicado: (2026)
por: Devine, Peter, et al.
Publicado: (2026)
Curriculum Learning for Small Code Language Models
por: Naïr, Marwa, et al.
Publicado: (2024)
por: Naïr, Marwa, et al.
Publicado: (2024)
Small Language Models: Survey, Measurements, and Insights
por: Lu, Zhenyan, et al.
Publicado: (2024)
por: Lu, Zhenyan, et al.
Publicado: (2024)
Towards Reasoning Ability of Small Language Models
por: Srivastava, Gaurav, et al.
Publicado: (2025)
por: Srivastava, Gaurav, et al.
Publicado: (2025)
Squat: Quant Small Language Models on the Edge
por: Shen, Xuan, et al.
Publicado: (2024)
por: Shen, Xuan, et al.
Publicado: (2024)
Fine-Grained Uncertainty Quantification for Long-Form Language Model Outputs: A Comparative Study
por: Bouchard, Dylan, et al.
Publicado: (2026)
por: Bouchard, Dylan, et al.
Publicado: (2026)
A Voter-Based Stochastic Rejection-Method Framework for Asymptotically Safe Language Model Outputs
por: Watts, Jake R., et al.
Publicado: (2024)
por: Watts, Jake R., et al.
Publicado: (2024)
Improving Large Models with Small models: Lower Costs and Better Performance
por: Chen, Dong, et al.
Publicado: (2024)
por: Chen, Dong, et al.
Publicado: (2024)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
por: Sinha, Neelabh, et al.
Publicado: (2024)
por: Sinha, Neelabh, et al.
Publicado: (2024)
GFlowPO: Generative Flow Network as a Language Model Prompt Optimizer
por: Cho, Junmo, et al.
Publicado: (2026)
por: Cho, Junmo, et al.
Publicado: (2026)
When Attention Sink Emerges in Language Models: An Empirical View
por: Gu, Xiangming, et al.
Publicado: (2024)
por: Gu, Xiangming, et al.
Publicado: (2024)
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
por: Feizi, Aarash, et al.
Publicado: (2025)
por: Feizi, Aarash, et al.
Publicado: (2025)
Phase Transitions in the Output Distribution of Large Language Models
por: Arnold, Julian, et al.
Publicado: (2024)
por: Arnold, Julian, et al.
Publicado: (2024)
When Should LLMs Be Less Specific? Selective Abstraction for Reliable Long-Form Text Generation
por: Goren, Shani, et al.
Publicado: (2026)
por: Goren, Shani, et al.
Publicado: (2026)
Ejemplares similares
-
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
por: Xu, Xiaoyu, et al.
Publicado: (2025) -
When Privacy Isn't Synthetic: Hidden Data Leakage in Generative AI Models
por: Mustaqim, S. M., et al.
Publicado: (2025) -
Inverse Scaling: When Bigger Isn't Better
por: McKenzie, Ian R., et al.
Publicado: (2023) -
Why Isn't Relational Learning Taking Over the World?
por: Poole, David
Publicado: (2025) -
SLOT: Structuring the Output of Large Language Models
por: Wang, Darren Yow-Bang, et al.
Publicado: (2025)