Evaluating LLMs' Mathematical and Coding Competency through Ontology-guided Interventions
Fuente:
arXiv
Guardado en:
| Autores principales: | Hong, Pengfei, Majumder, Navonil, Ghosal, Deepanway, Aditya, Somak, Mihalcea, Rada, Poria, Soujanya |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
por: Majumder, Navonil, et al.
Publicado: (2024)
por: Majumder, Navonil, et al.
Publicado: (2024)
Not All Votes Count! Programs as Verifiers Improve Self-Consistency of Language Models for Math Reasoning
por: Toh, Vernon Y. H., et al.
Publicado: (2024)
por: Toh, Vernon Y. H., et al.
Publicado: (2024)
Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse
por: Song, Maojia, et al.
Publicado: (2024)
por: Song, Maojia, et al.
Publicado: (2024)
Improving Text-To-Audio Models with Synthetic Captions
por: Kong, Zhifeng, et al.
Publicado: (2024)
por: Kong, Zhifeng, et al.
Publicado: (2024)
Inference Time Alignment with Reward-Guided Tree Search
por: Hung, Chia-Yu, et al.
Publicado: (2024)
por: Hung, Chia-Yu, et al.
Publicado: (2024)
NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks
por: Dutta, Aritra, et al.
Publicado: (2025)
por: Dutta, Aritra, et al.
Publicado: (2025)
Mustango: Toward Controllable Text-to-Music Generation
por: Melechovsky, Jan, et al.
Publicado: (2023)
por: Melechovsky, Jan, et al.
Publicado: (2023)
The Jumping Reasoning Curve? Tracking the Evolution of Reasoning Performance in GPT-[n] and o-[n] Models on Multimodal Puzzles
por: Toh, Vernon Y. H., et al.
Publicado: (2025)
por: Toh, Vernon Y. H., et al.
Publicado: (2025)
Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
por: Sun, Qi, et al.
Publicado: (2024)
por: Sun, Qi, et al.
Publicado: (2024)
Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense
por: Shen, Siqi, et al.
Publicado: (2024)
por: Shen, Siqi, et al.
Publicado: (2024)
Lessons from Training Grounded LLMs with Verifiable Rewards
por: Sim, Shang Hong, et al.
Publicado: (2025)
por: Sim, Shang Hong, et al.
Publicado: (2025)
LOGICPO: Efficient Translation of NL-based Logical Problems to FOL using LLMs and Preference Optimization
por: Viswanadha, Koushik, et al.
Publicado: (2025)
por: Viswanadha, Koushik, et al.
Publicado: (2025)
Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning
por: Ghosal, Deepanway, et al.
Publicado: (2024)
por: Ghosal, Deepanway, et al.
Publicado: (2024)
The Curious Case of Curiosity across Human Cultures and LLMs
por: Borah, Angana, et al.
Publicado: (2025)
por: Borah, Angana, et al.
Publicado: (2025)
DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling
por: Deep, Pala Tej, et al.
Publicado: (2024)
por: Deep, Pala Tej, et al.
Publicado: (2024)
Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMs
por: Puerto, Haritz, et al.
Publicado: (2024)
por: Puerto, Haritz, et al.
Publicado: (2024)
Towards Region-aware Bias Evaluation Metrics
por: Borah, Angana, et al.
Publicado: (2024)
por: Borah, Angana, et al.
Publicado: (2024)
Mind the (Belief) Gap: Group Identity in the World of LLMs
por: Borah, Angana, et al.
Publicado: (2025)
por: Borah, Angana, et al.
Publicado: (2025)
Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical Interactions
por: Deng, Naihao, et al.
Publicado: (2025)
por: Deng, Naihao, et al.
Publicado: (2025)
Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions
por: Borah, Angana, et al.
Publicado: (2024)
por: Borah, Angana, et al.
Publicado: (2024)
Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
por: Bhardwaj, Rishabh, et al.
Publicado: (2024)
por: Bhardwaj, Rishabh, et al.
Publicado: (2024)
PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns
por: Chia, Yew Ken, et al.
Publicado: (2024)
por: Chia, Yew Ken, et al.
Publicado: (2024)
[WIP] Jailbreak Paradox: The Achilles' Heel of LLMs
por: Rao, Abhinav, et al.
Publicado: (2024)
por: Rao, Abhinav, et al.
Publicado: (2024)
Rethinking Table Instruction Tuning
por: Deng, Naihao, et al.
Publicado: (2025)
por: Deng, Naihao, et al.
Publicado: (2025)
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
por: Hung, Chia-Yu, et al.
Publicado: (2024)
por: Hung, Chia-Yu, et al.
Publicado: (2024)
Towards Robust Instruction Tuning on Multimodal Large Language Models
por: Han, Wei, et al.
Publicado: (2024)
por: Han, Wei, et al.
Publicado: (2024)
MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning
por: Das, Debrup, et al.
Publicado: (2024)
por: Das, Debrup, et al.
Publicado: (2024)
PREMISE: Matching-based Prediction for Accurate Review Recommendation
por: Han, Wei, et al.
Publicado: (2025)
por: Han, Wei, et al.
Publicado: (2025)
The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models
por: Arif, Samee, et al.
Publicado: (2026)
por: Arif, Samee, et al.
Publicado: (2026)
Ruby Teaming: Improving Quality Diversity Search with Memory for Automated Red Teaming
por: Han, Vernon Toh Yan, et al.
Publicado: (2024)
por: Han, Vernon Toh Yan, et al.
Publicado: (2024)
Sowing the Wind, Reaping the Whirlwind: The Impact of Editing Language Models
por: Hazra, Rima, et al.
Publicado: (2024)
por: Hazra, Rima, et al.
Publicado: (2024)
Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations
por: Hazra, Rima, et al.
Publicado: (2024)
por: Hazra, Rima, et al.
Publicado: (2024)
NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
por: Hung, Chia-Yu, et al.
Publicado: (2025)
por: Hung, Chia-Yu, et al.
Publicado: (2025)
Whose wife is it anyway? Assessing bias against same-gender relationships in machine translation
por: Stewart, Ian, et al.
Publicado: (2024)
por: Stewart, Ian, et al.
Publicado: (2024)
Consistency Guided Knowledge Retrieval and Denoising in LLMs for Zero-shot Document-level Relation Triplet Extraction
por: Sun, Qi, et al.
Publicado: (2024)
por: Sun, Qi, et al.
Publicado: (2024)
MuCRASP: Multimodal Chain-of-thought Reasoning aware Structured Pruning
por: Dutta, Aritra, et al.
Publicado: (2026)
por: Dutta, Aritra, et al.
Publicado: (2026)
MAiDE-up: Multilingual Deception Detection of GPT-generated Hotel Reviews
por: Ignat, Oana, et al.
Publicado: (2024)
por: Ignat, Oana, et al.
Publicado: (2024)
Towards Dog Bark Decoding: Leveraging Human Speech Processing for Automated Bark Classification
por: Abzaliev, Artem, et al.
Publicado: (2024)
por: Abzaliev, Artem, et al.
Publicado: (2024)
Patient-Centered RAG for Oncology Visit Aid Following the Ottawa Decision Guide
por: Liu, Siyang, et al.
Publicado: (2025)
por: Liu, Siyang, et al.
Publicado: (2025)
Persuasion at Play: Understanding Misinformation Dynamics in Demographic-Aware Human-LLM Interactions
por: Borah, Angana, et al.
Publicado: (2025)
por: Borah, Angana, et al.
Publicado: (2025)
Ejemplares similares
-
Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
por: Majumder, Navonil, et al.
Publicado: (2024) -
Not All Votes Count! Programs as Verifiers Improve Self-Consistency of Language Models for Math Reasoning
por: Toh, Vernon Y. H., et al.
Publicado: (2024) -
Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse
por: Song, Maojia, et al.
Publicado: (2024) -
Improving Text-To-Audio Models with Synthetic Captions
por: Kong, Zhifeng, et al.
Publicado: (2024) -
Inference Time Alignment with Reward-Guided Tree Search
por: Hung, Chia-Yu, et al.
Publicado: (2024)