StructTest: Benchmarking LLMs' Reasoning through Compositional Structured Outputs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Hailin, Jiao, Fangkai, Ravaut, Mathieu, Farruque, Nawshad, Nguyen, Xuan Phi, Qin, Chengwei, Dey, Manan, Ding, Bosheng, Xiong, Caiming, Joty, Shafiq, Zhou, Yingbo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Comprehensive Survey of Contamination Detection Methods in Large Language Models
von: Ravaut, Mathieu, et al.
Veröffentlicht: (2024)
von: Ravaut, Mathieu, et al.
Veröffentlicht: (2024)
ChatGPT's One-year Anniversary: Are Open-Source Large Language Models Catching up?
von: Chen, Hailin, et al.
Veröffentlicht: (2023)
von: Chen, Hailin, et al.
Veröffentlicht: (2023)
Demystifying Domain-adaptive Post-training for Financial LLMs
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
Beyond Output Matching: Bidirectional Alignment for Enhanced In-Context Learning
von: Qin, Chengwei, et al.
Veröffentlicht: (2023)
von: Qin, Chengwei, et al.
Veröffentlicht: (2023)
Unsupervised Summarization Re-ranking
von: Ravaut, Mathieu, et al.
Veröffentlicht: (2022)
von: Ravaut, Mathieu, et al.
Veröffentlicht: (2022)
Relevant or Random: Can LLMs Truly Perform Analogical Reasoning?
von: Qin, Chengwei, et al.
Veröffentlicht: (2024)
von: Qin, Chengwei, et al.
Veröffentlicht: (2024)
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2026)
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2026)
A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
On Context Utilization in Summarization with Large Language Models
von: Ravaut, Mathieu, et al.
Veröffentlicht: (2023)
von: Ravaut, Mathieu, et al.
Veröffentlicht: (2023)
Exploring Self-supervised Logic-enhanced Training for Large Language Models
von: Jiao, Fangkai, et al.
Veröffentlicht: (2023)
von: Jiao, Fangkai, et al.
Veröffentlicht: (2023)
Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing
von: Jiao, Fangkai, et al.
Veröffentlicht: (2024)
von: Jiao, Fangkai, et al.
Veröffentlicht: (2024)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
Democratizing LLMs for Low-Resource Languages by Leveraging their English Dominant Abilities with Linguistically-Diverse Prompts
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2023)
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2023)
SFR-RAG: Towards Contextually Faithful LLMs
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2024)
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2024)
MAS-ZERO: Designing Multi-Agent Systems with Zero Supervision
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"
von: Ming, Yifei, et al.
Veröffentlicht: (2024)
von: Ming, Yifei, et al.
Veröffentlicht: (2024)
Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction
von: Shi, Zhenmei, et al.
Veröffentlicht: (2024)
von: Shi, Zhenmei, et al.
Veröffentlicht: (2024)
Modeling Uncertainty and Using Post-fusion as Fallback Improves Retrieval Augmented Generation with LLMs
von: Liu, Ye, et al.
Veröffentlicht: (2023)
von: Liu, Ye, et al.
Veröffentlicht: (2023)
ParaICL: Towards Parallel In-Context Learning
von: Li, Xingxuan, et al.
Veröffentlicht: (2024)
von: Li, Xingxuan, et al.
Veröffentlicht: (2024)
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2025)
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2025)
JudgeRank: Leveraging Large Language Models for Reasoning-Intensive Reranking
von: Niu, Tong, et al.
Veröffentlicht: (2024)
von: Niu, Tong, et al.
Veröffentlicht: (2024)
MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks
von: Ke, Zixuan, et al.
Veröffentlicht: (2026)
von: Ke, Zixuan, et al.
Veröffentlicht: (2026)
Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks
von: Li, Xingxuan, et al.
Veröffentlicht: (2024)
von: Li, Xingxuan, et al.
Veröffentlicht: (2024)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
SweRank: Software Issue Localization with Code Ranking
von: Reddy, Revanth Gangi, et al.
Veröffentlicht: (2025)
von: Reddy, Revanth Gangi, et al.
Veröffentlicht: (2025)
Personalised Distillation: Empowering Open-Sourced LLMs with Adaptive Learning for Code Generation
von: Chen, Hailin, et al.
Veröffentlicht: (2023)
von: Chen, Hailin, et al.
Veröffentlicht: (2023)
Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows
von: Ming, Yifei, et al.
Veröffentlicht: (2025)
von: Ming, Yifei, et al.
Veröffentlicht: (2025)
CodeXEmbed: A Generalist Embedding Model Family for Multiligual and Multi-task Code Retrieval
von: Liu, Ye, et al.
Veröffentlicht: (2024)
von: Liu, Ye, et al.
Veröffentlicht: (2024)
Efficiently Aligned Cross-Lingual Transfer Learning for Conversational Tasks using Prompt-Tuning
von: Tu, Lifu, et al.
Veröffentlicht: (2023)
von: Tu, Lifu, et al.
Veröffentlicht: (2023)
LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs
von: Long, Do Xuan, et al.
Veröffentlicht: (2024)
von: Long, Do Xuan, et al.
Veröffentlicht: (2024)
VIBEPASS: Can Vibe Coders Really Pass the Vibe Check?
von: Bansal, Srijan, et al.
Veröffentlicht: (2026)
von: Bansal, Srijan, et al.
Veröffentlicht: (2026)
Direct Judgement Preference Optimization
von: Wang, Peifeng, et al.
Veröffentlicht: (2024)
von: Wang, Peifeng, et al.
Veröffentlicht: (2024)
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
von: Yang, Jialin, et al.
Veröffentlicht: (2025)
von: Yang, Jialin, et al.
Veröffentlicht: (2025)
Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement Learning
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
Preference Optimization for Reasoning with Pseudo Feedback
von: Jiao, Fangkai, et al.
Veröffentlicht: (2024)
von: Jiao, Fangkai, et al.
Veröffentlicht: (2024)
Investigating Factuality in Long-Form Text Generation: The Roles of Self-Known and Self-Unknown
von: Tu, Lifu, et al.
Veröffentlicht: (2024)
von: Tu, Lifu, et al.
Veröffentlicht: (2024)
Lifelong Event Detection with Embedding Space Separation and Compaction
von: Qin, Chengwei, et al.
Veröffentlicht: (2024)
von: Qin, Chengwei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Comprehensive Survey of Contamination Detection Methods in Large Language Models
von: Ravaut, Mathieu, et al.
Veröffentlicht: (2024) -
ChatGPT's One-year Anniversary: Are Open-Source Large Language Models Catching up?
von: Chen, Hailin, et al.
Veröffentlicht: (2023) -
Demystifying Domain-adaptive Post-training for Financial LLMs
von: Ke, Zixuan, et al.
Veröffentlicht: (2025) -
Beyond Output Matching: Bidirectional Alignment for Enhanced In-Context Learning
von: Qin, Chengwei, et al.
Veröffentlicht: (2023) -
Unsupervised Summarization Re-ranking
von: Ravaut, Mathieu, et al.
Veröffentlicht: (2022)