Benchmarking Large Language Models on Controllable Generation under Diversified Instructions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yihan, Xu, Benfeng, Wang, Quan, Liu, Yi, Mao, Zhendong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding
von: Zhu, Chiwei, et al.
Veröffentlicht: (2025)
von: Zhu, Chiwei, et al.
Veröffentlicht: (2025)
An Index-based Approach for Efficient and Effective Web Content Extraction
von: Chen, Yihan, et al.
Veröffentlicht: (2025)
von: Chen, Yihan, et al.
Veröffentlicht: (2025)
Training LLM-Based Agents with Synthetic Self-Reflected Trajectories and Partial Masking
von: Chen, Yihan, et al.
Veröffentlicht: (2025)
von: Chen, Yihan, et al.
Veröffentlicht: (2025)
ExpertPrompting: Instructing Large Language Models to be Distinguished Experts
von: Xu, Benfeng, et al.
Veröffentlicht: (2023)
von: Xu, Benfeng, et al.
Veröffentlicht: (2023)
Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach
von: Li, Ruizhe, et al.
Veröffentlicht: (2025)
von: Li, Ruizhe, et al.
Veröffentlicht: (2025)
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
von: Du, Mingxuan, et al.
Veröffentlicht: (2025)
von: Du, Mingxuan, et al.
Veröffentlicht: (2025)
Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability
von: Zhu, Chiwei, et al.
Veröffentlicht: (2025)
von: Zhu, Chiwei, et al.
Veröffentlicht: (2025)
WildGraphBench: Benchmarking GraphRAG with Wild-Source Corpora
von: Wang, Pengyu, et al.
Veröffentlicht: (2026)
von: Wang, Pengyu, et al.
Veröffentlicht: (2026)
A-RAG: Scaling Agentic Retrieval-Augmented Generation via Hierarchical Retrieval Interfaces
von: Du, Mingxuan, et al.
Veröffentlicht: (2026)
von: Du, Mingxuan, et al.
Veröffentlicht: (2026)
Benchmarking and Improving Compositional Generalization of Multi-aspect Controllable Text Generation
von: Zhong, Tianqi, et al.
Veröffentlicht: (2024)
von: Zhong, Tianqi, et al.
Veröffentlicht: (2024)
MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
von: Guo, Zikang, et al.
Veröffentlicht: (2025)
von: Guo, Zikang, et al.
Veröffentlicht: (2025)
DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Report
von: Li, Ruizhe, et al.
Veröffentlicht: (2026)
von: Li, Ruizhe, et al.
Veröffentlicht: (2026)
Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization
von: Liu, Yixin, et al.
Veröffentlicht: (2023)
von: Liu, Yixin, et al.
Veröffentlicht: (2023)
FS-Researcher: Test-Time Scaling for Long-Horizon Research Tasks with File-System-Based Agents
von: Zhu, Chiwei, et al.
Veröffentlicht: (2026)
von: Zhu, Chiwei, et al.
Veröffentlicht: (2026)
Mitigating Biases in Language Models via Bias Unlearning
von: Liu, Dianqing, et al.
Veröffentlicht: (2025)
von: Liu, Dianqing, et al.
Veröffentlicht: (2025)
Instruction Mining: Instruction Data Selection for Tuning Large Language Models
von: Cao, Yihan, et al.
Veröffentlicht: (2023)
von: Cao, Yihan, et al.
Veröffentlicht: (2023)
Wiki Live Challenge: Challenging Deep Research Agents with Expert-Level Wikipedia Articles
von: Wang, Shaohan, et al.
Veröffentlicht: (2026)
von: Wang, Shaohan, et al.
Veröffentlicht: (2026)
Leveraging Importance Sampling to Detach Alignment Modules from Large Language Models
von: Liu, Yi, et al.
Veröffentlicht: (2025)
von: Liu, Yi, et al.
Veröffentlicht: (2025)
TypedThinker: Diversify Large Language Model Reasoning with Typed Thinking
von: Wang, Danqing, et al.
Veröffentlicht: (2024)
von: Wang, Danqing, et al.
Veröffentlicht: (2024)
AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios
von: Qi, Yunjia, et al.
Veröffentlicht: (2025)
von: Qi, Yunjia, et al.
Veröffentlicht: (2025)
FlipGuard: Defending Preference Alignment against Update Regression with Constrained Optimization
von: Zhu, Mingye, et al.
Veröffentlicht: (2024)
von: Zhu, Mingye, et al.
Veröffentlicht: (2024)
TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis
von: Wu, Xiaorui, et al.
Veröffentlicht: (2025)
von: Wu, Xiaorui, et al.
Veröffentlicht: (2025)
OMGEval: An Open Multilingual Generative Evaluation Benchmark for Large Language Models
von: Liu, Yang, et al.
Veröffentlicht: (2024)
von: Liu, Yang, et al.
Veröffentlicht: (2024)
Benchmarking Political Persuasion Risks Across Frontier Large Language Models
von: Chen, Zhongren, et al.
Veröffentlicht: (2026)
von: Chen, Zhongren, et al.
Veröffentlicht: (2026)
Align Documents to Questions: Question-Oriented Document Rewriting for Retrieval-Augmented Generation
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
Leveraging Robust Optimization for LLM Alignment under Distribution Shifts
von: Zhu, Mingye, et al.
Veröffentlicht: (2025)
von: Zhu, Mingye, et al.
Veröffentlicht: (2025)
Resilience of Large Language Models for Noisy Instructions
von: Wang, Bin, et al.
Veröffentlicht: (2024)
von: Wang, Bin, et al.
Veröffentlicht: (2024)
Enhancing Large Language Models Against Inductive Instructions with Dual-critique Prompting
von: Wang, Rui, et al.
Veröffentlicht: (2023)
von: Wang, Rui, et al.
Veröffentlicht: (2023)
SimpleStrat: Diversifying Language Model Generation with Stratification
von: Wong, Justin, et al.
Veröffentlicht: (2024)
von: Wong, Justin, et al.
Veröffentlicht: (2024)
EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language Models
von: Zou, Tao, et al.
Veröffentlicht: (2025)
von: Zou, Tao, et al.
Veröffentlicht: (2025)
PLANNER: Generating Diversified Paragraph via Latent Language Diffusion Model
von: Zhang, Yizhe, et al.
Veröffentlicht: (2023)
von: Zhang, Yizhe, et al.
Veröffentlicht: (2023)
EconLogicQA: A Question-Answering Benchmark for Evaluating Large Language Models in Economic Sequential Reasoning
von: Quan, Yinzhu, et al.
Veröffentlicht: (2024)
von: Quan, Yinzhu, et al.
Veröffentlicht: (2024)
The SIFo Benchmark: Investigating the Sequential Instruction Following Ability of Large Language Models
von: Chen, Xinyi, et al.
Veröffentlicht: (2024)
von: Chen, Xinyi, et al.
Veröffentlicht: (2024)
LexGenius: An Expert-Level Benchmark for Large Language Models in Legal General Intelligence
von: Liu, Wenjin, et al.
Veröffentlicht: (2025)
von: Liu, Wenjin, et al.
Veröffentlicht: (2025)
MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models
von: Yan, Qiao, et al.
Veröffentlicht: (2025)
von: Yan, Qiao, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models
von: Zhang, Huatian, et al.
Veröffentlicht: (2026)
von: Zhang, Huatian, et al.
Veröffentlicht: (2026)
KCS: Diversify Multi-hop Question Generation with Knowledge Composition Sampling
von: Wang, Yangfan, et al.
Veröffentlicht: (2025)
von: Wang, Yangfan, et al.
Veröffentlicht: (2025)
Improved Unbiased Watermark for Large Language Models
von: Chen, Ruibo, et al.
Veröffentlicht: (2025)
von: Chen, Ruibo, et al.
Veröffentlicht: (2025)
IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models
von: Gao, Yiming, et al.
Veröffentlicht: (2025)
von: Gao, Yiming, et al.
Veröffentlicht: (2025)
PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script Generation
von: Xu, Chenning, et al.
Veröffentlicht: (2026)
von: Xu, Chenning, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding
von: Zhu, Chiwei, et al.
Veröffentlicht: (2025) -
An Index-based Approach for Efficient and Effective Web Content Extraction
von: Chen, Yihan, et al.
Veröffentlicht: (2025) -
Training LLM-Based Agents with Synthetic Self-Reflected Trajectories and Partial Masking
von: Chen, Yihan, et al.
Veröffentlicht: (2025) -
ExpertPrompting: Instructing Large Language Models to be Distinguished Experts
von: Xu, Benfeng, et al.
Veröffentlicht: (2023) -
Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach
von: Li, Ruizhe, et al.
Veröffentlicht: (2025)