DiNeR: a Large Realistic Dataset for Evaluating Compositional Generalization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Chengang, Liu, Xiao, Feng, Yansong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Teaching Large Language Models an Unseen Language on the Fly
von: Zhang, Chen, et al.
Veröffentlicht: (2024)
von: Zhang, Chen, et al.
Veröffentlicht: (2024)
DiMA: An LLM-Powered Ride-Hailing Assistant at DiDi
von: Ning, Yansong, et al.
Veröffentlicht: (2025)
von: Ning, Yansong, et al.
Veröffentlicht: (2025)
Realistic Evaluation of Model Merging for Compositional Generalization
von: Tam, Derek, et al.
Veröffentlicht: (2024)
von: Tam, Derek, et al.
Veröffentlicht: (2024)
Haste Makes Waste: Evaluating Planning Abilities of LLMs for Efficient and Feasible Multitasking with Time Constraints Between Actions
von: Wu, Zirui, et al.
Veröffentlicht: (2025)
von: Wu, Zirui, et al.
Veröffentlicht: (2025)
CASA: Causality-driven Argument Sufficiency Assessment
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
Automating Legal Interpretation with LLMs: Retrieval, Generation, and Evaluation
von: Luo, Kangcheng, et al.
Veröffentlicht: (2025)
von: Luo, Kangcheng, et al.
Veröffentlicht: (2025)
ELLA: Empowering LLMs for Interpretable, Accurate and Informative Legal Advice
von: Hu, Yutong, et al.
Veröffentlicht: (2024)
von: Hu, Yutong, et al.
Veröffentlicht: (2024)
Only One Relation Possible? Modeling the Ambiguity in Event Temporal Relation Extraction
von: Hu, Yutong, et al.
Veröffentlicht: (2024)
von: Hu, Yutong, et al.
Veröffentlicht: (2024)
Can Perplexity Reflect Large Language Model's Ability in Long Text Understanding?
von: Hu, Yutong, et al.
Veröffentlicht: (2024)
von: Hu, Yutong, et al.
Veröffentlicht: (2024)
ScIRGen: Synthesize Realistic and Large-Scale RAG Dataset for Scientific Research
von: Lin, Junyong, et al.
Veröffentlicht: (2025)
von: Lin, Junyong, et al.
Veröffentlicht: (2025)
Realistic Evaluation of Toxicity in Large Language Models
von: Luong, Tinh Son, et al.
Veröffentlicht: (2024)
von: Luong, Tinh Son, et al.
Veröffentlicht: (2024)
Motion Generation from Fine-grained Textual Descriptions
von: Li, Kunhang, et al.
Veröffentlicht: (2024)
von: Li, Kunhang, et al.
Veröffentlicht: (2024)
ProTrix: Building Models for Planning and Reasoning over Tables with Sentence Context
von: Wu, Zirui, et al.
Veröffentlicht: (2024)
von: Wu, Zirui, et al.
Veröffentlicht: (2024)
RefTool: Reference-Guided Tool Creation for Knowledge-Intensive Reasoning
von: Liu, Xiao, et al.
Veröffentlicht: (2025)
von: Liu, Xiao, et al.
Veröffentlicht: (2025)
The Chronicles of RiDiC: Generating Datasets with Controlled Popularity Distribution for Long-form Factuality Evaluation
von: Braslavski, Pavel, et al.
Veröffentlicht: (2026)
von: Braslavski, Pavel, et al.
Veröffentlicht: (2026)
DiPaCo: Distributed Path Composition
von: Douillard, Arthur, et al.
Veröffentlicht: (2024)
von: Douillard, Arthur, et al.
Veröffentlicht: (2024)
Read it in Two Steps: Translating Extremely Low-Resource Languages with Code-Augmented Grammar Books
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
NeSyCoCo: A Neuro-Symbolic Concept Composer for Compositional Generalization
von: Kamali, Danial, et al.
Veröffentlicht: (2024)
von: Kamali, Danial, et al.
Veröffentlicht: (2024)
What Kinds of Tokens Benefit from Distant Text? An Analysis on Long Context Language Modeling
von: Hu, Yutong, et al.
Veröffentlicht: (2024)
von: Hu, Yutong, et al.
Veröffentlicht: (2024)
Automated Annotation of Evolving Corpora for Augmenting Longitudinal Network Data: A Framework Integrating Large Language Models and Expert Knowledge
von: Liu, Xiao, et al.
Veröffentlicht: (2025)
von: Liu, Xiao, et al.
Veröffentlicht: (2025)
Towards Realistic Few-Shot Relation Extraction: A New Meta Dataset and Evaluation
von: Alam, Fahmida, et al.
Veröffentlicht: (2024)
von: Alam, Fahmida, et al.
Veröffentlicht: (2024)
Enhancing Speech Large Language Models through Reinforced Behavior Alignment
von: Liu, Yansong, et al.
Veröffentlicht: (2025)
von: Liu, Yansong, et al.
Veröffentlicht: (2025)
Evaluating Morphological Compositional Generalization in Large Language Models
von: Ismayilzada, Mete, et al.
Veröffentlicht: (2024)
von: Ismayilzada, Mete, et al.
Veröffentlicht: (2024)
Chain of Condition: Construct, Verify and Solve Conditions for Conditional Question Answering
von: Lin, Jiuheng, et al.
Veröffentlicht: (2024)
von: Lin, Jiuheng, et al.
Veröffentlicht: (2024)
D$^2$Plan: Dual-Agent Dynamic Global Planning for Complex Retrieval-Augmented Reasoning
von: Luo, Kangcheng, et al.
Veröffentlicht: (2026)
von: Luo, Kangcheng, et al.
Veröffentlicht: (2026)
Cross-Lingual Transfer of Cultural Knowledge: An Asymmetric Phenomenon
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
AutoHall: Automated Factuality Hallucination Dataset Generation for Large Language Models
von: Cao, Zouying, et al.
Veröffentlicht: (2023)
von: Cao, Zouying, et al.
Veröffentlicht: (2023)
Augmenting Research Ideation with Data: An Empirical Investigation in Social Science
von: Liu, Xiao, et al.
Veröffentlicht: (2025)
von: Liu, Xiao, et al.
Veröffentlicht: (2025)
WikiFactDiff: A Large, Realistic, and Temporally Adaptable Dataset for Atomic Factual Knowledge Update in Causal Language Models
von: Khodja, Hichem Ammar, et al.
Veröffentlicht: (2024)
von: Khodja, Hichem Ammar, et al.
Veröffentlicht: (2024)
DiLM: Distilling Dataset into Language Model for Text-level Dataset Distillation
von: Maekawa, Aru, et al.
Veröffentlicht: (2024)
von: Maekawa, Aru, et al.
Veröffentlicht: (2024)
EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generation
von: Wang, Xinda, et al.
Veröffentlicht: (2025)
von: Wang, Xinda, et al.
Veröffentlicht: (2025)
JuniperLiu at CoMeDi Shared Task: Models as Annotators in Lexical Semantics Disagreements
von: Liu, Zhu, et al.
Veröffentlicht: (2024)
von: Liu, Zhu, et al.
Veröffentlicht: (2024)
Synthetic Dataset for Evaluating Complex Compositional Knowledge for Natural Language Inference
von: Akoju, Sushma Anand, et al.
Veröffentlicht: (2023)
von: Akoju, Sushma Anand, et al.
Veröffentlicht: (2023)
On Leakage of Code Generation Evaluation Datasets
von: Matton, Alexandre, et al.
Veröffentlicht: (2024)
von: Matton, Alexandre, et al.
Veröffentlicht: (2024)
Subject-level Inference for Realistic Text Anonymization Evaluation
von: Oh, Myeong Seok, et al.
Veröffentlicht: (2026)
von: Oh, Myeong Seok, et al.
Veröffentlicht: (2026)
Evaluating Text Creativity across Diverse Domains: A Dataset and Large Language Model Evaluator
von: Cao, Qian, et al.
Veröffentlicht: (2025)
von: Cao, Qian, et al.
Veröffentlicht: (2025)
LFED: A Literary Fiction Evaluation Dataset for Large Language Models
von: Yu, Linhao, et al.
Veröffentlicht: (2024)
von: Yu, Linhao, et al.
Veröffentlicht: (2024)
RECKON: Large-scale Reference-based Efficient Knowledge Evaluation for Large Language Model
von: Zhang, Lin, et al.
Veröffentlicht: (2025)
von: Zhang, Lin, et al.
Veröffentlicht: (2025)
Respond Beyond Language: A Benchmark for Video Generation in Response to Realistic User Intents
von: Wang, Shuting, et al.
Veröffentlicht: (2025)
von: Wang, Shuting, et al.
Veröffentlicht: (2025)
Privacy-Preserving Parameter-Efficient Fine-Tuning for Large Language Model Services
von: Li, Yansong, et al.
Veröffentlicht: (2023)
von: Li, Yansong, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Teaching Large Language Models an Unseen Language on the Fly
von: Zhang, Chen, et al.
Veröffentlicht: (2024) -
DiMA: An LLM-Powered Ride-Hailing Assistant at DiDi
von: Ning, Yansong, et al.
Veröffentlicht: (2025) -
Realistic Evaluation of Model Merging for Compositional Generalization
von: Tam, Derek, et al.
Veröffentlicht: (2024) -
Haste Makes Waste: Evaluating Planning Abilities of LLMs for Efficient and Feasible Multitasking with Time Constraints Between Actions
von: Wu, Zirui, et al.
Veröffentlicht: (2025) -
CASA: Causality-driven Argument Sufficiency Assessment
von: Liu, Xiao, et al.
Veröffentlicht: (2024)