Let LLMs Take on the Latest Challenges! A Chinese Dynamic Question Answering Benchmark
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Zhikun, Li, Yinghui, Ding, Ruixue, Wang, Xinyu, Chen, Boli, Jiang, Yong, Zheng, Hai-Tao, Lu, Wenlian, Xie, Pengjun, Huang, Fei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ProductAgent: Benchmarking Conversational Product Search Agent with Asking Clarification Questions
di: Ye, Jingheng, et al.
Pubblicazione: (2024)
di: Ye, Jingheng, et al.
Pubblicazione: (2024)
Geo-Encoder: A Chunk-Argument Bi-Encoder Framework for Chinese Geographic Re-Ranking
di: Cao, Yong, et al.
Pubblicazione: (2023)
di: Cao, Yong, et al.
Pubblicazione: (2023)
Improving Retrieval Augmented Open-Domain Question-Answering with Vectorized Contexts
di: Chen, Zhuo, et al.
Pubblicazione: (2024)
di: Chen, Zhuo, et al.
Pubblicazione: (2024)
Efficient Multimodal Planning Agent for Visual Question-Answering
di: Chen, Zhuo, et al.
Pubblicazione: (2026)
di: Chen, Zhuo, et al.
Pubblicazione: (2026)
Unfolding the Headline: Iterative Self-Questioning for News Retrieval and Timeline Summarization
di: Wu, Weiqi, et al.
Pubblicazione: (2025)
di: Wu, Weiqi, et al.
Pubblicazione: (2025)
Bidirectional End-to-End Learning of Retriever-Reader Paradigm for Entity Linking
di: Li, Yinghui, et al.
Pubblicazione: (2023)
di: Li, Yinghui, et al.
Pubblicazione: (2023)
Accurate Table Question Answering with Accessible LLMs
di: Jiang, Yangfan, et al.
Pubblicazione: (2026)
di: Jiang, Yangfan, et al.
Pubblicazione: (2026)
Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent
di: Li, Yangning, et al.
Pubblicazione: (2024)
di: Li, Yangning, et al.
Pubblicazione: (2024)
RaFe: Ranking Feedback Improves Query Rewriting for RAG
di: Mao, Shengyu, et al.
Pubblicazione: (2024)
di: Mao, Shengyu, et al.
Pubblicazione: (2024)
InsQABench: Benchmarking Chinese Insurance Domain Question Answering with Large Language Models
di: Ding, Jing, et al.
Pubblicazione: (2025)
di: Ding, Jing, et al.
Pubblicazione: (2025)
LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs -- No Silver Bullet for LC or RAG Routing
di: Li, Kuan, et al.
Pubblicazione: (2025)
di: Li, Kuan, et al.
Pubblicazione: (2025)
ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
di: Nie, Yuxiang, et al.
Pubblicazione: (2025)
di: Nie, Yuxiang, et al.
Pubblicazione: (2025)
One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMs
di: Li, Yinghui, et al.
Pubblicazione: (2025)
di: Li, Yinghui, et al.
Pubblicazione: (2025)
Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan
di: Yao, Jui-Ming, et al.
Pubblicazione: (2025)
di: Yao, Jui-Ming, et al.
Pubblicazione: (2025)
CoFE-RAG: A Comprehensive Full-chain Evaluation Framework for Retrieval-Augmented Generation with Enhanced Data Diversity
di: Liu, Jintao, et al.
Pubblicazione: (2024)
di: Liu, Jintao, et al.
Pubblicazione: (2024)
TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese Medicine
di: Xie, Jiacheng, et al.
Pubblicazione: (2025)
di: Xie, Jiacheng, et al.
Pubblicazione: (2025)
WebWalker: Benchmarking LLMs in Web Traversal
di: Wu, Jialong, et al.
Pubblicazione: (2025)
di: Wu, Jialong, et al.
Pubblicazione: (2025)
ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents
di: Wang, Qiuchen, et al.
Pubblicazione: (2025)
di: Wang, Qiuchen, et al.
Pubblicazione: (2025)
On the (In)Effectiveness of Large Language Models for Chinese Text Correction
di: Li, Yinghui, et al.
Pubblicazione: (2023)
di: Li, Yinghui, et al.
Pubblicazione: (2023)
Combining Knowledge Graph and LLMs for Enhanced Zero-shot Visual Question Answering
di: Tao, Qian, et al.
Pubblicazione: (2025)
di: Tao, Qian, et al.
Pubblicazione: (2025)
Detecting Knowledge Boundary of Vision Large Language Models by Sampling-Based Inference
di: Chen, Zhuo, et al.
Pubblicazione: (2025)
di: Chen, Zhuo, et al.
Pubblicazione: (2025)
WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent
di: Geng, Xinyu, et al.
Pubblicazione: (2025)
di: Geng, Xinyu, et al.
Pubblicazione: (2025)
Letting Teens Take the Lead.
di: Braun, Linda W.
Pubblicazione: (2001)
di: Braun, Linda W.
Pubblicazione: (2001)
LatEval: An Interactive LLMs Evaluation Benchmark with Incomplete Information from Lateral Thinking Puzzles
di: Huang, Shulin, et al.
Pubblicazione: (2023)
di: Huang, Shulin, et al.
Pubblicazione: (2023)
EXCGEC: A Benchmark for Edit-Wise Explainable Chinese Grammatical Error Correction
di: Ye, Jingheng, et al.
Pubblicazione: (2024)
di: Ye, Jingheng, et al.
Pubblicazione: (2024)
Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions
di: Chen, Hanjie, et al.
Pubblicazione: (2024)
di: Chen, Hanjie, et al.
Pubblicazione: (2024)
Beyond Code Snippets: Benchmarking LLMs on Repository-Level Question Answering
di: Alebachew, Yoseph Berhanu, et al.
Pubblicazione: (2026)
di: Alebachew, Yoseph Berhanu, et al.
Pubblicazione: (2026)
LLMs Meet Long Video: Advancing Long Video Question Answering with An Interactive Visual Adapter in LLMs
di: Li, Yunxin, et al.
Pubblicazione: (2024)
di: Li, Yunxin, et al.
Pubblicazione: (2024)
Traceable Cross-Source RAG for Chinese Tibetan Medicine Question Answering
di: Chen, Fengxian, et al.
Pubblicazione: (2026)
di: Chen, Fengxian, et al.
Pubblicazione: (2026)
On Time-Varying Delayed Stochastic Differential Systems with Non-Markovian Switching Parameters
di: Wu, Xinyu, et al.
Pubblicazione: (2024)
di: Wu, Xinyu, et al.
Pubblicazione: (2024)
Benchmarking Agentic Workflow Generation
di: Qiao, Shuofei, et al.
Pubblicazione: (2024)
di: Qiao, Shuofei, et al.
Pubblicazione: (2024)
CL$^2$GEC: A Multi-Discipline Benchmark for Continual Learning in Chinese Literature Grammatical Error Correction
di: Qin, Shang, et al.
Pubblicazione: (2025)
di: Qin, Shang, et al.
Pubblicazione: (2025)
ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering
di: Masry, Ahmed, et al.
Pubblicazione: (2025)
di: Masry, Ahmed, et al.
Pubblicazione: (2025)
DEXTER: A Benchmark for open-domain Complex Question Answering using LLMs
di: Prabhu, Venktesh V. Deepali, et al.
Pubblicazione: (2024)
di: Prabhu, Venktesh V. Deepali, et al.
Pubblicazione: (2024)
How Good LLMs Are at Answering Bangla Medical Visual Questions? Dataset and Benchmarking
di: Ahmed, Rafid, et al.
Pubblicazione: (2026)
di: Ahmed, Rafid, et al.
Pubblicazione: (2026)
When LLMs Meet Cunning Texts: A Fallacy Understanding Benchmark for Large Language Models
di: Li, Yinghui, et al.
Pubblicazione: (2024)
di: Li, Yinghui, et al.
Pubblicazione: (2024)
Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
di: Dong, Kuicai, et al.
Pubblicazione: (2025)
di: Dong, Kuicai, et al.
Pubblicazione: (2025)
Let the Agent Search: Autonomous Exploration Beats Rigid Workflows in Temporal Question Answering
di: Lv, Xufei, et al.
Pubblicazione: (2026)
di: Lv, Xufei, et al.
Pubblicazione: (2026)
Coal Mining Question Answering with LLMs
di: Rivera, Antonio Carlos, et al.
Pubblicazione: (2024)
di: Rivera, Antonio Carlos, et al.
Pubblicazione: (2024)
On the Calibration of Multilingual Question Answering LLMs
di: Yang, Yahan, et al.
Pubblicazione: (2023)
di: Yang, Yahan, et al.
Pubblicazione: (2023)
Documenti analoghi
-
ProductAgent: Benchmarking Conversational Product Search Agent with Asking Clarification Questions
di: Ye, Jingheng, et al.
Pubblicazione: (2024) -
Geo-Encoder: A Chunk-Argument Bi-Encoder Framework for Chinese Geographic Re-Ranking
di: Cao, Yong, et al.
Pubblicazione: (2023) -
Improving Retrieval Augmented Open-Domain Question-Answering with Vectorized Contexts
di: Chen, Zhuo, et al.
Pubblicazione: (2024) -
Efficient Multimodal Planning Agent for Visual Question-Answering
di: Chen, Zhuo, et al.
Pubblicazione: (2026) -
Unfolding the Headline: Iterative Self-Questioning for News Retrieval and Timeline Summarization
di: Wu, Weiqi, et al.
Pubblicazione: (2025)