BenTo: Benchmark Task Reduction with In-Context Transferability
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Hongyu, Li, Ming, Sun, Lichao, Zhou, Tianyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
SpecHub: Provable Acceleration to Multi-Draft Speculative Decoding
von: Sun, Ryan, et al.
Veröffentlicht: (2024)
von: Sun, Ryan, et al.
Veröffentlicht: (2024)
TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning
von: Yu, Fangxu, et al.
Veröffentlicht: (2025)
von: Yu, Fangxu, et al.
Veröffentlicht: (2025)
GR-Ben: A General Reasoning Benchmark for Evaluating Process Reward Models
von: Sun, Zhouhao, et al.
Veröffentlicht: (2026)
von: Sun, Zhouhao, et al.
Veröffentlicht: (2026)
In-Context Transfer Learning: Demonstration Synthesis by Transferring Similar Tasks
von: Wang, Dingzirui, et al.
Veröffentlicht: (2024)
von: Wang, Dingzirui, et al.
Veröffentlicht: (2024)
1+1>2: Can Large Language Models Serve as Cross-Lingual Knowledge Aggregators?
von: Huang, Yue, et al.
Veröffentlicht: (2024)
von: Huang, Yue, et al.
Veröffentlicht: (2024)
MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application
von: Peng, Xueqing, et al.
Veröffentlicht: (2025)
von: Peng, Xueqing, et al.
Veröffentlicht: (2025)
Neuro-Symbolic Synergy for Interactive World Modeling
von: Zhao, Hongyu, et al.
Veröffentlicht: (2026)
von: Zhao, Hongyu, et al.
Veröffentlicht: (2026)
Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
Submodular Context Partitioning and Compression for In-Context Learning
von: Zheng, Shaoyi, et al.
Veröffentlicht: (2025)
von: Zheng, Shaoyi, et al.
Veröffentlicht: (2025)
CodeIP: A Grammar-Guided Multi-Bit Watermark for Large Language Models of Code
von: Guan, Batu, et al.
Veröffentlicht: (2024)
von: Guan, Batu, et al.
Veröffentlicht: (2024)
Mosaic-IT: Cost-Free Compositional Data Synthesis for Instruction Tuning
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
FakeGPT: Fake News Generation, Explanation and Detection of Large Language Models
von: Huang, Yue, et al.
Veröffentlicht: (2023)
von: Huang, Yue, et al.
Veröffentlicht: (2023)
MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs
von: Zeng, Zhongshen, et al.
Veröffentlicht: (2024)
von: Zeng, Zhongshen, et al.
Veröffentlicht: (2024)
I Think, Therefore I am: Benchmarking Awareness of Large Language Models Using AwareBench
von: Li, Yuan, et al.
Veröffentlicht: (2024)
von: Li, Yuan, et al.
Veröffentlicht: (2024)
FinBen: A Holistic Financial Benchmark for Large Language Models
von: Xie, Qianqian, et al.
Veröffentlicht: (2024)
von: Xie, Qianqian, et al.
Veröffentlicht: (2024)
Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook
von: Li, Ming, et al.
Veröffentlicht: (2026)
von: Li, Ming, et al.
Veröffentlicht: (2026)
When AI Navigates the Fog of War
von: Li, Ming, et al.
Veröffentlicht: (2026)
von: Li, Ming, et al.
Veröffentlicht: (2026)
Where to show Demos in Your Prompt: A Positional Bias of In-Context Learning
von: Cobbina, Kwesi, et al.
Veröffentlicht: (2025)
von: Cobbina, Kwesi, et al.
Veröffentlicht: (2025)
Label Words as Local Task Vectors in In-Context Learning
von: Zheng, Bowen, et al.
Veröffentlicht: (2024)
von: Zheng, Bowen, et al.
Veröffentlicht: (2024)
CrossICL: Cross-Task In-Context Learning via Unsupervised Demonstration Transfer
von: Gao, Jinglong, et al.
Veröffentlicht: (2025)
von: Gao, Jinglong, et al.
Veröffentlicht: (2025)
BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law
von: Nagl, Sebastian, et al.
Veröffentlicht: (2026)
von: Nagl, Sebastian, et al.
Veröffentlicht: (2026)
TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice
von: Hu, Gang, et al.
Veröffentlicht: (2026)
von: Hu, Gang, et al.
Veröffentlicht: (2026)
PharmaShip: An Entity-Centric, Reading-Order-Supervised Benchmark for Chinese Pharmaceutical Shipping Documents
von: Xie, Tingwei, et al.
Veröffentlicht: (2025)
von: Xie, Tingwei, et al.
Veröffentlicht: (2025)
RMTBench: Benchmarking LLMs Through Multi-Turn User-Centric Role-Playing
von: Xiang, Hao, et al.
Veröffentlicht: (2025)
von: Xiang, Hao, et al.
Veröffentlicht: (2025)
LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contexts
von: Chen, Junhao, et al.
Veröffentlicht: (2025)
von: Chen, Junhao, et al.
Veröffentlicht: (2025)
TaskComplexity: A Dataset for Task Complexity Classification with In-Context Learning, FLAN-T5 and GPT-4o Benchmarks
von: Rasheed, Areeg Fahad, et al.
Veröffentlicht: (2024)
von: Rasheed, Areeg Fahad, et al.
Veröffentlicht: (2024)
What Happened in LLMs Layers when Trained for Fast vs. Slow Thinking: A Gradient Perspective
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
von: Adib, Shefayat E Shams, et al.
Veröffentlicht: (2026)
von: Adib, Shefayat E Shams, et al.
Veröffentlicht: (2026)
Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
von: Tie, Guiyao, et al.
Veröffentlicht: (2025)
von: Tie, Guiyao, et al.
Veröffentlicht: (2025)
Can Large Language Models Automatically Jailbreak GPT-4V?
von: Wu, Yuanwei, et al.
Veröffentlicht: (2024)
von: Wu, Yuanwei, et al.
Veröffentlicht: (2024)
M4LE: A Multi-Ability Multi-Range Multi-Task Multi-Domain Long-Context Evaluation Benchmark for Large Language Models
von: Kwan, Wai-Chung, et al.
Veröffentlicht: (2023)
von: Kwan, Wai-Chung, et al.
Veröffentlicht: (2023)
EPPCMinerBen: A Novel Benchmark for Evaluating Large Language Models on Electronic Patient-Provider Communication via the Patient Portal
von: Fodeh, Samah, et al.
Veröffentlicht: (2026)
von: Fodeh, Samah, et al.
Veröffentlicht: (2026)
CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era
von: Shi, Kaiwen, et al.
Veröffentlicht: (2026)
von: Shi, Kaiwen, et al.
Veröffentlicht: (2026)
MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models
von: Zhou, Keyan, et al.
Veröffentlicht: (2025)
von: Zhou, Keyan, et al.
Veröffentlicht: (2025)
CL-bench: A Benchmark for Context Learning
von: Dou, Shihan, et al.
Veröffentlicht: (2026)
von: Dou, Shihan, et al.
Veröffentlicht: (2026)
CCF: A Context Compression Framework for Efficient Long-Sequence Language Modeling
von: Li, Wenhao, et al.
Veröffentlicht: (2025)
von: Li, Wenhao, et al.
Veröffentlicht: (2025)
Open Grounded Planning: Challenges and Benchmark Construction
von: Guo, Shiguang, et al.
Veröffentlicht: (2024)
von: Guo, Shiguang, et al.
Veröffentlicht: (2024)
Does the Generator Mind its Contexts? An Analysis of Generative Model Faithfulness under Context Transfer
von: Hu, Xinshuo, et al.
Veröffentlicht: (2024)
von: Hu, Xinshuo, et al.
Veröffentlicht: (2024)
LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
von: Wu, Yuhao, et al.
Veröffentlicht: (2024)
von: Wu, Yuhao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?
von: Fan, Chenrui, et al.
Veröffentlicht: (2025) -
SpecHub: Provable Acceleration to Multi-Draft Speculative Decoding
von: Sun, Ryan, et al.
Veröffentlicht: (2024) -
TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning
von: Yu, Fangxu, et al.
Veröffentlicht: (2025) -
GR-Ben: A General Reasoning Benchmark for Evaluating Process Reward Models
von: Sun, Zhouhao, et al.
Veröffentlicht: (2026) -
In-Context Transfer Learning: Demonstration Synthesis by Transferring Similar Tasks
von: Wang, Dingzirui, et al.
Veröffentlicht: (2024)