Salvato in:
| Autori principali: | Gao, Lirong, Wang, Zeqing, Cai, Yuyan, Deng, Jiayi, Gu, Yanmei, Zhang, Yiming, Zhou, Jia, Zhang, Yanfei, Zhao, Junbo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2604.24690 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DORY: Deliberative Prompt Recovery for LLM
di: Gao, Lirong, et al.
Pubblicazione: (2024)
di: Gao, Lirong, et al.
Pubblicazione: (2024)
Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
di: Dong, Peijie, et al.
Pubblicazione: (2025)
di: Dong, Peijie, et al.
Pubblicazione: (2025)
Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
di: Zhou, Ziya, et al.
Pubblicazione: (2024)
di: Zhou, Ziya, et al.
Pubblicazione: (2024)
FLoE: Fisher-Based Layer Selection for Efficient Sparse Adaptation of Low-Rank Experts
di: Wang, Xinyi, et al.
Pubblicazione: (2025)
di: Wang, Xinyi, et al.
Pubblicazione: (2025)
ABench-Physics: Benchmarking Physical Reasoning in LLMs via High-Difficulty and Dynamic Physics Problems
di: Zhang, Yiming, et al.
Pubblicazione: (2025)
di: Zhang, Yiming, et al.
Pubblicazione: (2025)
3D-PreMise: Can Large Language Models Generate 3D Shapes with Sharp Features and Parametric Control?
di: Yuan, Zeqing, et al.
Pubblicazione: (2024)
di: Yuan, Zeqing, et al.
Pubblicazione: (2024)
A Thorough Examination of Decoding Methods in the Era of LLMs
di: Shi, Chufan, et al.
Pubblicazione: (2024)
di: Shi, Chufan, et al.
Pubblicazione: (2024)
Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?
di: Li, Xiang, et al.
Pubblicazione: (2025)
di: Li, Xiang, et al.
Pubblicazione: (2025)
Evaluating LLMs on Chinese Idiom Translation
di: Yang, Cai, et al.
Pubblicazione: (2025)
di: Yang, Cai, et al.
Pubblicazione: (2025)
Exploring the Capability Boundaries of LLMs in Mastering of Chinese Chouxiang Language
di: Lin, Dianqing, et al.
Pubblicazione: (2026)
di: Lin, Dianqing, et al.
Pubblicazione: (2026)
PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research
di: Miao, Tingjia, et al.
Pubblicazione: (2026)
di: Miao, Tingjia, et al.
Pubblicazione: (2026)
Collaborative QA using Interacting LLMs. Impact of Network Structure, Node Capability and Distributed Data
di: Jain, Adit, et al.
Pubblicazione: (2025)
di: Jain, Adit, et al.
Pubblicazione: (2025)
When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition
di: Li, Karen Jia-Hui, et al.
Pubblicazione: (2025)
di: Li, Karen Jia-Hui, et al.
Pubblicazione: (2025)
How Well Can Modern LLMs Act as Agent Cores in Radiology Environments?
di: Zheng, Qiaoyu, et al.
Pubblicazione: (2024)
di: Zheng, Qiaoyu, et al.
Pubblicazione: (2024)
Are Your LLMs Capable of Stable Reasoning?
di: Liu, Junnan, et al.
Pubblicazione: (2024)
di: Liu, Junnan, et al.
Pubblicazione: (2024)
Enough Coin Flips Can Make LLMs Act Bayesian
di: Gupta, Ritwik, et al.
Pubblicazione: (2025)
di: Gupta, Ritwik, et al.
Pubblicazione: (2025)
Towards Automatic Evaluation for LLMs' Clinical Capabilities: Metric, Data, and Algorithm
di: Liu, Lei, et al.
Pubblicazione: (2024)
di: Liu, Lei, et al.
Pubblicazione: (2024)
LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
di: Li, Chloe, et al.
Pubblicazione: (2025)
di: Li, Chloe, et al.
Pubblicazione: (2025)
Can LLMs Classify CVEs? Investigating LLMs Capabilities in Computing CVSS Vectors
di: Marchiori, Francesco, et al.
Pubblicazione: (2025)
di: Marchiori, Francesco, et al.
Pubblicazione: (2025)
Typestate via Revocable Capabilities
di: Jia, Songlin, et al.
Pubblicazione: (2025)
di: Jia, Songlin, et al.
Pubblicazione: (2025)
OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety
di: Liu, Chuang, et al.
Pubblicazione: (2024)
di: Liu, Chuang, et al.
Pubblicazione: (2024)
Can LLMs be Fooled? Investigating Vulnerabilities in LLMs
di: Abdali, Sara, et al.
Pubblicazione: (2024)
di: Abdali, Sara, et al.
Pubblicazione: (2024)
Evaluating Developmental Cognition Capabilities of LLMs
di: Xiao, Xiao, et al.
Pubblicazione: (2026)
di: Xiao, Xiao, et al.
Pubblicazione: (2026)
Olapa-MCoT: Enhancing the Chinese Mathematical Reasoning Capability of LLMs
di: Zhu, Shaojie, et al.
Pubblicazione: (2023)
di: Zhu, Shaojie, et al.
Pubblicazione: (2023)
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities
di: Li, Haoming, et al.
Pubblicazione: (2025)
di: Li, Haoming, et al.
Pubblicazione: (2025)
D.Va: Validate Your Demonstration First Before You Use It
di: Zhang, Qi, et al.
Pubblicazione: (2025)
di: Zhang, Qi, et al.
Pubblicazione: (2025)
Satisfiability Solving with LLMs: A Matched-Pair Evaluation of Reasoning Capability
di: Zhang, Leizhen, et al.
Pubblicazione: (2026)
di: Zhang, Leizhen, et al.
Pubblicazione: (2026)
An Empirical Study on the Capability of LLMs in Decomposing Bug Reports
di: Chen, Zhiyuan, et al.
Pubblicazione: (2025)
di: Chen, Zhiyuan, et al.
Pubblicazione: (2025)
Learners as Historians: Making History Come Alive through Historical Inquiry
di: Pappas, Marjorie L.
Pubblicazione: (2007)
di: Pappas, Marjorie L.
Pubblicazione: (2007)
Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
di: Chen, Zejian, et al.
Pubblicazione: (2026)
di: Chen, Zejian, et al.
Pubblicazione: (2026)
Can Editing LLMs Inject Harm?
di: Chen, Canyu, et al.
Pubblicazione: (2024)
di: Chen, Canyu, et al.
Pubblicazione: (2024)
Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
di: Tie, Guiyao, et al.
Pubblicazione: (2025)
di: Tie, Guiyao, et al.
Pubblicazione: (2025)
CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs
di: Yao, Zhiyuan, et al.
Pubblicazione: (2026)
di: Yao, Zhiyuan, et al.
Pubblicazione: (2026)
Can Prompts Rewind Time for LLMs? Evaluating the Effectiveness of Prompted Knowledge Cutoffs
di: Gao, Xin, et al.
Pubblicazione: (2025)
di: Gao, Xin, et al.
Pubblicazione: (2025)
SCAN: Structured Capability Assessment and Navigation for LLMs
di: Wang, Zongqi, et al.
Pubblicazione: (2025)
di: Wang, Zongqi, et al.
Pubblicazione: (2025)
Are LLMs Effective Negotiators? Systematic Evaluation of the Multifaceted Capabilities of LLMs in Negotiation Dialogues
di: Kwon, Deuksin, et al.
Pubblicazione: (2024)
di: Kwon, Deuksin, et al.
Pubblicazione: (2024)
LLMAID: Identifying AI Capabilities in Android Apps with LLMs
di: Liu, Pei, et al.
Pubblicazione: (2025)
di: Liu, Pei, et al.
Pubblicazione: (2025)
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
di: Yang, Jialin, et al.
Pubblicazione: (2025)
di: Yang, Jialin, et al.
Pubblicazione: (2025)
Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning
di: Wu, Yuyang, et al.
Pubblicazione: (2026)
di: Wu, Yuyang, et al.
Pubblicazione: (2026)
The College Archivist as College Historian: Baruch College Celebrates Its Historical Roots
di: Roff, Sandra
Pubblicazione: (2010)
di: Roff, Sandra
Pubblicazione: (2010)
Documenti analoghi
-
DORY: Deliberative Prompt Recovery for LLM
di: Gao, Lirong, et al.
Pubblicazione: (2024) -
Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
di: Dong, Peijie, et al.
Pubblicazione: (2025) -
Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
di: Zhou, Ziya, et al.
Pubblicazione: (2024) -
FLoE: Fisher-Based Layer Selection for Efficient Sparse Adaptation of Low-Rank Experts
di: Wang, Xinyi, et al.
Pubblicazione: (2025) -
ABench-Physics: Benchmarking Physical Reasoning in LLMs via High-Difficulty and Dynamic Physics Problems
di: Zhang, Yiming, et al.
Pubblicazione: (2025)