SciDA: Scientific Dynamic Assessor of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Junting, Miao, Tingjia, Liao, Yiyan, Wang, Qichao, Wen, Zhoufutu, Wang, Yanqin, Huang, Yunjie, Yan, Ge, Wang, Leqi, Xia, Yucheng, Gao, Hongwan, Zeng, Yuansong, Zheng, Renjie, Dun, Chen, Liang, Yitao, Yang, Tong, Huang, Wenhao, Zhang, Ge |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ScholarSearch: Benchmarking Scholar Searching Ability of LLMs
by: Zhou, Junting, et al.
Published: (2025)
by: Zhou, Junting, et al.
Published: (2025)
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
by: Wu, Siwei, et al.
Published: (2024)
by: Wu, Siwei, et al.
Published: (2024)
Assessor Personality Traits Are Not Educationally Important Drivers of Assessor Stringency/Leniency
by: Sebastian Dewhirst, et al.
Published: (2025)
by: Sebastian Dewhirst, et al.
Published: (2025)
Evaluating Library Assessors.
by: Credaro, Amanda
Published: (2003)
by: Credaro, Amanda
Published: (2003)
Time‐Resolved Small‐Angle X‐Ray Studies of Spherical Micelle Formation and Growth During Polymerization‐Induced Self‐Assembly in Polar Solvents
by: Zhiqing Mei, et al.
Published: (2025)
by: Zhiqing Mei, et al.
Published: (2025)
SciCapenter: Supporting Caption Composition for Scientific Figures with Machine-Generated Captions and Ratings
by: Hsu, Ting-Yao, et al.
Published: (2024)
by: Hsu, Ting-Yao, et al.
Published: (2024)
SciIF: Benchmarking Scientific Instruction Following Towards Rigorous Scientific Intelligence
by: Su, Encheng, et al.
Published: (2026)
by: Su, Encheng, et al.
Published: (2026)
AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMM
by: Wang, Jiarui, et al.
Published: (2024)
by: Wang, Jiarui, et al.
Published: (2024)
TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs
by: Wang, Juntong, et al.
Published: (2025)
by: Wang, Juntong, et al.
Published: (2025)
KOR-Bench: Benchmarking Language Models on Knowledge-Orthogonal Reasoning Tasks
by: Ma, Kaijing, et al.
Published: (2024)
by: Ma, Kaijing, et al.
Published: (2024)
Efficacy and Potential Mechanisms of Fire Needle Therapy for Moderate Acne: An Assessor‐Blinded, Randomized Controlled Trial
by: Ruie Wang, et al.
Published: (2026)
by: Ruie Wang, et al.
Published: (2026)
BABE: Biology Arena BEnchmark
by: Zhou, Junting, et al.
Published: (2026)
by: Zhou, Junting, et al.
Published: (2026)
Integrated lithium niobate microwave photonics: Driving next-generation wireless technologies
by: Feng, Hanke, et al.
Published: (2026)
by: Feng, Hanke, et al.
Published: (2026)
Beyond Single Prompts: Synergistic Fusion and Arrangement for VICL
by: Liao, Wenwen, et al.
Published: (2026)
by: Liao, Wenwen, et al.
Published: (2026)
Enhancing Visual In-Context Learning by Multi-Faceted Fusion
by: Liao, Wenwen, et al.
Published: (2026)
by: Liao, Wenwen, et al.
Published: (2026)
Medical Report Generation: A Hierarchical Task Structure-Based Cross-Modal Causal Intervention Framework
by: Song, Yucheng, et al.
Published: (2025)
by: Song, Yucheng, et al.
Published: (2025)
Trade-Offs Between Assessor Team Size and Assessor Expertise in Affecting Rating Accuracy in Assessment Centers
by: Andreja Wirz
Published: (2013)
by: Andreja Wirz
Published: (2013)
CoTJudger: A Graph-Driven Framework for Automatic Evaluation of Chain-of-Thought Efficiency and Redundancy in LRMs
by: Li, Siyi, et al.
Published: (2026)
by: Li, Siyi, et al.
Published: (2026)
SciEGQA: A Dataset for Scientific Evidence-Grounded Question Answering and Reasoning
by: Yu, Wenhan, et al.
Published: (2025)
by: Yu, Wenhan, et al.
Published: (2025)
UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow Models
by: Jiao, Guanlong, et al.
Published: (2025)
by: Jiao, Guanlong, et al.
Published: (2025)
Dark Photons in the Radio Sky: II. Resonant Conversions in the Intergalactic Medium
by: Baker, Ethan, et al.
Published: (2025)
by: Baker, Ethan, et al.
Published: (2025)
Dark Photons in the Radio Sky: I. Resonant Conversions in Halos
by: Baker, Ethan, et al.
Published: (2025)
by: Baker, Ethan, et al.
Published: (2025)
MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluation
by: Yang, Chenghao, et al.
Published: (2025)
by: Yang, Chenghao, et al.
Published: (2025)
SciMON: Scientific Inspiration Machines Optimized for Novelty
by: Wang, Qingyun, et al.
Published: (2023)
by: Wang, Qingyun, et al.
Published: (2023)
LLMs as Assessors: Right for the Right Reason?
by: Saha, Sourav, et al.
Published: (2026)
by: Saha, Sourav, et al.
Published: (2026)
SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding
by: Li, Sihang, et al.
Published: (2024)
by: Li, Sihang, et al.
Published: (2024)
SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis
by: Cai, Hengxing, et al.
Published: (2024)
by: Cai, Hengxing, et al.
Published: (2024)
Sci-CoE: Co-evolving Scientific Reasoning LLMs via Geometric Consensus with Sparse Supervision
by: He, Xiaohan, et al.
Published: (2026)
by: He, Xiaohan, et al.
Published: (2026)
Poisson Flow Joint Model for Multiphase contrast-enhanced CT
by: Ge, Rongjun, et al.
Published: (2025)
by: Ge, Rongjun, et al.
Published: (2025)
Understanding and Detecting Flaky Builds in GitHub Actions
by: Ge, Wenhao, et al.
Published: (2026)
by: Ge, Wenhao, et al.
Published: (2026)
MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical Language
by: Wang, Shun, et al.
Published: (2024)
by: Wang, Shun, et al.
Published: (2024)
Association between Glycosylated Hemoglobin and Serum Uric Acid: A US NHANES 2011–2020
by: Huan Li, et al.
Published: (2024)
by: Huan Li, et al.
Published: (2024)
“Back‐to‐Back” Radial Layered Skeleton Converging Heat Flow to Assist in Thermal Conduction of Aramid Nanofibers/Graphene Phase Change Composite Materials
by: Jun Tong, et al.
Published: (2024)
by: Jun Tong, et al.
Published: (2024)
SciGPT: A Large Language Model for Scientific Literature Understanding and Knowledge Discovery
by: She, Fengyu, et al.
Published: (2025)
by: She, Fengyu, et al.
Published: (2025)
SciMDR: Advancing Scientific Multimodal Document Reasoning
by: Chen, Ziyu, et al.
Published: (2026)
by: Chen, Ziyu, et al.
Published: (2026)
CryptoX : Compositional Reasoning Evaluation of Large Language Models
by: Shi, Jiajun, et al.
Published: (2025)
by: Shi, Jiajun, et al.
Published: (2025)
The origins of normativity: Assessor teaching and the emergence of norms
by: Laureano Castro
Published: (2023)
by: Laureano Castro
Published: (2023)
Epidemiological Characteristics, Target Distribution, and Clinical Value of Targeted Anticancer Drugs Approved in China: A Cross‐Sectional Study
by: Ting Zhu, et al.
Published: (2026)
by: Ting Zhu, et al.
Published: (2026)
WildSci: Advancing Scientific Reasoning from In-the-Wild Literature
by: Liu, Tengxiao, et al.
Published: (2026)
by: Liu, Tengxiao, et al.
Published: (2026)
Adaptive Donor‐Acceptor Modulation in Conjugated Polymers for Sustainable Solar‐Driven Water‐Electricity Cogeneration
by: Tong Liu, et al.
Published: (2025)
by: Tong Liu, et al.
Published: (2025)
Similar Items
-
ScholarSearch: Benchmarking Scholar Searching Ability of LLMs
by: Zhou, Junting, et al.
Published: (2025) -
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
by: Wu, Siwei, et al.
Published: (2024) -
Assessor Personality Traits Are Not Educationally Important Drivers of Assessor Stringency/Leniency
by: Sebastian Dewhirst, et al.
Published: (2025) -
Evaluating Library Assessors.
by: Credaro, Amanda
Published: (2003) -
Time‐Resolved Small‐Angle X‐Ray Studies of Spherical Micelle Formation and Growth During Polymerization‐Induced Self‐Assembly in Polar Solvents
by: Zhiqing Mei, et al.
Published: (2025)