AutoSCORE: Enhancing Automated Scoring with Multi-Agent Large Language Models via Structured Component Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yun, Ding, Zhaojun, Wu, Xuansheng, Sun, Siyue, Liu, Ninghao, Zhai, Xiaoming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
by: Wang, Yun, et al.
Published: (2026)
by: Wang, Yun, et al.
Published: (2026)
BRIDGE the Gap: Mitigating Bias Amplification in Automated Scoring of English Language Learners via Inter-group Data Augmentation
by: Wang, Yun, et al.
Published: (2026)
by: Wang, Yun, et al.
Published: (2026)
Applying Large Language Models and Chain-of-Thought for Automatic Scoring
by: Lee, Gyeong-Geon, et al.
Published: (2023)
by: Lee, Gyeong-Geon, et al.
Published: (2023)
Self-Regularization with Sparse Autoencoders for Controllable LLM-based Classification
by: Wu, Xuansheng, et al.
Published: (2025)
by: Wu, Xuansheng, et al.
Published: (2025)
Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring
by: Wu, Xuansheng, et al.
Published: (2024)
by: Wu, Xuansheng, et al.
Published: (2024)
Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders
by: Wu, Xuansheng, et al.
Published: (2025)
by: Wu, Xuansheng, et al.
Published: (2025)
Artificial Intelligence Bias on English Language Learners in Automatic Scoring
by: Guo, Shuchen, et al.
Published: (2025)
by: Guo, Shuchen, et al.
Published: (2025)
Foundation Models for Low-Resource Language Education (Vision Paper)
by: Ding, Zhaojun, et al.
Published: (2024)
by: Ding, Zhaojun, et al.
Published: (2024)
Retrieval-enhanced Knowledge Editing in Language Models for Multi-Hop Question Answering
by: Shi, Yucheng, et al.
Published: (2024)
by: Shi, Yucheng, et al.
Published: (2024)
Investigating CoT Monitorability in Large Reasoning Models
by: Yang, Shu, et al.
Published: (2025)
by: Yang, Shu, et al.
Published: (2025)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
by: Shu, Dong, et al.
Published: (2025)
by: Shu, Dong, et al.
Published: (2025)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
by: Zhao, Haiyan, et al.
Published: (2025)
by: Zhao, Haiyan, et al.
Published: (2025)
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
by: Shu, Dong, et al.
Published: (2025)
by: Shu, Dong, et al.
Published: (2025)
Auto-TA: Towards Scalable Automated Thematic Analysis (TA) via Multi-Agent Large Language Models with Reinforcement Learning
by: Yi, Seungjun, et al.
Published: (2025)
by: Yi, Seungjun, et al.
Published: (2025)
Could Small Language Models Serve as Recommenders? Towards Data-centric Cold-start Recommendations
by: Wu, Xuansheng, et al.
Published: (2023)
by: Wu, Xuansheng, et al.
Published: (2023)
Rank-Then-Score: Enhancing Large Language Models for Automated Essay Scoring
by: Cai, Yida, et al.
Published: (2025)
by: Cai, Yida, et al.
Published: (2025)
Efficient Multi-Task Inferencing with a Shared Backbone and Lightweight Task-Specific Adapters for Automatic Scoring
by: Latif, Ehsan, et al.
Published: (2024)
by: Latif, Ehsan, et al.
Published: (2024)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
by: Wu, Xuansheng, et al.
Published: (2023)
by: Wu, Xuansheng, et al.
Published: (2023)
AutoFlow: Automated Workflow Generation for Large Language Model Agents
by: Li, Zelong, et al.
Published: (2024)
by: Li, Zelong, et al.
Published: (2024)
Automating Dataset Updates Towards Reliable and Timely Evaluation of Large Language Models
by: Ying, Jiahao, et al.
Published: (2024)
by: Ying, Jiahao, et al.
Published: (2024)
Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era
by: Wu, Xuansheng, et al.
Published: (2024)
by: Wu, Xuansheng, et al.
Published: (2024)
SCORE: Systematic COnsistency and Robustness Evaluation for Large Language Models
by: Nalbandyan, Grigor, et al.
Published: (2025)
by: Nalbandyan, Grigor, et al.
Published: (2025)
AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents
by: Fu, Yao, et al.
Published: (2024)
by: Fu, Yao, et al.
Published: (2024)
AutoPCR: Automated Phenotype Concept Recognition by Prompting
by: Tao, Yicheng, et al.
Published: (2025)
by: Tao, Yicheng, et al.
Published: (2025)
Auto-Intent: Automated Intent Discovery and Self-Exploration for Large Language Model Web Agents
by: Kim, Jaekyeom, et al.
Published: (2024)
by: Kim, Jaekyeom, et al.
Published: (2024)
MASTER: Enhancing Large Language Model via Multi-Agent Simulated Teaching
by: Yue, Liang, et al.
Published: (2025)
by: Yue, Liang, et al.
Published: (2025)
Using Generative AI and Multi-Agents to Provide Automatic Feedback
by: Guo, Shuchen, et al.
Published: (2024)
by: Guo, Shuchen, et al.
Published: (2024)
Agent-Enhanced Large Language Models for Researching Political Institutions
by: Loffredo, Joseph R., et al.
Published: (2025)
by: Loffredo, Joseph R., et al.
Published: (2025)
HalluSAE: Detecting Hallucinations in Large Language Models via Sparse Auto-Encoders
by: Chen, Boshui, et al.
Published: (2026)
by: Chen, Boshui, et al.
Published: (2026)
Knowledge Distillation of LLM for Automatic Scoring of Science Education Assessments
by: Latif, Ehsan, et al.
Published: (2023)
by: Latif, Ehsan, et al.
Published: (2023)
AutoHall: Automated Factuality Hallucination Dataset Generation for Large Language Models
by: Cao, Zouying, et al.
Published: (2023)
by: Cao, Zouying, et al.
Published: (2023)
Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
by: Wang, Anyi, et al.
Published: (2025)
by: Wang, Anyi, et al.
Published: (2025)
Using GPT-4 to Augment Unbalanced Data for Automatic Scoring
by: Fang, Luyang, et al.
Published: (2023)
by: Fang, Luyang, et al.
Published: (2023)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
by: Reddy, Aashray, et al.
Published: (2025)
by: Reddy, Aashray, et al.
Published: (2025)
Fine-tuning ChatGPT for Automatic Scoring of Written Scientific Explanations in Chinese
by: Yang, Jie, et al.
Published: (2025)
by: Yang, Jie, et al.
Published: (2025)
Mil-SCORE: Benchmarking Long-Context Geospatial Reasoning and Planning in Large Language Models
by: Palnitkar, Aadi, et al.
Published: (2026)
by: Palnitkar, Aadi, et al.
Published: (2026)
Automated Genre-Aware Article Scoring and Feedback Using Large Language Models
by: Wang, Chihang, et al.
Published: (2024)
by: Wang, Chihang, et al.
Published: (2024)
AutoToM: Scaling Model-based Mental Inference via Automated Agent Modeling
by: Zhang, Zhining, et al.
Published: (2025)
by: Zhang, Zhining, et al.
Published: (2025)
AutoLogi: Automated Generation of Logic Puzzles for Evaluating Reasoning Abilities of Large Language Models
by: Zhu, Qin, et al.
Published: (2025)
by: Zhu, Qin, et al.
Published: (2025)
Automating Structural Engineering Workflows with Large Language Model Agents
by: Liang, Haoran, et al.
Published: (2025)
by: Liang, Haoran, et al.
Published: (2025)
Similar Items
-
Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
by: Wang, Yun, et al.
Published: (2026) -
BRIDGE the Gap: Mitigating Bias Amplification in Automated Scoring of English Language Learners via Inter-group Data Augmentation
by: Wang, Yun, et al.
Published: (2026) -
Applying Large Language Models and Chain-of-Thought for Automatic Scoring
by: Lee, Gyeong-Geon, et al.
Published: (2023) -
Self-Regularization with Sparse Autoencoders for Controllable LLM-based Classification
by: Wu, Xuansheng, et al.
Published: (2025) -
Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring
by: Wu, Xuansheng, et al.
Published: (2024)