SCAN: Structured Capability Assessment and Navigation for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zongqi, Gu, Tianle, Gong, Chen, Tian, Xin, Bao, Siqi, Yang, Yujiu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
by: Sun, Shaoning, et al.
Published: (2025)
by: Sun, Shaoning, et al.
Published: (2025)
Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking
by: Gu, Tianle, et al.
Published: (2025)
by: Gu, Tianle, et al.
Published: (2025)
Robust and Minimally Invasive Watermarking for EaaS
by: Wang, Zongqi, et al.
Published: (2024)
by: Wang, Zongqi, et al.
Published: (2024)
From Sparse Decisions to Dense Reasoning: A Multi-attribute Trajectory Paradigm for Multimodal Moderation
by: Gu, Tianle, et al.
Published: (2026)
by: Gu, Tianle, et al.
Published: (2026)
Chain of History: Learning and Forecasting with LLMs for Temporal Knowledge Graph Completion
by: Luo, Ruilin, et al.
Published: (2024)
by: Luo, Ruilin, et al.
Published: (2024)
Probing the Robustness of Large Language Models Safety to Latent Perturbations
by: Gu, Tianle, et al.
Published: (2025)
by: Gu, Tianle, et al.
Published: (2025)
MorphMark: Flexible Adaptive Watermarking for Large Language Models
by: Wang, Zongqi, et al.
Published: (2025)
by: Wang, Zongqi, et al.
Published: (2025)
Reward Modeling from Natural Language Human Feedback
by: Wang, Zongqi, et al.
Published: (2026)
by: Wang, Zongqi, et al.
Published: (2026)
MEOW: MEMOry Supervised LLM Unlearning Via Inverted Facts
by: Gu, Tianle, et al.
Published: (2024)
by: Gu, Tianle, et al.
Published: (2024)
Can Constructions "SCAN" Compositionality ?
by: Katrapati, Ganesh, et al.
Published: (2025)
by: Katrapati, Ganesh, et al.
Published: (2025)
Word Form Matters: LLMs' Semantic Reconstruction under Typoglycemia
by: Wang, Chenxi, et al.
Published: (2025)
by: Wang, Chenxi, et al.
Published: (2025)
A Thorough Examination of Decoding Methods in the Era of LLMs
by: Shi, Chufan, et al.
Published: (2024)
by: Shi, Chufan, et al.
Published: (2024)
Distributional Clarity: The Hidden Driver of RL-Friendliness in Large Language Models
by: Sun, Shaoning, et al.
Published: (2026)
by: Sun, Shaoning, et al.
Published: (2026)
PTD-SQL: Partitioning and Targeted Drilling with LLMs in Text-to-SQL
by: Luo, Ruilin, et al.
Published: (2024)
by: Luo, Ruilin, et al.
Published: (2024)
Distilling Bayesian Belief States into Language Models for Auditable Negotiation
by: Cui, Zongqi, et al.
Published: (2026)
by: Cui, Zongqi, et al.
Published: (2026)
INFERENCEDYNAMICS: Efficient Routing Across LLMs through Structured Capability and Knowledge Profiling
by: Shi, Haochen, et al.
Published: (2025)
by: Shi, Haochen, et al.
Published: (2025)
Evaluating the Capabilities of LLMs for Supporting Anticipatory Impact Assessment
by: Allaham, Mowafak, et al.
Published: (2024)
by: Allaham, Mowafak, et al.
Published: (2024)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
by: Lin, Zicheng, et al.
Published: (2024)
by: Lin, Zicheng, et al.
Published: (2024)
STEM: Efficient Relative Capability Evaluation of LLMs through Structured Transition Samples
by: Hu, Haiquan, et al.
Published: (2025)
by: Hu, Haiquan, et al.
Published: (2025)
Be Cautious When Merging Unfamiliar LLMs: A Phishing Model Capable of Stealing Privacy
by: Guo, Zhenyuan, et al.
Published: (2025)
by: Guo, Zhenyuan, et al.
Published: (2025)
How Well Do LLMs Handle Cantonese? Benchmarking Cantonese Capabilities of Large Language Models
by: Jiang, Jiyue, et al.
Published: (2024)
by: Jiang, Jiyue, et al.
Published: (2024)
HIPPO: Enhancing the Table Understanding Capability of LLMs through Hybrid-Modal Preference Optimization
by: Wang, Haolan, et al.
Published: (2025)
by: Wang, Haolan, et al.
Published: (2025)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
by: Jiang, Zhuohang, et al.
Published: (2025)
by: Jiang, Zhuohang, et al.
Published: (2025)
ShifCon: Enhancing Non-Dominant Language Capabilities with a Shift-based Multilingual Contrastive Framework
by: Zhang, Hengyuan, et al.
Published: (2024)
by: Zhang, Hengyuan, et al.
Published: (2024)
Structured Prompt Language: Declarative Context Management for LLMs
by: Gong, Wen G.
Published: (2026)
by: Gong, Wen G.
Published: (2026)
Exploring the Capability Boundaries of LLMs in Mastering of Chinese Chouxiang Language
by: Lin, Dianqing, et al.
Published: (2026)
by: Lin, Dianqing, et al.
Published: (2026)
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
by: Yang, Jialin, et al.
Published: (2025)
by: Yang, Jialin, et al.
Published: (2025)
Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination
by: Gao, Lirong, et al.
Published: (2026)
by: Gao, Lirong, et al.
Published: (2026)
LiFi: Lightweight Controlled Text Generation with Fine-Grained Control Codes
by: Shi, Chufan, et al.
Published: (2024)
by: Shi, Chufan, et al.
Published: (2024)
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
by: Lin, Zicheng, et al.
Published: (2024)
by: Lin, Zicheng, et al.
Published: (2024)
Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance
by: Yu, Jiachen, et al.
Published: (2026)
by: Yu, Jiachen, et al.
Published: (2026)
Exploring the Mystery of Influential Data for Mathematical Reasoning
by: Ni, Xinzhe, et al.
Published: (2024)
by: Ni, Xinzhe, et al.
Published: (2024)
PROST-LLM: Progressively Enhancing the Speech-to-Speech Translation Capability in LLMs
by: Xu, Jing, et al.
Published: (2026)
by: Xu, Jing, et al.
Published: (2026)
FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability
by: Xia, Congying, et al.
Published: (2024)
by: Xia, Congying, et al.
Published: (2024)
Towards Automatic Evaluation for LLMs' Clinical Capabilities: Metric, Data, and Algorithm
by: Liu, Lei, et al.
Published: (2024)
by: Liu, Lei, et al.
Published: (2024)
The SCAN Statistical Model Checker
by: Ghiorzi, Enrico, et al.
Published: (2026)
by: Ghiorzi, Enrico, et al.
Published: (2026)
Reasoning Under Uncertainty: Exploring Probabilistic Reasoning Capabilities of LLMs
by: Pournemat, Mobina, et al.
Published: (2025)
by: Pournemat, Mobina, et al.
Published: (2025)
Are Your LLMs Capable of Stable Reasoning?
by: Liu, Junnan, et al.
Published: (2024)
by: Liu, Junnan, et al.
Published: (2024)
How Difficulty-Aware Staged Reinforcement Learning Enhances LLMs' Reasoning Capabilities: A Preliminary Experimental Study
by: Ji, Yunjie, et al.
Published: (2025)
by: Ji, Yunjie, et al.
Published: (2025)
MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding
by: Deng, Jingyuan, et al.
Published: (2025)
by: Deng, Jingyuan, et al.
Published: (2025)
Similar Items
-
S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
by: Sun, Shaoning, et al.
Published: (2025) -
Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking
by: Gu, Tianle, et al.
Published: (2025) -
Robust and Minimally Invasive Watermarking for EaaS
by: Wang, Zongqi, et al.
Published: (2024) -
From Sparse Decisions to Dense Reasoning: A Multi-attribute Trajectory Paradigm for Multimodal Moderation
by: Gu, Tianle, et al.
Published: (2026) -
Chain of History: Learning and Forecasting with LLMs for Temporal Knowledge Graph Completion
by: Luo, Ruilin, et al.
Published: (2024)