Towards Robust Knowledge Tracing Models via k-Sparse Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Shuyan, Liu, Zitao, Zhao, Xiangyu, Luo, Weiqi, Weng, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Low-Resource Knowledge Tracing Tasks by Supervised Pre-training and Importance Mechanism Fine-tuning
by: Zhang, Hengyuan, et al.
Published: (2024)
by: Zhang, Hengyuan, et al.
Published: (2024)
What Are Step-Level Reward Models Rewarding? Counterintuitive Findings from MCTS-Boosted Mathematical Reasoning
by: Ma, Yiran, et al.
Published: (2024)
by: Ma, Yiran, et al.
Published: (2024)
Personalized Knowledge Tracing through Student Representation Reconstruction and Class Imbalance Mitigation
by: Chen, Zhiyu, et al.
Published: (2024)
by: Chen, Zhiyu, et al.
Published: (2024)
A Question-centric Multi-experts Contrastive Learning Framework for Improving the Accuracy and Interpretability of Deep Sequential Knowledge Tracing Models
by: Zhang, Hengyuan, et al.
Published: (2024)
by: Zhang, Hengyuan, et al.
Published: (2024)
A Mathematical Theory of Top-$k$ Sparse Attention via Total Variation Distance
by: Tzachristas, Georgios, et al.
Published: (2025)
by: Tzachristas, Georgios, et al.
Published: (2025)
Domain Generalizable Knowledge Tracing via Concept Aggregation and Relation-Based Attention
by: Xie, Yuquan, et al.
Published: (2024)
by: Xie, Yuquan, et al.
Published: (2024)
Sparse Attention Decomposition Applied to Circuit Tracing
by: Franco, Gabriel, et al.
Published: (2024)
by: Franco, Gabriel, et al.
Published: (2024)
Improving Sparse Autoencoder with Dynamic Attention
by: Wang, Dongsheng, et al.
Published: (2026)
by: Wang, Dongsheng, et al.
Published: (2026)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
by: Yuan, Jingyang, et al.
Published: (2025)
by: Yuan, Jingyang, et al.
Published: (2025)
NOSA: Native and Offloadable Sparse Attention
by: Huang, Yuxiang, et al.
Published: (2025)
by: Huang, Yuxiang, et al.
Published: (2025)
GRASP: group-Shapley feature selection for patients
by: Luo, Yuheng, et al.
Published: (2026)
by: Luo, Yuheng, et al.
Published: (2026)
S2O: Early Stopping for Sparse Attention via Online Permutation
by: Zhang, Yu, et al.
Published: (2026)
by: Zhang, Yu, et al.
Published: (2026)
Nonlinearity, Feedback and Uniform Consistency in Causal Structural Learning
by: Wang, Shuyan
Published: (2023)
by: Wang, Shuyan
Published: (2023)
vAttention: Verified Sparse Attention
by: Desai, Aditya, et al.
Published: (2025)
by: Desai, Aditya, et al.
Published: (2025)
Intrinsically Interpretable Attention via Sparse Post-Training
by: Draye, Florent, et al.
Published: (2025)
by: Draye, Florent, et al.
Published: (2025)
Advancing Personalized Learning Analysis via an Innovative Domain Knowledge Informed Attention-based Knowledge Tracing Method
by: Kose, Shubham, et al.
Published: (2025)
by: Kose, Shubham, et al.
Published: (2025)
SparseDM: Toward Sparse Efficient Diffusion Models
by: Wang, Kafeng, et al.
Published: (2024)
by: Wang, Kafeng, et al.
Published: (2024)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
Towards Robust Multi-Modal Reasoning via Model Selection
by: Liu, Xiangyan, et al.
Published: (2023)
by: Liu, Xiangyan, et al.
Published: (2023)
Stem: Rethinking Causal Information Flow in Sparse Attention
by: Niu, Lin, et al.
Published: (2026)
by: Niu, Lin, et al.
Published: (2026)
Trainable Dynamic Mask Sparse Attention
by: Shi, Jingze, et al.
Published: (2025)
by: Shi, Jingze, et al.
Published: (2025)
Cumulative Distribution Function based General Temporal Point Processes
by: Wang, Maolin, et al.
Published: (2024)
by: Wang, Maolin, et al.
Published: (2024)
FedSKD: Aggregation-free Model-heterogeneous Federated Learning via Multi-dimensional Similarity Knowledge Distillation for Medical Image Classification
by: Weng, Ziqiao, et al.
Published: (2025)
by: Weng, Ziqiao, et al.
Published: (2025)
Improving Robustness In Sparse Autoencoders via Masked Regularization
by: Narayanaswamy, Vivek, et al.
Published: (2026)
by: Narayanaswamy, Vivek, et al.
Published: (2026)
Robust and Efficient Zeroth-Order LLM Fine-Tuning via Adaptive Bayesian Subspace Optimizer
by: Feng, Jian, et al.
Published: (2026)
by: Feng, Jian, et al.
Published: (2026)
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
by: Gao, Yizhao, et al.
Published: (2025)
by: Gao, Yizhao, et al.
Published: (2025)
Self-Indexing KVCache: Predicting Sparse Attention from Compressed Keys
by: Yang, Xu, et al.
Published: (2026)
by: Yang, Xu, et al.
Published: (2026)
LLM DNA: Tracing Model Evolution via Functional Representations
by: Wu, Zhaomin, et al.
Published: (2025)
by: Wu, Zhaomin, et al.
Published: (2025)
ZETA: Leveraging Z-order Curves for Efficient Top-k Attention
by: Zeng, Qiuhao, et al.
Published: (2025)
by: Zeng, Qiuhao, et al.
Published: (2025)
Double-P: Hierarchical Top-P Sparse Attention for Long-Context LLMs
by: Ni, Wentao, et al.
Published: (2026)
by: Ni, Wentao, et al.
Published: (2026)
HSR-Enhanced Sparse Attention Acceleration
by: Chen, Bo, et al.
Published: (2024)
by: Chen, Bo, et al.
Published: (2024)
Robust Knowledge Transfer in Tiered Reinforcement Learning
by: Huang, Jiawei, et al.
Published: (2023)
by: Huang, Jiawei, et al.
Published: (2023)
Towards Interpretable Adversarial Examples via Sparse Adversarial Attack
by: Lin, Fudong, et al.
Published: (2025)
by: Lin, Fudong, et al.
Published: (2025)
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
by: Huang, Yuxiang, et al.
Published: (2026)
by: Huang, Yuxiang, et al.
Published: (2026)
Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework
by: Weng, Jiaqi, et al.
Published: (2025)
by: Weng, Jiaqi, et al.
Published: (2025)
Towards Graph Foundation Models: Training on Knowledge Graphs Enables Transferability to General Graphs
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention
by: Xu, Hongtao, et al.
Published: (2026)
by: Xu, Hongtao, et al.
Published: (2026)
Circuit Complexity of Hierarchical Knowledge Tracing and Implications for Log-Precision Transformers
by: Liu, Naiming, et al.
Published: (2026)
by: Liu, Naiming, et al.
Published: (2026)
Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models
by: Yuan, Hongbang, et al.
Published: (2024)
by: Yuan, Hongbang, et al.
Published: (2024)
From Sparsity to Simplicity: Enabling Simpler Sequential Replacements via Sparse Attention Distillation
by: Ren, Yuxin, et al.
Published: (2026)
by: Ren, Yuxin, et al.
Published: (2026)
Similar Items
-
Improving Low-Resource Knowledge Tracing Tasks by Supervised Pre-training and Importance Mechanism Fine-tuning
by: Zhang, Hengyuan, et al.
Published: (2024) -
What Are Step-Level Reward Models Rewarding? Counterintuitive Findings from MCTS-Boosted Mathematical Reasoning
by: Ma, Yiran, et al.
Published: (2024) -
Personalized Knowledge Tracing through Student Representation Reconstruction and Class Imbalance Mitigation
by: Chen, Zhiyu, et al.
Published: (2024) -
A Question-centric Multi-experts Contrastive Learning Framework for Improving the Accuracy and Interpretability of Deep Sequential Knowledge Tracing Models
by: Zhang, Hengyuan, et al.
Published: (2024) -
A Mathematical Theory of Top-$k$ Sparse Attention via Total Variation Distance
by: Tzachristas, Georgios, et al.
Published: (2025)