Selecting Auxiliary Data via Neural Tangent Kernels for Low-Resource Domains
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Pingjie, Liu, Hongcheng, Liao, Yusheng, Fan, Ziqing, Du, Yaxin, Tang, Shuo, Wang, Yanfeng, Wang, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
M2K-VDG: Model-Adaptive Multimodal Knowledge Anchor Enhanced Video-grounded Dialogue Generation
by: Liu, Hongcheng, et al.
Published: (2024)
by: Liu, Hongcheng, et al.
Published: (2024)
When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
by: Liu, Hongcheng, et al.
Published: (2025)
by: Liu, Hongcheng, et al.
Published: (2025)
Leveraging Diverse Modeling Contexts with Collaborating Learning for Neural Machine Translation
by: Liao, Yusheng, et al.
Published: (2024)
by: Liao, Yusheng, et al.
Published: (2024)
Towards Omni-RAG: Comprehensive Retrieval-Augmented Generation for Large Language Models in Medical Applications
by: Chen, Zhe, et al.
Published: (2025)
by: Chen, Zhe, et al.
Published: (2025)
Automatic Interactive Evaluation for Large Language Models with State Aware Patient Simulator
by: Liao, Yusheng, et al.
Published: (2024)
by: Liao, Yusheng, et al.
Published: (2024)
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMs
by: Liu, Hongcheng, et al.
Published: (2026)
by: Liu, Hongcheng, et al.
Published: (2026)
Med-PMC: Medical Personalized Multi-modal Consultation with a Proactive Ask-First-Observe-Next Paradigm
by: Liu, Hongcheng, et al.
Published: (2024)
by: Liu, Hongcheng, et al.
Published: (2024)
Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning
by: Ou, Siqu, et al.
Published: (2025)
by: Ou, Siqu, et al.
Published: (2025)
MING-MOE: Enhancing Medical Multi-Task Learning in Large Language Models with Sparse Mixture of Low-Rank Adapter Experts
by: Liao, Yusheng, et al.
Published: (2024)
by: Liao, Yusheng, et al.
Published: (2024)
DSVD: Dynamic Self-Verify Decoding for Faithful Generation in Large Language Models
by: Guo, YiQiu, et al.
Published: (2025)
by: Guo, YiQiu, et al.
Published: (2025)
HeteroRAG: A Heterogeneous Retrieval-Augmented Generation Framework for Medical Vision Language Tasks
by: Chen, Zhe, et al.
Published: (2025)
by: Chen, Zhe, et al.
Published: (2025)
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Combatting Dimensional Collapse in LLM Pre-Training Data via Diversified File Selection
by: Fan, Ziqing, et al.
Published: (2025)
by: Fan, Ziqing, et al.
Published: (2025)
TAIA: Large Language Models are Out-of-Distribution Data Learners
by: Jiang, Shuyang, et al.
Published: (2024)
by: Jiang, Shuyang, et al.
Published: (2024)
ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical Agents
by: Liao, Yusheng, et al.
Published: (2024)
by: Liao, Yusheng, et al.
Published: (2024)
Overthinking Reduction with Decoupled Rewards and Curriculum Data Scheduling
by: Jiang, Shuyang, et al.
Published: (2025)
by: Jiang, Shuyang, et al.
Published: (2025)
Decoding Linguistic Representations of Human Brain
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
MedCare: Advancing Medical LLMs through Decoupling Clinical Alignment and Knowledge Aggregation
by: Liao, Yusheng, et al.
Published: (2024)
by: Liao, Yusheng, et al.
Published: (2024)
Reconstruct the Pruned Model without Any Retraining
by: Wang, Pingjie, et al.
Published: (2024)
by: Wang, Pingjie, et al.
Published: (2024)
DICE: Structured Reasoning in LLMs through SLM-Guided Chain-of-Thought Correction
by: Li, Yiqi, et al.
Published: (2025)
by: Li, Yiqi, et al.
Published: (2025)
AgentEHR: Advancing Autonomous Clinical Decision-Making via Retrospective Summarization
by: Liao, Yusheng, et al.
Published: (2026)
by: Liao, Yusheng, et al.
Published: (2026)
VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
by: Liu, Hongcheng, et al.
Published: (2025)
by: Liu, Hongcheng, et al.
Published: (2025)
MedS$^3$: Towards Medical Slow Thinking with Self-Evolved Soft Dual-sided Process Supervision
by: Jiang, Shuyang, et al.
Published: (2025)
by: Jiang, Shuyang, et al.
Published: (2025)
Joint Selection for Large-Scale Pre-Training Data via Policy Gradient-based Mask Learning
by: Fan, Ziqing, et al.
Published: (2025)
by: Fan, Ziqing, et al.
Published: (2025)
MetaMT,a MetaLearning Method Leveraging Multiple Domain Data for Low Resource Machine Translation
by: Li, Rumeng, et al.
Published: (2019)
by: Li, Rumeng, et al.
Published: (2019)
VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context
by: Liu, Heyang, et al.
Published: (2025)
by: Liu, Heyang, et al.
Published: (2025)
A Unified Data Augmentation Framework for Low-Resource Multi-Domain Dialogue Generation
by: Liu, Yongkang, et al.
Published: (2024)
by: Liu, Yongkang, et al.
Published: (2024)
M$^3$AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset
by: Chen, Zhe, et al.
Published: (2024)
by: Chen, Zhe, et al.
Published: (2024)
OpenFedLLM: Training Large Language Models on Decentralized Private Data via Federated Learning
by: Ye, Rui, et al.
Published: (2024)
by: Ye, Rui, et al.
Published: (2024)
A*-Thought: Efficient Reasoning via Bidirectional Compression for Low-Resource Settings
by: Xu, Xiaoang, et al.
Published: (2025)
by: Xu, Xiaoang, et al.
Published: (2025)
Learning to Adapt to Low-Resource Paraphrase Generation
by: Li, Zhigen, et al.
Published: (2024)
by: Li, Zhigen, et al.
Published: (2024)
VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
by: Liu, Heyang, et al.
Published: (2025)
by: Liu, Heyang, et al.
Published: (2025)
Resolving Knowledge Conflicts in Domain-specific Data Selection: A Case Study on Medical Instruction-tuning
by: Zhong, Qihuang, et al.
Published: (2025)
by: Zhong, Qihuang, et al.
Published: (2025)
ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification
by: Fan, Ziqing, et al.
Published: (2025)
by: Fan, Ziqing, et al.
Published: (2025)
Model Reprogramming Demystified: A Neural Tangent Kernel Perspective
by: Chung, Ming-Yu, et al.
Published: (2025)
by: Chung, Ming-Yu, et al.
Published: (2025)
VocalNet-M2: Advancing Low-Latency Spoken Language Modeling via Integrated Multi-Codebook Tokenization and Multi-Token Prediction
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Fast Graph Condensation with Structure-based Neural Tangent Kernel
by: Wang, Lin, et al.
Published: (2023)
by: Wang, Lin, et al.
Published: (2023)
Synthesizing Post-Training Data for LLMs through Multi-Agent Simulation
by: Tang, Shuo, et al.
Published: (2024)
by: Tang, Shuo, et al.
Published: (2024)
An Effective Deployment of Diffusion LM for Data Augmentation in Low-Resource Sentiment Classification
by: Chen, Zhuowei, et al.
Published: (2024)
by: Chen, Zhuowei, et al.
Published: (2024)
Similar Items
-
M2K-VDG: Model-Adaptive Multimodal Knowledge Anchor Enhanced Video-grounded Dialogue Generation
by: Liu, Hongcheng, et al.
Published: (2024) -
When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
by: Liu, Hongcheng, et al.
Published: (2025) -
Leveraging Diverse Modeling Contexts with Collaborating Learning for Neural Machine Translation
by: Liao, Yusheng, et al.
Published: (2024) -
Towards Omni-RAG: Comprehensive Retrieval-Augmented Generation for Large Language Models in Medical Applications
by: Chen, Zhe, et al.
Published: (2025) -
Automatic Interactive Evaluation for Large Language Models with State Aware Patient Simulator
by: Liao, Yusheng, et al.
Published: (2024)