LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Jian, Yu, Hang, Liu, Bingchang, Yang, Wenjie, Di, Peng, Li, Jianguo, Zhang, Yue |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CoBa: Convergence Balancer for Multitask Finetuning of Large Language Models
von: Gong, Zi, et al.
Veröffentlicht: (2024)
von: Gong, Zi, et al.
Veröffentlicht: (2024)
Unifying the Perspectives of NLP and Software Engineering: A Survey on Language Models for Code
von: Zhang, Ziyin, et al.
Veröffentlicht: (2023)
von: Zhang, Ziyin, et al.
Veröffentlicht: (2023)
Unified Data Selection for LLM Reasoning
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2026)
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2026)
InstructDiff: Domain-Adaptive Data Selection via Differential Entropy for Efficient LLM Fine-Tuning
von: Su, Junyou, et al.
Veröffentlicht: (2026)
von: Su, Junyou, et al.
Veröffentlicht: (2026)
Pruning as a Domain-specific LLM Extractor
von: Zhang, Nan, et al.
Veröffentlicht: (2024)
von: Zhang, Nan, et al.
Veröffentlicht: (2024)
D2LLM: Decomposed and Distilled Large Language Models for Semantic Search
von: Liao, Zihan, et al.
Veröffentlicht: (2024)
von: Liao, Zihan, et al.
Veröffentlicht: (2024)
GALLa: Graph Aligned Large Language Models for Improved Source Code Understanding
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
LLM as Graph Kernel: Rethinking Message Passing on Text-Rich Graphs
von: Zhang, Ying, et al.
Veröffentlicht: (2026)
von: Zhang, Ying, et al.
Veröffentlicht: (2026)
F2LLM Technical Report: Matching SOTA Embedding Performance with 6 Million Open-Source Data
von: Zhang, Ziyin, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyin, et al.
Veröffentlicht: (2025)
Resolving Knowledge Conflicts in Domain-specific Data Selection: A Case Study on Medical Instruction-tuning
von: Zhong, Qihuang, et al.
Veröffentlicht: (2025)
von: Zhong, Qihuang, et al.
Veröffentlicht: (2025)
E2LLM: Encoder Elongated Large Language Models for Long-Context Understanding and Reasoning
von: Liao, Zihan, et al.
Veröffentlicht: (2024)
von: Liao, Zihan, et al.
Veröffentlicht: (2024)
A Data Synthesis Method Driven by Large Language Models for Proactive Mining of Implicit User Intentions in Tourism
von: Wang, Jinqiang, et al.
Veröffentlicht: (2025)
von: Wang, Jinqiang, et al.
Veröffentlicht: (2025)
ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
EIoU-EMC: A Novel Loss for Domain-specific Nested Entity Recognition
von: Zhang, Jian, et al.
Veröffentlicht: (2025)
von: Zhang, Jian, et al.
Veröffentlicht: (2025)
Greedy Information Projection for LLM Data Selection
von: Dong, Victor Ye, et al.
Veröffentlicht: (2026)
von: Dong, Victor Ye, et al.
Veröffentlicht: (2026)
DavIR: Data Selection via Implicit Reward for Large Language Models
von: Zhou, Haotian, et al.
Veröffentlicht: (2023)
von: Zhou, Haotian, et al.
Veröffentlicht: (2023)
F2LLM-v2: Inclusive, Performant, and Efficient Embeddings for a Multilingual World
von: Zhang, Ziyin, et al.
Veröffentlicht: (2026)
von: Zhang, Ziyin, et al.
Veröffentlicht: (2026)
Incubating Text Classifiers Following User Instruction with Nothing but LLM
von: Peng, Letian, et al.
Veröffentlicht: (2024)
von: Peng, Letian, et al.
Veröffentlicht: (2024)
DALLMi: Domain Adaption for LLM-based Multi-label Classifier
von: Beţianu, Miruna, et al.
Veröffentlicht: (2024)
von: Beţianu, Miruna, et al.
Veröffentlicht: (2024)
ImF: Implicit Fingerprint for Large Language Models
von: Wu, Jiaxuan, et al.
Veröffentlicht: (2025)
von: Wu, Jiaxuan, et al.
Veröffentlicht: (2025)
QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining
von: Liu, Fengze, et al.
Veröffentlicht: (2025)
von: Liu, Fengze, et al.
Veröffentlicht: (2025)
COMAP: Co-Evolving World Models and Agent Policies for LLM Agents
von: Liu, Youwei, et al.
Veröffentlicht: (2026)
von: Liu, Youwei, et al.
Veröffentlicht: (2026)
NITP: Next Implicit Token Prediction for LLM Pre-training
von: Zhang, Xiangdong, et al.
Veröffentlicht: (2026)
von: Zhang, Xiangdong, et al.
Veröffentlicht: (2026)
LLM with Relation Classifier for Document-Level Relation Extraction
von: Li, Xingzuo, et al.
Veröffentlicht: (2024)
von: Li, Xingzuo, et al.
Veröffentlicht: (2024)
Autonomous Data Selection with Zero-shot Generative Classifiers for Mathematical Texts
von: Zhang, Yifan, et al.
Veröffentlicht: (2024)
von: Zhang, Yifan, et al.
Veröffentlicht: (2024)
Selection of LLM Fine-Tuning Data based on Orthogonal Rules
von: Li, Xiaomin, et al.
Veröffentlicht: (2024)
von: Li, Xiaomin, et al.
Veröffentlicht: (2024)
Domain-specific Guided Summarization for Mental Health Posts
von: Qian, Lu, et al.
Veröffentlicht: (2024)
von: Qian, Lu, et al.
Veröffentlicht: (2024)
DiSRouter: Distributed Self-Routing for LLM Selections
von: Zheng, Hang, et al.
Veröffentlicht: (2025)
von: Zheng, Hang, et al.
Veröffentlicht: (2025)
Evaluating and Enhancing Large Language Models Performance in Domain-specific Medicine: Osteoarthritis Management with DocOA
von: Chen, Xi, et al.
Veröffentlicht: (2024)
von: Chen, Xi, et al.
Veröffentlicht: (2024)
Explicit and Implicit Data Augmentation for Social Event Detection
von: Ma, Congbo, et al.
Veröffentlicht: (2025)
von: Ma, Congbo, et al.
Veröffentlicht: (2025)
Token Cleaning: Fine-Grained Data Selection for LLM Supervised Fine-Tuning
von: Pang, Jinlong, et al.
Veröffentlicht: (2025)
von: Pang, Jinlong, et al.
Veröffentlicht: (2025)
Rodimus*: Breaking the Accuracy-Efficiency Trade-Off with Efficient Attentions
von: He, Zhihao, et al.
Veröffentlicht: (2024)
von: He, Zhihao, et al.
Veröffentlicht: (2024)
SALP-CG: Standard-Aligned LLM Pipeline for Classifying and Grading Large Volumes of Online Conversational Health Data
von: Yan, Yiwei, et al.
Veröffentlicht: (2025)
von: Yan, Yiwei, et al.
Veröffentlicht: (2025)
AugTriever: Unsupervised Dense Retrieval and Domain Adaptation by Scalable Data Augmentation
von: Meng, Rui, et al.
Veröffentlicht: (2022)
von: Meng, Rui, et al.
Veröffentlicht: (2022)
ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning
von: Wu, Yang, et al.
Veröffentlicht: (2024)
von: Wu, Yang, et al.
Veröffentlicht: (2024)
From Myopic Selection to Long-Horizon Awareness: Sequential LLM Routing for Multi-Turn Dialogue
von: Zhang, Jiarui, et al.
Veröffentlicht: (2026)
von: Zhang, Jiarui, et al.
Veröffentlicht: (2026)
C2LLM Technical Report: A New Frontier in Code Retrieval via Adaptive Cross-Attention Pooling
von: Qin, Jin, et al.
Veröffentlicht: (2025)
von: Qin, Jin, et al.
Veröffentlicht: (2025)
Selecting Auxiliary Data via Neural Tangent Kernels for Low-Resource Domains
von: Wang, Pingjie, et al.
Veröffentlicht: (2025)
von: Wang, Pingjie, et al.
Veröffentlicht: (2025)
Don't Say No: Jailbreaking LLM by Suppressing Refusal
von: Zhou, Yukai, et al.
Veröffentlicht: (2024)
von: Zhou, Yukai, et al.
Veröffentlicht: (2024)
Internalizing ASR with Implicit Chain of Thought for Efficient Speech-to-Speech Conversational LLM
von: Yuen, Robin Shing-Hei, et al.
Veröffentlicht: (2024)
von: Yuen, Robin Shing-Hei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CoBa: Convergence Balancer for Multitask Finetuning of Large Language Models
von: Gong, Zi, et al.
Veröffentlicht: (2024) -
Unifying the Perspectives of NLP and Software Engineering: A Survey on Language Models for Code
von: Zhang, Ziyin, et al.
Veröffentlicht: (2023) -
Unified Data Selection for LLM Reasoning
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2026) -
InstructDiff: Domain-Adaptive Data Selection via Differential Entropy for Efficient LLM Fine-Tuning
von: Su, Junyou, et al.
Veröffentlicht: (2026) -
Pruning as a Domain-specific LLM Extractor
von: Zhang, Nan, et al.
Veröffentlicht: (2024)