Position: Let's Develop Data Probes to Fundamentally Understand How Data Affects LLM Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Shiqiang, Woisetschläger, Herbert, Jacobsen, Hans Arno, Ji, Mingyue |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MAR-FL: A Communication Efficient Peer-to-Peer Federated Learning System
by: Mulitze, Felix, et al.
Published: (2025)
by: Mulitze, Felix, et al.
Published: (2025)
MESS+: Dynamically Learned Inference-Time LLM Routing in Model Zoos with Service Level Guarantees
by: Woisetschläger, Herbert, et al.
Published: (2025)
by: Woisetschläger, Herbert, et al.
Published: (2025)
Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining
by: Sow, Daouda, et al.
Published: (2025)
by: Sow, Daouda, et al.
Published: (2025)
An Extremely Data-efficient and Generative LLM-based Reinforcement Learning Agent for Recommenders
by: Feng, Shuang, et al.
Published: (2024)
by: Feng, Shuang, et al.
Published: (2024)
No One Left Behind: How to Exploit the Incomplete and Skewed Multi-Label Data for Conversion Rate Prediction
by: Jia, Qinglin, et al.
Published: (2025)
by: Jia, Qinglin, et al.
Published: (2025)
Let It Go? Not Quite: Addressing Item Cold Start in Sequential Recommendations with Content-Based Initialization
by: Pembek, Anton, et al.
Published: (2025)
by: Pembek, Anton, et al.
Published: (2025)
How Many Tools Should an LLM Agent See? A Chance-Corrected Answer
by: Repantis, Vyzantinos, et al.
Published: (2026)
by: Repantis, Vyzantinos, et al.
Published: (2026)
A Survey on Efficient Federated Learning Methods for Foundation Model Training
by: Woisetschläger, Herbert, et al.
Published: (2024)
by: Woisetschläger, Herbert, et al.
Published: (2024)
A Data Science Approach to Calcutta High Court Judgments: An Efficient LLM and RAG-powered Framework for Summarization and Similar Cases Retrieval
by: Banerjee, Puspendu, et al.
Published: (2025)
by: Banerjee, Puspendu, et al.
Published: (2025)
To LLM, or Not to LLM: How Designers and Developers Navigate LLMs as Tools or Teammates
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
What Data is Really Necessary? A Feasibility Study of Inference Data Minimization for Recommender Systems
by: Leysen, Jens, et al.
Published: (2025)
by: Leysen, Jens, et al.
Published: (2025)
TOBUGraph: Knowledge Graph-Based Retrieval for Enhanced LLM Performance Beyond RAG
by: Kashmira, Savini, et al.
Published: (2024)
by: Kashmira, Savini, et al.
Published: (2024)
Prompt Optimization with Logged Bandit Data
by: Kiyohara, Haruka, et al.
Published: (2025)
by: Kiyohara, Haruka, et al.
Published: (2025)
WellFactor: Patient Profiling using Integrative Embedding of Healthcare Data
by: Choi, Dongjin, et al.
Published: (2023)
by: Choi, Dongjin, et al.
Published: (2023)
CADC: Encoding User-Item Interactions for Compressing Recommendation Model Training Data
by: Zarch, Hossein Entezari, et al.
Published: (2024)
by: Zarch, Hossein Entezari, et al.
Published: (2024)
Position-Aware Sequential Attention for Accurate Next Item Recommendations
by: Nabiev, Timur, et al.
Published: (2026)
by: Nabiev, Timur, et al.
Published: (2026)
Lightweight yet Efficient: An External Attentive Graph Convolutional Network with Positional Prompts for Sequential Recommendation
by: Zhang, Jinyu, et al.
Published: (2025)
by: Zhang, Jinyu, et al.
Published: (2025)
CirrusBench: Evaluating LLM-based Agents Beyond Correctness in Real-World Cloud Service Environments
by: Yu, Yi, et al.
Published: (2026)
by: Yu, Yi, et al.
Published: (2026)
Lightweight Adaptation for LLM-based Technical Service Agent: Latent Logic Augmentation and Robust Noise Reduction
by: Yu, Yi, et al.
Published: (2026)
by: Yu, Yi, et al.
Published: (2026)
LLM-assisted Vector Similarity Search
by: Riyadh, Md, et al.
Published: (2024)
by: Riyadh, Md, et al.
Published: (2024)
LiDDA: Data Driven Attribution at LinkedIn
by: Bencina, John, et al.
Published: (2025)
by: Bencina, John, et al.
Published: (2025)
LLM-based Schema-Guided Extraction and Validation of Missing-Person Intelligence from Heterogeneous Data Sources
by: Castillo, Joshua, et al.
Published: (2026)
by: Castillo, Joshua, et al.
Published: (2026)
Preference Discerning with LLM-Enhanced Generative Retrieval
by: Paischer, Fabian, et al.
Published: (2024)
by: Paischer, Fabian, et al.
Published: (2024)
Exploring How Fair Model Representations Relate to Fair Recommendations
by: Vassøy, Bjørnar, et al.
Published: (2026)
by: Vassøy, Bjørnar, et al.
Published: (2026)
R1-Ranker: Teaching LLM Rankers to Reason
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
Aligning Dense Retrievers with LLM Utility via Distillation
by: Sandhu, Rajinder, et al.
Published: (2026)
by: Sandhu, Rajinder, et al.
Published: (2026)
Vision-Guided Chunking Is All You Need: Enhancing RAG with Multimodal Document Understanding
by: Tripathi, Vishesh, et al.
Published: (2025)
by: Tripathi, Vishesh, et al.
Published: (2025)
Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge
by: Fabbri, Francesco, et al.
Published: (2025)
by: Fabbri, Francesco, et al.
Published: (2025)
Contextual Attention-Based Multimodal Fusion of LLM and CNN for Sentiment Analysis
by: Zerkouk, Meriem, et al.
Published: (2025)
by: Zerkouk, Meriem, et al.
Published: (2025)
MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
by: Ai, Qingyao, et al.
Published: (2025)
by: Ai, Qingyao, et al.
Published: (2025)
NextMem: Towards Latent Factual Memory for LLM-based Agents
by: Zhang, Zeyu, et al.
Published: (2026)
by: Zhang, Zeyu, et al.
Published: (2026)
Trust by Design: Skill Profiles for Transparent, Cost-Aware LLM Routing
by: Okamoto, Mika, et al.
Published: (2026)
by: Okamoto, Mika, et al.
Published: (2026)
Simpson's Paradox in Behavioral Curves: How Aggregation Distorts Parametric Models of User Dynamics
by: Zhou, Chao
Published: (2026)
by: Zhou, Chao
Published: (2026)
CONE: Embeddings for Complex Numerical Data Preserving Unit and Variable Semantics
by: Shrestha, Gyanendra, et al.
Published: (2026)
by: Shrestha, Gyanendra, et al.
Published: (2026)
Auto-FP: An Experimental Study of Automated Feature Preprocessing for Tabular Data
by: Qi, Danrui, et al.
Published: (2023)
by: Qi, Danrui, et al.
Published: (2023)
TutorLLM: Customizing Learning Recommendations with Knowledge Tracing and Retrieval-Augmented Generation
by: Li, Zhaoxing, et al.
Published: (2025)
by: Li, Zhaoxing, et al.
Published: (2025)
Scaling Search Relevance: Augmenting App Store Ranking with LLM-Generated Judgments
by: Christakopoulou, Evangelia, et al.
Published: (2026)
by: Christakopoulou, Evangelia, et al.
Published: (2026)
To Judge or not to Judge: Using LLM Judgements for Advertiser Keyphrase Relevance at eBay
by: Dey, Soumik, et al.
Published: (2025)
by: Dey, Soumik, et al.
Published: (2025)
ALERT: Zero-shot LLM Jailbreak Detection via Internal Discrepancy Amplification
by: Lin, Xiao, et al.
Published: (2026)
by: Lin, Xiao, et al.
Published: (2026)
RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
by: Li, Shijun, et al.
Published: (2026)
by: Li, Shijun, et al.
Published: (2026)
Similar Items
-
MAR-FL: A Communication Efficient Peer-to-Peer Federated Learning System
by: Mulitze, Felix, et al.
Published: (2025) -
MESS+: Dynamically Learned Inference-Time LLM Routing in Model Zoos with Service Level Guarantees
by: Woisetschläger, Herbert, et al.
Published: (2025) -
Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining
by: Sow, Daouda, et al.
Published: (2025) -
An Extremely Data-efficient and Generative LLM-based Reinforcement Learning Agent for Recommenders
by: Feng, Shuang, et al.
Published: (2024) -
No One Left Behind: How to Exploit the Incomplete and Skewed Multi-Label Data for Conversion Rate Prediction
by: Jia, Qinglin, et al.
Published: (2025)