FLUX: Data Worth Training On
Fuente:
arXiv
Saved in:
| Main Authors: | Gowtham, Rupesh, Sai, Kumar, Sanjay, Saravanan, Chaithanya, Venkata |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Blu-WERP (Web Extraction and Refinement Pipeline): A Scalable Pipeline for Preprocessing Large Language Model Datasets
by: Gowtham, et al.
Published: (2025)
by: Gowtham, et al.
Published: (2025)
Developing an AI Assistant for Knowledge Management and Workforce Training in State DOTs
by: Amaram, Divija, et al.
Published: (2026)
by: Amaram, Divija, et al.
Published: (2026)
Not All Personas Are Worth It: Culture-Reflective Persona Data Augmentation
by: Han, Ji-Eun, et al.
Published: (2025)
by: Han, Ji-Eun, et al.
Published: (2025)
Large Language Models aren't all that you need
by: Holla, Kiran Voderhobli, et al.
Published: (2024)
by: Holla, Kiran Voderhobli, et al.
Published: (2024)
Is Peer-Reviewing Worth the Effort?
by: Church, Kenneth, et al.
Published: (2024)
by: Church, Kenneth, et al.
Published: (2024)
AI Hallucinations: A Misnomer Worth Clarifying
by: Maleki, Negar, et al.
Published: (2024)
by: Maleki, Negar, et al.
Published: (2024)
Teaching Transformers Causal Reasoning through Axiomatic Training
by: Vashishtha, Aniket, et al.
Published: (2024)
by: Vashishtha, Aniket, et al.
Published: (2024)
What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions
by: Choe, Sang Keun, et al.
Published: (2024)
by: Choe, Sang Keun, et al.
Published: (2024)
Zero-Shot Confidence Estimation for Small LLMs: When Supervised Baselines Aren't Worth Training
by: Nguyen, Luong N.
Published: (2026)
by: Nguyen, Luong N.
Published: (2026)
How does a Multilingual LM Handle Multiple Languages?
by: Kakarla, Santhosh, et al.
Published: (2025)
by: Kakarla, Santhosh, et al.
Published: (2025)
Causal Order: The Key to Leveraging Imperfect Experts in Causal Inference
by: Vashishtha, Aniket, et al.
Published: (2023)
by: Vashishtha, Aniket, et al.
Published: (2023)
When AI Speaks, Whose Values Does It Express? A Cross-Cultural Audit of Individualism-Collectivism Bias in Large Language Models
by: Venkata, Pruthvinath Jeripity
Published: (2026)
by: Venkata, Pruthvinath Jeripity
Published: (2026)
An Expert is Worth One Token: Synergizing Multiple Expert LLMs as Generalist via Expert Token Routing
by: Chai, Ziwei, et al.
Published: (2024)
by: Chai, Ziwei, et al.
Published: (2024)
A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone
by: Hao, Jitai, et al.
Published: (2025)
by: Hao, Jitai, et al.
Published: (2025)
Is Escalation Worth It? A Decision-Theoretic Characterization of LLM Cascades
by: Bouchard, Dylan
Published: (2026)
by: Bouchard, Dylan
Published: (2026)
FastFLUX: Pruning FLUX with Block-wise Replacement and Sandwich Training
by: Cai, Fuhan, et al.
Published: (2025)
by: Cai, Fuhan, et al.
Published: (2025)
Turn Waste into Worth: Rectifying Top-$k$ Router of MoE
by: Zeng, Zhiyuan, et al.
Published: (2024)
by: Zeng, Zhiyuan, et al.
Published: (2024)
Programming by Backprop: An Instruction is Worth 100 Examples When Finetuning LLMs
by: Cook, Jonathan, et al.
Published: (2025)
by: Cook, Jonathan, et al.
Published: (2025)
Rewarding Intellectual Humility Learning When Not To Answer In Large Language Models
by: Jha, Abha, et al.
Published: (2026)
by: Jha, Abha, et al.
Published: (2026)
Data-Constrained Synthesis of Training Data for De-Identification
by: Vakili, Thomas, et al.
Published: (2025)
by: Vakili, Thomas, et al.
Published: (2025)
Detecting and Mitigating Bias in LLMs through Knowledge Graph-Augmented Training
by: Kumar, Rajeev, et al.
Published: (2025)
by: Kumar, Rajeev, et al.
Published: (2025)
Unifying Structured Data as Graph for Data-to-Text Pre-Training
by: Li, Shujie, et al.
Published: (2024)
by: Li, Shujie, et al.
Published: (2024)
Improving the Reliability of LLMs: Combining CoT, RAG, Self-Consistency, and Self-Verification
by: Kumar, Adarsh, et al.
Published: (2025)
by: Kumar, Adarsh, et al.
Published: (2025)
PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts
by: An, Bang, et al.
Published: (2023)
by: An, Bang, et al.
Published: (2023)
The Chameleon Nature of LLMs: Quantifying Multi-Turn Stance Instability in Search-Enabled Language Models
by: Ratnakar, Shivam, et al.
Published: (2025)
by: Ratnakar, Shivam, et al.
Published: (2025)
Training with Pseudo-Code for Instruction Following
by: Kumar, Prince, et al.
Published: (2025)
by: Kumar, Prince, et al.
Published: (2025)
A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding
by: Lu, Jinghui, et al.
Published: (2024)
by: Lu, Jinghui, et al.
Published: (2024)
Unlearning Traces the Influential Training Data of Language Models
by: Isonuma, Masaru, et al.
Published: (2024)
by: Isonuma, Masaru, et al.
Published: (2024)
Balanced Data Sampling for Language Model Training with Clustering
by: Shao, Yunfan, et al.
Published: (2024)
by: Shao, Yunfan, et al.
Published: (2024)
Beyond Public Access in LLM Pre-Training Data
by: Rosenblat, Sruly, et al.
Published: (2025)
by: Rosenblat, Sruly, et al.
Published: (2025)
The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories
by: Shah, Raj Sanjay, et al.
Published: (2025)
by: Shah, Raj Sanjay, et al.
Published: (2025)
Closing the Data Loop: Using OpenDataArena to Engineer Superior Training Datasets
by: Gao, Xin, et al.
Published: (2025)
by: Gao, Xin, et al.
Published: (2025)
Detecting RLVR Training Data via Structural Convergence of Reasoning
by: Zhang, Hongbo, et al.
Published: (2026)
by: Zhang, Hongbo, et al.
Published: (2026)
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
by: Ge, Albert, et al.
Published: (2025)
by: Ge, Albert, et al.
Published: (2025)
SlimPajama-DC: Understanding Data Combinations for LLM Training
by: Shen, Zhiqiang, et al.
Published: (2023)
by: Shen, Zhiqiang, et al.
Published: (2023)
Data Management For Training Large Language Models: A Survey
by: Wang, Zige, et al.
Published: (2023)
by: Wang, Zige, et al.
Published: (2023)
Sentiment Analysis of Cyberbullying Data in Social Media
by: Susmitha, Arvapalli Sai, et al.
Published: (2024)
by: Susmitha, Arvapalli Sai, et al.
Published: (2024)
Can LLMs Detect Their Confabulations? Estimating Reliability in Uncertainty-Aware Language Models
by: Zhou, Tianyi, et al.
Published: (2025)
by: Zhou, Tianyi, et al.
Published: (2025)
READ: Reinforcement-based Adversarial Learning for Text Classification with Limited Labeled Data
by: Sharma, Rohit, et al.
Published: (2025)
by: Sharma, Rohit, et al.
Published: (2025)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
by: Mirtaheri, Parsa, et al.
Published: (2025)
by: Mirtaheri, Parsa, et al.
Published: (2025)
Similar Items
-
Blu-WERP (Web Extraction and Refinement Pipeline): A Scalable Pipeline for Preprocessing Large Language Model Datasets
by: Gowtham, et al.
Published: (2025) -
Developing an AI Assistant for Knowledge Management and Workforce Training in State DOTs
by: Amaram, Divija, et al.
Published: (2026) -
Not All Personas Are Worth It: Culture-Reflective Persona Data Augmentation
by: Han, Ji-Eun, et al.
Published: (2025) -
Large Language Models aren't all that you need
by: Holla, Kiran Voderhobli, et al.
Published: (2024) -
Is Peer-Reviewing Worth the Effort?
by: Church, Kenneth, et al.
Published: (2024)