ClusterFusion: Hybrid Clustering with Embedding Guidance and LLM Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Yiming, Yuan, Yuan, Viswanathan, Vijay, Neubig, Graham |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ClusterFusion++: Expanding Cluster-Level Fusion to Full Transformer-Block Decoding
by: Jin, ChiHeng, et al.
Published: (2026)
by: Jin, ChiHeng, et al.
Published: (2026)
Better Synthetic Data by Retrieving and Transforming Existing Datasets
by: Gandhi, Saumya, et al.
Published: (2024)
by: Gandhi, Saumya, et al.
Published: (2024)
Training Task Experts through Retrieval Based Distillation
by: Ge, Jiaxin, et al.
Published: (2024)
by: Ge, Jiaxin, et al.
Published: (2024)
ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive
by: Luo, Xinhao, et al.
Published: (2025)
by: Luo, Xinhao, et al.
Published: (2025)
SELF-GUIDE: Better Task-Specific Instruction Following via Self-Synthetic Finetuning
by: Zhao, Chenyang, et al.
Published: (2024)
by: Zhao, Chenyang, et al.
Published: (2024)
Checklists Are Better Than Reward Models For Aligning Language Models
by: Viswanathan, Vijay, et al.
Published: (2025)
by: Viswanathan, Vijay, et al.
Published: (2025)
Synthetic Multimodal Question Generation
by: Wu, Ian, et al.
Published: (2024)
by: Wu, Ian, et al.
Published: (2024)
An Improved Deep Learning Model for Word Embeddings Based Clustering for Large Text Datasets
by: Sutrakar, Vijay Kumar, et al.
Published: (2025)
by: Sutrakar, Vijay Kumar, et al.
Published: (2025)
Go-Browse: Training Web Agents with Structured Exploration
by: Gandhi, Apurva, et al.
Published: (2025)
by: Gandhi, Apurva, et al.
Published: (2025)
BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language Models
by: Tjuatja, Lindia, et al.
Published: (2025)
by: Tjuatja, Lindia, et al.
Published: (2025)
Evaluating Language Models as Synthetic Data Generators
by: Kim, Seungone, et al.
Published: (2024)
by: Kim, Seungone, et al.
Published: (2024)
Effective Strategies for Asynchronous Software Engineering Agents
by: Geng, Jiayi, et al.
Published: (2026)
by: Geng, Jiayi, et al.
Published: (2026)
Scalable Parameter-Light Spectral Method for Clustering Short Text Embeddings with a Cohesion-Based Evaluation Metric
by: Neveditsin, Nikita, et al.
Published: (2025)
by: Neveditsin, Nikita, et al.
Published: (2025)
Repetition Improves Language Model Embeddings
by: Springer, Jacob Mitchell, et al.
Published: (2024)
by: Springer, Jacob Mitchell, et al.
Published: (2024)
LLM-MemCluster: Empowering Large Language Models with Dynamic Memory for Text Clustering
by: Zhu, Yuanjie, et al.
Published: (2025)
by: Zhu, Yuanjie, et al.
Published: (2025)
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
by: Zhang, Charlie, et al.
Published: (2025)
by: Zhang, Charlie, et al.
Published: (2025)
Midtraining Bridges Pretraining and Posttraining Distributions
by: Liu, Emmy, et al.
Published: (2025)
by: Liu, Emmy, et al.
Published: (2025)
An Incomplete Loop: Instruction Inference, Instruction Following, and In-context Learning in Language Models
by: Liu, Emmy, et al.
Published: (2024)
by: Liu, Emmy, et al.
Published: (2024)
What Is Missing in Multilingual Visual Reasoning and How to Fix It
by: Song, Yueqi, et al.
Published: (2024)
by: Song, Yueqi, et al.
Published: (2024)
Solving NLP Problems through Human-System Collaboration: A Discussion-based Approach
by: Kaneko, Masahiro, et al.
Published: (2023)
by: Kaneko, Masahiro, et al.
Published: (2023)
Hybrid Topic-Semantic Labeling and Graph Embeddings for Unsupervised Legal Document Clustering
by: Bastola, Deepak, et al.
Published: (2025)
by: Bastola, Deepak, et al.
Published: (2025)
Efficient Long-Context LLM Inference via KV Cache Clustering
by: Hu, Jie, et al.
Published: (2025)
by: Hu, Jie, et al.
Published: (2025)
Training Proactive and Personalized LLM Agents
by: Sun, Weiwei, et al.
Published: (2025)
by: Sun, Weiwei, et al.
Published: (2025)
Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks
by: Vijayvargiya, Sanidhya, et al.
Published: (2026)
by: Vijayvargiya, Sanidhya, et al.
Published: (2026)
SDEC: Semantic Deep Embedded Clustering
by: Rahman, Mohammad Wali Ur, et al.
Published: (2025)
by: Rahman, Mohammad Wali Ur, et al.
Published: (2025)
German Text Embedding Clustering Benchmark
by: Wehrli, Silvan, et al.
Published: (2024)
by: Wehrli, Silvan, et al.
Published: (2024)
Detecting LLM Hallucinations via Embedding Cluster Geometry: A Three-Type Taxonomy with Measurable Signatures
by: Korun, Matic
Published: (2026)
by: Korun, Matic
Published: (2026)
Beyond Browsing: API-Based Web Agents
by: Song, Yueqi, et al.
Published: (2024)
by: Song, Yueqi, et al.
Published: (2024)
Normalization through Fine-tuning: Understanding Wav2vec 2.0 Embeddings for Phonetic Analysis
by: Wang, Yiming, et al.
Published: (2025)
by: Wang, Yiming, et al.
Published: (2025)
RAGGED: Towards Informed Design of Scalable and Stable RAG Systems
by: Hsia, Jennifer, et al.
Published: (2024)
by: Hsia, Jennifer, et al.
Published: (2024)
What Goes Into a LM Acceptability Judgment? Rethinking the Impact of Frequency and Length
by: Tjuatja, Lindia, et al.
Published: (2024)
by: Tjuatja, Lindia, et al.
Published: (2024)
Adaptation of Embedding Models to Financial Filings via LLM Distillation
by: Brenner, Eliot, et al.
Published: (2025)
by: Brenner, Eliot, et al.
Published: (2025)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
by: Huan, Maggie, et al.
Published: (2025)
by: Huan, Maggie, et al.
Published: (2025)
Training Versatile Coding Agents in Synthetic Environments
by: Zhu, Yiqi, et al.
Published: (2025)
by: Zhu, Yiqi, et al.
Published: (2025)
Rethinking LLM Language Adaptation: A Case Study on Chinese Mixtral
by: Cui, Yiming, et al.
Published: (2024)
by: Cui, Yiming, et al.
Published: (2024)
Inducing Programmatic Skills for Agentic Tasks
by: Wang, Zora Zhiruo, et al.
Published: (2025)
by: Wang, Zora Zhiruo, et al.
Published: (2025)
Agent Workflow Memory
by: Wang, Zora Zhiruo, et al.
Published: (2024)
by: Wang, Zora Zhiruo, et al.
Published: (2024)
Mistake Notebook Learning: Batch-Clustered Failures for Training-Free Agent Adaptation
by: Su, Xuanbo, et al.
Published: (2025)
by: Su, Xuanbo, et al.
Published: (2025)
ChameleonLLM: Batch-Aware Dynamic Low-Rank Adaptation via Inference-Time Clusters
by: Yuksel, Kamer Ali, et al.
Published: (2025)
by: Yuksel, Kamer Ali, et al.
Published: (2025)
Classification and Clustering of Sentence-Level Embeddings of Scientific Articles Generated by Contrastive Learning
by: Guedes, Gustavo Bartz, et al.
Published: (2024)
by: Guedes, Gustavo Bartz, et al.
Published: (2024)
Similar Items
-
ClusterFusion++: Expanding Cluster-Level Fusion to Full Transformer-Block Decoding
by: Jin, ChiHeng, et al.
Published: (2026) -
Better Synthetic Data by Retrieving and Transforming Existing Datasets
by: Gandhi, Saumya, et al.
Published: (2024) -
Training Task Experts through Retrieval Based Distillation
by: Ge, Jiaxin, et al.
Published: (2024) -
ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive
by: Luo, Xinhao, et al.
Published: (2025) -
SELF-GUIDE: Better Task-Specific Instruction Following via Self-Synthetic Finetuning
by: Zhao, Chenyang, et al.
Published: (2024)