Text Clustering with Large Language Model Embeddings
Fuente:
arXiv
Salvato in:
| Autori principali: | Petukhova, Alina, Matos-Carvalho, João P., Fachada, Nuno |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Whisper-LM: Improving ASR Models with Language Models for Low-Resource Languages
di: de Zuazo, Xabier, et al.
Pubblicazione: (2025)
di: de Zuazo, Xabier, et al.
Pubblicazione: (2025)
Rule Extraction in Machine Learning: Chat Incremental Pattern Constructor
di: Nwokocha, Caleb Princewill
Pubblicazione: (2022)
di: Nwokocha, Caleb Princewill
Pubblicazione: (2022)
Do Reasoning Models Enhance Embedding Models?
di: Chan, Wun Yu, et al.
Pubblicazione: (2026)
di: Chan, Wun Yu, et al.
Pubblicazione: (2026)
Linguistic Collapse: Neural Collapse in (Large) Language Models
di: Wu, Robert, et al.
Pubblicazione: (2024)
di: Wu, Robert, et al.
Pubblicazione: (2024)
Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
di: Gupta, Aayush
Pubblicazione: (2025)
di: Gupta, Aayush
Pubblicazione: (2025)
Research on a hybrid LSTM-CNN-Attention model for text-based web content classification
di: Kuz, Mykola, et al.
Pubblicazione: (2025)
di: Kuz, Mykola, et al.
Pubblicazione: (2025)
LLM-supported document separation for printed reviews from zbMATH Open
di: Pluzhnikov, Ivan, et al.
Pubblicazione: (2026)
di: Pluzhnikov, Ivan, et al.
Pubblicazione: (2026)
Fine-tuning of Large Language Models for Constituency Parsing Using a Sequence to Sequence Approach
di: Delgado, Francisco Jose Cortes, et al.
Pubblicazione: (2025)
di: Delgado, Francisco Jose Cortes, et al.
Pubblicazione: (2025)
How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework
di: Nieth, Björn, et al.
Pubblicazione: (2026)
di: Nieth, Björn, et al.
Pubblicazione: (2026)
Approaches to Semantic Textual Similarity in Slovak Language: From Algorithms to Transformers
di: Radosky, Lukas, et al.
Pubblicazione: (2026)
di: Radosky, Lukas, et al.
Pubblicazione: (2026)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
di: Kim, Heejun, et al.
Pubblicazione: (2026)
di: Kim, Heejun, et al.
Pubblicazione: (2026)
DeepSeek-V3, GPT-4, Phi-4, and LLaMA-3.3 generate correct code for LoRaWAN-related engineering tasks
di: Fernandes, Daniel, et al.
Pubblicazione: (2025)
di: Fernandes, Daniel, et al.
Pubblicazione: (2025)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
di: Breneur, Oleksandr Marchenko, et al.
Pubblicazione: (2026)
di: Breneur, Oleksandr Marchenko, et al.
Pubblicazione: (2026)
Mapping the Web of Science, a large-scale graph and text-based dataset with LLM embeddings
di: Kunt, Tim, et al.
Pubblicazione: (2026)
di: Kunt, Tim, et al.
Pubblicazione: (2026)
IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following
di: Sun, Mingrui, et al.
Pubblicazione: (2026)
di: Sun, Mingrui, et al.
Pubblicazione: (2026)
Doğal Dil İşlemede Tokenizasyon Standartları ve Ölçümü: Türkçe Üzerinden Büyük Dil Modellerinin Karşılaştırmalı Analizi
di: Bayram, M. Ali, et al.
Pubblicazione: (2025)
di: Bayram, M. Ali, et al.
Pubblicazione: (2025)
Büyük Dil Modelleri için TR-MMLU Benchmarkı: Performans Değerlendirmesi, Zorluklar ve İyileştirme Fırsatları
di: Bayram, M. Ali, et al.
Pubblicazione: (2025)
di: Bayram, M. Ali, et al.
Pubblicazione: (2025)
Targeted Lexical Injection: Unlocking Latent Cross-Lingual Alignment in Lugha-Llama via Early-Layer LoRA Fine-Tuning
di: Ngugi, Stanley
Pubblicazione: (2025)
di: Ngugi, Stanley
Pubblicazione: (2025)
CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
di: Badshah, Sher, et al.
Pubblicazione: (2025)
di: Badshah, Sher, et al.
Pubblicazione: (2025)
MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition
di: Abtahi, Farhad, et al.
Pubblicazione: (2026)
di: Abtahi, Farhad, et al.
Pubblicazione: (2026)
Claim Automation using Large Language Model
di: Mo, Zhengda, et al.
Pubblicazione: (2026)
di: Mo, Zhengda, et al.
Pubblicazione: (2026)
Optimizing Retrieval-Augmented Generation (RAG) for Colloquial Cantonese: A LoRA-Based Systematic Review
di: Calonge, David Santandreu, et al.
Pubblicazione: (2025)
di: Calonge, David Santandreu, et al.
Pubblicazione: (2025)
The Syntactic Acceptability Dataset (Preview): A Resource for Machine Learning and Linguistic Analysis of English
di: Juzek, Tom S
Pubblicazione: (2025)
di: Juzek, Tom S
Pubblicazione: (2025)
Tokens with Meaning: A Hybrid Tokenization Approach for Turkish
di: Bayram, M. Ali, et al.
Pubblicazione: (2025)
di: Bayram, M. Ali, et al.
Pubblicazione: (2025)
EdgeJury: Cross-Reviewed Small-Model Ensembles for Truthful Question Answering on Serverless Edge Inference
di: Kumar, Aayush
Pubblicazione: (2025)
di: Kumar, Aayush
Pubblicazione: (2025)
Beyond Subtokens: A Rich Character Embedding for Low-resource and Morphologically Complex Languages
di: Schneider, Felix, et al.
Pubblicazione: (2026)
di: Schneider, Felix, et al.
Pubblicazione: (2026)
Semantic Retention and Extreme Compression in LLMs: Can We Have Both?
di: Laborde, Stanislas, et al.
Pubblicazione: (2025)
di: Laborde, Stanislas, et al.
Pubblicazione: (2025)
Smotrom tvoja pa ander drogoj verden! Resurrecting Dead Pidgin with Generative Models: Russenorsk Case Study
di: Tikhonov, Alexey, et al.
Pubblicazione: (2025)
di: Tikhonov, Alexey, et al.
Pubblicazione: (2025)
Conformal Prediction Sets for Next-Token Prediction in Large Language Models: Balancing Coverage Guarantees with Set Efficiency
di: Kotla, Yoshith Roy, et al.
Pubblicazione: (2025)
di: Kotla, Yoshith Roy, et al.
Pubblicazione: (2025)
ProactBench: Beyond What The User Asked For
di: Harfi, Sepehr, et al.
Pubblicazione: (2026)
di: Harfi, Sepehr, et al.
Pubblicazione: (2026)
How much do LLMs learn from negative examples?
di: Hamdan, Shadi, et al.
Pubblicazione: (2025)
di: Hamdan, Shadi, et al.
Pubblicazione: (2025)
Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
di: Fan, Jingxing, et al.
Pubblicazione: (2025)
di: Fan, Jingxing, et al.
Pubblicazione: (2025)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
di: Basu, Abhinaba
Pubblicazione: (2026)
di: Basu, Abhinaba
Pubblicazione: (2026)
Rethinking the Multilingual Reasoning Gap with Layer Swap
di: Lasbordes, Maxence, et al.
Pubblicazione: (2026)
di: Lasbordes, Maxence, et al.
Pubblicazione: (2026)
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark
di: Bayram, M. Ali, et al.
Pubblicazione: (2025)
di: Bayram, M. Ali, et al.
Pubblicazione: (2025)
Advancements in Machine Learning and Deep Learning for Early Detection and Management of Mental Health Disorder
di: Kannan, Kamala Devi, et al.
Pubblicazione: (2024)
di: Kannan, Kamala Devi, et al.
Pubblicazione: (2024)
Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback
di: Juzek, Tom S., et al.
Pubblicazione: (2025)
di: Juzek, Tom S., et al.
Pubblicazione: (2025)
Towards Ontology-Enhanced Representation Learning for Large Language Models
di: Ronzano, Francesco, et al.
Pubblicazione: (2024)
di: Ronzano, Francesco, et al.
Pubblicazione: (2024)
Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis
di: Zanbaghi, Shahin, et al.
Pubblicazione: (2025)
di: Zanbaghi, Shahin, et al.
Pubblicazione: (2025)
Constitution or Collapse? Exploring Constitutional AI with Llama 3-8B
di: Zhang, Xue
Pubblicazione: (2025)
di: Zhang, Xue
Pubblicazione: (2025)
Documenti analoghi
-
Whisper-LM: Improving ASR Models with Language Models for Low-Resource Languages
di: de Zuazo, Xabier, et al.
Pubblicazione: (2025) -
Rule Extraction in Machine Learning: Chat Incremental Pattern Constructor
di: Nwokocha, Caleb Princewill
Pubblicazione: (2022) -
Do Reasoning Models Enhance Embedding Models?
di: Chan, Wun Yu, et al.
Pubblicazione: (2026) -
Linguistic Collapse: Neural Collapse in (Large) Language Models
di: Wu, Robert, et al.
Pubblicazione: (2024) -
Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
di: Gupta, Aayush
Pubblicazione: (2025)