Generative Deduplication For Socia Media Data Selection
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Xianming, Li, Jing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
BeLLM: Backward Dependency Enhanced Large Language Model for Sentence Embeddings
di: Li, Xianming, et al.
Pubblicazione: (2023)
di: Li, Xianming, et al.
Pubblicazione: (2023)
AnglE-optimized Text Embeddings
di: Li, Xianming, et al.
Pubblicazione: (2023)
di: Li, Xianming, et al.
Pubblicazione: (2023)
SEDD: Scalable and Efficient Dataset Deduplication with GPUs
di: Son, Youngjun, et al.
Pubblicazione: (2025)
di: Son, Youngjun, et al.
Pubblicazione: (2025)
Deduplicating and Ranking Solution Programs for Suggesting Reference Solutions
di: Shirafuji, Atsushi, et al.
Pubblicazione: (2023)
di: Shirafuji, Atsushi, et al.
Pubblicazione: (2023)
2D Matryoshka Sentence Embeddings
di: Li, Xianming, et al.
Pubblicazione: (2024)
di: Li, Xianming, et al.
Pubblicazione: (2024)
Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks
di: Schelpe, Sietse
Pubblicazione: (2026)
di: Schelpe, Sietse
Pubblicazione: (2026)
BaichuanSEED: Sharing the Potential of ExtensivE Data Collection and Deduplication by Introducing a Competitive Large Language Model Baseline
di: Dong, Guosheng, et al.
Pubblicazione: (2024)
di: Dong, Guosheng, et al.
Pubblicazione: (2024)
Merlin: Deterministic Byte-Exact Deduplication for Lossless Context Optimization in Large Language Model Inference
di: Schelpe, Sietse
Pubblicazione: (2026)
di: Schelpe, Sietse
Pubblicazione: (2026)
Privacy-Preserving Data Deduplication for Enhancing Federated Learning of Language Models (Extended Version)
di: Abadi, Aydin, et al.
Pubblicazione: (2024)
di: Abadi, Aydin, et al.
Pubblicazione: (2024)
UniPoll: A Unified Social Media Poll Generation Framework via Multi-Objective Optimization
di: Li, Yixia, et al.
Pubblicazione: (2023)
di: Li, Yixia, et al.
Pubblicazione: (2023)
ShadowPEFT: Shadow Network for Parameter-Efficient Fine-Tuning
di: Li, Xianming, et al.
Pubblicazione: (2026)
di: Li, Xianming, et al.
Pubblicazione: (2026)
Evaluating Deduplication Techniques for Economic Research Paper Titles with a Focus on Semantic Similarity using NLP and LLMs
di: You, Doohee, et al.
Pubblicazione: (2024)
di: You, Doohee, et al.
Pubblicazione: (2024)
PopALM: Popularity-Aligned Language Models for Social Media Trendy Response Prediction
di: Yu, Erxin, et al.
Pubblicazione: (2024)
di: Yu, Erxin, et al.
Pubblicazione: (2024)
STARE at the Structure: Steering ICL Exemplar Selection with Structural Alignment
di: Li, Jiaqian, et al.
Pubblicazione: (2025)
di: Li, Jiaqian, et al.
Pubblicazione: (2025)
From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning
di: Li, Ming, et al.
Pubblicazione: (2023)
di: Li, Ming, et al.
Pubblicazione: (2023)
Unified Data Selection for LLM Reasoning
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
ProRank: Prompt Warmup via Reinforcement Learning for Small Language Models Reranking
di: Li, Xianming, et al.
Pubblicazione: (2025)
di: Li, Xianming, et al.
Pubblicazione: (2025)
Entropy-Based Data Selection for Language Models
di: Li, Hongming, et al.
Pubblicazione: (2026)
di: Li, Hongming, et al.
Pubblicazione: (2026)
Instruction Data Selection via Answer Divergence
di: Li, Bo, et al.
Pubblicazione: (2026)
di: Li, Bo, et al.
Pubblicazione: (2026)
How Does Knowledge Selection Help Retrieval Augmented Generation?
di: Li, Xiangci, et al.
Pubblicazione: (2024)
di: Li, Xiangci, et al.
Pubblicazione: (2024)
OASIS: Order-Augmented Strategy for Improved Code Search
di: Gao, Zuchen, et al.
Pubblicazione: (2025)
di: Gao, Zuchen, et al.
Pubblicazione: (2025)
Selective Weak-to-Strong Generalization
di: Lang, Hao, et al.
Pubblicazione: (2025)
di: Lang, Hao, et al.
Pubblicazione: (2025)
CAP: Data Contamination Detection via Consistency Amplification
di: Zhao, Yi, et al.
Pubblicazione: (2024)
di: Zhao, Yi, et al.
Pubblicazione: (2024)
Data Selection via Optimal Control for Language Models
di: Gu, Yuxian, et al.
Pubblicazione: (2024)
di: Gu, Yuxian, et al.
Pubblicazione: (2024)
NaturalThoughts: Selecting and Distilling Reasoning Traces for General Reasoning Tasks
di: Li, Yang, et al.
Pubblicazione: (2025)
di: Li, Yang, et al.
Pubblicazione: (2025)
Showing LLM-Generated Code Selectively Based on Confidence of LLMs
di: Li, Jia, et al.
Pubblicazione: (2024)
di: Li, Jia, et al.
Pubblicazione: (2024)
Towards Universal Debiasing for Language Models-based Tabular Data Generation
di: Li, Tianchun, et al.
Pubblicazione: (2025)
di: Li, Tianchun, et al.
Pubblicazione: (2025)
HICL: Hashtag-Driven In-Context Learning for Social Media Natural Language Understanding
di: Tan, Hanzhuo, et al.
Pubblicazione: (2023)
di: Tan, Hanzhuo, et al.
Pubblicazione: (2023)
IndiVec: An Exploration of Leveraging Large Language Models for Media Bias Detection with Fine-Grained Bias Indicators
di: Lin, Luyang, et al.
Pubblicazione: (2024)
di: Lin, Luyang, et al.
Pubblicazione: (2024)
Two Directions for Clinical Data Generation with Large Language Models: Data-to-Label and Label-to-Data
di: Li, Rumeng, et al.
Pubblicazione: (2023)
di: Li, Rumeng, et al.
Pubblicazione: (2023)
LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection
di: Wu, Jian, et al.
Pubblicazione: (2025)
di: Wu, Jian, et al.
Pubblicazione: (2025)
Data Selection for Multi-turn Dialogue Instruction Tuning
di: Li, Bo, et al.
Pubblicazione: (2026)
di: Li, Bo, et al.
Pubblicazione: (2026)
Investigating Chain-of-thought with ChatGPT for Stance Detection on Social Media
di: Zhang, Bowen, et al.
Pubblicazione: (2023)
di: Zhang, Bowen, et al.
Pubblicazione: (2023)
Picking the Cream of the Crop: Visual-Centric Data Selection with Collaborative Agents
di: Liu, Zhenyu, et al.
Pubblicazione: (2025)
di: Liu, Zhenyu, et al.
Pubblicazione: (2025)
ToReMi: Topic-Aware Data Reweighting for Dynamic Pre-Training Data Selection
di: Zhu, Xiaoxuan, et al.
Pubblicazione: (2025)
di: Zhu, Xiaoxuan, et al.
Pubblicazione: (2025)
Data Whisperer: Efficient Data Selection for Task-Specific LLM Fine-Tuning via Few-Shot In-Context Learning
di: Wang, Shaobo, et al.
Pubblicazione: (2025)
di: Wang, Shaobo, et al.
Pubblicazione: (2025)
CrowdSelect: Synthetic Instruction Data Selection with Multi-LLM Wisdom
di: Li, Yisen, et al.
Pubblicazione: (2025)
di: Li, Yisen, et al.
Pubblicazione: (2025)
Predictive Data Selection: The Data That Predicts Is the Data That Teaches
di: Shum, Kashun, et al.
Pubblicazione: (2025)
di: Shum, Kashun, et al.
Pubblicazione: (2025)
FairDeDup: Detecting and Mitigating Vision-Language Fairness Disparities in Semantic Dataset Deduplication
di: Slyman, Eric, et al.
Pubblicazione: (2024)
di: Slyman, Eric, et al.
Pubblicazione: (2024)
Importance-Aware Data Selection for Efficient LLM Instruction Tuning
di: Jiang, Tingyu, et al.
Pubblicazione: (2025)
di: Jiang, Tingyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
BeLLM: Backward Dependency Enhanced Large Language Model for Sentence Embeddings
di: Li, Xianming, et al.
Pubblicazione: (2023) -
AnglE-optimized Text Embeddings
di: Li, Xianming, et al.
Pubblicazione: (2023) -
SEDD: Scalable and Efficient Dataset Deduplication with GPUs
di: Son, Youngjun, et al.
Pubblicazione: (2025) -
Deduplicating and Ranking Solution Programs for Suggesting Reference Solutions
di: Shirafuji, Atsushi, et al.
Pubblicazione: (2023) -
2D Matryoshka Sentence Embeddings
di: Li, Xianming, et al.
Pubblicazione: (2024)