Multilingual De-Duplication Strategies: Applying scalable similarity search with monolingual & multilingual embedding models
Fuente:
arXiv
Saved in:
| Main Authors: | Pasch, Stefan, Petridis, Dimitirios, Cutura, Jannic |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Applying Cognitive Design Patterns to General LLM Agents
by: Wray, Robert E., et al.
Published: (2025)
by: Wray, Robert E., et al.
Published: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
Predictive Simultaneous Interpretation: Harnessing Large Language Models for Democratizing Real-Time Multilingual Communication
by: Iida, Kurando, et al.
Published: (2024)
by: Iida, Kurando, et al.
Published: (2024)
Multilingual jailbreaking of LLMs using low-resource languages
by: Marx, Dylan, et al.
Published: (2026)
by: Marx, Dylan, et al.
Published: (2026)
mEdIT: Multilingual Text Editing via Instruction Tuning
by: Raheja, Vipul, et al.
Published: (2024)
by: Raheja, Vipul, et al.
Published: (2024)
Bielik 11B v3: Multilingual Large Language Model for European Languages
by: Ociepa, Krzysztof, et al.
Published: (2025)
by: Ociepa, Krzysztof, et al.
Published: (2025)
FLeX: Fourier-based Low-rank EXpansion for multilingual transfer
by: Narasimhan, Gaurav
Published: (2026)
by: Narasimhan, Gaurav
Published: (2026)
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
by: Hilasaca, Kenji, et al.
Published: (2026)
by: Hilasaca, Kenji, et al.
Published: (2026)
jina-embeddings-v3: Multilingual Embeddings With Task LoRA
by: Sturua, Saba, et al.
Published: (2024)
by: Sturua, Saba, et al.
Published: (2024)
jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval
by: Günther, Michael, et al.
Published: (2025)
by: Günther, Michael, et al.
Published: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
Unlocking the Wisdom of Large Language Models: An Introduction to The Path to Artificial General Intelligence
by: Chang, Edward Y.
Published: (2024)
by: Chang, Edward Y.
Published: (2024)
SciDMT: A Large-Scale Corpus for Detecting Scientific Mentions
by: Pan, Huitong, et al.
Published: (2024)
by: Pan, Huitong, et al.
Published: (2024)
Executing Arithmetic: Fine-Tuning Large Language Models as Turing Machines
by: Lai, Junyu, et al.
Published: (2024)
by: Lai, Junyu, et al.
Published: (2024)
Process Supervision-Guided Policy Optimization for Code Generation
by: Dai, Ning, et al.
Published: (2024)
by: Dai, Ning, et al.
Published: (2024)
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs
by: Rai, Daking, et al.
Published: (2024)
by: Rai, Daking, et al.
Published: (2024)
Intrinsic Evaluation of RAG Systems for Deep-Logic Questions
by: Hu, Junyi, et al.
Published: (2024)
by: Hu, Junyi, et al.
Published: (2024)
KemenkeuGPT: Leveraging a Large Language Model on Indonesia's Government Financial Data and Regulations to Enhance Decision Making
by: Febrian, Gilang Fajar, et al.
Published: (2024)
by: Febrian, Gilang Fajar, et al.
Published: (2024)
EVINCE: Optimizing Multi-LLM Dialogues Using Conditional Statistics and Information Theory
by: Chang, Edward Y.
Published: (2024)
by: Chang, Edward Y.
Published: (2024)
AI Predicts AGI: Leveraging AGI Forecasting and Peer Review to Explore LLMs' Complex Reasoning Capabilities
by: Davide, Fabrizio, et al.
Published: (2024)
by: Davide, Fabrizio, et al.
Published: (2024)
Ensuring Ground Truth Accuracy in Healthcare with the EVINCE framework
by: Chang, Edward Y.
Published: (2024)
by: Chang, Edward Y.
Published: (2024)
SOCIA-Nabla: Textual Gradient Meets Multi-Agent Orchestration for Automated Simulator Generation
by: Hua, Yuncheng, et al.
Published: (2025)
by: Hua, Yuncheng, et al.
Published: (2025)
ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning
by: Chang, Edward Y., et al.
Published: (2025)
by: Chang, Edward Y., et al.
Published: (2025)
Evaluating Steering Techniques using Human Similarity Judgments
by: Studdiford, Zach, et al.
Published: (2025)
by: Studdiford, Zach, et al.
Published: (2025)
Reasoning-Based AI for Startup Evaluation (R.A.I.S.E.): A Memory-Augmented, Multi-Step Decision Framework
by: Preuveneers, Jack, et al.
Published: (2025)
by: Preuveneers, Jack, et al.
Published: (2025)
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
by: Wang, Liang, et al.
Published: (2026)
by: Wang, Liang, et al.
Published: (2026)
RADD: Retrieval-Augmented Discrete Diffusion for Multi-Modal Knowledge Graph Completion
by: Niu, Guanglin, et al.
Published: (2026)
by: Niu, Guanglin, et al.
Published: (2026)
Quantifying Self-Preservation Bias in Large Language Models
by: Migliarini, Matteo, et al.
Published: (2026)
by: Migliarini, Matteo, et al.
Published: (2026)
Evaluating Relational Reasoning in LLMs with REL
by: Fesser, Lukas, et al.
Published: (2026)
by: Fesser, Lukas, et al.
Published: (2026)
Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms
by: Le, Linh, et al.
Published: (2026)
by: Le, Linh, et al.
Published: (2026)
KACE: Knowledge-Adaptive Context Engineering for Mathematical Reasoning
by: Parashar, Jayant, et al.
Published: (2026)
by: Parashar, Jayant, et al.
Published: (2026)
RAudit: A Blind Auditing Protocol for Large Language Model Reasoning
by: Chang, Edward Y., et al.
Published: (2026)
by: Chang, Edward Y., et al.
Published: (2026)
Towards ethical multimodal systems
by: Roger, Alexis, et al.
Published: (2023)
by: Roger, Alexis, et al.
Published: (2023)
AI-Powered Annotation Pipelines for Stabilizing Large Language Models: A Human-AI Synergy Approach
by: Pathak, Gangesh, et al.
Published: (2025)
by: Pathak, Gangesh, et al.
Published: (2025)
LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
by: Ivanov, Igor
Published: (2025)
by: Ivanov, Igor
Published: (2025)
ToolWeaver: Weaving Collaborative Semantics for Scalable Tool Use in Large Language Models
by: Fang, Bowen, et al.
Published: (2026)
by: Fang, Bowen, et al.
Published: (2026)
PersistBench: When Should Long-Term Memories Be Forgotten by LLMs?
by: Pulipaka, Sidharth, et al.
Published: (2026)
by: Pulipaka, Sidharth, et al.
Published: (2026)
Heimdall: test-time scaling on the generative verification
by: Shi, Wenlei, et al.
Published: (2025)
by: Shi, Wenlei, et al.
Published: (2025)
Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints
by: Peng, Songping, et al.
Published: (2026)
by: Peng, Songping, et al.
Published: (2026)
Similar Items
-
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025) -
Applying Cognitive Design Patterns to General LLM Agents
by: Wray, Robert E., et al.
Published: (2025) -
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025) -
Predictive Simultaneous Interpretation: Harnessing Large Language Models for Democratizing Real-Time Multilingual Communication
by: Iida, Kurando, et al.
Published: (2024) -
Multilingual jailbreaking of LLMs using low-resource languages
by: Marx, Dylan, et al.
Published: (2026)