Getting Your Indices in a Row: Full-Text Search for LLM Training Data for Real World
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Marinas, Ines Altemir, Kucherenko, Anastasiia, Sternfeld, Alexander, Kucharavy, Andrei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Going over Fine Web with a Fine-Tooth Comb: Technical Report of Indexing Fine Web for Problematic Content Search and Retrieval
von: Marinas, Inés Altemir, et al.
Veröffentlicht: (2025)
von: Marinas, Inés Altemir, et al.
Veröffentlicht: (2025)
Low-Perplexity LLM-Generated Sequences and Where To Find Them
von: Wuhrmann, Arthur, et al.
Veröffentlicht: (2025)
von: Wuhrmann, Arthur, et al.
Veröffentlicht: (2025)
TypePilot: Leveraging the Scala Type System for Secure LLM-generated Code
von: Sternfeld, Alexander, et al.
Veröffentlicht: (2025)
von: Sternfeld, Alexander, et al.
Veröffentlicht: (2025)
From Model to Breach: Towards Actionable LLM-Generated Vulnerabilities Reporting
von: Vallez, Cyril, et al.
Veröffentlicht: (2025)
von: Vallez, Cyril, et al.
Veröffentlicht: (2025)
Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs
von: Sternfeld, Alexander, et al.
Veröffentlicht: (2026)
von: Sternfeld, Alexander, et al.
Veröffentlicht: (2026)
Monitoring Transformative Technological Convergence Through LLM-Extracted Semantic Entity Triple Graphs
von: Sternfeld, Alexander, et al.
Veröffentlicht: (2025)
von: Sternfeld, Alexander, et al.
Veröffentlicht: (2025)
LLM Detectors Still Fall Short of Real World: Case of LLM-Generated Short News-Like Posts
von: Gameiro, Henrique Da Silva, et al.
Veröffentlicht: (2024)
von: Gameiro, Henrique Da Silva, et al.
Veröffentlicht: (2024)
How to Get Your LLM to Generate Challenging Problems for Evaluation
von: Patel, Arkil, et al.
Veröffentlicht: (2025)
von: Patel, Arkil, et al.
Veröffentlicht: (2025)
DetectRL-X: Towards Reliable Multilingual and Real-World LLM-Generated Text Detection
von: Wu, Junchao, et al.
Veröffentlicht: (2026)
von: Wu, Junchao, et al.
Veröffentlicht: (2026)
Don't Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration Pitfalls
von: Wang, Ante, et al.
Veröffentlicht: (2025)
von: Wang, Ante, et al.
Veröffentlicht: (2025)
DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
Evaluating LLM Alignment on Personality Inference from Real-World Interview Data
von: Zhu, Jianfeng, et al.
Veröffentlicht: (2025)
von: Zhu, Jianfeng, et al.
Veröffentlicht: (2025)
WIKIGENBENCH: Exploring Full-length Wikipedia Generation under Real-World Scenario
von: Zhang, Jiebin, et al.
Veröffentlicht: (2024)
von: Zhang, Jiebin, et al.
Veröffentlicht: (2024)
LLM-to-Speech: A Synthetic Data Pipeline for Training Dialectal Text-to-Speech Models
von: Khamis, Ahmed Khaled, et al.
Veröffentlicht: (2026)
von: Khamis, Ahmed Khaled, et al.
Veröffentlicht: (2026)
Who Gets the Reward, Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents
von: Yang, Chih-Hsuan, et al.
Veröffentlicht: (2025)
von: Yang, Chih-Hsuan, et al.
Veröffentlicht: (2025)
Retracing the Past: LLMs Emit Training Data When They Get Lost
von: Ko, Myeongseob, et al.
Veröffentlicht: (2025)
von: Ko, Myeongseob, et al.
Veröffentlicht: (2025)
SearchLLM: Detecting LLM Paraphrased Text by Measuring the Similarity with Regeneration of the Candidate Source via Search Engine
von: Nguyen-Son, Hoang-Quoc, et al.
Veröffentlicht: (2026)
von: Nguyen-Son, Hoang-Quoc, et al.
Veröffentlicht: (2026)
Perturb Your Data: Paraphrase-Guided Training Data Watermarking
von: Shetty, Pranav, et al.
Veröffentlicht: (2025)
von: Shetty, Pranav, et al.
Veröffentlicht: (2025)
Getting More Juice Out of Your Data: Hard Pair Refinement Enhances Visual-Language Models Without Extra Data
von: Wang, Haonan, et al.
Veröffentlicht: (2023)
von: Wang, Haonan, et al.
Veröffentlicht: (2023)
How to Train Your Latent Diffusion Language Model Jointly With the Latent Space
von: Meshchaninov, Viacheslav, et al.
Veröffentlicht: (2026)
von: Meshchaninov, Viacheslav, et al.
Veröffentlicht: (2026)
How to Train Your Deep Research Agent? Prompt, Reward, and Policy Optimization in Search-R1
von: Xu, Yinuo, et al.
Veröffentlicht: (2026)
von: Xu, Yinuo, et al.
Veröffentlicht: (2026)
Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
von: Islam, Tunazzina
Veröffentlicht: (2026)
von: Islam, Tunazzina
Veröffentlicht: (2026)
Crown, Frame, Reverse: Layer-Wise Scaling Variants for LLM Pre-Training
von: Baroian, Andrei, et al.
Veröffentlicht: (2025)
von: Baroian, Andrei, et al.
Veröffentlicht: (2025)
Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations
von: Acquaye, Christabel, et al.
Veröffentlicht: (2026)
von: Acquaye, Christabel, et al.
Veröffentlicht: (2026)
Rep2Text: Decoding Full Text from a Single LLM Token Representation
von: Zhao, Haiyan, et al.
Veröffentlicht: (2025)
von: Zhao, Haiyan, et al.
Veröffentlicht: (2025)
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios
von: Meng, Jinxiang, et al.
Veröffentlicht: (2026)
von: Meng, Jinxiang, et al.
Veröffentlicht: (2026)
How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?
von: Papi, Sara, et al.
Veröffentlicht: (2024)
von: Papi, Sara, et al.
Veröffentlicht: (2024)
Using Natural Language Processing to find Indication for Burnout with Text Classification: From Online Data to Real-World Data
von: Kurpicz-Briki, Mascha, et al.
Veröffentlicht: (2024)
von: Kurpicz-Briki, Mascha, et al.
Veröffentlicht: (2024)
Small Language Models in the Real World: Insights from Industrial Text Classification
von: Li, Lujun, et al.
Veröffentlicht: (2025)
von: Li, Lujun, et al.
Veröffentlicht: (2025)
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
von: Wang, Kun, et al.
Veröffentlicht: (2025)
von: Wang, Kun, et al.
Veröffentlicht: (2025)
Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review
von: Yu, Sungduk, et al.
Veröffentlicht: (2025)
von: Yu, Sungduk, et al.
Veröffentlicht: (2025)
Is Your Paper Being Reviewed by an LLM? Investigating AI Text Detectability in Peer Review
von: Yu, Sungduk, et al.
Veröffentlicht: (2024)
von: Yu, Sungduk, et al.
Veröffentlicht: (2024)
Teaching People LLM's Errors and Getting it Right
von: Stringham, Nathan, et al.
Veröffentlicht: (2025)
von: Stringham, Nathan, et al.
Veröffentlicht: (2025)
BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
von: Chizhov, Pavel, et al.
Veröffentlicht: (2024)
von: Chizhov, Pavel, et al.
Veröffentlicht: (2024)
Deep Anomaly Detection in Text
von: Manolache, Andrei
Veröffentlicht: (2023)
von: Manolache, Andrei
Veröffentlicht: (2023)
Syntriever: How to Train Your Retriever with Synthetic Data from LLMs
von: Kim, Minsang, et al.
Veröffentlicht: (2025)
von: Kim, Minsang, et al.
Veröffentlicht: (2025)
Squrve: A Unified and Modular Framework for Complex Real-World Text-to-SQL Tasks
von: Wang, Yihan, et al.
Veröffentlicht: (2025)
von: Wang, Yihan, et al.
Veröffentlicht: (2025)
Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw
von: Wang, Zijun, et al.
Veröffentlicht: (2026)
von: Wang, Zijun, et al.
Veröffentlicht: (2026)
FastDraft: How to Train Your Draft
von: Zafrir, Ofir, et al.
Veröffentlicht: (2024)
von: Zafrir, Ofir, et al.
Veröffentlicht: (2024)
Training-free LLM-generated Text Detection by Mining Token Probability Sequences
von: Xu, Yihuai, et al.
Veröffentlicht: (2024)
von: Xu, Yihuai, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Going over Fine Web with a Fine-Tooth Comb: Technical Report of Indexing Fine Web for Problematic Content Search and Retrieval
von: Marinas, Inés Altemir, et al.
Veröffentlicht: (2025) -
Low-Perplexity LLM-Generated Sequences and Where To Find Them
von: Wuhrmann, Arthur, et al.
Veröffentlicht: (2025) -
TypePilot: Leveraging the Scala Type System for Secure LLM-generated Code
von: Sternfeld, Alexander, et al.
Veröffentlicht: (2025) -
From Model to Breach: Towards Actionable LLM-Generated Vulnerabilities Reporting
von: Vallez, Cyril, et al.
Veröffentlicht: (2025) -
Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs
von: Sternfeld, Alexander, et al.
Veröffentlicht: (2026)