Overcoming Copyright Barriers in Corpus Distribution Through Non-Reversible Hashing
Fuente:
arXiv
Saved in:
| Main Authors: | Amalvy, Arthur, Labatut, Vincent, Bost, Xavier, Huang, Hen-Hsen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Known Facts: Generating Unseen Temporal Knowledge to Address Data Contamination in LLM Evaluation
by: Amalvy, Arthur, et al.
Published: (2026)
by: Amalvy, Arthur, et al.
Published: (2026)
Annotation Guidelines for Corpus Novelties: Part 1 -- Named Entity Recognition
by: Amalvy, Arthur, et al.
Published: (2024)
by: Amalvy, Arthur, et al.
Published: (2024)
Annotation Guidelines for Corpus Novelties: Part 2 -- Alias Resolution Version 1.0
by: Amalvy, Arthur, et al.
Published: (2024)
by: Amalvy, Arthur, et al.
Published: (2024)
Learning to Rank Context for Named Entity Recognition Using a Synthetic Dataset
by: Amalvy, Arthur, et al.
Published: (2023)
by: Amalvy, Arthur, et al.
Published: (2023)
The Role of Natural Language Processing Tasks in Automatic Literary Character Network Construction
by: Amalvy, Arthur, et al.
Published: (2024)
by: Amalvy, Arthur, et al.
Published: (2024)
Renard: A Modular Pipeline for Extracting Character Networks from Narrative Texts
by: Amalvy, Arthur, et al.
Published: (2024)
by: Amalvy, Arthur, et al.
Published: (2024)
The Role of Global and Local Context in Named Entity Recognition
by: Amalvy, Arthur, et al.
Published: (2023)
by: Amalvy, Arthur, et al.
Published: (2023)
Democratizing LLM Efficiency: From Hyperscale Optimizations to Universal Deployability
by: Huang, Hen-Hsen
Published: (2025)
by: Huang, Hen-Hsen
Published: (2025)
Interconnected Kingdoms: Comparing 'A Song of Ice and Fire' Adaptations Across Media Using Complex Networks
by: Amalvy, Arthur, et al.
Published: (2024)
by: Amalvy, Arthur, et al.
Published: (2024)
Persistent Homology of Topic Networks for the Prediction of Reader Curiosity
by: Hopp, Manuel D. S., et al.
Published: (2025)
by: Hopp, Manuel D. S., et al.
Published: (2025)
Diverge to Induce Prompting: Multi-Rationale Induction for Zero-Shot Reasoning
by: Chen, Po-Chun, et al.
Published: (2026)
by: Chen, Po-Chun, et al.
Published: (2026)
Strategy-Induct: Task-Level Strategy Induction for Instruction Generation
by: Chen, Po-Chun, et al.
Published: (2026)
by: Chen, Po-Chun, et al.
Published: (2026)
DINA: A Dual Defense Framework Against Internal Noise and External Attacks in Natural Language Processing
by: Chuang, Ko-Wei, et al.
Published: (2025)
by: Chuang, Ko-Wei, et al.
Published: (2025)
Personalized Graph-Empowered Large Language Model for Proactive Information Access
by: Chang, Chia Cheng, et al.
Published: (2026)
by: Chang, Chia Cheng, et al.
Published: (2026)
Co-Trained Retriever-Generator Framework for Question Generation in Earnings Calls
by: Juan, Yining, et al.
Published: (2024)
by: Juan, Yining, et al.
Published: (2024)
No One Fits All: From Fixed Prompting to Learned Routing in Multilingual LLMs
by: Wu, Wei-Chi, et al.
Published: (2026)
by: Wu, Wei-Chi, et al.
Published: (2026)
Pre-Finetuning with Impact Duration Awareness for Stock Movement Prediction
by: Chiu, Chr-Jr, et al.
Published: (2024)
by: Chiu, Chr-Jr, et al.
Published: (2024)
"Why" Has the Least Side Effect on Model Editing
by: Pan, Tsung-Hsuan, et al.
Published: (2024)
by: Pan, Tsung-Hsuan, et al.
Published: (2024)
Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks
by: Chan, Brian J, et al.
Published: (2024)
by: Chan, Brian J, et al.
Published: (2024)
Diagnosing Model Editing via Knowledge Spectrum
by: Pan, Tsung-Hsuan, et al.
Published: (2025)
by: Pan, Tsung-Hsuan, et al.
Published: (2025)
Unveiling Selection Biases: Exploring Order and Token Sensitivity in Large Language Models
by: Wei, Sheng-Lun, et al.
Published: (2024)
by: Wei, Sheng-Lun, et al.
Published: (2024)
Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
by: Lin, Wei-Hsiang, et al.
Published: (2025)
by: Lin, Wei-Hsiang, et al.
Published: (2025)
Visual Lifelog Retrieval through Captioning-Enhanced Interpretation
by: Shih, Yu-Fei, et al.
Published: (2025)
by: Shih, Yu-Fei, et al.
Published: (2025)
Efficient Beam Search for Large Language Models Using Trie-Based Decoding
by: Chan, Brian J, et al.
Published: (2025)
by: Chan, Brian J, et al.
Published: (2025)
Overcoming Low-Resource Barriers in Tulu: Neural Models and Corpus Creation for OffensiveLanguage Identification
by: D, Anusha M, et al.
Published: (2025)
by: D, Anusha M, et al.
Published: (2025)
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models
by: Yeh, Cheng-Kai, et al.
Published: (2025)
by: Yeh, Cheng-Kai, et al.
Published: (2025)
Enhancing Investment Opinion Ranking through Argument-Based Sentiment Analysis
by: Chen, Chung-Chi, et al.
Published: (2024)
by: Chen, Chung-Chi, et al.
Published: (2024)
Using Contextually Aligned Online Reviews to Measure LLMs' Performance Disparities Across Language Varieties
by: Tang, Zixin, et al.
Published: (2025)
by: Tang, Zixin, et al.
Published: (2025)
Bias in the Ear of the Listener: Assessing Sensitivity in Audio Language Models Across Linguistic, Demographic, and Positional Variations
by: Wei, Sheng-Lun, et al.
Published: (2026)
by: Wei, Sheng-Lun, et al.
Published: (2026)
From Opinion Mining to Financial Argument Mining
by: Chen, Chung-Chi, et al.
Published: (2021)
by: Chen, Chung-Chi, et al.
Published: (2021)
Barriers to Universal Reasoning With Transformers (And How to Overcome Them)
by: Kraus, Oliver, et al.
Published: (2026)
by: Kraus, Oliver, et al.
Published: (2026)
SmartSpatial: Enhancing the 3D Spatial Arrangement Capabilities of Stable Diffusion Models and Introducing a Novel 3D Spatial Evaluation Framework
by: Huang, Mao Xun, et al.
Published: (2025)
by: Huang, Mao Xun, et al.
Published: (2025)
Copyright Detective: A Forensic System to Evidence LLMs Flickering Copyright Leakage Risks
by: Zhang, Guangwei, et al.
Published: (2026)
by: Zhang, Guangwei, et al.
Published: (2026)
Beyond English: Unveiling Multilingual Bias in LLM Copyright Compliance
by: Chen, Yupeng, et al.
Published: (2025)
by: Chen, Yupeng, et al.
Published: (2025)
"Wait, did you mean the doctor?": Collecting a Dialogue Corpus for Topical Analysis
by: Decker, Amandine, et al.
Published: (2025)
by: Decker, Amandine, et al.
Published: (2025)
DarijaBanking: A New Resource for Overcoming Language Barriers in Banking Intent Detection for Moroccan Arabic Speakers
by: Skiredj, Abderrahman, et al.
Published: (2024)
by: Skiredj, Abderrahman, et al.
Published: (2024)
Spotlight Attention: Towards Efficient LLM Generation via Non-linear Hashing-based KV Cache Retrieval
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
Overcoming Non-monotonicity in Transducer-based Streaming Generation
by: Ma, Zhengrui, et al.
Published: (2024)
by: Ma, Zhengrui, et al.
Published: (2024)
The Language of Interoception: Examining Embodiment and Emotion Through a Corpus of Body Part Mentions
by: Wu, Sophie, et al.
Published: (2025)
by: Wu, Sophie, et al.
Published: (2025)
Evaluating Copyright Takedown Methods for Language Models
by: Wei, Boyi, et al.
Published: (2024)
by: Wei, Boyi, et al.
Published: (2024)
Similar Items
-
Beyond Known Facts: Generating Unseen Temporal Knowledge to Address Data Contamination in LLM Evaluation
by: Amalvy, Arthur, et al.
Published: (2026) -
Annotation Guidelines for Corpus Novelties: Part 1 -- Named Entity Recognition
by: Amalvy, Arthur, et al.
Published: (2024) -
Annotation Guidelines for Corpus Novelties: Part 2 -- Alias Resolution Version 1.0
by: Amalvy, Arthur, et al.
Published: (2024) -
Learning to Rank Context for Named Entity Recognition Using a Synthetic Dataset
by: Amalvy, Arthur, et al.
Published: (2023) -
The Role of Natural Language Processing Tasks in Automatic Literary Character Network Construction
by: Amalvy, Arthur, et al.
Published: (2024)