WikiSplit++: Easy Data Refinement for Split and Rephrase
Fuente:
arXiv
Saved in:
| Main Authors: | Tsukagoshi, Hayato, Hirao, Tsutomu, Morishita, Makoto, Chousa, Katsuki, Sasano, Ryohei, Takeda, Koichi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Simplifying Translations for Children: Iterative Simplification Considering Age of Acquisition with LLMs
by: Oshika, Masashi, et al.
Published: (2024)
by: Oshika, Masashi, et al.
Published: (2024)
Improving Sentence Embeddings with Automatic Generation of Training Data Using Few-shot Examples
by: Sato, Soma, et al.
Published: (2024)
by: Sato, Soma, et al.
Published: (2024)
Sentence Representations via Gaussian Embedding
by: Yoda, Shohei, et al.
Published: (2023)
by: Yoda, Shohei, et al.
Published: (2023)
Ruri: Japanese General Text Embeddings
by: Tsukagoshi, Hayato, et al.
Published: (2024)
by: Tsukagoshi, Hayato, et al.
Published: (2024)
Redundancy, Isotropy, and Intrinsic Dimensionality of Prompt-based Text Embeddings
by: Tsukagoshi, Hayato, et al.
Published: (2025)
by: Tsukagoshi, Hayato, et al.
Published: (2025)
FrameEOL: Semantic Frame Induction using Causal Language Models
by: Yano, Chihiro, et al.
Published: (2025)
by: Yano, Chihiro, et al.
Published: (2025)
When Is 0.1% Enough? Analyzing the Combined Effects of Dimensionality Reduction and Quantization on Text Embedding Compression
by: Kisako, Riku, et al.
Published: (2026)
by: Kisako, Riku, et al.
Published: (2026)
A Japanese-Chinese Parallel Corpus Using Crowdsourcing for Web Mining
by: Nagata, Masaaki, et al.
Published: (2024)
by: Nagata, Masaaki, et al.
Published: (2024)
CiMaTe: Citation Count Prediction Effectively Leveraging the Main Text
by: Hirako, Jun, et al.
Published: (2024)
by: Hirako, Jun, et al.
Published: (2024)
Verifying Claims About Metaphors with Large-Scale Automatic Metaphor Identification
by: Aono, Kotaro, et al.
Published: (2024)
by: Aono, Kotaro, et al.
Published: (2024)
Are Social Sentiments Inherent in LLMs? An Empirical Study on Extraction of Inter-demographic Sentiments
by: Tanaka, Kunitomo, et al.
Published: (2024)
by: Tanaka, Kunitomo, et al.
Published: (2024)
Split and Rephrase with Large Language Models
by: Ponce, David, et al.
Published: (2023)
by: Ponce, David, et al.
Published: (2023)
JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus
by: Nagata, Masaaki, et al.
Published: (2025)
by: Nagata, Masaaki, et al.
Published: (2025)
Hacking Neural Evaluation Metrics with Single Hub Text
by: Deguchi, Hiroyuki, et al.
Published: (2025)
by: Deguchi, Hiroyuki, et al.
Published: (2025)
One Single Hub Text Breaks CLIP: Identifying Vulnerabilities in Cross-Modal Encoders via Hubness
by: Deguchi, Hiroyuki, et al.
Published: (2026)
by: Deguchi, Hiroyuki, et al.
Published: (2026)
How Do Language Models Acquire Character-Level Information?
by: Sato, Soma, et al.
Published: (2026)
by: Sato, Soma, et al.
Published: (2026)
Can We Still Hear the Accent? Investigating the Resilience of Native Language Signals in the LLM Era
by: Utami, Nabelanita, et al.
Published: (2026)
by: Utami, Nabelanita, et al.
Published: (2026)
On Representational Dissociation of Language and Arithmetic in Large Language Models
by: Kisako, Riku, et al.
Published: (2025)
by: Kisako, Riku, et al.
Published: (2025)
Can we obtain significant success in RST discourse parsing by using Large Language Models?
by: Maekawa, Aru, et al.
Published: (2024)
by: Maekawa, Aru, et al.
Published: (2024)
Argument Mining as a Text-to-Text Generation Task
by: Kawarada, Masayuki, et al.
Published: (2026)
by: Kawarada, Masayuki, et al.
Published: (2026)
Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling
by: Maini, Pratyush, et al.
Published: (2024)
by: Maini, Pratyush, et al.
Published: (2024)
WiCER: Wiki-memory Compile, Evaluate, Refine Iterative Knowledge Compilation for LLM Wiki Systems
by: Huerta, Juan M.
Published: (2026)
by: Huerta, Juan M.
Published: (2026)
To Drop or Not to Drop? Predicting Argument Ellipsis Judgments: A Case Study in Japanese
by: Ishizuki, Yukiko, et al.
Published: (2024)
by: Ishizuki, Yukiko, et al.
Published: (2024)
Do LLMs and Humans Find the Same Questions Difficult? A Case Study on Japanese Quiz Answering
by: Sugiura, Naoya, et al.
Published: (2025)
by: Sugiura, Naoya, et al.
Published: (2025)
Chinese Spelling Correction as Rephrasing Language Model
by: Liu, Linfeng, et al.
Published: (2023)
by: Liu, Linfeng, et al.
Published: (2023)
Quantifying the Risks of Tool-assisted Rephrasing to Linguistic Diversity
by: Wang, Mengying, et al.
Published: (2024)
by: Wang, Mengying, et al.
Published: (2024)
Tokenization with Split Trees
by: Schmidt, Craig W., et al.
Published: (2026)
by: Schmidt, Craig W., et al.
Published: (2026)
Rephrasing Electronic Health Records for Pretraining Clinical Language Models
by: Liu, Jinghui, et al.
Published: (2024)
by: Liu, Jinghui, et al.
Published: (2024)
Rethinking Cross-Subject Data Splitting for Brain-to-Text Decoding
by: Yin, Congchi, et al.
Published: (2023)
by: Yin, Congchi, et al.
Published: (2023)
DART: An AIGT Detector using AMR of Rephrased Text
by: Park, Hyeonchu, et al.
Published: (2024)
by: Park, Hyeonchu, et al.
Published: (2024)
Generating Diverse Translation with Perturbed kNN-MT
by: Nishida, Yuto, et al.
Published: (2024)
by: Nishida, Yuto, et al.
Published: (2024)
EchoPrompt: Instructing the Model to Rephrase Queries for Improved In-context Learning
by: Mekala, Rajasekhar Reddy, et al.
Published: (2023)
by: Mekala, Rajasekhar Reddy, et al.
Published: (2023)
Wiki-Quantities and Wiki-Measurements: Datasets of Quantities and their Measurement Context from Wikipedia
by: Göpfert, Jan, et al.
Published: (2025)
by: Göpfert, Jan, et al.
Published: (2025)
SplitReason: Learning To Offload Reasoning
by: Akhauri, Yash, et al.
Published: (2025)
by: Akhauri, Yash, et al.
Published: (2025)
How Good is Post-Hoc Watermarking With Language Model Rephrasing?
by: Fernandez, Pierre, et al.
Published: (2025)
by: Fernandez, Pierre, et al.
Published: (2025)
Rephrase and Contrast: Fine-Tuning Language Models for Enhanced Understanding of Communication and Computer Networks
by: Wang, Liujianfu, et al.
Published: (2024)
by: Wang, Liujianfu, et al.
Published: (2024)
TimeMachine-bench: A Benchmark for Evaluating Model Capabilities in Repository-Level Migration Tasks
by: Fujii, Ryo, et al.
Published: (2026)
by: Fujii, Ryo, et al.
Published: (2026)
Editing Knowledge Representation of Language Model via Rephrased Prefix Prompts
by: Cai, Yuchen, et al.
Published: (2024)
by: Cai, Yuchen, et al.
Published: (2024)
TATRA: Training-Free Instance-Adaptive Prompting Through Rephrasing and Aggregation
by: Dziuba, Bartosz, et al.
Published: (2026)
by: Dziuba, Bartosz, et al.
Published: (2026)
DynSplit-KV: Dynamic Semantic Splitting for KVCache Compression in Efficient Long-Context LLM Inference
by: Ye, Jiancai, et al.
Published: (2026)
by: Ye, Jiancai, et al.
Published: (2026)
Similar Items
-
Simplifying Translations for Children: Iterative Simplification Considering Age of Acquisition with LLMs
by: Oshika, Masashi, et al.
Published: (2024) -
Improving Sentence Embeddings with Automatic Generation of Training Data Using Few-shot Examples
by: Sato, Soma, et al.
Published: (2024) -
Sentence Representations via Gaussian Embedding
by: Yoda, Shohei, et al.
Published: (2023) -
Ruri: Japanese General Text Embeddings
by: Tsukagoshi, Hayato, et al.
Published: (2024) -
Redundancy, Isotropy, and Intrinsic Dimensionality of Prompt-based Text Embeddings
by: Tsukagoshi, Hayato, et al.
Published: (2025)