Saved in:
| Main Authors: | Suzuki, Hisami, Katsumata, Satoru, Kodama, Takashi, Takahashi, Tetsuro, Nakayama, Kouta, Sekine, Satoshi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.02372 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
JAPAGEN: Efficient Few/Zero-shot Learning via Japanese Training Dataset Generation with LLM
by: Fujii, Takuro, et al.
Published: (2024)
by: Fujii, Takuro, et al.
Published: (2024)
Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?
by: Furuhashi, Momoka, et al.
Published: (2025)
by: Furuhashi, Momoka, et al.
Published: (2025)
Which Feedback Works for Whom? Differential Effects of LLM-Generated Feedback Elements Across Learner Profiles
by: Furuhashi, Momoka, et al.
Published: (2026)
by: Furuhashi, Momoka, et al.
Published: (2026)
llm-jp-modernbert: A ModernBERT Model Trained on a Large-Scale Japanese Corpus with Long Context Length
by: Sugiura, Issa, et al.
Published: (2025)
by: Sugiura, Issa, et al.
Published: (2025)
WikiSQE: A Large-Scale Dataset for Sentence Quality Estimation in Wikipedia
by: Ando, Kenichiro, et al.
Published: (2023)
by: Ando, Kenichiro, et al.
Published: (2023)
LLM-jp: A Cross-organizational Project for the Research and Development of Fully Open Japanese LLMs
by: LLM-jp, et al.
Published: (2024)
by: LLM-jp, et al.
Published: (2024)
RecMind: Japanese Movie Recommendation Dialogue with Seeker's Internal State
by: Kodama, Takashi, et al.
Published: (2024)
by: Kodama, Takashi, et al.
Published: (2024)
LLM as a Scorer: The Impact of Output Order on Dialogue Evaluation
by: Chen, Yi-Pei, et al.
Published: (2024)
by: Chen, Yi-Pei, et al.
Published: (2024)
A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities
by: Imamura, Kenji, et al.
Published: (2026)
by: Imamura, Kenji, et al.
Published: (2026)
Exclusive Unlearning
by: Sasaki, Mutsumi, et al.
Published: (2026)
by: Sasaki, Mutsumi, et al.
Published: (2026)
Should We Respect LLMs? A Cross-Lingual Study on the Influence of Prompt Politeness on LLM Performance
by: Yin, Ziqi, et al.
Published: (2024)
by: Yin, Ziqi, et al.
Published: (2024)
A Better LLM Evaluator for Text Generation: The Impact of Prompt Output Sequencing and Optimization
by: Chu, KuanChao, et al.
Published: (2024)
by: Chu, KuanChao, et al.
Published: (2024)
When Choices Become Risks: Safety Failures of Large Language Models under Multiple-Choice Constraints
by: Chen, Yuheng, et al.
Published: (2026)
by: Chen, Yuheng, et al.
Published: (2026)
JaFIn: Japanese Financial Instruction Dataset
by: Tanabe, Kota, et al.
Published: (2024)
by: Tanabe, Kota, et al.
Published: (2024)
Economy Watchers Survey Provides Datasets and Tasks for Japanese Financial Domain
by: Suzuki, Masahiro, et al.
Published: (2024)
by: Suzuki, Masahiro, et al.
Published: (2024)
Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability
by: Takami, Kyosuke, et al.
Published: (2026)
by: Takami, Kyosuke, et al.
Published: (2026)
Correct Chains, Wrong Answers: Dissociating Reasoning from Output in LLM Logic
by: Rao, Abinav, et al.
Published: (2026)
by: Rao, Abinav, et al.
Published: (2026)
Overcoming the "Impracticality" of RAG: Proposing a Real-World Benchmark and Multi-Dimensional Diagnostic Framework
by: Narita, Kenichirou, et al.
Published: (2026)
by: Narita, Kenichirou, et al.
Published: (2026)
Datasets for Multilingual Answer Sentence Selection
by: Gabburo, Matteo, et al.
Published: (2024)
by: Gabburo, Matteo, et al.
Published: (2024)
Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems
by: Agrawal, Aakriti, et al.
Published: (2025)
by: Agrawal, Aakriti, et al.
Published: (2025)
JRadiEvo: A Japanese Radiology Report Generation Model Enhanced by Evolutionary Optimization of Model Merging
by: Baba, Kaito, et al.
Published: (2024)
by: Baba, Kaito, et al.
Published: (2024)
100 Books for Teachers of English as a Second Language: An Annotated Bibliography.
by: Springer, Hisami K., Comp.
Published: (1971)
by: Springer, Hisami K., Comp.
Published: (1971)
Automated Long Answer Grading with RiceChem Dataset
by: Sonkar, Shashank, et al.
Published: (2024)
by: Sonkar, Shashank, et al.
Published: (2024)
Cross-Context Review: Improving LLM Output Quality by Separating Production and Review Sessions
by: Song, Tae-Eun
Published: (2026)
by: Song, Tae-Eun
Published: (2026)
JFinTEB: Japanese Financial Text Embedding Benchmark
by: Suzuki, Masahiro, et al.
Published: (2026)
by: Suzuki, Masahiro, et al.
Published: (2026)
Rethinking Supervised Fine-Tuning: Emphasizing Key Answer Tokens for Improved LLM Accuracy
by: Shi, Xiaofeng, et al.
Published: (2025)
by: Shi, Xiaofeng, et al.
Published: (2025)
Multilingual KokoroChat: A Multi-LLM Ensemble Translation Method for Creating a Multilingual Counseling Dialogue Dataset
by: Suzuki, Ryoma, et al.
Published: (2026)
by: Suzuki, Ryoma, et al.
Published: (2026)
70B-parameter large language models in Japanese medical question-answering
by: Sukeda, Issey, et al.
Published: (2024)
by: Sukeda, Issey, et al.
Published: (2024)
Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling
by: Cappelletti, Silvia, et al.
Published: (2025)
by: Cappelletti, Silvia, et al.
Published: (2025)
Cancer-Answer: Empowering Cancer Care with Advanced Large Language Models
by: Deroy, Aniket, et al.
Published: (2024)
by: Deroy, Aniket, et al.
Published: (2024)
Span-Level Hallucination Detection for LLM-Generated Answers
by: Elchafei, Passant, et al.
Published: (2025)
by: Elchafei, Passant, et al.
Published: (2025)
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
by: Cencerrado, Iván Vicente Moreno, et al.
Published: (2025)
by: Cencerrado, Iván Vicente Moreno, et al.
Published: (2025)
Pretraining and Updates of Domain-Specific LLM: A Case Study in the Japanese Business Domain
by: Takahashi, Kosuke, et al.
Published: (2024)
by: Takahashi, Kosuke, et al.
Published: (2024)
Wrong Answers Can Also Be Useful: PlausibleQA -- A Large-Scale QA Dataset with Answer Plausibility Scores
by: Mozafari, Jamshid, et al.
Published: (2025)
by: Mozafari, Jamshid, et al.
Published: (2025)
Integrated Framework for LLM Evaluation with Answer Generation
by: Lee, Sujeong, et al.
Published: (2025)
by: Lee, Sujeong, et al.
Published: (2025)
Refining Answer Distributions for Improved Large Language Model Reasoning
by: Pal, Soumyasundar, et al.
Published: (2024)
by: Pal, Soumyasundar, et al.
Published: (2024)
A Dataset of Open-Domain Question Answering with Multiple-Span Answers
by: Luo, Zhiyi, et al.
Published: (2024)
by: Luo, Zhiyi, et al.
Published: (2024)
Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models
by: Dam, Harvey, et al.
Published: (2025)
by: Dam, Harvey, et al.
Published: (2025)
Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers
by: Wang, Yuhan, et al.
Published: (2026)
by: Wang, Yuhan, et al.
Published: (2026)
When Safety Fails Before the Answer: Benchmarking Harmful Behavior Detection in Reasoning Chains
by: Kakkar, Ishita, et al.
Published: (2026)
by: Kakkar, Ishita, et al.
Published: (2026)
Similar Items
-
JAPAGEN: Efficient Few/Zero-shot Learning via Japanese Training Dataset Generation with LLM
by: Fujii, Takuro, et al.
Published: (2024) -
Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?
by: Furuhashi, Momoka, et al.
Published: (2025) -
Which Feedback Works for Whom? Differential Effects of LLM-Generated Feedback Elements Across Learner Profiles
by: Furuhashi, Momoka, et al.
Published: (2026) -
llm-jp-modernbert: A ModernBERT Model Trained on a Large-Scale Japanese Corpus with Long Context Length
by: Sugiura, Issa, et al.
Published: (2025) -
WikiSQE: A Large-Scale Dataset for Sentence Quality Estimation in Wikipedia
by: Ando, Kenichiro, et al.
Published: (2023)