Rectifying Belief Space via Unlearning to Harness LLMs' Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Niwa, Ayana, Kaneko, Masahiro, Inui, Kentaro |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
JailNewsBench: Multi-Lingual and Regional Benchmark for Fake News Generation under Jailbreak Attacks
di: Kaneko, Masahiro, et al.
Pubblicazione: (2026)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2026)
ExaGPT: Example-Based Machine-Generated Text Detection for Human Interpretability
di: Koike, Ryuto, et al.
Pubblicazione: (2025)
di: Koike, Ryuto, et al.
Pubblicazione: (2025)
JUBAKU: An Adversarial Benchmark for Exposing Culturally Grounded Stereotypes in Japanese LLMs
di: Shiotani, Taihei, et al.
Pubblicazione: (2026)
di: Shiotani, Taihei, et al.
Pubblicazione: (2026)
AmbigNLG: Addressing Task Ambiguity in Instruction for NLG
di: Niwa, Ayana, et al.
Pubblicazione: (2024)
di: Niwa, Ayana, et al.
Pubblicazione: (2024)
Spelling-out is not Straightforward: LLMs' Capability of Tokenization from Token to Characters
di: Hiraoka, Tatsuya, et al.
Pubblicazione: (2025)
di: Hiraoka, Tatsuya, et al.
Pubblicazione: (2025)
Repetition Neurons: How Do Language Models Produce Repetitions?
di: Hiraoka, Tatsuya, et al.
Pubblicazione: (2024)
di: Hiraoka, Tatsuya, et al.
Pubblicazione: (2024)
Monotonic Representation of Numeric Properties in Language Models
di: Heinzerling, Benjamin, et al.
Pubblicazione: (2024)
di: Heinzerling, Benjamin, et al.
Pubblicazione: (2024)
ACORN: Aspect-wise Commonsense Reasoning Explanation Evaluation
di: Brassard, Ana, et al.
Pubblicazione: (2024)
di: Brassard, Ana, et al.
Pubblicazione: (2024)
Cell-Based Representation of Relational Binding in Language Models
di: Dai, Qin, et al.
Pubblicazione: (2026)
di: Dai, Qin, et al.
Pubblicazione: (2026)
Representational Analysis of Binding in Language Models
di: Dai, Qin, et al.
Pubblicazione: (2024)
di: Dai, Qin, et al.
Pubblicazione: (2024)
A Japanese Benchmark for Evaluating Social Bias in Reasoning Based on Attribution Theory
di: Shiotani, Taihei, et al.
Pubblicazione: (2026)
di: Shiotani, Taihei, et al.
Pubblicazione: (2026)
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
LLMs Can Compensate for Deficiencies in Visual Representations
di: Takishita, Sho, et al.
Pubblicazione: (2025)
di: Takishita, Sho, et al.
Pubblicazione: (2025)
Understanding and Controlling Repetition Neurons and Induction Heads in In-Context Learning
di: Doan, Nhi Hoai, et al.
Pubblicazione: (2025)
di: Doan, Nhi Hoai, et al.
Pubblicazione: (2025)
Measuring AI Reasoning: A Guide for Researchers
di: Nwadike, Munachiso Samuel, et al.
Pubblicazione: (2026)
di: Nwadike, Munachiso Samuel, et al.
Pubblicazione: (2026)
LLM Unlearning with LLM Beliefs
di: Li, Kemou, et al.
Pubblicazione: (2025)
di: Li, Kemou, et al.
Pubblicazione: (2025)
Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
di: Kaneko, Masahiro
Pubblicazione: (2026)
di: Kaneko, Masahiro
Pubblicazione: (2026)
Rectify Evaluation Preference: Improving LLMs' Critique on Math Reasoning via Perplexity-aware Reinforcement Learning
di: Tian, Changyuan, et al.
Pubblicazione: (2025)
di: Tian, Changyuan, et al.
Pubblicazione: (2025)
Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
The Geometry of Numerical Reasoning: Language Models Compare Numeric Properties in Linear Subspaces
di: El-Shangiti, Ahmed Oumar, et al.
Pubblicazione: (2024)
di: El-Shangiti, Ahmed Oumar, et al.
Pubblicazione: (2024)
A Little Leak Will Sink a Great Ship: Survey of Transparency for Large Language Models from Start to Finish
di: Kaneko, Masahiro, et al.
Pubblicazione: (2024)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2024)
Syntactic Learnability of Echo State Neural Language Models at Scale
di: Ueda, Ryo, et al.
Pubblicazione: (2025)
di: Ueda, Ryo, et al.
Pubblicazione: (2025)
TopK Language Models
di: Takahashi, Ryosuke, et al.
Pubblicazione: (2025)
di: Takahashi, Ryosuke, et al.
Pubblicazione: (2025)
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge
di: Zhang, Ying, et al.
Pubblicazione: (2025)
di: Zhang, Ying, et al.
Pubblicazione: (2025)
J-UniMorph: Japanese Morphological Annotation through the Universal Feature Schema
di: Matsuzaki, Kosuke, et al.
Pubblicazione: (2024)
di: Matsuzaki, Kosuke, et al.
Pubblicazione: (2024)
A Large Collection of Model-generated Contradictory Responses for Consistency-aware Dialogue Systems
di: Sato, Shiki, et al.
Pubblicazione: (2024)
di: Sato, Shiki, et al.
Pubblicazione: (2024)
Analyzing Feed-Forward Blocks in Transformers through the Lens of Attention Maps
di: Kobayashi, Goro, et al.
Pubblicazione: (2023)
di: Kobayashi, Goro, et al.
Pubblicazione: (2023)
First Heuristic Then Rational: Dynamic Use of Heuristics in Language Model Reasoning
di: Aoki, Yoichi, et al.
Pubblicazione: (2024)
di: Aoki, Yoichi, et al.
Pubblicazione: (2024)
FinchGPT: a Transformer based language model for birdsong analysis
di: Kobayashi, Kosei, et al.
Pubblicazione: (2025)
di: Kobayashi, Kosei, et al.
Pubblicazione: (2025)
An Audit and Analysis of LLM-Assisted Health Misinformation Jailbreaks Against LLMs
di: Hussain, Ayana, et al.
Pubblicazione: (2025)
di: Hussain, Ayana, et al.
Pubblicazione: (2025)
How Individual Traits and Language Styles Shape Preferences In Open-ended User-LLM Interaction: A Preliminary Study
di: Chevi, Rendi, et al.
Pubblicazione: (2025)
di: Chevi, Rendi, et al.
Pubblicazione: (2025)
Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting
di: Kaneko, Masahiro, et al.
Pubblicazione: (2024)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2024)
LLMs Faithfully and Iteratively Compute Answers During CoT: A Systematic Analysis With Multi-step Arithmetics
di: Kudo, Keito, et al.
Pubblicazione: (2024)
di: Kudo, Keito, et al.
Pubblicazione: (2024)
Emergence of Primacy and Recency Effect in Mamba: A Mechanistic Point of View
di: Airlangga, Muhammad Cendekia, et al.
Pubblicazione: (2025)
di: Airlangga, Muhammad Cendekia, et al.
Pubblicazione: (2025)
On Entity Identification in Language Models
di: Sakata, Masaki, et al.
Pubblicazione: (2025)
di: Sakata, Masaki, et al.
Pubblicazione: (2025)
Linear Representations of Hierarchical Concepts in Language Models
di: Sakata, Masaki, et al.
Pubblicazione: (2026)
di: Sakata, Masaki, et al.
Pubblicazione: (2026)
Why Mean Pooling Works: Quantifying Second-Order Collapse in Text Embeddings
di: Hara, Tomomasa, et al.
Pubblicazione: (2026)
di: Hara, Tomomasa, et al.
Pubblicazione: (2026)
Reducing the Cost: Cross-Prompt Pre-Finetuning for Short Answer Scoring
di: Funayama, Hiroaki, et al.
Pubblicazione: (2024)
di: Funayama, Hiroaki, et al.
Pubblicazione: (2024)
To Drop or Not to Drop? Predicting Argument Ellipsis Judgments: A Case Study in Japanese
di: Ishizuki, Yukiko, et al.
Pubblicazione: (2024)
di: Ishizuki, Yukiko, et al.
Pubblicazione: (2024)
MQM-Chat: Multidimensional Quality Metrics for Chat Translation
di: Li, Yunmeng, et al.
Pubblicazione: (2024)
di: Li, Yunmeng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
JailNewsBench: Multi-Lingual and Regional Benchmark for Fake News Generation under Jailbreak Attacks
di: Kaneko, Masahiro, et al.
Pubblicazione: (2026) -
ExaGPT: Example-Based Machine-Generated Text Detection for Human Interpretability
di: Koike, Ryuto, et al.
Pubblicazione: (2025) -
JUBAKU: An Adversarial Benchmark for Exposing Culturally Grounded Stereotypes in Japanese LLMs
di: Shiotani, Taihei, et al.
Pubblicazione: (2026) -
AmbigNLG: Addressing Task Ambiguity in Instruction for NLG
di: Niwa, Ayana, et al.
Pubblicazione: (2024) -
Spelling-out is not Straightforward: LLMs' Capability of Tokenization from Token to Characters
di: Hiraoka, Tatsuya, et al.
Pubblicazione: (2025)