Zipfian Whitening
Fuente:
arXiv
Saved in:
| Main Authors: | Yokoi, Sho, Bao, Han, Kurita, Hiroto, Shimodaira, Hidetoshi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quantifying Lexical Semantic Shift via Unbalanced Optimal Transport
by: Kishino, Ryo, et al.
Published: (2024)
by: Kishino, Ryo, et al.
Published: (2024)
DeLTa: A Decoding Strategy based on Logit Trajectory Prediction Improves Factuality and Reasoning Ability
by: He, Yunzhen, et al.
Published: (2025)
by: He, Yunzhen, et al.
Published: (2025)
TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
by: Shing, Makoto, et al.
Published: (2025)
by: Shing, Makoto, et al.
Published: (2025)
Why Mean Pooling Works: Quantifying Second-Order Collapse in Text Embeddings
by: Hara, Tomomasa, et al.
Published: (2026)
by: Hara, Tomomasa, et al.
Published: (2026)
Subspace Representations for Soft Set Operations and Sentence Similarities
by: Ishibashi, Yoichi, et al.
Published: (2022)
by: Ishibashi, Yoichi, et al.
Published: (2022)
SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale Corpora
by: Yoneda, Masataka, et al.
Published: (2026)
by: Yoneda, Masataka, et al.
Published: (2026)
Whitening Not Recommended for Classification Tasks in LLMs
by: Forooghi, Ali, et al.
Published: (2024)
by: Forooghi, Ali, et al.
Published: (2024)
Norm of Mean Contextualized Embeddings Determines their Variance
by: Yamagiwa, Hiroaki, et al.
Published: (2024)
by: Yamagiwa, Hiroaki, et al.
Published: (2024)
Knowledge Sanitization of Large Language Models
by: Ishibashi, Yoichi, et al.
Published: (2023)
by: Ishibashi, Yoichi, et al.
Published: (2023)
3D Rotation and Translation for Hyperbolic Knowledge Graph Embedding
by: Zhu, Yihua, et al.
Published: (2023)
by: Zhu, Yihua, et al.
Published: (2023)
Block-Diagonal Orthogonal Relation and Matrix Entity for Knowledge Graph Embedding
by: Zhu, Yihua, et al.
Published: (2024)
by: Zhu, Yihua, et al.
Published: (2024)
Shimo Lab at "Discharge Me!": Discharge Summarization by Prompt-Driven Concatenation of Electronic Health Record Sections
by: He, Yunzhen, et al.
Published: (2024)
by: He, Yunzhen, et al.
Published: (2024)
Revisiting Cosine Similarity via Normalized ICA-transformed Embeddings
by: Yamagiwa, Hiroaki, et al.
Published: (2024)
by: Yamagiwa, Hiroaki, et al.
Published: (2024)
Understanding Higher-Order Correlations Among Semantic Components in Embeddings
by: Oyama, Momose, et al.
Published: (2024)
by: Oyama, Momose, et al.
Published: (2024)
Axis Tour: Word Tour Determines the Order of Axes in ICA-transformed Embeddings
by: Yamagiwa, Hiroaki, et al.
Published: (2024)
by: Yamagiwa, Hiroaki, et al.
Published: (2024)
Measuring Affinity between Attention-Head Weight Subspaces via the Projection Kernel
by: Yamagiwa, Hiroaki, et al.
Published: (2026)
by: Yamagiwa, Hiroaki, et al.
Published: (2026)
Language Model Maps for Prompt-Response Distributions via Log-Likelihood Vectors
by: Takase, Yusuke, et al.
Published: (2026)
by: Takase, Yusuke, et al.
Published: (2026)
RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding
by: Hoshino, Yuichiro, et al.
Published: (2025)
by: Hoshino, Yuichiro, et al.
Published: (2025)
Scaling Laws for Upcycling Mixture-of-Experts Language Models
by: Liew, Seng Pei, et al.
Published: (2025)
by: Liew, Seng Pei, et al.
Published: (2025)
Mapping 1,000+ Language Models via the Log-Likelihood Vector
by: Oyama, Momose, et al.
Published: (2025)
by: Oyama, Momose, et al.
Published: (2025)
Likelihood Variance as Text Importance for Resampling Texts to Map Language Models
by: Oyama, Momose, et al.
Published: (2025)
by: Oyama, Momose, et al.
Published: (2025)
Improving Prediction Accuracy of Semantic Segmentation Methods Using Convolutional Autoencoder Based Pre-processing Layers
by: Shimodaira, Hisashi
Published: (2024)
by: Shimodaira, Hisashi
Published: (2024)
Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-Tuning
by: Yano, Kazuki, et al.
Published: (2026)
by: Yano, Kazuki, et al.
Published: (2026)
Non-Zipfian Distribution of Stopwords and Subset Selection Models
by: Li, Wentian, et al.
Published: (2026)
by: Li, Wentian, et al.
Published: (2026)
DOS: Distilling Observable Softmaps of Zipfian Prototypes for Self-Supervised Point Representation
by: Abdelsamad, Mohamed, et al.
Published: (2025)
by: Abdelsamad, Mohamed, et al.
Published: (2025)
Revisiting the Capacity Gap in Chain-of-Thought Distillation from a Practical Perspective
by: Kajitsuka, Tokio, et al.
Published: (2026)
by: Kajitsuka, Tokio, et al.
Published: (2026)
Beyond Chains: Bridging Large Language Models and Knowledge Bases in Complex Question Answering
by: Zhu, Yihua, et al.
Published: (2025)
by: Zhu, Yihua, et al.
Published: (2025)
Necessary and Sufficient Watermark for Large Language Models
by: Takezawa, Yuki, et al.
Published: (2023)
by: Takezawa, Yuki, et al.
Published: (2023)
Exploring Multilingual Large Language Models for Enhanced TNM classification of Radiology Report in lung cancer staging
by: Matsuo, Hidetoshi, et al.
Published: (2024)
by: Matsuo, Hidetoshi, et al.
Published: (2024)
Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization
by: Zhang, Zheyuan, et al.
Published: (2026)
by: Zhang, Zheyuan, et al.
Published: (2026)
A Single Linear Layer Yields Task-Adapted Low-Rank Matrices
by: Kim, Hwichan, et al.
Published: (2024)
by: Kim, Hwichan, et al.
Published: (2024)
Establishing a Scale for Kullback-Leibler Divergence in Language Models Across Various Settings
by: Kishino, Ryo, et al.
Published: (2025)
by: Kishino, Ryo, et al.
Published: (2025)
Domain Mixture Design via Log-Likelihood Differences for Aligning Language Models with a Target Model
by: Kishino, Ryo, et al.
Published: (2026)
by: Kishino, Ryo, et al.
Published: (2026)
FAAST: Forward-Only Associative Learning via Closed-Form Fast Weights for Test-Time Supervised Adaptation
by: Bao, Guangsheng, et al.
Published: (2026)
by: Bao, Guangsheng, et al.
Published: (2026)
Lean Formalization of Generalization Error Bound by Rademacher Complexity and Dudley's Entropy Integral
by: Sonoda, Sho, et al.
Published: (2025)
by: Sonoda, Sho, et al.
Published: (2025)
Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs
by: Huang, Yue, et al.
Published: (2026)
by: Huang, Yue, et al.
Published: (2026)
Structured Pruning for Diverse Best-of-N Reasoning Optimization
by: Nguyen, Hieu Trung, et al.
Published: (2025)
by: Nguyen, Hieu Trung, et al.
Published: (2025)
Analyzing Feed-Forward Blocks in Transformers through the Lens of Attention Maps
by: Kobayashi, Goro, et al.
Published: (2023)
by: Kobayashi, Goro, et al.
Published: (2023)
Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model
by: Hong, Yuzhong, et al.
Published: (2024)
by: Hong, Yuzhong, et al.
Published: (2024)
Task-driven Layerwise Additive Activation Intervention
by: Nguyen, Hieu Trung, et al.
Published: (2025)
by: Nguyen, Hieu Trung, et al.
Published: (2025)
Similar Items
-
Quantifying Lexical Semantic Shift via Unbalanced Optimal Transport
by: Kishino, Ryo, et al.
Published: (2024) -
DeLTa: A Decoding Strategy based on Logit Trajectory Prediction Improves Factuality and Reasoning Ability
by: He, Yunzhen, et al.
Published: (2025) -
TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
by: Shing, Makoto, et al.
Published: (2025) -
Why Mean Pooling Works: Quantifying Second-Order Collapse in Text Embeddings
by: Hara, Tomomasa, et al.
Published: (2026) -
Subspace Representations for Soft Set Operations and Sentence Similarities
by: Ishibashi, Yoichi, et al.
Published: (2022)