An Information-theoretic Multi-task Representation Learning Framework for Natural Language Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Hu, Dou, Wei, Lingwei, Zhou, Wei, Hu, Songlin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Representation Learning with Conditional Information Flow Maximization
di: Hu, Dou, et al.
Pubblicazione: (2024)
di: Hu, Dou, et al.
Pubblicazione: (2024)
Structured Probabilistic Coding
di: Hu, Dou, et al.
Pubblicazione: (2023)
di: Hu, Dou, et al.
Pubblicazione: (2023)
Transferring Structure Knowledge: A New Task to Fake news Detection Towards Cold-Start Propagation
di: Wei, Lingwei, et al.
Pubblicazione: (2024)
di: Wei, Lingwei, et al.
Pubblicazione: (2024)
An Information-theoretic Propagation Denoising and Fusion Framework for Fake News Detection
di: Chen, Mengyang, et al.
Pubblicazione: (2026)
di: Chen, Mengyang, et al.
Pubblicazione: (2026)
Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models
di: Nagle, Alliot, et al.
Pubblicazione: (2024)
di: Nagle, Alliot, et al.
Pubblicazione: (2024)
Enhancing Multi-Hop Fact Verification with Structured Knowledge-Augmented Large Language Models
di: Cao, Han, et al.
Pubblicazione: (2025)
di: Cao, Han, et al.
Pubblicazione: (2025)
A Mathematical Theory for Learning Semantic Languages by Abstract Learners
di: Liao, Kuo-Yu, et al.
Pubblicazione: (2024)
di: Liao, Kuo-Yu, et al.
Pubblicazione: (2024)
Understanding Factual Recall in Transformers via Associative Memories
di: Nichani, Eshaan, et al.
Pubblicazione: (2024)
di: Nichani, Eshaan, et al.
Pubblicazione: (2024)
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
di: Yang, Tong, et al.
Pubblicazione: (2024)
di: Yang, Tong, et al.
Pubblicazione: (2024)
A Rate-Distortion Framework for Summarization
di: Arda, Enes, et al.
Pubblicazione: (2025)
di: Arda, Enes, et al.
Pubblicazione: (2025)
Self-Play Only Evolves When Self-Synthetic Pipeline Ensures Learnable Information Gain
di: Liu, Wei, et al.
Pubblicazione: (2026)
di: Liu, Wei, et al.
Pubblicazione: (2026)
The Information of Large Language Model Geometry
di: Tan, Zhiquan, et al.
Pubblicazione: (2024)
di: Tan, Zhiquan, et al.
Pubblicazione: (2024)
Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language Models
di: Wei, Lai, et al.
Pubblicazione: (2024)
di: Wei, Lai, et al.
Pubblicazione: (2024)
Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
di: Makkuva, Ashok Vardhan, et al.
Pubblicazione: (2024)
di: Makkuva, Ashok Vardhan, et al.
Pubblicazione: (2024)
Information-Theoretic Generative Clustering of Documents
di: Du, Xin, et al.
Pubblicazione: (2024)
di: Du, Xin, et al.
Pubblicazione: (2024)
NanoKnow: How to Know What Your Language Model Knows
di: Gu, Lingwei, et al.
Pubblicazione: (2026)
di: Gu, Lingwei, et al.
Pubblicazione: (2026)
Multi-turn Training with Basic Human Feedback Helps Little on LLM Reasoning
di: Liu, Qiang, et al.
Pubblicazione: (2025)
di: Liu, Qiang, et al.
Pubblicazione: (2025)
An Enhanced Text Compression Approach Using Transformer-based Language Models
di: Rahman, Chowdhury Mofizur, et al.
Pubblicazione: (2024)
di: Rahman, Chowdhury Mofizur, et al.
Pubblicazione: (2024)
MultiTok: Variable-Length Tokenization for Efficient LLMs Adapted from LZW Compression
di: Elias, Noel, et al.
Pubblicazione: (2024)
di: Elias, Noel, et al.
Pubblicazione: (2024)
Efficient Learned Data Compression via Dual-Stream Feature Decoupling
di: Ma, Huidong, et al.
Pubblicazione: (2026)
di: Ma, Huidong, et al.
Pubblicazione: (2026)
Filtering Beats Fine Tuning: A Bayesian Kalman View of In Context Learning in LLMs
di: Kiruluta, Andrew
Pubblicazione: (2026)
di: Kiruluta, Andrew
Pubblicazione: (2026)
RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory
di: Zuo, Fei, et al.
Pubblicazione: (2026)
di: Zuo, Fei, et al.
Pubblicazione: (2026)
Structure-aware Propagation Generation with Large Language Models for Fake News Detection
di: Chen, Mengyang, et al.
Pubblicazione: (2025)
di: Chen, Mengyang, et al.
Pubblicazione: (2025)
How Many Features Can a Language Model Store Under the Linear Representation Hypothesis?
di: Garg, Nikhil, et al.
Pubblicazione: (2026)
di: Garg, Nikhil, et al.
Pubblicazione: (2026)
Theoretical Limits of Language Model Alignment
di: Paes, Lucas Monteiro, et al.
Pubblicazione: (2026)
di: Paes, Lucas Monteiro, et al.
Pubblicazione: (2026)
Hierarchical Knowledge Distillation on Text Graph for Data-limited Attribute Inference
di: Li, Quan, et al.
Pubblicazione: (2024)
di: Li, Quan, et al.
Pubblicazione: (2024)
Multi-task retriever fine-tuning for domain-specific and efficient RAG
di: Béchard, Patrice, et al.
Pubblicazione: (2025)
di: Béchard, Patrice, et al.
Pubblicazione: (2025)
Information-Theoretic Framework for Understanding Modern Machine-Learning
di: Feder, Meir, et al.
Pubblicazione: (2025)
di: Feder, Meir, et al.
Pubblicazione: (2025)
Understanding Survey Paper Taxonomy about Large Language Models via Graph Representation Learning
di: Zhuang, Jun, et al.
Pubblicazione: (2024)
di: Zhuang, Jun, et al.
Pubblicazione: (2024)
Language Modeling Is Compression
di: Delétang, Grégoire, et al.
Pubblicazione: (2023)
di: Delétang, Grégoire, et al.
Pubblicazione: (2023)
Advancing Academic Knowledge Retrieval via LLM-enhanced Representation Similarity Fusion
di: Dai, Wei, et al.
Pubblicazione: (2024)
di: Dai, Wei, et al.
Pubblicazione: (2024)
An Information Theoretic Perspective on Agentic System Design
di: He, Shizhe, et al.
Pubblicazione: (2025)
di: He, Shizhe, et al.
Pubblicazione: (2025)
What Makes the Preferred Thinking Direction for LLMs in Multiple-choice Questions?
di: Zhang, Yizhe, et al.
Pubblicazione: (2025)
di: Zhang, Yizhe, et al.
Pubblicazione: (2025)
Cost-aware LLM-based Online Dataset Annotation
di: Elumar, Eray Can, et al.
Pubblicazione: (2025)
di: Elumar, Eray Can, et al.
Pubblicazione: (2025)
Iterative Counterfactual Data Augmentation
di: Plyler, Mitchell, et al.
Pubblicazione: (2025)
di: Plyler, Mitchell, et al.
Pubblicazione: (2025)
Proposal and study of statistical features for string similarity computation and classification
di: Rodrigues, E. O., et al.
Pubblicazione: (2026)
di: Rodrigues, E. O., et al.
Pubblicazione: (2026)
Theoretical guarantees on the best-of-n alignment policy
di: Beirami, Ahmad, et al.
Pubblicazione: (2024)
di: Beirami, Ahmad, et al.
Pubblicazione: (2024)
Effective Context in Transformers: An Analysis of Fragmentation and Tokenization
di: Fesharaki, Amirmehdi Jafari, et al.
Pubblicazione: (2026)
di: Fesharaki, Amirmehdi Jafari, et al.
Pubblicazione: (2026)
InfAlign: Inference-aware language model alignment
di: Balashankar, Ananth, et al.
Pubblicazione: (2024)
di: Balashankar, Ananth, et al.
Pubblicazione: (2024)
Speculative Decoding Scaling Laws (SDSL): Throughput Optimization Made Simple
di: Bozorgkhoo, Amirhossein, et al.
Pubblicazione: (2026)
di: Bozorgkhoo, Amirhossein, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Representation Learning with Conditional Information Flow Maximization
di: Hu, Dou, et al.
Pubblicazione: (2024) -
Structured Probabilistic Coding
di: Hu, Dou, et al.
Pubblicazione: (2023) -
Transferring Structure Knowledge: A New Task to Fake news Detection Towards Cold-Start Propagation
di: Wei, Lingwei, et al.
Pubblicazione: (2024) -
An Information-theoretic Propagation Denoising and Fusion Framework for Fake News Detection
di: Chen, Mengyang, et al.
Pubblicazione: (2026) -
Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models
di: Nagle, Alliot, et al.
Pubblicazione: (2024)