Does Representation Matter? Exploring Intermediate Layers in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Skean, Oscar, Arefin, Md Rifat, LeCun, Yann, Shwartz-Ziv, Ravid |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Layer by Layer: Uncovering Hidden Representations in Language Models
von: Skean, Oscar, et al.
Veröffentlicht: (2025)
von: Skean, Oscar, et al.
Veröffentlicht: (2025)
Seq-VCR: Preventing Collapse in Intermediate Transformer Representations for Enhanced Reasoning
von: Arefin, Md Rifat, et al.
Veröffentlicht: (2024)
von: Arefin, Md Rifat, et al.
Veröffentlicht: (2024)
Variance-Covariance Regularization Improves Representation Learning
von: Zhu, Jiachen, et al.
Veröffentlicht: (2023)
von: Zhu, Jiachen, et al.
Veröffentlicht: (2023)
Video Representation Learning with Joint-Embedding Predictive Architectures
von: Drozdov, Katrina, et al.
Veröffentlicht: (2024)
von: Drozdov, Katrina, et al.
Veröffentlicht: (2024)
On Training in Imagination
von: Timor, Nadav, et al.
Veröffentlicht: (2026)
von: Timor, Nadav, et al.
Veröffentlicht: (2026)
Rate-In: Information-Driven Adaptive Dropout Rates for Improved Inference-Time Uncertainty Estimation
von: Zeevi, Tal, et al.
Veröffentlicht: (2024)
von: Zeevi, Tal, et al.
Veröffentlicht: (2024)
Antislop: A Comprehensive Framework for Identifying and Eliminating Repetitive Patterns in Language Models
von: Paech, Samuel, et al.
Veröffentlicht: (2025)
von: Paech, Samuel, et al.
Veröffentlicht: (2025)
From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
von: Shani, Chen, et al.
Veröffentlicht: (2025)
von: Shani, Chen, et al.
Veröffentlicht: (2025)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
von: Sanyal, Sunny, et al.
Veröffentlicht: (2024)
von: Sanyal, Sunny, et al.
Veröffentlicht: (2024)
The Entropy Enigma: Success and Failure of Entropy Minimization
von: Press, Ori, et al.
Veröffentlicht: (2024)
von: Press, Ori, et al.
Veröffentlicht: (2024)
JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
von: Ioannides, Georgios, et al.
Veröffentlicht: (2025)
von: Ioannides, Georgios, et al.
Veröffentlicht: (2025)
Just How Flexible are Neural Networks in Practice?
von: Shwartz-Ziv, Ravid, et al.
Veröffentlicht: (2024)
von: Shwartz-Ziv, Ravid, et al.
Veröffentlicht: (2024)
Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures
von: Ioannides, Georgios, et al.
Veröffentlicht: (2026)
von: Ioannides, Georgios, et al.
Veröffentlicht: (2026)
Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
von: Queipo-de-Llano, Enrique, et al.
Veröffentlicht: (2025)
von: Queipo-de-Llano, Enrique, et al.
Veröffentlicht: (2025)
AI Must Embrace Specialization via Superhuman Adaptable Intelligence
von: Goldfeder, Judah, et al.
Veröffentlicht: (2026)
von: Goldfeder, Judah, et al.
Veröffentlicht: (2026)
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
von: Huang, Hai, et al.
Veröffentlicht: (2025)
von: Huang, Hai, et al.
Veröffentlicht: (2025)
The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
von: Balestriero, Randall, et al.
Veröffentlicht: (2025)
von: Balestriero, Randall, et al.
Veröffentlicht: (2025)
Learning by Reconstruction Produces Uninformative Features For Perception
von: Balestriero, Randall, et al.
Veröffentlicht: (2024)
von: Balestriero, Randall, et al.
Veröffentlicht: (2024)
Transformers without Normalization
von: Zhu, Jiachen, et al.
Veröffentlicht: (2025)
von: Zhu, Jiachen, et al.
Veröffentlicht: (2025)
URLOST: Unsupervised Representation Learning without Stationarity or Topology
von: Yun, Zeyu, et al.
Veröffentlicht: (2023)
von: Yun, Zeyu, et al.
Veröffentlicht: (2023)
Learning to Compress: Local Rank and Information Compression in Deep Neural Networks
von: Patel, Niket, et al.
Veröffentlicht: (2024)
von: Patel, Niket, et al.
Veröffentlicht: (2024)
Variance Covariance Regularization Enforces Pairwise Independence in Self-Supervised Representations
von: Mialon, Grégoire, et al.
Veröffentlicht: (2022)
von: Mialon, Grégoire, et al.
Veröffentlicht: (2022)
LiveBench: A Challenging, Contamination-Limited LLM Benchmark
von: White, Colin, et al.
Veröffentlicht: (2024)
von: White, Colin, et al.
Veröffentlicht: (2024)
Introduction to Latent Variable Energy-Based Models: A Path Towards Autonomous Machine Intelligence
von: Dawid, Anna, et al.
Veröffentlicht: (2023)
von: Dawid, Anna, et al.
Veröffentlicht: (2023)
A hierarchical loss and its problems when classifying non-hierarchically
von: Wu, Cinna, et al.
Veröffentlicht: (2017)
von: Wu, Cinna, et al.
Veröffentlicht: (2017)
Fast and Exact Enumeration of Deep Networks Partitions Regions
von: Balestriero, Randall, et al.
Veröffentlicht: (2024)
von: Balestriero, Randall, et al.
Veröffentlicht: (2024)
Towards an Improved Understanding and Utilization of Maximum Manifold Capacity Representations
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2024)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2024)
Short Data, Long Context: Distilling Positional Knowledge in Transformers
von: Huber, Patrick, et al.
Veröffentlicht: (2026)
von: Huber, Patrick, et al.
Veröffentlicht: (2026)
Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
von: Zhai, Yuexiang, et al.
Veröffentlicht: (2024)
von: Zhai, Yuexiang, et al.
Veröffentlicht: (2024)
Can Large Language Models Understand Intermediate Representations in Compilers?
von: Jiang, Hailong, et al.
Veröffentlicht: (2025)
von: Jiang, Hailong, et al.
Veröffentlicht: (2025)
Sorted LLaMA: Unlocking the Potential of Intermediate Layers of Large Language Models for Dynamic Inference
von: Kavehzadeh, Parsa, et al.
Veröffentlicht: (2023)
von: Kavehzadeh, Parsa, et al.
Veröffentlicht: (2023)
ILRe: Intermediate Layer Retrieval for Context Compression in Causal Language Models
von: Liang, Manlai, et al.
Veröffentlicht: (2025)
von: Liang, Manlai, et al.
Veröffentlicht: (2025)
Learning and Leveraging World Models in Visual Representation Learning
von: Garrido, Quentin, et al.
Veröffentlicht: (2024)
von: Garrido, Quentin, et al.
Veröffentlicht: (2024)
Semantic Tube Prediction: Beating LLM Data Efficiency with JEPA
von: Huang, Hai, et al.
Veröffentlicht: (2026)
von: Huang, Hai, et al.
Veröffentlicht: (2026)
OpenDebateEvidence: A Massive-Scale Argument Mining and Summarization Dataset
von: Roush, Allen, et al.
Veröffentlicht: (2024)
von: Roush, Allen, et al.
Veröffentlicht: (2024)
You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
von: LeVi, Amit, et al.
Veröffentlicht: (2025)
von: LeVi, Amit, et al.
Veröffentlicht: (2025)
Rectified LpJEPA: Joint-Embedding Predictive Architectures with Sparse and Maximum-Entropy Representations
von: Kuang, Yilun, et al.
Veröffentlicht: (2026)
von: Kuang, Yilun, et al.
Veröffentlicht: (2026)
Exploring the Potential of the Large Language Models (LLMs) in Identifying Misleading News Headlines
von: Rony, Md Main Uddin, et al.
Veröffentlicht: (2024)
von: Rony, Md Main Uddin, et al.
Veröffentlicht: (2024)
Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
von: Balestriero, Randall, et al.
Veröffentlicht: (2025)
von: Balestriero, Randall, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Layer by Layer: Uncovering Hidden Representations in Language Models
von: Skean, Oscar, et al.
Veröffentlicht: (2025) -
Seq-VCR: Preventing Collapse in Intermediate Transformer Representations for Enhanced Reasoning
von: Arefin, Md Rifat, et al.
Veröffentlicht: (2024) -
Variance-Covariance Regularization Improves Representation Learning
von: Zhu, Jiachen, et al.
Veröffentlicht: (2023) -
Video Representation Learning with Joint-Embedding Predictive Architectures
von: Drozdov, Katrina, et al.
Veröffentlicht: (2024) -
On Training in Imagination
von: Timor, Nadav, et al.
Veröffentlicht: (2026)