How much do contextualized representations encode long-range context?
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Sun, Simeng, Hsieh, Cheng-Ping |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
L0-Reasoning Bench: Evaluating Procedural Correctness in Language Models via Simple Program Execution
par: Sun, Simeng, et autres
Publié: (2025)
par: Sun, Simeng, et autres
Publié: (2025)
An empirical study on the limitation of Transformers in program trace generation
par: Sun, Simeng
Publié: (2025)
par: Sun, Simeng
Publié: (2025)
How much do language models memorize?
par: Morris, John X., et autres
Publié: (2025)
par: Morris, John X., et autres
Publié: (2025)
Clinical ModernBERT: An efficient and long context encoder for biomedical text
par: Lee, Simon A., et autres
Publié: (2025)
par: Lee, Simon A., et autres
Publié: (2025)
RULER: What's the Real Context Size of Your Long-Context Language Models?
par: Hsieh, Cheng-Ping, et autres
Publié: (2024)
par: Hsieh, Cheng-Ping, et autres
Publié: (2024)
Injecting Wiktionary to improve token-level contextual representations using contrastive learning
par: Mosolova, Anna, et autres
Publié: (2024)
par: Mosolova, Anna, et autres
Publié: (2024)
Cartridges: Lightweight and general-purpose long context representations via self-study
par: Eyuboglu, Sabri, et autres
Publié: (2025)
par: Eyuboglu, Sabri, et autres
Publié: (2025)
Exploring the encoding of linguistic representations in the Fully-Connected Layer of generative CNNs for Speech
par: Šegedin, Bruno Ferenc, et autres
Publié: (2025)
par: Šegedin, Bruno Ferenc, et autres
Publié: (2025)
SWAN-GPT: An Efficient and Scalable Approach for Long-Context Language Modeling
par: Puvvada, Krishna C., et autres
Publié: (2025)
par: Puvvada, Krishna C., et autres
Publié: (2025)
Suri: Multi-constraint Instruction Following for Long-form Text Generation
par: Pham, Chau Minh, et autres
Publié: (2024)
par: Pham, Chau Minh, et autres
Publié: (2024)
How much reliable is ChatGPT's prediction on Information Extraction under Input Perturbations?
par: Mondal, Ishani, et autres
Publié: (2024)
par: Mondal, Ishani, et autres
Publié: (2024)
Interpreting the structure of multi-object representations in vision encoders
par: Khajuria, Tarun, et autres
Publié: (2024)
par: Khajuria, Tarun, et autres
Publié: (2024)
How much speech data is necessary for ASR in African languages? An evaluation of data scaling in Kinyarwanda and Kikuyu
par: Akera, Benjamin, et autres
Publié: (2025)
par: Akera, Benjamin, et autres
Publié: (2025)
How much do LLMs learn from negative examples?
par: Hamdan, Shadi, et autres
Publié: (2025)
par: Hamdan, Shadi, et autres
Publié: (2025)
Turbulence-like 5/3 spectral scaling in contextual representations of language as a complex system
par: Yang, Zhongxin, et autres
Publié: (2026)
par: Yang, Zhongxin, et autres
Publié: (2026)
CLIPPER: Compression enables long-context synthetic data generation
par: Pham, Chau Minh, et autres
Publié: (2025)
par: Pham, Chau Minh, et autres
Publié: (2025)
Nationality encoding in language model hidden states: Probing culturally differentiated representations in persona-conditioned academic text
par: Jackson, Paul, et autres
Publié: (2026)
par: Jackson, Paul, et autres
Publié: (2026)
Can LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval and haystacks
par: Hengle, Amey, et autres
Publié: (2025)
par: Hengle, Amey, et autres
Publié: (2025)
Comparing representations of long clinical texts for the task of patient note-identification
par: Alsaidi, Safa, et autres
Publié: (2025)
par: Alsaidi, Safa, et autres
Publié: (2025)
TopicGPT: A Prompt-based Topic Modeling Framework
par: Pham, Chau Minh, et autres
Publié: (2023)
par: Pham, Chau Minh, et autres
Publié: (2023)
BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
par: Almeida, Thales Sales, et autres
Publié: (2025)
par: Almeida, Thales Sales, et autres
Publié: (2025)
Does quantization affect models' performance on long-context tasks?
par: Mekala, Anmol, et autres
Publié: (2025)
par: Mekala, Anmol, et autres
Publié: (2025)
One ruler to measure them all: Benchmarking multilingual long-context language models
par: Kim, Yekyung, et autres
Publié: (2025)
par: Kim, Yekyung, et autres
Publié: (2025)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
par: Wang, Chonghua, et autres
Publié: (2024)
par: Wang, Chonghua, et autres
Publié: (2024)
Guideline Learning for In-context Information Extraction
par: Pang, Chaoxu, et autres
Publié: (2023)
par: Pang, Chaoxu, et autres
Publié: (2023)
Machine learning and emoji prediction: How much accuracy can MARBERT achieve?
par: Shormani, Mohammed Q., et autres
Publié: (2026)
par: Shormani, Mohammed Q., et autres
Publié: (2026)
so much depends / upon / a whitespace: Why Whitespace Matters for Poets and LLMs
par: Bhyravajjula, Sriharsh, et autres
Publié: (2025)
par: Bhyravajjula, Sriharsh, et autres
Publié: (2025)
Semantics or spelling? Probing contextual word embeddings with orthographic noise
par: Matthews, Jacob A., et autres
Publié: (2024)
par: Matthews, Jacob A., et autres
Publié: (2024)
Extracting domain-specific terms using contextual word embeddings
par: Repar, Andraž, et autres
Publié: (2025)
par: Repar, Andraž, et autres
Publié: (2025)
Large language models reorganize representational geometry during in-context learning
par: Xiong, Hua-Dong, et autres
Publié: (2026)
par: Xiong, Hua-Dong, et autres
Publié: (2026)
IndicSentEval: How Effectively do Multilingual Transformer Models encode Linguistic Properties for Indic Languages?
par: Aravapalli, Akhilesh, et autres
Publié: (2024)
par: Aravapalli, Akhilesh, et autres
Publié: (2024)
Failure of contextual invariance in large language models
par: Kumar, Sagar, et autres
Publié: (2026)
par: Kumar, Sagar, et autres
Publié: (2026)
One Thousand and One Pairs: A "novel" challenge for long-context language models
par: Karpinska, Marzena, et autres
Publié: (2024)
par: Karpinska, Marzena, et autres
Publié: (2024)
AIC CTU@FEVER 8: On-premise fact checking through long context RAG
par: Ullrich, Herbert, et autres
Publié: (2025)
par: Ullrich, Herbert, et autres
Publié: (2025)
HYBRIDMIND: Meta Selection of Natural Language and Symbolic Language for Enhanced LLM Reasoning
par: Han, Simeng, et autres
Publié: (2024)
par: Han, Simeng, et autres
Publié: (2024)
How much does context affect the accuracy of AI health advice?
par: Garg, Prashant, et autres
Publié: (2025)
par: Garg, Prashant, et autres
Publié: (2025)
Adjoint sharding for very long context training of state space models
par: Xu, Xingzi, et autres
Publié: (2025)
par: Xu, Xingzi, et autres
Publié: (2025)
Phase transition on a context-sensitive random language model with short range interactions
par: Toji, Yuma, et autres
Publié: (2026)
par: Toji, Yuma, et autres
Publié: (2026)
Contextual morphologically-guided tokenization for Latin encoder models
par: Hudspeth, Marisa, et autres
Publié: (2025)
par: Hudspeth, Marisa, et autres
Publié: (2025)
A graph-based analysis of semantic types and coercion in contextualized word embeddings
par: Chen, Long, et autres
Publié: (2026)
par: Chen, Long, et autres
Publié: (2026)
Documents similaires
-
L0-Reasoning Bench: Evaluating Procedural Correctness in Language Models via Simple Program Execution
par: Sun, Simeng, et autres
Publié: (2025) -
An empirical study on the limitation of Transformers in program trace generation
par: Sun, Simeng
Publié: (2025) -
How much do language models memorize?
par: Morris, John X., et autres
Publié: (2025) -
Clinical ModernBERT: An efficient and long context encoder for biomedical text
par: Lee, Simon A., et autres
Publié: (2025) -
RULER: What's the Real Context Size of Your Long-Context Language Models?
par: Hsieh, Cheng-Ping, et autres
Publié: (2024)