Code Pretraining Improves Entity Tracking Abilities of Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Najoung, Schuster, Sebastian, Toshniwal, Shubham |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Do Language Models Track Entities Across State Changes?
di: Tang, Zilu, et al.
Pubblicazione: (2026)
di: Tang, Zilu, et al.
Pubblicazione: (2026)
Major Entity Identification: A Generalizable Alternative to Coreference Resolution
di: Manikantan, Kawshik, et al.
Pubblicazione: (2024)
di: Manikantan, Kawshik, et al.
Pubblicazione: (2024)
A systematic framework for generating novel experimental hypotheses from language models
di: Misra, Kanishka, et al.
Pubblicazione: (2024)
di: Misra, Kanishka, et al.
Pubblicazione: (2024)
Emergent Abilities of Large Language Models under Continued Pretraining for Language Adaptation
di: Elhady, Ahmed, et al.
Pubblicazione: (2025)
di: Elhady, Ahmed, et al.
Pubblicazione: (2025)
Personas as a Way to Model Truthfulness in Language Models
di: Joshi, Nitish, et al.
Pubblicazione: (2023)
di: Joshi, Nitish, et al.
Pubblicazione: (2023)
Improving Multi-Domain Task-Oriented Dialogue System with Offline Reinforcement Learning
di: Prajapat, Dharmendra, et al.
Pubblicazione: (2024)
di: Prajapat, Dharmendra, et al.
Pubblicazione: (2024)
Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks
di: Wu, Zhaofeng, et al.
Pubblicazione: (2023)
di: Wu, Zhaofeng, et al.
Pubblicazione: (2023)
Exploring the Implicit Semantic Ability of Multimodal Large Language Models: A Pilot Study on Entity Set Expansion
di: Wang, Hebin, et al.
Pubblicazione: (2024)
di: Wang, Hebin, et al.
Pubblicazione: (2024)
Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter It
di: Qin, Yulu, et al.
Pubblicazione: (2025)
di: Qin, Yulu, et al.
Pubblicazione: (2025)
Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models
di: Arad, Dana, et al.
Pubblicazione: (2025)
di: Arad, Dana, et al.
Pubblicazione: (2025)
CodePMP: Scalable Preference Model Pretraining for Large Language Model Reasoning
di: Yu, Huimu, et al.
Pubblicazione: (2024)
di: Yu, Huimu, et al.
Pubblicazione: (2024)
Has Your Pretrained Model Improved? A Multi-head Posterior Based Approach
di: Aboagye, Prince, et al.
Pubblicazione: (2024)
di: Aboagye, Prince, et al.
Pubblicazione: (2024)
A Causal Language Modeling Detour Improves Encoder Continued Pretraining
di: Touchent, Rian, et al.
Pubblicazione: (2026)
di: Touchent, Rian, et al.
Pubblicazione: (2026)
Scope Ambiguities in Large Language Models
di: Kamath, Gaurav, et al.
Pubblicazione: (2024)
di: Kamath, Gaurav, et al.
Pubblicazione: (2024)
Death of the Novel(ty): Beyond n-Gram Novelty as a Metric for Textual Creativity
di: Saakyan, Arkadiy, et al.
Pubblicazione: (2025)
di: Saakyan, Arkadiy, et al.
Pubblicazione: (2025)
SAAS: Solving Ability Amplification Strategy for Enhanced Mathematical Reasoning in Large Language Models
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages
di: Africa, David Demitri, et al.
Pubblicazione: (2025)
di: Africa, David Demitri, et al.
Pubblicazione: (2025)
Entity-Aware Biaffine Attention Model for Improved Constituent Parsing with Reduced Entity Violations
di: Bai, Xinyi
Pubblicazione: (2024)
di: Bai, Xinyi
Pubblicazione: (2024)
Exploring Language Model's Code Generation Ability with Auxiliary Functions
di: Lee, Seonghyeon, et al.
Pubblicazione: (2024)
di: Lee, Seonghyeon, et al.
Pubblicazione: (2024)
IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs
di: Manikantan, Kawshik, et al.
Pubblicazione: (2024)
di: Manikantan, Kawshik, et al.
Pubblicazione: (2024)
Code-driven Number Sequence Calculation: Enhancing the inductive Reasoning Abilities of Large Language Models
di: Chen, Kedi, et al.
Pubblicazione: (2025)
di: Chen, Kedi, et al.
Pubblicazione: (2025)
Crystal: Illuminating LLM Abilities on Language and Code
di: Tao, Tianhua, et al.
Pubblicazione: (2024)
di: Tao, Tianhua, et al.
Pubblicazione: (2024)
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
di: Toshniwal, Shubham, et al.
Pubblicazione: (2024)
di: Toshniwal, Shubham, et al.
Pubblicazione: (2024)
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
di: Toshniwal, Shubham, et al.
Pubblicazione: (2024)
di: Toshniwal, Shubham, et al.
Pubblicazione: (2024)
LTNER: Large Language Model Tagging for Named Entity Recognition with Contextualized Entity Marking
di: Yan, Faren, et al.
Pubblicazione: (2024)
di: Yan, Faren, et al.
Pubblicazione: (2024)
The Tokenization Bottleneck: How Vocabulary Extension Improves Chemistry Representation Learning in Pretrained Language Models
di: Kalamkar, Prathamesh, et al.
Pubblicazione: (2025)
di: Kalamkar, Prathamesh, et al.
Pubblicazione: (2025)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
di: Kim, Dongjun, et al.
Pubblicazione: (2025)
di: Kim, Dongjun, et al.
Pubblicazione: (2025)
Multilingual Pretraining for Pixel Language Models
di: Kesen, Ilker, et al.
Pubblicazione: (2025)
di: Kesen, Ilker, et al.
Pubblicazione: (2025)
Improving Arithmetic Reasoning Ability of Large Language Models through Relation Tuples, Verification and Dynamic Feedback
di: Miao, Zhongtao, et al.
Pubblicazione: (2024)
di: Miao, Zhongtao, et al.
Pubblicazione: (2024)
Leveraging Large Language Models for Entity Matching
di: Huang, Qianyu, et al.
Pubblicazione: (2024)
di: Huang, Qianyu, et al.
Pubblicazione: (2024)
CodeNER: Code Prompting for Named Entity Recognition
di: Han, Sungwoo, et al.
Pubblicazione: (2025)
di: Han, Sungwoo, et al.
Pubblicazione: (2025)
Improving LLM Abilities in Idiomatic Translation
di: Donthi, Sundesh, et al.
Pubblicazione: (2024)
di: Donthi, Sundesh, et al.
Pubblicazione: (2024)
BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages
di: Manoj, Guduru, et al.
Pubblicazione: (2025)
di: Manoj, Guduru, et al.
Pubblicazione: (2025)
AnyMatch -- Efficient Zero-Shot Entity Matching with a Small Language Model
di: Zhang, Zeyu, et al.
Pubblicazione: (2024)
di: Zhang, Zeyu, et al.
Pubblicazione: (2024)
AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset
di: Moshkov, Ivan, et al.
Pubblicazione: (2025)
di: Moshkov, Ivan, et al.
Pubblicazione: (2025)
Mix of Experts Language Model for Named Entity Recognition
di: Chen, Xinwei, et al.
Pubblicazione: (2024)
di: Chen, Xinwei, et al.
Pubblicazione: (2024)
Unlocking the Power of Large Language Models for Entity Alignment
di: Jiang, Xuhui, et al.
Pubblicazione: (2024)
di: Jiang, Xuhui, et al.
Pubblicazione: (2024)
On the Representations of Entities in Auto-regressive Large Language Models
di: Morand, Victor, et al.
Pubblicazione: (2025)
di: Morand, Victor, et al.
Pubblicazione: (2025)
Counting Ability of Large Language Models and Impact of Tokenization
di: Zhang, Xiang, et al.
Pubblicazione: (2024)
di: Zhang, Xiang, et al.
Pubblicazione: (2024)
Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models
di: Sun, Haoran, et al.
Pubblicazione: (2024)
di: Sun, Haoran, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Do Language Models Track Entities Across State Changes?
di: Tang, Zilu, et al.
Pubblicazione: (2026) -
Major Entity Identification: A Generalizable Alternative to Coreference Resolution
di: Manikantan, Kawshik, et al.
Pubblicazione: (2024) -
A systematic framework for generating novel experimental hypotheses from language models
di: Misra, Kanishka, et al.
Pubblicazione: (2024) -
Emergent Abilities of Large Language Models under Continued Pretraining for Language Adaptation
di: Elhady, Ahmed, et al.
Pubblicazione: (2025) -
Personas as a Way to Model Truthfulness in Language Models
di: Joshi, Nitish, et al.
Pubblicazione: (2023)