On the Semantic and Syntactic Information Encoded in Proto-Tokens for One-Step Text Reconstruction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bondarenko, Ivan, Palkin, Egor, Tikunov, Fedor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Neural Proto-Language Reconstruction
von: Cui, Chenxuan, et al.
Veröffentlicht: (2024)
von: Cui, Chenxuan, et al.
Veröffentlicht: (2024)
Exploring the Hidden Capacity of LLMs for One-Step Text Generation
von: Mezentsev, Gleb, et al.
Veröffentlicht: (2025)
von: Mezentsev, Gleb, et al.
Veröffentlicht: (2025)
Logit-KL Flow Matching: Non-Autoregressive Text Generation via Sampling-Hybrid Inference
von: Sevriugov, Egor, et al.
Veröffentlicht: (2024)
von: Sevriugov, Egor, et al.
Veröffentlicht: (2024)
LOLgorithm: Integrating Semantic,Syntactic and Contextual Elements for Humor Classification
von: Khurana, Tanisha, et al.
Veröffentlicht: (2024)
von: Khurana, Tanisha, et al.
Veröffentlicht: (2024)
RFBES at SemEval-2024 Task 8: Investigating Syntactic and Semantic Features for Distinguishing AI-Generated and Human-Written Texts
von: Rad, Mohammad Heydari, et al.
Veröffentlicht: (2024)
von: Rad, Mohammad Heydari, et al.
Veröffentlicht: (2024)
Circuit Fingerprints: How Answer Tokens Encode Their Geometrical Path
von: Saurez, Andres, et al.
Veröffentlicht: (2026)
von: Saurez, Andres, et al.
Veröffentlicht: (2026)
SSCAE -- Semantic, Syntactic, and Context-aware natural language Adversarial Examples generator
von: Asl, Javad Rafiei, et al.
Veröffentlicht: (2024)
von: Asl, Javad Rafiei, et al.
Veröffentlicht: (2024)
Dual Encoder: Exploiting the Potential of Syntactic and Semantic for Aspect Sentiment Triplet Extraction
von: Zhao, Xiaowei, et al.
Veröffentlicht: (2024)
von: Zhao, Xiaowei, et al.
Veröffentlicht: (2024)
Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo
von: Loula, João, et al.
Veröffentlicht: (2025)
von: Loula, João, et al.
Veröffentlicht: (2025)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
DLM-One: Diffusion Language Models for One-Step Sequence Generation
von: Chen, Tianqi, et al.
Veröffentlicht: (2025)
von: Chen, Tianqi, et al.
Veröffentlicht: (2025)
Understanding Token Probability Encoding in Output Embeddings
von: Cho, Hakaze, et al.
Veröffentlicht: (2024)
von: Cho, Hakaze, et al.
Veröffentlicht: (2024)
One Token to Fool LLM-as-a-Judge
von: Zhao, Yulai, et al.
Veröffentlicht: (2025)
von: Zhao, Yulai, et al.
Veröffentlicht: (2025)
Interpretable Syntactic Representations Enable Hierarchical Word Vectors
von: Silwal, Biraj
Veröffentlicht: (2024)
von: Silwal, Biraj
Veröffentlicht: (2024)
Towards Token-Level Text Anomaly Detection
von: Cao, Yang, et al.
Veröffentlicht: (2026)
von: Cao, Yang, et al.
Veröffentlicht: (2026)
Do LLMs Encode Functional Importance of Reasoning Tokens?
von: Singh, Janvijay, et al.
Veröffentlicht: (2026)
von: Singh, Janvijay, et al.
Veröffentlicht: (2026)
SAP: Syntactic Attention Pruning for Transformer-based Language Models
von: Lee, Tzu-Yun, et al.
Veröffentlicht: (2025)
von: Lee, Tzu-Yun, et al.
Veröffentlicht: (2025)
Large Language Models Encode Semantics and Alignment in Linearly Separable Representations
von: Saglam, Baturay, et al.
Veröffentlicht: (2025)
von: Saglam, Baturay, et al.
Veröffentlicht: (2025)
Urdu Dependency Parsing and Treebank Development: A Syntactic and Morphological Perspective
von: Habib, Nudrat
Veröffentlicht: (2024)
von: Habib, Nudrat
Veröffentlicht: (2024)
Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs
von: Nakkiran, Preetum, et al.
Veröffentlicht: (2025)
von: Nakkiran, Preetum, et al.
Veröffentlicht: (2025)
Beyond Early-Token Bias: Model-Specific and Language-Specific Position Effects in Multilingual LLMs
von: Menschikov, Mikhail, et al.
Veröffentlicht: (2025)
von: Menschikov, Mikhail, et al.
Veröffentlicht: (2025)
SENTRA: Selected-Next-Token Transformer for LLM Text Detection
von: Plyler, Mitchell, et al.
Veröffentlicht: (2025)
von: Plyler, Mitchell, et al.
Veröffentlicht: (2025)
RaguTeam at SemEval-2026 Task 8: Meno and Friends in a Judge-Orchestrated LLM Ensemble for Faithful Multi-Turn Response Generation
von: Bondarenko, Ivan, et al.
Veröffentlicht: (2026)
von: Bondarenko, Ivan, et al.
Veröffentlicht: (2026)
ParaScopes: What do Language Models Activations Encode About Future Text?
von: Pochinkov, Nicky, et al.
Veröffentlicht: (2025)
von: Pochinkov, Nicky, et al.
Veröffentlicht: (2025)
Syntactic Control of Language Models by Posterior Inference
von: Xefteri, Vicky, et al.
Veröffentlicht: (2025)
von: Xefteri, Vicky, et al.
Veröffentlicht: (2025)
Step Rejection Fine-Tuning: A Practical Distillation Recipe
von: Slinko, Igor, et al.
Veröffentlicht: (2026)
von: Slinko, Igor, et al.
Veröffentlicht: (2026)
Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
One Pass Streaming Algorithm for Super Long Token Attention Approximation in Sublinear Space
von: Addanki, Raghav, et al.
Veröffentlicht: (2023)
von: Addanki, Raghav, et al.
Veröffentlicht: (2023)
Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca
von: Cui, Yiming, et al.
Veröffentlicht: (2023)
von: Cui, Yiming, et al.
Veröffentlicht: (2023)
Guided Score identity Distillation for Data-Free One-Step Text-to-Image Generation
von: Zhou, Mingyuan, et al.
Veröffentlicht: (2024)
von: Zhou, Mingyuan, et al.
Veröffentlicht: (2024)
Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
von: Mao, Yu, et al.
Veröffentlicht: (2025)
von: Mao, Yu, et al.
Veröffentlicht: (2025)
Not All Denoising Steps Are Equal: Model Scheduling for Faster Masked Diffusion Language Models
von: Sedykh, Ivan, et al.
Veröffentlicht: (2026)
von: Sedykh, Ivan, et al.
Veröffentlicht: (2026)
AutoJudge: Judge Decoding Without Manual Annotation
von: Garipov, Roman, et al.
Veröffentlicht: (2025)
von: Garipov, Roman, et al.
Veröffentlicht: (2025)
In-Context Learning of a Linear Transformer Block: Benefits of the MLP Component and One-Step GD Initialization
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2024)
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2024)
HawkesLLM: Semantic Uncertainty Propagation in Agentic Text Simulation
von: Deng, Zewei, et al.
Veröffentlicht: (2026)
von: Deng, Zewei, et al.
Veröffentlicht: (2026)
Advancing Anomaly Detection: Non-Semantic Financial Data Encoding with LLMs
von: Bakumenko, Alexander, et al.
Veröffentlicht: (2024)
von: Bakumenko, Alexander, et al.
Veröffentlicht: (2024)
$K$-MSHC: Unmasking Minimally Sufficient Head Circuits in Large Language Models with Experiments on Syntactic Classification Tasks
von: Chowdhary, Pratim, et al.
Veröffentlicht: (2025)
von: Chowdhary, Pratim, et al.
Veröffentlicht: (2025)
Token Homogenization under Positional Bias
von: Yusupov, Viacheslav, et al.
Veröffentlicht: (2025)
von: Yusupov, Viacheslav, et al.
Veröffentlicht: (2025)
IAPO: Information-Aware Policy Optimization for Token-Efficient Reasoning
von: He, Yinhan, et al.
Veröffentlicht: (2026)
von: He, Yinhan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Neural Proto-Language Reconstruction
von: Cui, Chenxuan, et al.
Veröffentlicht: (2024) -
Exploring the Hidden Capacity of LLMs for One-Step Text Generation
von: Mezentsev, Gleb, et al.
Veröffentlicht: (2025) -
Logit-KL Flow Matching: Non-Autoregressive Text Generation via Sampling-Hybrid Inference
von: Sevriugov, Egor, et al.
Veröffentlicht: (2024) -
LOLgorithm: Integrating Semantic,Syntactic and Contextual Elements for Humor Classification
von: Khurana, Tanisha, et al.
Veröffentlicht: (2024) -
RFBES at SemEval-2024 Task 8: Investigating Syntactic and Semantic Features for Distinguishing AI-Generated and Human-Written Texts
von: Rad, Mohammad Heydari, et al.
Veröffentlicht: (2024)