On the Semantic and Syntactic Information Encoded in Proto-Tokens for One-Step Text Reconstruction
Fuente:
arXiv
Guardado en:
| Autores principales: | Bondarenko, Ivan, Palkin, Egor, Tikunov, Fedor |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Neural Proto-Language Reconstruction
por: Cui, Chenxuan, et al.
Publicado: (2024)
por: Cui, Chenxuan, et al.
Publicado: (2024)
Exploring the Hidden Capacity of LLMs for One-Step Text Generation
por: Mezentsev, Gleb, et al.
Publicado: (2025)
por: Mezentsev, Gleb, et al.
Publicado: (2025)
Logit-KL Flow Matching: Non-Autoregressive Text Generation via Sampling-Hybrid Inference
por: Sevriugov, Egor, et al.
Publicado: (2024)
por: Sevriugov, Egor, et al.
Publicado: (2024)
LOLgorithm: Integrating Semantic,Syntactic and Contextual Elements for Humor Classification
por: Khurana, Tanisha, et al.
Publicado: (2024)
por: Khurana, Tanisha, et al.
Publicado: (2024)
RFBES at SemEval-2024 Task 8: Investigating Syntactic and Semantic Features for Distinguishing AI-Generated and Human-Written Texts
por: Rad, Mohammad Heydari, et al.
Publicado: (2024)
por: Rad, Mohammad Heydari, et al.
Publicado: (2024)
Circuit Fingerprints: How Answer Tokens Encode Their Geometrical Path
por: Saurez, Andres, et al.
Publicado: (2026)
por: Saurez, Andres, et al.
Publicado: (2026)
SSCAE -- Semantic, Syntactic, and Context-aware natural language Adversarial Examples generator
por: Asl, Javad Rafiei, et al.
Publicado: (2024)
por: Asl, Javad Rafiei, et al.
Publicado: (2024)
Dual Encoder: Exploiting the Potential of Syntactic and Semantic for Aspect Sentiment Triplet Extraction
por: Zhao, Xiaowei, et al.
Publicado: (2024)
por: Zhao, Xiaowei, et al.
Publicado: (2024)
Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo
por: Loula, João, et al.
Publicado: (2025)
por: Loula, João, et al.
Publicado: (2025)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
por: Kim, Eunji, et al.
Publicado: (2024)
por: Kim, Eunji, et al.
Publicado: (2024)
DLM-One: Diffusion Language Models for One-Step Sequence Generation
por: Chen, Tianqi, et al.
Publicado: (2025)
por: Chen, Tianqi, et al.
Publicado: (2025)
Understanding Token Probability Encoding in Output Embeddings
por: Cho, Hakaze, et al.
Publicado: (2024)
por: Cho, Hakaze, et al.
Publicado: (2024)
One Token to Fool LLM-as-a-Judge
por: Zhao, Yulai, et al.
Publicado: (2025)
por: Zhao, Yulai, et al.
Publicado: (2025)
Interpretable Syntactic Representations Enable Hierarchical Word Vectors
por: Silwal, Biraj
Publicado: (2024)
por: Silwal, Biraj
Publicado: (2024)
Towards Token-Level Text Anomaly Detection
por: Cao, Yang, et al.
Publicado: (2026)
por: Cao, Yang, et al.
Publicado: (2026)
Do LLMs Encode Functional Importance of Reasoning Tokens?
por: Singh, Janvijay, et al.
Publicado: (2026)
por: Singh, Janvijay, et al.
Publicado: (2026)
SAP: Syntactic Attention Pruning for Transformer-based Language Models
por: Lee, Tzu-Yun, et al.
Publicado: (2025)
por: Lee, Tzu-Yun, et al.
Publicado: (2025)
Large Language Models Encode Semantics and Alignment in Linearly Separable Representations
por: Saglam, Baturay, et al.
Publicado: (2025)
por: Saglam, Baturay, et al.
Publicado: (2025)
Urdu Dependency Parsing and Treebank Development: A Syntactic and Morphological Perspective
por: Habib, Nudrat
Publicado: (2024)
por: Habib, Nudrat
Publicado: (2024)
Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs
por: Nakkiran, Preetum, et al.
Publicado: (2025)
por: Nakkiran, Preetum, et al.
Publicado: (2025)
Beyond Early-Token Bias: Model-Specific and Language-Specific Position Effects in Multilingual LLMs
por: Menschikov, Mikhail, et al.
Publicado: (2025)
por: Menschikov, Mikhail, et al.
Publicado: (2025)
SENTRA: Selected-Next-Token Transformer for LLM Text Detection
por: Plyler, Mitchell, et al.
Publicado: (2025)
por: Plyler, Mitchell, et al.
Publicado: (2025)
RaguTeam at SemEval-2026 Task 8: Meno and Friends in a Judge-Orchestrated LLM Ensemble for Faithful Multi-Turn Response Generation
por: Bondarenko, Ivan, et al.
Publicado: (2026)
por: Bondarenko, Ivan, et al.
Publicado: (2026)
ParaScopes: What do Language Models Activations Encode About Future Text?
por: Pochinkov, Nicky, et al.
Publicado: (2025)
por: Pochinkov, Nicky, et al.
Publicado: (2025)
Syntactic Control of Language Models by Posterior Inference
por: Xefteri, Vicky, et al.
Publicado: (2025)
por: Xefteri, Vicky, et al.
Publicado: (2025)
Step Rejection Fine-Tuning: A Practical Distillation Recipe
por: Slinko, Igor, et al.
Publicado: (2026)
por: Slinko, Igor, et al.
Publicado: (2026)
Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
por: Foroutan, Negar, et al.
Publicado: (2025)
por: Foroutan, Negar, et al.
Publicado: (2025)
One Pass Streaming Algorithm for Super Long Token Attention Approximation in Sublinear Space
por: Addanki, Raghav, et al.
Publicado: (2023)
por: Addanki, Raghav, et al.
Publicado: (2023)
Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca
por: Cui, Yiming, et al.
Publicado: (2023)
por: Cui, Yiming, et al.
Publicado: (2023)
Guided Score identity Distillation for Data-Free One-Step Text-to-Image Generation
por: Zhou, Mingyuan, et al.
Publicado: (2024)
por: Zhou, Mingyuan, et al.
Publicado: (2024)
Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
por: Kim, Minchan, et al.
Publicado: (2024)
por: Kim, Minchan, et al.
Publicado: (2024)
Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
por: Mao, Yu, et al.
Publicado: (2025)
por: Mao, Yu, et al.
Publicado: (2025)
Not All Denoising Steps Are Equal: Model Scheduling for Faster Masked Diffusion Language Models
por: Sedykh, Ivan, et al.
Publicado: (2026)
por: Sedykh, Ivan, et al.
Publicado: (2026)
AutoJudge: Judge Decoding Without Manual Annotation
por: Garipov, Roman, et al.
Publicado: (2025)
por: Garipov, Roman, et al.
Publicado: (2025)
In-Context Learning of a Linear Transformer Block: Benefits of the MLP Component and One-Step GD Initialization
por: Zhang, Ruiqi, et al.
Publicado: (2024)
por: Zhang, Ruiqi, et al.
Publicado: (2024)
HawkesLLM: Semantic Uncertainty Propagation in Agentic Text Simulation
por: Deng, Zewei, et al.
Publicado: (2026)
por: Deng, Zewei, et al.
Publicado: (2026)
Advancing Anomaly Detection: Non-Semantic Financial Data Encoding with LLMs
por: Bakumenko, Alexander, et al.
Publicado: (2024)
por: Bakumenko, Alexander, et al.
Publicado: (2024)
$K$-MSHC: Unmasking Minimally Sufficient Head Circuits in Large Language Models with Experiments on Syntactic Classification Tasks
por: Chowdhary, Pratim, et al.
Publicado: (2025)
por: Chowdhary, Pratim, et al.
Publicado: (2025)
Token Homogenization under Positional Bias
por: Yusupov, Viacheslav, et al.
Publicado: (2025)
por: Yusupov, Viacheslav, et al.
Publicado: (2025)
IAPO: Information-Aware Policy Optimization for Token-Efficient Reasoning
por: He, Yinhan, et al.
Publicado: (2026)
por: He, Yinhan, et al.
Publicado: (2026)
Ejemplares similares
-
Neural Proto-Language Reconstruction
por: Cui, Chenxuan, et al.
Publicado: (2024) -
Exploring the Hidden Capacity of LLMs for One-Step Text Generation
por: Mezentsev, Gleb, et al.
Publicado: (2025) -
Logit-KL Flow Matching: Non-Autoregressive Text Generation via Sampling-Hybrid Inference
por: Sevriugov, Egor, et al.
Publicado: (2024) -
LOLgorithm: Integrating Semantic,Syntactic and Contextual Elements for Humor Classification
por: Khurana, Tanisha, et al.
Publicado: (2024) -
RFBES at SemEval-2024 Task 8: Investigating Syntactic and Semantic Features for Distinguishing AI-Generated and Human-Written Texts
por: Rad, Mohammad Heydari, et al.
Publicado: (2024)