Dissociating Decodability and Causal Use in Bracket-Sequence Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Sharma, Aryan, Dawes, Cutter, Raval, Shivam |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
H-Probes: Extracting Hierarchical Structures From Latent Representations of Language Models
di: Dawes, Cutter, et al.
Pubblicazione: (2026)
di: Dawes, Cutter, et al.
Pubblicazione: (2026)
Linear probes rely on textual evidence: Results from leakage mitigation studies in language models
di: Boxo, Gerard, et al.
Pubblicazione: (2025)
di: Boxo, Gerard, et al.
Pubblicazione: (2025)
Causal Reflection with Language Models
di: Aryan, Abi, et al.
Pubblicazione: (2025)
di: Aryan, Abi, et al.
Pubblicazione: (2025)
Transforming Chatbot Text: A Sequence-to-Sequence Approach
di: Reddy, Natesh, et al.
Pubblicazione: (2025)
di: Reddy, Natesh, et al.
Pubblicazione: (2025)
Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models
di: Roy, Amartya, et al.
Pubblicazione: (2025)
di: Roy, Amartya, et al.
Pubblicazione: (2025)
On The Adaptation of Unlimiformer for Decoder-Only Transformers
di: Ahrabian, Kian, et al.
Pubblicazione: (2024)
di: Ahrabian, Kian, et al.
Pubblicazione: (2024)
TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
di: Sun, Hanshi, et al.
Pubblicazione: (2024)
di: Sun, Hanshi, et al.
Pubblicazione: (2024)
Teaching Transformers Causal Reasoning through Axiomatic Training
di: Vashishtha, Aniket, et al.
Pubblicazione: (2024)
di: Vashishtha, Aniket, et al.
Pubblicazione: (2024)
LayerNorm Induces Recency Bias in Transformer Decoders
di: Kim, Junu, et al.
Pubblicazione: (2025)
di: Kim, Junu, et al.
Pubblicazione: (2025)
How Powerful are Decoder-Only Transformer Neural Models?
di: Roberts, Jesse
Pubblicazione: (2023)
di: Roberts, Jesse
Pubblicazione: (2023)
Learning Extrapolative Sequence Transformations from Markov Chains
di: Hager, Sophia, et al.
Pubblicazione: (2025)
di: Hager, Sophia, et al.
Pubblicazione: (2025)
Faster Transformer Decoding: N-gram Masked Self-Attention
di: Chelba, Ciprian, et al.
Pubblicazione: (2020)
di: Chelba, Ciprian, et al.
Pubblicazione: (2020)
Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs
di: Venkatesh, Sohan
Pubblicazione: (2026)
di: Venkatesh, Sohan
Pubblicazione: (2026)
Task Schema and Binding: A Double Dissociation Study of In-Context Learning
di: Kim, Chaeha
Pubblicazione: (2025)
di: Kim, Chaeha
Pubblicazione: (2025)
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
di: Yang, Songlin, et al.
Pubblicazione: (2024)
di: Yang, Songlin, et al.
Pubblicazione: (2024)
Illuminate: A novel approach for depression detection with explainable analysis and proactive therapy using prompt engineering
di: Agrawal, Aryan
Pubblicazione: (2024)
di: Agrawal, Aryan
Pubblicazione: (2024)
A Group Theoretic Analysis of the Symmetries Underlying Base Addition and Their Learnability by Neural Networks
di: Dawes, Cutter, et al.
Pubblicazione: (2025)
di: Dawes, Cutter, et al.
Pubblicazione: (2025)
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
di: Dong, Yihe, et al.
Pubblicazione: (2025)
di: Dong, Yihe, et al.
Pubblicazione: (2025)
Decoding Speculative Decoding
di: Yan, Minghao, et al.
Pubblicazione: (2024)
di: Yan, Minghao, et al.
Pubblicazione: (2024)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
di: Collins, Liam, et al.
Pubblicazione: (2024)
di: Collins, Liam, et al.
Pubblicazione: (2024)
TabSHAP
di: Chaudhary, Aryan, et al.
Pubblicazione: (2026)
di: Chaudhary, Aryan, et al.
Pubblicazione: (2026)
Automatic Differential Diagnosis using Transformer-Based Multi-Label Sequence Classification
di: Sadi, Abu Adnan, et al.
Pubblicazione: (2024)
di: Sadi, Abu Adnan, et al.
Pubblicazione: (2024)
Dissociating model architectures from inference computations
di: Sajid, Noor, et al.
Pubblicazione: (2025)
di: Sajid, Noor, et al.
Pubblicazione: (2025)
PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration
di: Wu, Songhao, et al.
Pubblicazione: (2025)
di: Wu, Songhao, et al.
Pubblicazione: (2025)
CMLFormer: A Dual Decoder Transformer with Switching Point Learning for Code-Mixed Language Modeling
di: Baral, Aditeya, et al.
Pubblicazione: (2025)
di: Baral, Aditeya, et al.
Pubblicazione: (2025)
A2SF: Accumulative Attention Scoring with Forgetting Factor for Token Pruning in Transformer Decoder
di: Jo, Hyun-rae, et al.
Pubblicazione: (2024)
di: Jo, Hyun-rae, et al.
Pubblicazione: (2024)
Better Think Thrice: Learning to Reason Causally with Double Counterfactual Consistency
di: Lin, Victoria, et al.
Pubblicazione: (2026)
di: Lin, Victoria, et al.
Pubblicazione: (2026)
Encoding Agent Trajectories as Representations with Sequence Transformers
di: Tsiligkaridis, Athanasios, et al.
Pubblicazione: (2024)
di: Tsiligkaridis, Athanasios, et al.
Pubblicazione: (2024)
FourierNAT: A Fourier-Mixing-Based Non-Autoregressive Transformer for Parallel Sequence Generation
di: Kiruluta, Andrew, et al.
Pubblicazione: (2025)
di: Kiruluta, Andrew, et al.
Pubblicazione: (2025)
Accelerating Transformer Inference for Translation via Parallel Decoding
di: Santilli, Andrea, et al.
Pubblicazione: (2023)
di: Santilli, Andrea, et al.
Pubblicazione: (2023)
Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer
di: Lu, Wenquan, et al.
Pubblicazione: (2025)
di: Lu, Wenquan, et al.
Pubblicazione: (2025)
Position Information Emerges in Causal Transformers Without Positional Encodings via Similarity of Nearby Embeddings
di: Zuo, Chunsheng, et al.
Pubblicazione: (2024)
di: Zuo, Chunsheng, et al.
Pubblicazione: (2024)
Scaling Test-Time Inference with Policy-Optimized, Dynamic Retrieval-Augmented Generation via KV Caching and Decoding
di: Srinivas, Sakhinana Sagar, et al.
Pubblicazione: (2025)
di: Srinivas, Sakhinana Sagar, et al.
Pubblicazione: (2025)
PARSE: LLM Driven Schema Optimization for Reliable Entity Extraction
di: Shrimal, Anubhav, et al.
Pubblicazione: (2025)
di: Shrimal, Anubhav, et al.
Pubblicazione: (2025)
Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generation
di: Ren, Liliang, et al.
Pubblicazione: (2025)
di: Ren, Liliang, et al.
Pubblicazione: (2025)
Development of Pre-Trained Transformer-based Models for the Nepali Language
di: Thapa, Prajwal, et al.
Pubblicazione: (2024)
di: Thapa, Prajwal, et al.
Pubblicazione: (2024)
Can Post-Training Transform LLMs into Causal Reasoners?
di: Chen, Junqi, et al.
Pubblicazione: (2026)
di: Chen, Junqi, et al.
Pubblicazione: (2026)
Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task
di: Curth, Alicia, et al.
Pubblicazione: (2026)
di: Curth, Alicia, et al.
Pubblicazione: (2026)
Decoding Decoded: Understanding Hyperparameter Effects in Open-Ended Text Generation
di: Arias, Esteban Garces, et al.
Pubblicazione: (2024)
di: Arias, Esteban Garces, et al.
Pubblicazione: (2024)
Multiscale Byte Language Models -- A Hierarchical Architecture for Causal Million-Length Sequence Modeling
di: Egli, Eric, et al.
Pubblicazione: (2025)
di: Egli, Eric, et al.
Pubblicazione: (2025)
Documenti analoghi
-
H-Probes: Extracting Hierarchical Structures From Latent Representations of Language Models
di: Dawes, Cutter, et al.
Pubblicazione: (2026) -
Linear probes rely on textual evidence: Results from leakage mitigation studies in language models
di: Boxo, Gerard, et al.
Pubblicazione: (2025) -
Causal Reflection with Language Models
di: Aryan, Abi, et al.
Pubblicazione: (2025) -
Transforming Chatbot Text: A Sequence-to-Sequence Approach
di: Reddy, Natesh, et al.
Pubblicazione: (2025) -
Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models
di: Roy, Amartya, et al.
Pubblicazione: (2025)