Explicitly Encoding Structural Symmetry is Key to Length Generalization in Arithmetic Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | Sabbaghi, Mahdi, Pappas, George, Hassani, Hamed, Goel, Surbhi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
InfoSFT: Learn More and Forget Less with Information-Aware Token Weighting
di: Sabbaghi, Mahdi, et al.
Pubblicazione: (2026)
di: Sabbaghi, Mahdi, et al.
Pubblicazione: (2026)
Robust Policy Optimization to Prevent Catastrophic Forgetting
di: Sabbaghi, Mahdi, et al.
Pubblicazione: (2026)
di: Sabbaghi, Mahdi, et al.
Pubblicazione: (2026)
Adversarial Reasoning at Jailbreaking Time
di: Sabbaghi, Mahdi, et al.
Pubblicazione: (2025)
di: Sabbaghi, Mahdi, et al.
Pubblicazione: (2025)
Length Optimization in Conformal Prediction
di: Kiyani, Shayan, et al.
Pubblicazione: (2024)
di: Kiyani, Shayan, et al.
Pubblicazione: (2024)
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
di: Cho, Hanseul, et al.
Pubblicazione: (2024)
di: Cho, Hanseul, et al.
Pubblicazione: (2024)
From Interpolation to Extrapolation: Complete Length Generalization for Arithmetic Transformers
di: Duan, Shaoxiong, et al.
Pubblicazione: (2023)
di: Duan, Shaoxiong, et al.
Pubblicazione: (2023)
In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
di: Panigrahi, Abhishek, et al.
Pubblicazione: (2025)
di: Panigrahi, Abhishek, et al.
Pubblicazione: (2025)
Conformal Language Model Reasoning with Coherent Factuality
di: Rubin-Toles, Maxon, et al.
Pubblicazione: (2025)
di: Rubin-Toles, Maxon, et al.
Pubblicazione: (2025)
Task Structure Reverses Layerwise State Encoding in Sequence Models
di: Jiang, Yuhang
Pubblicazione: (2026)
di: Jiang, Yuhang
Pubblicazione: (2026)
Human-AI Collaborative Uncertainty Quantification
di: Noorani, Sima, et al.
Pubblicazione: (2025)
di: Noorani, Sima, et al.
Pubblicazione: (2025)
Evaluating the Performance of Large Language Models via Debates
di: Moniri, Behrad, et al.
Pubblicazione: (2024)
di: Moniri, Behrad, et al.
Pubblicazione: (2024)
Conformal Prediction with Learned Features
di: Kiyani, Shayan, et al.
Pubblicazione: (2024)
di: Kiyani, Shayan, et al.
Pubblicazione: (2024)
Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge
di: Chen, Yan-Lun, et al.
Pubblicazione: (2025)
di: Chen, Yan-Lun, et al.
Pubblicazione: (2025)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
di: Gambardella, Andrew, et al.
Pubblicazione: (2024)
di: Gambardella, Andrew, et al.
Pubblicazione: (2024)
Fusion Matters: Length-Aware Analysis of Positional-Encoding Fusion in Transformers
di: Hallam, Mohamed Amine, et al.
Pubblicazione: (2026)
di: Hallam, Mohamed Amine, et al.
Pubblicazione: (2026)
Conformal Prediction Beyond the Seen: A Missing Mass Perspective for Uncertainty Quantification in Generative Models
di: Noorani, Sima, et al.
Pubblicazione: (2025)
di: Noorani, Sima, et al.
Pubblicazione: (2025)
Automated Black-box Prompt Engineering for Personalized Text-to-Image Generation
di: He, Yutong, et al.
Pubblicazione: (2024)
di: He, Yutong, et al.
Pubblicazione: (2024)
Watermarking Language Models with Error Correcting Codes
di: Chao, Patrick, et al.
Pubblicazione: (2024)
di: Chao, Patrick, et al.
Pubblicazione: (2024)
Multi-Round Human-AI Collaboration with User-Specified Requirements
di: Noorani, Sima, et al.
Pubblicazione: (2026)
di: Noorani, Sima, et al.
Pubblicazione: (2026)
Exploring Ordinality in Text Classification: A Comparative Study of Explicit and Implicit Techniques
di: Kasa, Siva Rajesh, et al.
Pubblicazione: (2024)
di: Kasa, Siva Rajesh, et al.
Pubblicazione: (2024)
Logicbreaks: A Framework for Understanding Subversion of Rule-based Inference
di: Xue, Anton, et al.
Pubblicazione: (2024)
di: Xue, Anton, et al.
Pubblicazione: (2024)
Unraveling Arithmetic in Large Language Models: The Role of Algebraic Structures
di: Chang, Fu-Chieh, et al.
Pubblicazione: (2024)
di: Chang, Fu-Chieh, et al.
Pubblicazione: (2024)
IGC: Integrating a Gated Calculator into an LLM to Solve Arithmetic Tasks Reliably and Efficiently
di: Dietz, Florian, et al.
Pubblicazione: (2025)
di: Dietz, Florian, et al.
Pubblicazione: (2025)
Investigating Task Arithmetic for Zero-Shot Information Retrieval
di: Braga, Marco, et al.
Pubblicazione: (2025)
di: Braga, Marco, et al.
Pubblicazione: (2025)
Trellis: Learning to Compress Key-Value Memory in Attention Models
di: Karami, Mahdi, et al.
Pubblicazione: (2025)
di: Karami, Mahdi, et al.
Pubblicazione: (2025)
On Provable Length and Compositional Generalization
di: Ahuja, Kartik, et al.
Pubblicazione: (2024)
di: Ahuja, Kartik, et al.
Pubblicazione: (2024)
Decision Theoretic Foundations for Conformal Prediction: Optimal Uncertainty Quantification for Risk-Averse Agents
di: Kiyani, Shayan, et al.
Pubblicazione: (2025)
di: Kiyani, Shayan, et al.
Pubblicazione: (2025)
When to Trust the Cheap Check: Weak and Strong Verification for Reasoning
di: Kiyani, Shayan, et al.
Pubblicazione: (2026)
di: Kiyani, Shayan, et al.
Pubblicazione: (2026)
Robust Decision Making with Partially Calibrated Forecasts
di: Kiyani, Shayan, et al.
Pubblicazione: (2025)
di: Kiyani, Shayan, et al.
Pubblicazione: (2025)
Mining Mental Health Signals: A Comparative Study of Four Machine Learning Methods for Depression Detection from Social Media Posts in Sorani Kurdish
di: Mohammed, Idrees, et al.
Pubblicazione: (2025)
di: Mohammed, Idrees, et al.
Pubblicazione: (2025)
Improving Variable-Length Generation in Diffusion Language Models via Length Regularization
di: Cheng, Zicong, et al.
Pubblicazione: (2026)
di: Cheng, Zicong, et al.
Pubblicazione: (2026)
Temporal Difference Learning with Compressed Updates: Error-Feedback meets Reinforcement Learning
di: Mitra, Aritra, et al.
Pubblicazione: (2023)
di: Mitra, Aritra, et al.
Pubblicazione: (2023)
StrAE: Autoencoding for Pre-Trained Embeddings using Explicit Structure
di: Opper, Mattia, et al.
Pubblicazione: (2023)
di: Opper, Mattia, et al.
Pubblicazione: (2023)
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation
di: He, Zhenyu, et al.
Pubblicazione: (2024)
di: He, Zhenyu, et al.
Pubblicazione: (2024)
Adaptively profiling models with task elicitation
di: Brown, Davis, et al.
Pubblicazione: (2025)
di: Brown, Davis, et al.
Pubblicazione: (2025)
Multilingual Language Models Encode Script Over Linguistic Structure
di: Verma, Aastha A K, et al.
Pubblicazione: (2026)
di: Verma, Aastha A K, et al.
Pubblicazione: (2026)
Localize-and-Stitch: Efficient Model Merging via Sparse Task Arithmetic
di: He, Yifei, et al.
Pubblicazione: (2024)
di: He, Yifei, et al.
Pubblicazione: (2024)
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
di: Robey, Alexander, et al.
Pubblicazione: (2023)
di: Robey, Alexander, et al.
Pubblicazione: (2023)
Language Models are Symbolic Learners in Arithmetic
di: Deng, Chunyuan, et al.
Pubblicazione: (2024)
di: Deng, Chunyuan, et al.
Pubblicazione: (2024)
Steering Language Models with Weight Arithmetic
di: Fierro, Constanza, et al.
Pubblicazione: (2025)
di: Fierro, Constanza, et al.
Pubblicazione: (2025)
Documenti analoghi
-
InfoSFT: Learn More and Forget Less with Information-Aware Token Weighting
di: Sabbaghi, Mahdi, et al.
Pubblicazione: (2026) -
Robust Policy Optimization to Prevent Catastrophic Forgetting
di: Sabbaghi, Mahdi, et al.
Pubblicazione: (2026) -
Adversarial Reasoning at Jailbreaking Time
di: Sabbaghi, Mahdi, et al.
Pubblicazione: (2025) -
Length Optimization in Conformal Prediction
di: Kiyani, Shayan, et al.
Pubblicazione: (2024) -
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
di: Cho, Hanseul, et al.
Pubblicazione: (2024)