Inducing Systematicity in Transformers by Attending to Structurally Quantized Embeddings
Fuente:
arXiv
Guardado en:
| Autores principales: | Jiang, Yichen, Zhou, Xiang, Bansal, Mohit |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
por: Yadav, Prateek, et al.
Publicado: (2023)
por: Yadav, Prateek, et al.
Publicado: (2023)
Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
por: Xiao, Hanqi, et al.
Publicado: (2025)
por: Xiao, Hanqi, et al.
Publicado: (2025)
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
por: Sung, Yi-Lin, et al.
Publicado: (2025)
por: Sung, Yi-Lin, et al.
Publicado: (2025)
UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning
por: Patil, Vaidehi, et al.
Publicado: (2025)
por: Patil, Vaidehi, et al.
Publicado: (2025)
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
por: Hase, Peter, et al.
Publicado: (2024)
por: Hase, Peter, et al.
Publicado: (2024)
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
por: Chen, Justin Chih-Yao, et al.
Publicado: (2023)
por: Chen, Justin Chih-Yao, et al.
Publicado: (2023)
Soft Self-Consistency Improves Language Model Agents
por: Wang, Han, et al.
Publicado: (2024)
por: Wang, Han, et al.
Publicado: (2024)
DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback
por: Khan, Zaid, et al.
Publicado: (2024)
por: Khan, Zaid, et al.
Publicado: (2024)
Multi-Attribute Steering of Language Models via Targeted Intervention
por: Nguyen, Duy, et al.
Publicado: (2025)
por: Nguyen, Duy, et al.
Publicado: (2025)
EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents
por: Zala, Abhay, et al.
Publicado: (2024)
por: Zala, Abhay, et al.
Publicado: (2024)
Effective Reasoning Chains Reduce Intrinsic Dimensionality
por: Prasad, Archiki, et al.
Publicado: (2026)
por: Prasad, Archiki, et al.
Publicado: (2026)
Confidence-aware Self-Semantic Distillation on Knowledge Graph Embedding
por: Liu, Yichen, et al.
Publicado: (2022)
por: Liu, Yichen, et al.
Publicado: (2022)
Executable Functional Abstractions: Inferring Generative Programs for Advanced Math Problems
por: Khan, Zaid, et al.
Publicado: (2025)
por: Khan, Zaid, et al.
Publicado: (2025)
One Life to Learn: Inferring Symbolic World Models for Stochastic Environments from Unguided Exploration
por: Khan, Zaid, et al.
Publicado: (2025)
por: Khan, Zaid, et al.
Publicado: (2025)
ReGAL: Refactoring Programs to Discover Generalizable Abstractions
por: Stengel-Eskin, Elias, et al.
Publicado: (2024)
por: Stengel-Eskin, Elias, et al.
Publicado: (2024)
Zero-Training Temporal Drift Detection for Transformer Sentiment Models: A Comprehensive Analysis on Authentic Social Media Streams
por: Bansal, Aayam, et al.
Publicado: (2025)
por: Bansal, Aayam, et al.
Publicado: (2025)
Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
por: Saha, Swarnadeep, et al.
Publicado: (2023)
por: Saha, Swarnadeep, et al.
Publicado: (2023)
Evaluating Very Long-Term Conversational Memory of LLM Agents
por: Maharana, Adyasha, et al.
Publicado: (2024)
por: Maharana, Adyasha, et al.
Publicado: (2024)
Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills
por: Chen, Justin Chih-Yao, et al.
Publicado: (2025)
por: Chen, Justin Chih-Yao, et al.
Publicado: (2025)
Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind
por: Xiao, Hanqi, et al.
Publicado: (2026)
por: Xiao, Hanqi, et al.
Publicado: (2026)
Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models
por: Prasad, Archiki, et al.
Publicado: (2023)
por: Prasad, Archiki, et al.
Publicado: (2023)
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
por: Sung, Yi-Lin, et al.
Publicado: (2023)
por: Sung, Yi-Lin, et al.
Publicado: (2023)
What Matters for Model Merging at Scale?
por: Yadav, Prateek, et al.
Publicado: (2024)
por: Yadav, Prateek, et al.
Publicado: (2024)
MixCE: Training Autoregressive Language Models by Mixing Forward and Reverse Cross-Entropies
por: Zhang, Shiyue, et al.
Publicado: (2023)
por: Zhang, Shiyue, et al.
Publicado: (2023)
ADaPT: As-Needed Decomposition and Planning with Language Models
por: Prasad, Archiki, et al.
Publicado: (2023)
por: Prasad, Archiki, et al.
Publicado: (2023)
System-1.x: Learning to Balance Fast and Slow Planning with Language Models
por: Saha, Swarnadeep, et al.
Publicado: (2024)
por: Saha, Swarnadeep, et al.
Publicado: (2024)
Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression
por: Singh, Joykirat, et al.
Publicado: (2025)
por: Singh, Joykirat, et al.
Publicado: (2025)
Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy
por: Li, Pingzhi, et al.
Publicado: (2023)
por: Li, Pingzhi, et al.
Publicado: (2023)
PRInTS: Reward Modeling for Long-Horizon Information Seeking
por: Lee, Jaewoo, et al.
Publicado: (2025)
por: Lee, Jaewoo, et al.
Publicado: (2025)
DiagrammerGPT: Generating Open-Domain, Open-Platform Diagrams via LLM Planning
por: Zala, Abhay, et al.
Publicado: (2023)
por: Zala, Abhay, et al.
Publicado: (2023)
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
por: Lin, Han, et al.
Publicado: (2023)
por: Lin, Han, et al.
Publicado: (2023)
Grouped Sequency-arranged Rotation: Optimizing Rotation Transformation for Quantization for Free
por: Choi, Euntae, et al.
Publicado: (2025)
por: Choi, Euntae, et al.
Publicado: (2025)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
por: Zandieh, Amir, et al.
Publicado: (2024)
por: Zandieh, Amir, et al.
Publicado: (2024)
Resolving Conflicts in Lifelong Learning via Aligning Updates in Subspaces
por: Zhou, Yueer, et al.
Publicado: (2025)
por: Zhou, Yueer, et al.
Publicado: (2025)
Should We Attend More or Less? Modulating Attention for Fairness
por: Zayed, Abdelrahman, et al.
Publicado: (2023)
por: Zayed, Abdelrahman, et al.
Publicado: (2023)
CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending
por: Zhu, Shiyi, et al.
Publicado: (2023)
por: Zhu, Shiyi, et al.
Publicado: (2023)
Cog-DRIFT: Exploration on Adaptively Reformulated Instances Enables Learning from Hard Reasoning Problems
por: Chen, Justin Chih-Yao, et al.
Publicado: (2026)
por: Chen, Justin Chih-Yao, et al.
Publicado: (2026)
Task-Stratified Knowledge Scaling Laws for Post-Training Quantized Large Language Models
por: Zhou, Chenxi, et al.
Publicado: (2025)
por: Zhou, Chenxi, et al.
Publicado: (2025)
From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization
por: Zhou, Chenxi, et al.
Publicado: (2026)
por: Zhou, Chenxi, et al.
Publicado: (2026)
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
por: Wan, David, et al.
Publicado: (2024)
por: Wan, David, et al.
Publicado: (2024)
Ejemplares similares
-
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
por: Yadav, Prateek, et al.
Publicado: (2023) -
Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
por: Xiao, Hanqi, et al.
Publicado: (2025) -
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
por: Sung, Yi-Lin, et al.
Publicado: (2025) -
UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning
por: Patil, Vaidehi, et al.
Publicado: (2025) -
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
por: Hase, Peter, et al.
Publicado: (2024)