Salvato in:
| Autori principali: | Jiang, Yichen, Zhou, Xiang, Bansal, Mohit |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2402.06492 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
di: Yadav, Prateek, et al.
Pubblicazione: (2023)
di: Yadav, Prateek, et al.
Pubblicazione: (2023)
Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
di: Xiao, Hanqi, et al.
Pubblicazione: (2025)
di: Xiao, Hanqi, et al.
Pubblicazione: (2025)
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
di: Sung, Yi-Lin, et al.
Pubblicazione: (2025)
di: Sung, Yi-Lin, et al.
Pubblicazione: (2025)
UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning
di: Patil, Vaidehi, et al.
Pubblicazione: (2025)
di: Patil, Vaidehi, et al.
Pubblicazione: (2025)
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
di: Hase, Peter, et al.
Pubblicazione: (2024)
di: Hase, Peter, et al.
Pubblicazione: (2024)
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2023)
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2023)
Soft Self-Consistency Improves Language Model Agents
di: Wang, Han, et al.
Pubblicazione: (2024)
di: Wang, Han, et al.
Pubblicazione: (2024)
DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback
di: Khan, Zaid, et al.
Pubblicazione: (2024)
di: Khan, Zaid, et al.
Pubblicazione: (2024)
Multi-Attribute Steering of Language Models via Targeted Intervention
di: Nguyen, Duy, et al.
Pubblicazione: (2025)
di: Nguyen, Duy, et al.
Pubblicazione: (2025)
ReGAL: Refactoring Programs to Discover Generalizable Abstractions
di: Stengel-Eskin, Elias, et al.
Pubblicazione: (2024)
di: Stengel-Eskin, Elias, et al.
Pubblicazione: (2024)
EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents
di: Zala, Abhay, et al.
Pubblicazione: (2024)
di: Zala, Abhay, et al.
Pubblicazione: (2024)
Effective Reasoning Chains Reduce Intrinsic Dimensionality
di: Prasad, Archiki, et al.
Pubblicazione: (2026)
di: Prasad, Archiki, et al.
Pubblicazione: (2026)
Executable Functional Abstractions: Inferring Generative Programs for Advanced Math Problems
di: Khan, Zaid, et al.
Pubblicazione: (2025)
di: Khan, Zaid, et al.
Pubblicazione: (2025)
One Life to Learn: Inferring Symbolic World Models for Stochastic Environments from Unguided Exploration
di: Khan, Zaid, et al.
Pubblicazione: (2025)
di: Khan, Zaid, et al.
Pubblicazione: (2025)
Confidence-aware Self-Semantic Distillation on Knowledge Graph Embedding
di: Liu, Yichen, et al.
Pubblicazione: (2022)
di: Liu, Yichen, et al.
Pubblicazione: (2022)
Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models
di: Prasad, Archiki, et al.
Pubblicazione: (2023)
di: Prasad, Archiki, et al.
Pubblicazione: (2023)
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
di: Sung, Yi-Lin, et al.
Pubblicazione: (2023)
di: Sung, Yi-Lin, et al.
Pubblicazione: (2023)
Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
di: Saha, Swarnadeep, et al.
Pubblicazione: (2023)
di: Saha, Swarnadeep, et al.
Pubblicazione: (2023)
Zero-Training Temporal Drift Detection for Transformer Sentiment Models: A Comprehensive Analysis on Authentic Social Media Streams
di: Bansal, Aayam, et al.
Pubblicazione: (2025)
di: Bansal, Aayam, et al.
Pubblicazione: (2025)
Evaluating Very Long-Term Conversational Memory of LLM Agents
di: Maharana, Adyasha, et al.
Pubblicazione: (2024)
di: Maharana, Adyasha, et al.
Pubblicazione: (2024)
Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2025)
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2025)
Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind
di: Xiao, Hanqi, et al.
Pubblicazione: (2026)
di: Xiao, Hanqi, et al.
Pubblicazione: (2026)
DiagrammerGPT: Generating Open-Domain, Open-Platform Diagrams via LLM Planning
di: Zala, Abhay, et al.
Pubblicazione: (2023)
di: Zala, Abhay, et al.
Pubblicazione: (2023)
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
di: Lin, Han, et al.
Pubblicazione: (2023)
di: Lin, Han, et al.
Pubblicazione: (2023)
What Matters for Model Merging at Scale?
di: Yadav, Prateek, et al.
Pubblicazione: (2024)
di: Yadav, Prateek, et al.
Pubblicazione: (2024)
MixCE: Training Autoregressive Language Models by Mixing Forward and Reverse Cross-Entropies
di: Zhang, Shiyue, et al.
Pubblicazione: (2023)
di: Zhang, Shiyue, et al.
Pubblicazione: (2023)
ADaPT: As-Needed Decomposition and Planning with Language Models
di: Prasad, Archiki, et al.
Pubblicazione: (2023)
di: Prasad, Archiki, et al.
Pubblicazione: (2023)
System-1.x: Learning to Balance Fast and Slow Planning with Language Models
di: Saha, Swarnadeep, et al.
Pubblicazione: (2024)
di: Saha, Swarnadeep, et al.
Pubblicazione: (2024)
Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression
di: Singh, Joykirat, et al.
Pubblicazione: (2025)
di: Singh, Joykirat, et al.
Pubblicazione: (2025)
Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy
di: Li, Pingzhi, et al.
Pubblicazione: (2023)
di: Li, Pingzhi, et al.
Pubblicazione: (2023)
PRInTS: Reward Modeling for Long-Horizon Information Seeking
di: Lee, Jaewoo, et al.
Pubblicazione: (2025)
di: Lee, Jaewoo, et al.
Pubblicazione: (2025)
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
di: Wan, David, et al.
Pubblicazione: (2024)
di: Wan, David, et al.
Pubblicazione: (2024)
Exposing and Addressing Cross-Task Inconsistency in Unified Vision-Language Models
di: Maharana, Adyasha, et al.
Pubblicazione: (2023)
di: Maharana, Adyasha, et al.
Pubblicazione: (2023)
Cog-DRIFT: Exploration on Adaptively Reformulated Instances Enables Learning from Hard Reasoning Problems
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2026)
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2026)
Resolving Conflicts in Lifelong Learning via Aligning Updates in Subspaces
di: Zhou, Yueer, et al.
Pubblicazione: (2025)
di: Zhou, Yueer, et al.
Pubblicazione: (2025)
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
Should We Attend More or Less? Modulating Attention for Fairness
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2023)
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2023)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
di: Zandieh, Amir, et al.
Pubblicazione: (2024)
di: Zandieh, Amir, et al.
Pubblicazione: (2024)
Anyprefer: An Agentic Framework for Preference Data Synthesis
di: Zhou, Yiyang, et al.
Pubblicazione: (2025)
di: Zhou, Yiyang, et al.
Pubblicazione: (2025)
A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning
di: Yadav, Prateek, et al.
Pubblicazione: (2024)
di: Yadav, Prateek, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
di: Yadav, Prateek, et al.
Pubblicazione: (2023) -
Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
di: Xiao, Hanqi, et al.
Pubblicazione: (2025) -
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
di: Sung, Yi-Lin, et al.
Pubblicazione: (2025) -
UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning
di: Patil, Vaidehi, et al.
Pubblicazione: (2025) -
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
di: Hase, Peter, et al.
Pubblicazione: (2024)