Saved in:
| Main Authors: | Chevalier, Alexis, Geng, Jiayi, Wettig, Alexander, Chen, Howard, Mizera, Sebastian, Annala, Toni, Aragon, Max Jameson, Fanlo, Arturo Rodríguez, Frieder, Simon, Machado, Simon, Prabhakar, Akshara, Thieu, Ellie, Wang, Jiachen T., Wang, Zirui, Wu, Xindi, Xia, Mengzhou, Xia, Wenhan, Yu, Jiatong, Zhu, Jun-Jie, Ren, Zhiyong Jason, Arora, Sanjeev, Chen, Danqi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2402.11111 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LESS: Selecting Influential Data for Targeted Instruction Tuning
by: Xia, Mengzhou, et al.
Published: (2024)
by: Xia, Mengzhou, et al.
Published: (2024)
CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
by: Wang, Zirui, et al.
Published: (2024)
by: Wang, Zirui, et al.
Published: (2024)
Trainable Transformer in Transformer
by: Panigrahi, Abhishek, et al.
Published: (2023)
by: Panigrahi, Abhishek, et al.
Published: (2023)
SimPO: Simple Preference Optimization with a Reference-Free Reward
by: Meng, Yu, et al.
Published: (2024)
by: Meng, Yu, et al.
Published: (2024)
LitSearch: A Retrieval Benchmark for Scientific Literature Search
by: Ajith, Anirudh, et al.
Published: (2024)
by: Ajith, Anirudh, et al.
Published: (2024)
Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
by: Xia, Mengzhou, et al.
Published: (2023)
by: Xia, Mengzhou, et al.
Published: (2023)
Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training
by: Zhong, Zexuan, et al.
Published: (2024)
by: Zhong, Zexuan, et al.
Published: (2024)
Motivic Steenrod problem away from the characteristic
by: Annala, Toni, et al.
Published: (2024)
by: Annala, Toni, et al.
Published: (2024)
Motivic spectra and universality of $K$-theory
by: Annala, Toni, et al.
Published: (2022)
by: Annala, Toni, et al.
Published: (2022)
A note on weight filtrations at the characteristic
by: Annala, Toni, et al.
Published: (2025)
by: Annala, Toni, et al.
Published: (2025)
Motivic Steenrod operations at the characteristic via infinite ramification
by: Annala, Toni, et al.
Published: (2025)
by: Annala, Toni, et al.
Published: (2025)
Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving
by: Lin, Yong, et al.
Published: (2025)
by: Lin, Yong, et al.
Published: (2025)
Atiyah duality for motivic spectra
by: Annala, Toni, et al.
Published: (2024)
by: Annala, Toni, et al.
Published: (2024)
Algebraic cobordism and a Conner-Floyd isomorphism for algebraic K-theory
by: Annala, Toni, et al.
Published: (2023)
by: Annala, Toni, et al.
Published: (2023)
QuRating: Selecting High-Quality Data for Training Language Models
by: Wettig, Alexander, et al.
Published: (2024)
by: Wettig, Alexander, et al.
Published: (2024)
Finding Transformer Circuits with Edge Pruning
by: Bhaskar, Adithya, et al.
Published: (2024)
by: Bhaskar, Adithya, et al.
Published: (2024)
How to Train Long-Context Language Models (Effectively)
by: Gao, Tianyu, et al.
Published: (2024)
by: Gao, Tianyu, et al.
Published: (2024)
Extracting Rule-based Descriptions of Attention Features in Transformers
by: Friedman, Dan, et al.
Published: (2025)
by: Friedman, Dan, et al.
Published: (2025)
The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
by: Zhu, Xinyu, et al.
Published: (2025)
by: Zhu, Xinyu, et al.
Published: (2025)
On the Impossibility of Retrain Equivalence in Machine Unlearning
by: Yu, Jiatong, et al.
Published: (2025)
by: Yu, Jiatong, et al.
Published: (2025)
Unknown Aware AI-Generated Content Attribution
by: Thieu, Ellie, et al.
Published: (2026)
by: Thieu, Ellie, et al.
Published: (2026)
ICONS: Influence Consensus for Vision-Language Data Selection
by: Wu, Xindi, et al.
Published: (2024)
by: Wu, Xindi, et al.
Published: (2024)
How Does RL Post-training Induce Skill Composition? A Case Study on Countdown
by: Park, Simon, et al.
Published: (2025)
by: Park, Simon, et al.
Published: (2025)
Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search
by: Yen, Howard, et al.
Published: (2025)
by: Yen, Howard, et al.
Published: (2025)
Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?
by: Bhaskar, Adithya, et al.
Published: (2025)
by: Bhaskar, Adithya, et al.
Published: (2025)
ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty
by: Wu, Xindi, et al.
Published: (2024)
by: Wu, Xindi, et al.
Published: (2024)
Landau Singularities Revisited: Computational Algebraic Geometry for Feynman Integrals
by: Fevola, Claudia, et al.
Published: (2023)
by: Fevola, Claudia, et al.
Published: (2023)
Principal Landau Determinants
by: Fevola, Claudia, et al.
Published: (2023)
by: Fevola, Claudia, et al.
Published: (2023)
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
by: Razin, Noam, et al.
Published: (2024)
by: Razin, Noam, et al.
Published: (2024)
Detecting Pretraining Data from Large Language Models
by: Shi, Weijia, et al.
Published: (2023)
by: Shi, Weijia, et al.
Published: (2023)
No LLM Solved Yu Tsumura's 554th Problem
by: Frieder, Simon, et al.
Published: (2025)
by: Frieder, Simon, et al.
Published: (2025)
Afectos e intersticios. David Vilaseca y la autobiografía errática
by: Isaias Fanlo
Published: (2024)
by: Isaias Fanlo
Published: (2024)
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning
by: Kaur, Simran, et al.
Published: (2024)
by: Kaur, Simran, et al.
Published: (2024)
Organize the Web: Constructing Domains Enhances Pre-Training Data Curation
by: Wettig, Alexander, et al.
Published: (2025)
by: Wettig, Alexander, et al.
Published: (2025)
Metadata Conditioning Accelerates Language Model Pre-training
by: Gao, Tianyu, et al.
Published: (2025)
by: Gao, Tianyu, et al.
Published: (2025)
What is in Your Safe Data? Identifying Benign Data that Breaks Safety
by: He, Luxi, et al.
Published: (2024)
by: He, Luxi, et al.
Published: (2024)
A novel gauge-equivariant neural-network architecture for preconditioners in lattice QCD
by: Pfahler, Simon, et al.
Published: (2026)
by: Pfahler, Simon, et al.
Published: (2026)
Close-up-GS: Enhancing Close-Up View Synthesis in 3D Gaussian Splatting with Progressive Self-Training
by: Xia, Jiatong, et al.
Published: (2025)
by: Xia, Jiatong, et al.
Published: (2025)
Training-Free Instance-Aware 3D Scene Reconstruction and Diffusion-Based View Synthesis from Sparse Images
by: Xia, Jiatong, et al.
Published: (2026)
by: Xia, Jiatong, et al.
Published: (2026)
Deciphering the Factors Influencing the Efficacy of Chain-of-Thought: Probability, Memorization, and Noisy Reasoning
by: Prabhakar, Akshara, et al.
Published: (2024)
by: Prabhakar, Akshara, et al.
Published: (2024)
Similar Items
-
LESS: Selecting Influential Data for Targeted Instruction Tuning
by: Xia, Mengzhou, et al.
Published: (2024) -
CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
by: Wang, Zirui, et al.
Published: (2024) -
Trainable Transformer in Transformer
by: Panigrahi, Abhishek, et al.
Published: (2023) -
SimPO: Simple Preference Optimization with a Reference-Free Reward
by: Meng, Yu, et al.
Published: (2024) -
LitSearch: A Retrieval Benchmark for Scientific Literature Search
by: Ajith, Anirudh, et al.
Published: (2024)