The Impact of Element Ordering on LM Agent Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Chi, Wayne, Talwalkar, Ameet, Donahue, Chris |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UPS: Efficiently Building Foundation Models for PDE Solving via Cross-Modal Adaptation
by: Shen, Junhong, et al.
Published: (2024)
by: Shen, Junhong, et al.
Published: (2024)
CoMind: Towards Community-Driven Agents for Machine Learning Engineering
by: Li, Sijie, et al.
Published: (2025)
by: Li, Sijie, et al.
Published: (2025)
Agreement-Based Cascading for Efficient Inference
by: Kolawole, Steven, et al.
Published: (2024)
by: Kolawole, Steven, et al.
Published: (2024)
Provably tuning the ElasticNet across instances
by: Balcan, Maria-Florina, et al.
Published: (2022)
by: Balcan, Maria-Florina, et al.
Published: (2022)
FrontierCO: Real-World and Large-Scale Evaluation of Machine Learning Solvers for Combinatorial Optimization
by: Feng, Shengyu, et al.
Published: (2025)
by: Feng, Shengyu, et al.
Published: (2025)
Learning to Relax: Setting Solver Parameters Across a Sequence of Linear System Instances
by: Khodak, Mikhail, et al.
Published: (2023)
by: Khodak, Mikhail, et al.
Published: (2023)
Spend Less, Fit Better: Budget-Efficient Scaling Law Fitting via Active Experiment Selection
by: Li, Sijie, et al.
Published: (2026)
by: Li, Sijie, et al.
Published: (2026)
Multitask Learning Can Improve Worst-Group Outcomes
by: Kulkarni, Atharva, et al.
Published: (2023)
by: Kulkarni, Atharva, et al.
Published: (2023)
Where Does My Model Underperform? A Human Evaluation of Slice Discovery Algorithms
by: Johnson, Nari, et al.
Published: (2023)
by: Johnson, Nari, et al.
Published: (2023)
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
by: Kolawole, Steven, et al.
Published: (2024)
by: Kolawole, Steven, et al.
Published: (2024)
Pre-Generating Multi-Difficulty PDE Data for Few-Shot Neural PDE Solvers
by: Choudhary, Naman, et al.
Published: (2025)
by: Choudhary, Naman, et al.
Published: (2025)
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
by: Ruan, Yangjun, et al.
Published: (2023)
by: Ruan, Yangjun, et al.
Published: (2023)
Understanding Optimization in Deep Learning with Central Flows
by: Cohen, Jeremy M., et al.
Published: (2024)
by: Cohen, Jeremy M., et al.
Published: (2024)
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
by: Shen, Junhong, et al.
Published: (2025)
by: Shen, Junhong, et al.
Published: (2025)
Learn Hard Problems During RL with Reference Guided Fine-tuning
by: Wu, Yangzhen, et al.
Published: (2026)
by: Wu, Yangzhen, et al.
Published: (2026)
Comparing Developer and LLM Biases in Code Evaluation
by: Mittal, Aditya, et al.
Published: (2026)
by: Mittal, Aditya, et al.
Published: (2026)
CodePDE: An Inference Framework for LLM-driven PDE Solver Generation
by: Li, Shanda, et al.
Published: (2025)
by: Li, Shanda, et al.
Published: (2025)
Sample Complexity and Representation Ability of Test-time Scaling Paradigms
by: Huang, Baihe, et al.
Published: (2025)
by: Huang, Baihe, et al.
Published: (2025)
Specialized Foundation Models Struggle to Beat Supervised Baselines
by: Xu, Zongzhe, et al.
Published: (2024)
by: Xu, Zongzhe, et al.
Published: (2024)
RECODE: Reasoning Through Code Generation for Visual Question Answering
by: Shen, Junhong, et al.
Published: (2025)
by: Shen, Junhong, et al.
Published: (2025)
Learning Personalized Decision Support Policies
by: Bhatt, Umang, et al.
Published: (2023)
by: Bhatt, Umang, et al.
Published: (2023)
Just Label the Repeats for In-The-Wild Audio-to-Score Alignment
by: Bukey, Irmak, et al.
Published: (2024)
by: Bukey, Irmak, et al.
Published: (2024)
Rethinking Music Captioning with Music Metadata LLMs
by: Bukey, Irmak, et al.
Published: (2026)
by: Bukey, Irmak, et al.
Published: (2026)
Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
by: Long, Phillip, et al.
Published: (2026)
by: Long, Phillip, et al.
Published: (2026)
EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
by: Chi, Wayne, et al.
Published: (2025)
by: Chi, Wayne, et al.
Published: (2025)
ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response
by: Xie, Stephan, et al.
Published: (2026)
by: Xie, Stephan, et al.
Published: (2026)
Toto 2.0: Time Series Forecasting Enters the Scaling Era
by: Khwaja, Emaad, et al.
Published: (2026)
by: Khwaja, Emaad, et al.
Published: (2026)
SeisLM: a Foundation Model for Seismic Waveforms
by: Liu, Tianlin, et al.
Published: (2024)
by: Liu, Tianlin, et al.
Published: (2024)
Anticipatory Music Transformer
by: Thickstun, John, et al.
Published: (2023)
by: Thickstun, John, et al.
Published: (2023)
The Impact of Variable Ordering on Bayesian Network Structure Learning
by: Kitson, Neville K, et al.
Published: (2022)
by: Kitson, Neville K, et al.
Published: (2022)
Scalable Time-Series Causal Discovery with Approximate Causal Ordering
by: Jiao, Ziyang, et al.
Published: (2024)
by: Jiao, Ziyang, et al.
Published: (2024)
Modulating Language Model Experiences through Frictions
by: Collins, Katherine M., et al.
Published: (2024)
by: Collins, Katherine M., et al.
Published: (2024)
Copilot Arena: A Platform for Code LLM Evaluation in the Wild
by: Chi, Wayne, et al.
Published: (2025)
by: Chi, Wayne, et al.
Published: (2025)
ExLM: Rethinking the Impact of [MASK] Tokens in Masked Language Models
by: Zheng, Kangjie, et al.
Published: (2025)
by: Zheng, Kangjie, et al.
Published: (2025)
Local deployment of large-scale music AI models on commodity hardware
by: Zhou, Xun, et al.
Published: (2024)
by: Zhou, Xun, et al.
Published: (2024)
Imitation Learning for Multi-turn LM Agents via On-policy Expert Corrections
by: Lauffer, Niklas, et al.
Published: (2025)
by: Lauffer, Niklas, et al.
Published: (2025)
Impact of Decentralized Learning on Player Utilities in Stackelberg Games
by: Donahue, Kate, et al.
Published: (2024)
by: Donahue, Kate, et al.
Published: (2024)
Label-consistent clustering for evolving data
by: Gadekar, Ameet, et al.
Published: (2025)
by: Gadekar, Ameet, et al.
Published: (2025)
Do Music Generation Models Encode Music Theory?
by: Wei, Megan, et al.
Published: (2024)
by: Wei, Megan, et al.
Published: (2024)
GameDevBench: Evaluating Agentic Capabilities Through Game Development
by: Chi, Wayne, et al.
Published: (2026)
by: Chi, Wayne, et al.
Published: (2026)
Similar Items
-
UPS: Efficiently Building Foundation Models for PDE Solving via Cross-Modal Adaptation
by: Shen, Junhong, et al.
Published: (2024) -
CoMind: Towards Community-Driven Agents for Machine Learning Engineering
by: Li, Sijie, et al.
Published: (2025) -
Agreement-Based Cascading for Efficient Inference
by: Kolawole, Steven, et al.
Published: (2024) -
Provably tuning the ElasticNet across instances
by: Balcan, Maria-Florina, et al.
Published: (2022) -
FrontierCO: Real-World and Large-Scale Evaluation of Machine Learning Solvers for Combinatorial Optimization
by: Feng, Shengyu, et al.
Published: (2025)