Manifold Metric: A Loss Landscape Approach for Predicting Model Performance
Fuente:
arXiv
Salvato in:
| Autori principali: | Malviya, Pranshu, Huang, Jerry, Baratin, Aristide, Fournier, Quentin, Chandar, Sarath |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Lookbehind-SAM: k steps back, 1 step forward
di: Mordido, Gonçalo, et al.
Pubblicazione: (2023)
di: Mordido, Gonçalo, et al.
Pubblicazione: (2023)
CoPeP: Benchmarking Continual Pretraining for Protein Language Models
di: Patil, Darshan, et al.
Pubblicazione: (2026)
di: Patil, Darshan, et al.
Pubblicazione: (2026)
Promoting Exploration in Memory-Augmented Adam using Critical Momenta
di: Malviya, Pranshu, et al.
Pubblicazione: (2023)
di: Malviya, Pranshu, et al.
Pubblicazione: (2023)
Torque-Aware Momentum
di: Malviya, Pranshu, et al.
Pubblicazione: (2024)
di: Malviya, Pranshu, et al.
Pubblicazione: (2024)
Exploring Quantization for Efficient Pre-Training of Transformer Language Models
di: Chitsaz, Kamran, et al.
Pubblicazione: (2024)
di: Chitsaz, Kamran, et al.
Pubblicazione: (2024)
NovoMolGen: Rethinking Molecular Language Model Pretraining
di: Chitsaz, Kamran, et al.
Pubblicazione: (2025)
di: Chitsaz, Kamran, et al.
Pubblicazione: (2025)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
Lazy vs hasty: linearization in deep networks impacts learning schedule based on example difficulty
di: George, Thomas, et al.
Pubblicazione: (2022)
di: George, Thomas, et al.
Pubblicazione: (2022)
Towards Practical Tool Usage for Continually Learning LLMs
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
Do Large Language Models Know How Much They Know?
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing
di: Hashemzadeh, Maryam, et al.
Pubblicazione: (2026)
di: Hashemzadeh, Maryam, et al.
Pubblicazione: (2026)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
Faithfulness Measurable Masked Language Models
di: Madsen, Andreas, et al.
Pubblicazione: (2023)
di: Madsen, Andreas, et al.
Pubblicazione: (2023)
Why Don't Prompt-Based Fairness Metrics Correlate?
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2024)
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2024)
Small Encoders Can Rival Large Decoders in Detecting Groundedness
di: Abbes, Istabrak, et al.
Pubblicazione: (2025)
di: Abbes, Istabrak, et al.
Pubblicazione: (2025)
Are self-explanations from Large Language Models faithful?
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
Mastering Memory Tasks with World Models
di: Samsami, Mohammad Reza, et al.
Pubblicazione: (2024)
di: Samsami, Mohammad Reza, et al.
Pubblicazione: (2024)
Bias in Motion: Theoretical Insights into the Dynamics of Bias in SGD Training
di: Jain, Anchit, et al.
Pubblicazione: (2024)
di: Jain, Anchit, et al.
Pubblicazione: (2024)
Enhancing Joint Motion Prediction for Individuals with Limb Loss Through Model Reprogramming
di: Dey, Sharmita, et al.
Pubblicazione: (2024)
di: Dey, Sharmita, et al.
Pubblicazione: (2024)
Any-Property-Conditional Molecule Generation with Self-Criticism using Spanning Trees
di: Jolicoeur-Martineau, Alexia, et al.
Pubblicazione: (2024)
di: Jolicoeur-Martineau, Alexia, et al.
Pubblicazione: (2024)
Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
di: Nilaksh, et al.
Pubblicazione: (2026)
di: Nilaksh, et al.
Pubblicazione: (2026)
The Expressive Limits of Diagonal SSMs for State-Tracking
di: Shakerinava, Mehran, et al.
Pubblicazione: (2026)
di: Shakerinava, Mehran, et al.
Pubblicazione: (2026)
Generating $π$-Functional Molecules Using STGG+ with Active Learning
di: Jolicoeur-Martineau, Alexia, et al.
Pubblicazione: (2025)
di: Jolicoeur-Martineau, Alexia, et al.
Pubblicazione: (2025)
Parity Requires Unified Input Dependence and Negative Eigenvalues in SSMs
di: Khavari, Behnoush, et al.
Pubblicazione: (2025)
di: Khavari, Behnoush, et al.
Pubblicazione: (2025)
Investigating the Multilingual Calibration Effects of Language Model Instruction-Tuning
di: Huang, Jerry, et al.
Pubblicazione: (2026)
di: Huang, Jerry, et al.
Pubblicazione: (2026)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
Sub-goal Distillation: A Method to Improve Small Language Agents
di: Hashemzadeh, Maryam, et al.
Pubblicazione: (2024)
di: Hashemzadeh, Maryam, et al.
Pubblicazione: (2024)
Intelligent Switching for Reset-Free RL
di: Patil, Darshan, et al.
Pubblicazione: (2024)
di: Patil, Darshan, et al.
Pubblicazione: (2024)
Neural Coherence : Find higher performance to out-of-distribution tasks from few samples
di: Guiroy, Simon, et al.
Pubblicazione: (2025)
di: Guiroy, Simon, et al.
Pubblicazione: (2025)
Steering Large Language Model Activations in Sparse Spaces
di: Bayat, Reza, et al.
Pubblicazione: (2025)
di: Bayat, Reza, et al.
Pubblicazione: (2025)
Layerwise LQR for Geometry-Aware Optimization of Deep Networks
di: Dufort-Labbé, Simon, et al.
Pubblicazione: (2026)
di: Dufort-Labbé, Simon, et al.
Pubblicazione: (2026)
Navigating Potholes with Geometry-Aware Sharpness Minimization
di: Dufort-Labbé, Simon, et al.
Pubblicazione: (2026)
di: Dufort-Labbé, Simon, et al.
Pubblicazione: (2026)
Squeezing More from the Stream : Learning Representation Online for Streaming Reinforcement Learning
di: Nilaksh, et al.
Pubblicazione: (2026)
di: Nilaksh, et al.
Pubblicazione: (2026)
MST-R: Multi-Stage Tuning for Retrieval Systems and Metric Evaluation
di: Malviya, Yash, et al.
Pubblicazione: (2024)
di: Malviya, Yash, et al.
Pubblicazione: (2024)
CrystalGym: A New Benchmark for Materials Discovery Using Reinforcement Learning
di: Govindarajan, Prashant, et al.
Pubblicazione: (2025)
di: Govindarajan, Prashant, et al.
Pubblicazione: (2025)
BindGPT: A Scalable Framework for 3D Molecular Design via Language Modeling and Reinforcement Learning
di: Zholus, Artem, et al.
Pubblicazione: (2024)
di: Zholus, Artem, et al.
Pubblicazione: (2024)
Interpretability Needs a New Paradigm
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
Should We Attend More or Less? Modulating Attention for Fairness
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2023)
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2023)
GRPO-$λ$: Credit Assignment improves LLM Reasoning
di: Parthasarathi, Prasanna, et al.
Pubblicazione: (2025)
di: Parthasarathi, Prasanna, et al.
Pubblicazione: (2025)
On the Hyperparameter Loss Landscapes of Machine Learning Models: An Exploratory Study
di: Huang, Mingyu, et al.
Pubblicazione: (2023)
di: Huang, Mingyu, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Lookbehind-SAM: k steps back, 1 step forward
di: Mordido, Gonçalo, et al.
Pubblicazione: (2023) -
CoPeP: Benchmarking Continual Pretraining for Protein Language Models
di: Patil, Darshan, et al.
Pubblicazione: (2026) -
Promoting Exploration in Memory-Augmented Adam using Critical Momenta
di: Malviya, Pranshu, et al.
Pubblicazione: (2023) -
Torque-Aware Momentum
di: Malviya, Pranshu, et al.
Pubblicazione: (2024) -
Exploring Quantization for Efficient Pre-Training of Transformer Language Models
di: Chitsaz, Kamran, et al.
Pubblicazione: (2024)