Excision Score: Evaluating Edits with Surgical Precision
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gruzinov, Nikolai, Sycheva, Ksenia, Barr, Earl T., Bezzubov, Alex |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation
von: Bagirov, Farid, et al.
Veröffentlicht: (2025)
von: Bagirov, Farid, et al.
Veröffentlicht: (2025)
Diff-XYZ: A Benchmark for Evaluating Diff Understanding
von: Glukhov, Evgeniy, et al.
Veröffentlicht: (2025)
von: Glukhov, Evgeniy, et al.
Veröffentlicht: (2025)
Gaussian Process Tilted Nonparametric Density Estimation using Fisher Divergence Score Matching
von: Paisley, John, et al.
Veröffentlicht: (2025)
von: Paisley, John, et al.
Veröffentlicht: (2025)
How Robustly do LLMs Understand Execution Semantics?
von: Spiess, Claudio, et al.
Veröffentlicht: (2026)
von: Spiess, Claudio, et al.
Veröffentlicht: (2026)
Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings
von: Tsvetkov, Petr, et al.
Veröffentlicht: (2024)
von: Tsvetkov, Petr, et al.
Veröffentlicht: (2024)
Automatic Semantic Augmentation of Language Model Prompts (for Code Summarization)
von: Ahmed, Toufique, et al.
Veröffentlicht: (2023)
von: Ahmed, Toufique, et al.
Veröffentlicht: (2023)
QiMeng-PRepair: Precise Code Repair via Edit-Aware Reward Optimization
von: Ke, Changxin, et al.
Veröffentlicht: (2026)
von: Ke, Changxin, et al.
Veröffentlicht: (2026)
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits
von: Yosef, Ron, et al.
Veröffentlicht: (2025)
von: Yosef, Ron, et al.
Veröffentlicht: (2025)
Edit Flows: Flow Matching with Edit Operations
von: Havasi, Marton, et al.
Veröffentlicht: (2025)
von: Havasi, Marton, et al.
Veröffentlicht: (2025)
GeoEdit: Local Frames for Fast, Training-Free On-Manifold Editing in Diffusion Models
von: Zhang, Yiming, et al.
Veröffentlicht: (2026)
von: Zhang, Yiming, et al.
Veröffentlicht: (2026)
GEDAN: Learning the Edit Costs for Graph Edit Distance
von: Leonardi, Francesco, et al.
Veröffentlicht: (2025)
von: Leonardi, Francesco, et al.
Veröffentlicht: (2025)
EUGENE: Explainable Structure-aware Graph Edit Distance Estimation with Generalized Edit Costs
von: Bommakanti, Aditya, et al.
Veröffentlicht: (2024)
von: Bommakanti, Aditya, et al.
Veröffentlicht: (2024)
Challenge on Optimization of Context Collection for Code Completion
von: Ustalov, Dmitry, et al.
Veröffentlicht: (2025)
von: Ustalov, Dmitry, et al.
Veröffentlicht: (2025)
Denoising Score Matching with Random Features: Insights on Diffusion Models from Precise Learning Curves
von: George, Anand Jerry, et al.
Veröffentlicht: (2025)
von: George, Anand Jerry, et al.
Veröffentlicht: (2025)
OT Score: An OT based Confidence Score for Prototype-Assisted Source Free Unsupervised Domain Adaptation
von: Zhang, Yiming, et al.
Veröffentlicht: (2025)
von: Zhang, Yiming, et al.
Veröffentlicht: (2025)
SpotEdit: Evaluating Visually-Guided Image Editing Methods
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2025)
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2025)
Next Edit Prediction: Learning to Predict Code Edits from Context and Interaction History
von: Lu, Ruofan, et al.
Veröffentlicht: (2025)
von: Lu, Ruofan, et al.
Veröffentlicht: (2025)
CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization
von: Das, Debeshee, et al.
Veröffentlicht: (2025)
von: Das, Debeshee, et al.
Veröffentlicht: (2025)
Generalized Tree Edit Distance (GTED): A Faithful Evaluation Metric for Statement Autoformalization
von: Liu, Yuntian, et al.
Veröffentlicht: (2025)
von: Liu, Yuntian, et al.
Veröffentlicht: (2025)
Differentiable Optimization of Similarity Scores Between Models and Brains
von: Cloos, Nathan, et al.
Veröffentlicht: (2024)
von: Cloos, Nathan, et al.
Veröffentlicht: (2024)
Edit-Based Flow Matching for Temporal Point Processes
von: Lüdke, David, et al.
Veröffentlicht: (2025)
von: Lüdke, David, et al.
Veröffentlicht: (2025)
Metanetworks as Regulatory Operators: Learning to Edit for Requirement Compliance
von: Kalogeropoulos, Ioannis, et al.
Veröffentlicht: (2025)
von: Kalogeropoulos, Ioannis, et al.
Veröffentlicht: (2025)
Search over Self-Edit Strategies for LLM Adaptation
von: Cheong, Alistair, et al.
Veröffentlicht: (2026)
von: Cheong, Alistair, et al.
Veröffentlicht: (2026)
SCOPED: Score-Curvature Out-of-distribution Proximity Evaluator for Diffusion
von: Barkley, Brett, et al.
Veröffentlicht: (2025)
von: Barkley, Brett, et al.
Veröffentlicht: (2025)
Evaluation of Bagging Predictors with Kernel Density Estimation and Bagging Score
von: Seitz, Philipp, et al.
Veröffentlicht: (2026)
von: Seitz, Philipp, et al.
Veröffentlicht: (2026)
Tabular Diffusion Counterfactual Explanations
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
Stochastic Variational Inference with Tuneable Stochastic Annealing
von: Paisley, John, et al.
Veröffentlicht: (2025)
von: Paisley, John, et al.
Veröffentlicht: (2025)
Robust Learning of Diverse Code Edits
von: Aggarwal, Tushar, et al.
Veröffentlicht: (2025)
von: Aggarwal, Tushar, et al.
Veröffentlicht: (2025)
Resolving UnderEdit & OverEdit with Iterative & Neighbor-Assisted Model Editing
von: Baghel, Bhiman Kumar, et al.
Veröffentlicht: (2025)
von: Baghel, Bhiman Kumar, et al.
Veröffentlicht: (2025)
Pandora's Regret: A Proper Scoring Rule for Evaluating Sequential Search
von: Flores, Gerardo A., et al.
Veröffentlicht: (2026)
von: Flores, Gerardo A., et al.
Veröffentlicht: (2026)
Evaluating Posterior Probabilities: Decision Theory, Proper Scoring Rules, and Calibration
von: Ferrer, Luciana, et al.
Veröffentlicht: (2024)
von: Ferrer, Luciana, et al.
Veröffentlicht: (2024)
Bayesian-LoRA: LoRA based Parameter Efficient Fine-Tuning using Optimal Quantization levels and Rank Values trough Differentiable Bayesian Gates
von: Meo, Cristian, et al.
Veröffentlicht: (2024)
von: Meo, Cristian, et al.
Veröffentlicht: (2024)
Efficient Exploration in Deep Reinforcement Learning: A Novel Bayesian Actor-Critic Algorithm
von: Rozanov, Nikolai
Veröffentlicht: (2024)
von: Rozanov, Nikolai
Veröffentlicht: (2024)
PaintBench: Deterministic Evaluation of Precise Visual Editing
von: Xu, Kai, et al.
Veröffentlicht: (2026)
von: Xu, Kai, et al.
Veröffentlicht: (2026)
POME: Post Optimization Model Edit via Muon-style Projection
von: Liu, Yong, et al.
Veröffentlicht: (2025)
von: Liu, Yong, et al.
Veröffentlicht: (2025)
EvoFlows: Evolutionary Edit-Based Flow-Matching for Protein Engineering
von: Deutschmann, Nicolas, et al.
Veröffentlicht: (2026)
von: Deutschmann, Nicolas, et al.
Veröffentlicht: (2026)
Fighting Sampling Bias: A Framework for Training and Evaluating Credit Scoring Models
von: Kozodoi, Nikita, et al.
Veröffentlicht: (2024)
von: Kozodoi, Nikita, et al.
Veröffentlicht: (2024)
Norm Anchors Make Model Edits Last
von: Liu, Mingda, et al.
Veröffentlicht: (2026)
von: Liu, Mingda, et al.
Veröffentlicht: (2026)
A Hardware-Aware, Per-Layer Methodology for Post-Training Quantization of Large Language Models
von: Killian, Earl
Veröffentlicht: (2026)
von: Killian, Earl
Veröffentlicht: (2026)
Improving Summarization with Human Edits
von: Yao, Zonghai, et al.
Veröffentlicht: (2023)
von: Yao, Zonghai, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation
von: Bagirov, Farid, et al.
Veröffentlicht: (2025) -
Diff-XYZ: A Benchmark for Evaluating Diff Understanding
von: Glukhov, Evgeniy, et al.
Veröffentlicht: (2025) -
Gaussian Process Tilted Nonparametric Density Estimation using Fisher Divergence Score Matching
von: Paisley, John, et al.
Veröffentlicht: (2025) -
How Robustly do LLMs Understand Execution Semantics?
von: Spiess, Claudio, et al.
Veröffentlicht: (2026) -
Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings
von: Tsvetkov, Petr, et al.
Veröffentlicht: (2024)