Generalized Tree Edit Distance (GTED): A Faithful Evaluation Metric for Statement Autoformalization
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yuntian, Zhu, Tao, Liu, Xiaoyang, Chen, Yu, Liu, Zhaoxuan, Guo, Qingfeng, Zhang, Jiashuo, Bao, Kangjie, Luo, Tao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ASSESS: A Semantic and Structural Evaluation Framework for Statement Similarity
by: Liu, Xiaoyang, et al.
Published: (2025)
by: Liu, Xiaoyang, et al.
Published: (2025)
ATLAS: Autoformalizing Theorems through Lifting, Augmentation, and Synthesis of Data
by: Liu, Xiaoyang, et al.
Published: (2025)
by: Liu, Xiaoyang, et al.
Published: (2025)
Decompose, Structure, and Repair: A Neuro-Symbolic Framework for Autoformalization via Operator Trees
by: Liu, Xiaoyang, et al.
Published: (2026)
by: Liu, Xiaoyang, et al.
Published: (2026)
Reliable Evaluation and Benchmarks for Statement Autoformalization
by: Poiroux, Auguste, et al.
Published: (2024)
by: Poiroux, Auguste, et al.
Published: (2024)
Faithful Autoformalization via Roundtrip Verification and Repair
by: Amrollahi, Daneshvar, et al.
Published: (2026)
by: Amrollahi, Daneshvar, et al.
Published: (2026)
ProofFlow: A Dependency Graph Approach to Faithful Proof Autoformalization
by: Cabral, Rafael, et al.
Published: (2025)
by: Cabral, Rafael, et al.
Published: (2025)
FormalAlign: Automated Alignment Evaluation for Autoformalization
by: Lu, Jianqiao, et al.
Published: (2024)
by: Lu, Jianqiao, et al.
Published: (2024)
Towards Faithful and Controllable Personalization via Critique-Post-Edit Reinforcement Learning
by: Zhu, Chenghao, et al.
Published: (2025)
by: Zhu, Chenghao, et al.
Published: (2025)
BadEdit: Backdooring large language models by model editing
by: Li, Yanzhou, et al.
Published: (2024)
by: Li, Yanzhou, et al.
Published: (2024)
FormaRL: Enhancing Autoformalization with no Labeled Data
by: Huang, Yanxing, et al.
Published: (2025)
by: Huang, Yanxing, et al.
Published: (2025)
Autoformalizer with Tool Feedback
by: Guo, Qi, et al.
Published: (2025)
by: Guo, Qi, et al.
Published: (2025)
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
by: Liu, Chenxu, et al.
Published: (2026)
by: Liu, Chenxu, et al.
Published: (2026)
An Evaluation Benchmark for Autoformalization in Lean4
by: Gulati, Aryan, et al.
Published: (2024)
by: Gulati, Aryan, et al.
Published: (2024)
Autoformalization in the Era of Large Language Models: A Survey
by: Weng, Ke, et al.
Published: (2025)
by: Weng, Ke, et al.
Published: (2025)
A Causal Lens for Evaluating Faithfulness Metrics
by: Zaman, Kerem, et al.
Published: (2025)
by: Zaman, Kerem, et al.
Published: (2025)
GEDAN: Learning the Edit Costs for Graph Edit Distance
by: Leonardi, Francesco, et al.
Published: (2025)
by: Leonardi, Francesco, et al.
Published: (2025)
Generative Agents for Multi-Agent Autoformalization of Interaction Scenarios
by: Mensfelt, Agnieszka, et al.
Published: (2024)
by: Mensfelt, Agnieszka, et al.
Published: (2024)
Rethinking the Evaluation Protocol of Domain Generalization
by: Yu, Han, et al.
Published: (2023)
by: Yu, Han, et al.
Published: (2023)
Improving Autoformalization Using Direct Dependency Retrieval
by: Wang, Shaoqi, et al.
Published: (2025)
by: Wang, Shaoqi, et al.
Published: (2025)
Data Heterogeneity Modeling for Trustworthy Machine Learning
by: Liu, Jiashuo, et al.
Published: (2025)
by: Liu, Jiashuo, et al.
Published: (2025)
Munkres' General Topology Autoformalized in Isabelle/HOL
by: Bryant, Dustin, et al.
Published: (2026)
by: Bryant, Dustin, et al.
Published: (2026)
Autoformalizing Euclidean Geometry
by: Murphy, Logan, et al.
Published: (2024)
by: Murphy, Logan, et al.
Published: (2024)
Towards a Common Framework for Autoformalization
by: Mensfelt, Agnieszka, et al.
Published: (2025)
by: Mensfelt, Agnieszka, et al.
Published: (2025)
Are We Evaluating the Edit Locality of LLM Model Editing Properly?
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
Autoformalize Mathematical Statements by Symbolic Equivalence and Semantic Consistency
by: Li, Zenan, et al.
Published: (2024)
by: Li, Zenan, et al.
Published: (2024)
Graph Edit Distance with General Costs Using Neural Set Divergence
by: Jain, Eeshaan, et al.
Published: (2024)
by: Jain, Eeshaan, et al.
Published: (2024)
Counterfactual Edits for Generative Evaluation
by: Lymperaiou, Maria, et al.
Published: (2023)
by: Lymperaiou, Maria, et al.
Published: (2023)
Impact-driven Context Filtering For Cross-file Code Completion
by: Li, Yanzhou, et al.
Published: (2025)
by: Li, Yanzhou, et al.
Published: (2025)
F-Fidelity: A Robust Framework for Faithfulness Evaluation of Explainable AI
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization
by: Agarwal, Anmol, et al.
Published: (2026)
by: Agarwal, Anmol, et al.
Published: (2026)
A New Approach Towards Autoformalization
by: Patel, Nilay, et al.
Published: (2023)
by: Patel, Nilay, et al.
Published: (2023)
MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows
by: Li, Xiquan, et al.
Published: (2025)
by: Li, Xiquan, et al.
Published: (2025)
Transformer Circuit Faithfulness Metrics are not Robust
by: Miller, Joseph, et al.
Published: (2024)
by: Miller, Joseph, et al.
Published: (2024)
Norm Anchors Make Model Edits Last
by: Liu, Mingda, et al.
Published: (2026)
by: Liu, Mingda, et al.
Published: (2026)
Attention Distance: A Novel Metric for Directed Fuzzing with Large Language Models
by: Bin, Wang, et al.
Published: (2025)
by: Bin, Wang, et al.
Published: (2025)
Frechet Music Distance: A Metric For Generative Symbolic Music Evaluation
by: Retkowski, Jan, et al.
Published: (2024)
by: Retkowski, Jan, et al.
Published: (2024)
ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation
by: Huang, Pengcheng, et al.
Published: (2025)
by: Huang, Pengcheng, et al.
Published: (2025)
Fréchet Power-Scenario Distance: A Metric for Evaluating Generative AI Models across Multiple Time-Scales in Smart Grids
by: Cai, Yuting, et al.
Published: (2025)
by: Cai, Yuting, et al.
Published: (2025)
Towards Metric-Faithful Neural Graph Matching
by: Shivottam, Jyotirmaya, et al.
Published: (2026)
by: Shivottam, Jyotirmaya, et al.
Published: (2026)
RPTS: Tree-Structured Reasoning Process Scoring for Faithful Multimodal Evaluation
by: Wang, Haofeng, et al.
Published: (2025)
by: Wang, Haofeng, et al.
Published: (2025)
Similar Items
-
ASSESS: A Semantic and Structural Evaluation Framework for Statement Similarity
by: Liu, Xiaoyang, et al.
Published: (2025) -
ATLAS: Autoformalizing Theorems through Lifting, Augmentation, and Synthesis of Data
by: Liu, Xiaoyang, et al.
Published: (2025) -
Decompose, Structure, and Repair: A Neuro-Symbolic Framework for Autoformalization via Operator Trees
by: Liu, Xiaoyang, et al.
Published: (2026) -
Reliable Evaluation and Benchmarks for Statement Autoformalization
by: Poiroux, Auguste, et al.
Published: (2024) -
Faithful Autoformalization via Roundtrip Verification and Repair
by: Amrollahi, Daneshvar, et al.
Published: (2026)