HELEA: Hard-Negative Benchmark and LLM-based Reranking for Robust Entity Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jang, Yoonjin, Kim, Junwoo, Ko, Youngjoong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation
von: Koc, Vincent
Veröffentlicht: (2025)
von: Koc, Vincent
Veröffentlicht: (2025)
Retrieval and Augmentation of Domain Knowledge for Text-to-SQL Semantic Parsing
von: Patwardhan, Manasi, et al.
Veröffentlicht: (2025)
von: Patwardhan, Manasi, et al.
Veröffentlicht: (2025)
Both Ends Count! Just How Good are LLM Agents at "Text-to-Big SQL"?
von: Eizaguirre, Germán T., et al.
Veröffentlicht: (2026)
von: Eizaguirre, Germán T., et al.
Veröffentlicht: (2026)
ReaGeo: Reasoning-Enhanced End-to-End Geocoding with LLMs
von: Cui, Jian, et al.
Veröffentlicht: (2026)
von: Cui, Jian, et al.
Veröffentlicht: (2026)
GraphEval36K: Benchmarking Coding and Reasoning Capabilities of Large Language Models on Graph Datasets
von: Wu, Qiming, et al.
Veröffentlicht: (2024)
von: Wu, Qiming, et al.
Veröffentlicht: (2024)
Reviewing the Reviewer: Graph-Enhanced LLMs for E-commerce Appeal Adjudication
von: Du, Yuchen, et al.
Veröffentlicht: (2026)
von: Du, Yuchen, et al.
Veröffentlicht: (2026)
AutoBench: Automating LLM Evaluation through Reciprocal Peer Assessment
von: Loi, Dario, et al.
Veröffentlicht: (2025)
von: Loi, Dario, et al.
Veröffentlicht: (2025)
HOME-KGQA: A Benchmark Dataset for Multimodal Knowledge Graph Question Answering on Household Daily Activities
von: Egami, Shusaku, et al.
Veröffentlicht: (2026)
von: Egami, Shusaku, et al.
Veröffentlicht: (2026)
Free Access to World News: Reconstructing Full-Text Articles from GDELT
von: Colladon, A. Fronzetti, et al.
Veröffentlicht: (2025)
von: Colladon, A. Fronzetti, et al.
Veröffentlicht: (2025)
How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game Environment
von: Ansell, Rebecca, et al.
Veröffentlicht: (2026)
von: Ansell, Rebecca, et al.
Veröffentlicht: (2026)
Critical Insights into Leading Conversational AI Models
von: Kohli, Urja, et al.
Veröffentlicht: (2025)
von: Kohli, Urja, et al.
Veröffentlicht: (2025)
ChatGPT4PCG Competition: Character-like Level Generation for Science Birds
von: Taveekitworachai, Pittawat, et al.
Veröffentlicht: (2023)
von: Taveekitworachai, Pittawat, et al.
Veröffentlicht: (2023)
Comparison of Unsupervised Metrics for Evaluating Judicial Decision Extraction
von: Litvak, Ivan Leonidovich, et al.
Veröffentlicht: (2025)
von: Litvak, Ivan Leonidovich, et al.
Veröffentlicht: (2025)
EMR-AGENT: Automating Cohort and Feature Extraction from EMR Databases
von: Lee, Kwanhyung, et al.
Veröffentlicht: (2025)
von: Lee, Kwanhyung, et al.
Veröffentlicht: (2025)
BERTopic for Topic Modeling of Hindi Short Texts: A Comparative Study
von: Mutsaddi, Atharva, et al.
Veröffentlicht: (2025)
von: Mutsaddi, Atharva, et al.
Veröffentlicht: (2025)
DySK-Attn: A Framework for Efficient, Real-Time Knowledge Updating in Large Language Models via Dynamic Sparse Knowledge Attention
von: Khan, Kabir, et al.
Veröffentlicht: (2025)
von: Khan, Kabir, et al.
Veröffentlicht: (2025)
Falkor-IRAC: Graph-Constrained Generation for Verified Legal Reasoning in Indian Judicial AI
von: Bose, Joy
Veröffentlicht: (2026)
von: Bose, Joy
Veröffentlicht: (2026)
Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
von: Nakamura, Mason, et al.
Veröffentlicht: (2025)
von: Nakamura, Mason, et al.
Veröffentlicht: (2025)
ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
von: Chen, Hao, et al.
Veröffentlicht: (2025)
von: Chen, Hao, et al.
Veröffentlicht: (2025)
Survey Transfer Learning: Recycling Data with Silicon Responses
von: Amini, Ali
Veröffentlicht: (2025)
von: Amini, Ali
Veröffentlicht: (2025)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
Structured Prompting and Feedback-Guided Reasoning with LLMs for Data Interpretation
von: Rath, Amit
Veröffentlicht: (2025)
von: Rath, Amit
Veröffentlicht: (2025)
Primary Care Diagnoses as a Reliable Predictor for Orthopedic Surgical Interventions
von: Verma, Khushboo, et al.
Veröffentlicht: (2025)
von: Verma, Khushboo, et al.
Veröffentlicht: (2025)
Game of Thought: Robust Information Seeking with Large Language Models Using Game Theory
von: Cui, Langyuan, et al.
Veröffentlicht: (2026)
von: Cui, Langyuan, et al.
Veröffentlicht: (2026)
REPOT: Recoverable Program-of-Thought via Checkpoint Repair
von: Mazaheri, Parsa
Veröffentlicht: (2026)
von: Mazaheri, Parsa
Veröffentlicht: (2026)
Assisting humans in complex comparisons: automated information comparison at scale
von: Yuen, Truman, et al.
Veröffentlicht: (2024)
von: Yuen, Truman, et al.
Veröffentlicht: (2024)
Reinforced Language Models for Sequential Decision Making
von: Dilkes, Jim, et al.
Veröffentlicht: (2025)
von: Dilkes, Jim, et al.
Veröffentlicht: (2025)
REVOLVE: Optimizing AI Systems by Tracking Response Evolution in Textual Optimization
von: Zhang, Peiyan, et al.
Veröffentlicht: (2024)
von: Zhang, Peiyan, et al.
Veröffentlicht: (2024)
GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
von: Zhang, Peiyan, et al.
Veröffentlicht: (2025)
von: Zhang, Peiyan, et al.
Veröffentlicht: (2025)
Knowledge Graphs as the Missing Data Layer for LLM-Based Industrial Asset Operations
von: Mandarapu, Madhulatha, et al.
Veröffentlicht: (2026)
von: Mandarapu, Madhulatha, et al.
Veröffentlicht: (2026)
Evo-DKD: Dual-Knowledge Decoding for Autonomous Ontology Evolution in Large Language Models
von: Raman, Vishal, et al.
Veröffentlicht: (2025)
von: Raman, Vishal, et al.
Veröffentlicht: (2025)
Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts
von: Agarwal, Amit, et al.
Veröffentlicht: (2024)
von: Agarwal, Amit, et al.
Veröffentlicht: (2024)
Spark-LLM-Eval: A Distributed Framework for Statistically Rigorous Large Language Model Evaluation
von: Mitra, Subhadip
Veröffentlicht: (2026)
von: Mitra, Subhadip
Veröffentlicht: (2026)
HiFi-RAG: Hierarchical Content Filtering and Two-Pass Generation for Open-Domain RAG
von: Nuengsigkapian, Cattalyya
Veröffentlicht: (2025)
von: Nuengsigkapian, Cattalyya
Veröffentlicht: (2025)
What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics
von: Wachowiak, Lennart, et al.
Veröffentlicht: (2025)
von: Wachowiak, Lennart, et al.
Veröffentlicht: (2025)
Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents
von: Chua, Jaymari, et al.
Veröffentlicht: (2025)
von: Chua, Jaymari, et al.
Veröffentlicht: (2025)
COVID-19 on YouTube: A Data-Driven Analysis of Sentiment, Toxicity, and Content Recommendations
von: Su, Vanessa, et al.
Veröffentlicht: (2024)
von: Su, Vanessa, et al.
Veröffentlicht: (2024)
Content and Engagement Trends in COVID-19 YouTube Videos: Evidence from the Late Pandemic
von: Thakur, Nirmalya, et al.
Veröffentlicht: (2025)
von: Thakur, Nirmalya, et al.
Veröffentlicht: (2025)
Mpox Narrative on Instagram: A Labeled Multilingual Dataset of Instagram Posts on Mpox for Sentiment, Hate Speech, and Anxiety Analysis
von: Thakur, Nirmalya
Veröffentlicht: (2024)
von: Thakur, Nirmalya
Veröffentlicht: (2024)
Five Years of COVID-19 Discourse on Instagram: A Labeled Instagram Dataset of Over Half a Million Posts for Multilingual Sentiment Analysis
von: Thakur, Nirmalya
Veröffentlicht: (2024)
von: Thakur, Nirmalya
Veröffentlicht: (2024)
Ähnliche Einträge
-
Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation
von: Koc, Vincent
Veröffentlicht: (2025) -
Retrieval and Augmentation of Domain Knowledge for Text-to-SQL Semantic Parsing
von: Patwardhan, Manasi, et al.
Veröffentlicht: (2025) -
Both Ends Count! Just How Good are LLM Agents at "Text-to-Big SQL"?
von: Eizaguirre, Germán T., et al.
Veröffentlicht: (2026) -
ReaGeo: Reasoning-Enhanced End-to-End Geocoding with LLMs
von: Cui, Jian, et al.
Veröffentlicht: (2026) -
GraphEval36K: Benchmarking Coding and Reasoning Capabilities of Large Language Models on Graph Datasets
von: Wu, Qiming, et al.
Veröffentlicht: (2024)