arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Weiqi, Ou, Jiefu, Song, Yangqiu, Van Durme, Benjamin, Khashabi, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment
by: Ou, Jiefu, et al.
Published: (2024)
by: Ou, Jiefu, et al.
Published: (2024)
Removed by arXiv
by: arXiv, Removed by
Published: (2023)
by: arXiv, Removed by
Published: (2023)
MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset
by: Wang, Weiqi, et al.
Published: (2024)
by: Wang, Weiqi, et al.
Published: (2024)
Are Finer Citations Always Better? Rethinking Granularity for Attributed Generation
by: Wang, Hexuan, et al.
Published: (2026)
by: Wang, Hexuan, et al.
Published: (2026)
Crystal: Characterizing Relative Impact of Scholarly Publications
by: Collison, Hannah, et al.
Published: (2026)
by: Collison, Hannah, et al.
Published: (2026)
Certified Mitigation of Worst-Case LLM Copyright Infringement
by: Zhang, Jingyu, et al.
Published: (2025)
by: Zhang, Jingyu, et al.
Published: (2025)
X-raying the arXiv: A Large-Scale Analysis of arXiv Submissions' Source Files
by: Apruzzese, Giovanni, et al.
Published: (2026)
by: Apruzzese, Giovanni, et al.
Published: (2026)
Data for "arXiv:2507.15554"
by: Lin, Ssu-Chih, et al.
Published: (2025)
by: Lin, Ssu-Chih, et al.
Published: (2025)
CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?
by: Ou, Jiefu, et al.
Published: (2025)
by: Ou, Jiefu, et al.
Published: (2025)
RORA: Robust Free-Text Rationale Evaluation
by: Jiang, Zhengping, et al.
Published: (2024)
by: Jiang, Zhengping, et al.
Published: (2024)
Many-Tier Instruction Hierarchy in LLM Agents
by: Zhang, Jingyu, et al.
Published: (2026)
by: Zhang, Jingyu, et al.
Published: (2026)
Scaling Accessible Mathematics on arXiv: HTML Conversion and MathML 4
by: Ginev, Deyan, et al.
Published: (2026)
by: Ginev, Deyan, et al.
Published: (2026)
Data of figures in arXiv:2602.07941
by: Cho, Sungtae, et al.
Published: (2026)
by: Cho, Sungtae, et al.
Published: (2026)
WebArXiv: Evaluating Multimodal Agents on Time-Invariant arXiv Tasks
by: Sun, Zihao, et al.
Published: (2025)
by: Sun, Zihao, et al.
Published: (2025)
Text-Tuple-Table: Towards Information Integration in Text-to-Table Generation via Global Tuple Extraction
by: Deng, Zheye, et al.
Published: (2024)
by: Deng, Zheye, et al.
Published: (2024)
Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data
by: Zhang, Jingyu, et al.
Published: (2024)
by: Zhang, Jingyu, et al.
Published: (2024)
Topics, Authors, and Institutions in Large Language Model Research: Trends from 17K arXiv Papers
by: Movva, Rajiv, et al.
Published: (2023)
by: Movva, Rajiv, et al.
Published: (2023)
Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
by: Zhang, Jingyu, et al.
Published: (2024)
by: Zhang, Jingyu, et al.
Published: (2024)
NLLG Quarterly arXiv Report 09/24: What are the most influential current AI Papers?
by: Leiter, Christoph, et al.
Published: (2024)
by: Leiter, Christoph, et al.
Published: (2024)
SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
by: Jiang, Dongwei, et al.
Published: (2024)
by: Jiang, Dongwei, et al.
Published: (2024)
The extension of zbMATH Open by arXiv preprints
by: Beckenbach, Isabel, et al.
Published: (2024)
by: Beckenbach, Isabel, et al.
Published: (2024)
Comments on Symplectic bipotentials arXiv:2410.23122
by: Buliga, Marius
Published: (2026)
by: Buliga, Marius
Published: (2026)
Citegeist: Automated Generation of Related Work Analysis on the arXiv Corpus
by: Beger, Claas, et al.
Published: (2025)
by: Beger, Claas, et al.
Published: (2025)
Challenging the Evaluator: LLM Sycophancy Under User Rebuttal
by: Kim, Sungwon, et al.
Published: (2025)
by: Kim, Sungwon, et al.
Published: (2025)
Science Hierarchography: Hierarchical Organization of Science Literature
by: Gao, Muhan, et al.
Published: (2025)
by: Gao, Muhan, et al.
Published: (2025)
Dated Data: Tracing Knowledge Cutoffs in Large Language Models
by: Cheng, Jeffrey, et al.
Published: (2024)
by: Cheng, Jeffrey, et al.
Published: (2024)
Structuring the Unstructured: A Systematic Review of Text-to-Structure Generation for Agentic AI with a Universal Evaluation Framework
by: Deng, Zheye, et al.
Published: (2025)
by: Deng, Zheye, et al.
Published: (2025)
RealHiTBench: A Comprehensive Realistic Hierarchical Table Benchmark for Evaluating LLM-Based Table Analysis
by: Wu, Pengzuo, et al.
Published: (2025)
by: Wu, Pengzuo, et al.
Published: (2025)
Comment on "Multidimensional arrow of time" (arXiv:2601.14134)
by: Galiautdinov, Andrei
Published: (2026)
by: Galiautdinov, Andrei
Published: (2026)
Influence Prediction in Collaboration Networks: An Empirical Study on arXiv
by: Lin, Marina, et al.
Published: (2025)
by: Lin, Marina, et al.
Published: (2025)
"According to ...": Prompting Language Models Improves Quoting from Pre-Training Data
by: Weller, Orion, et al.
Published: (2023)
by: Weller, Orion, et al.
Published: (2023)
Jailbreak Distillation: Renewable Safety Benchmarking
by: Zhang, Jingyu, et al.
Published: (2025)
by: Zhang, Jingyu, et al.
Published: (2025)
Comment on arXiv:2511.21731v1: Identifying Quantum Structure in AI Language: Evidence for Evolutionary Convergence of Human and Artificial Cognition
by: Sienicki, Krzysztof
Published: (2026)
by: Sienicki, Krzysztof
Published: (2026)
Comments on Weinstein's comments arXiv: 2509.09361 & 2510.03793
by: Ginoux, Jean-Marc
Published: (2025)
by: Ginoux, Jean-Marc
Published: (2025)
Code and data to produce figures of the paper arXiv:2508.09417
by: Zhang, Jiaju
Published: (2026)
by: Zhang, Jiaju
Published: (2026)
asifalsuny/arXiv-2502.20539: v1.0.1
by: asifalsuny
Published: (2025)
by: asifalsuny
Published: (2025)
Anderson transition in high dimension: comments to arXiv:2403.01974
by: Suslov, I. M.
Published: (2025)
by: Suslov, I. M.
Published: (2025)
Comment on: "The future of the correlated electron problem", arXiv:2010.00584
by: Shaginyan, V. R., et al.
Published: (2025)
by: Shaginyan, V. R., et al.
Published: (2025)
Reply to 'Comments on Graphon Signal Processing' [arXiv:2310.14683]
by: Ruiz, Luana, et al.
Published: (2024)
by: Ruiz, Luana, et al.
Published: (2024)
Notes on Crowther and the "Interpretation" of Quantum Mechanics (arXiv:2512.14315)
by: Sienicki, Mikołaj, et al.
Published: (2025)
by: Sienicki, Mikołaj, et al.
Published: (2025)
Similar Items
-
WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment
by: Ou, Jiefu, et al.
Published: (2024) -
Removed by arXiv
by: arXiv, Removed by
Published: (2023) -
MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset
by: Wang, Weiqi, et al.
Published: (2024) -
Are Finer Citations Always Better? Rethinking Granularity for Attributed Generation
by: Wang, Hexuan, et al.
Published: (2026) -
Crystal: Characterizing Relative Impact of Scholarly Publications
by: Collison, Hannah, et al.
Published: (2026)