Saved in:
| Main Authors: | Strauss, Ilan, Yang, Jangho, O'Reilly, Tim, Rosenblat, Sruly, Moure, Isobel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2508.00838 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Public Access in LLM Pre-Training Data
by: Rosenblat, Sruly, et al.
Published: (2025)
by: Rosenblat, Sruly, et al.
Published: (2025)
Real-World Gaps in AI Governance Research
by: Strauss, Ilan, et al.
Published: (2025)
by: Strauss, Ilan, et al.
Published: (2025)
AI Blob! LLM-Driven Recontextualization of Italian Television Archives
by: Balestri, Roberto
Published: (2025)
by: Balestri, Roberto
Published: (2025)
Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs
by: Seo, Yongsik, et al.
Published: (2026)
by: Seo, Yongsik, et al.
Published: (2026)
Enhancing Research Idea Generation through Combinatorial Innovation and Multi-Agent Iterative Search Strategies
by: Chen, Shuai, et al.
Published: (2026)
by: Chen, Shuai, et al.
Published: (2026)
Large language models for automated scholarly paper review: A survey
by: Zhuang, Zhenzhen, et al.
Published: (2025)
by: Zhuang, Zhenzhen, et al.
Published: (2025)
Named Entity Recognition of Historical Texts via Large Language Model
by: Zhang, Shibingfeng, et al.
Published: (2025)
by: Zhang, Shibingfeng, et al.
Published: (2025)
Triples and Knowledge-Infused Embeddings for Clustering and Classification of Scientific Documents
by: Arcan, Mihael
Published: (2025)
by: Arcan, Mihael
Published: (2025)
LLMs4SchemaDiscovery: A Human-in-the-Loop Workflow for Scientific Schema Mining with Large Language Models
by: Sadruddin, Sameer, et al.
Published: (2025)
by: Sadruddin, Sameer, et al.
Published: (2025)
Mining for Species, Locations, Habitats, and Ecosystems from Scientific Papers in Invasion Biology: A Large-Scale Exploratory Study with Large Language Models
by: D'Souza, Jennifer, et al.
Published: (2025)
by: D'Souza, Jennifer, et al.
Published: (2025)
Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents
by: Greif, Gavin, et al.
Published: (2025)
by: Greif, Gavin, et al.
Published: (2025)
SciClaims: An End-to-End Generative System for Biomedical Claim Analysis
by: Ortega, Raúl, et al.
Published: (2025)
by: Ortega, Raúl, et al.
Published: (2025)
SC4ANM: Identifying Optimal Section Combinations for Automated Novelty Prediction in Academic Papers
by: Wu, Wenqing, et al.
Published: (2025)
by: Wu, Wenqing, et al.
Published: (2025)
Fairness Evaluation of Large Language Models in Academic Library Reference Services
by: Wang, Haining, et al.
Published: (2025)
by: Wang, Haining, et al.
Published: (2025)
Polarity Detection of Sustainable Development Goals in News Text
by: Cadeddu, Andrea, et al.
Published: (2025)
by: Cadeddu, Andrea, et al.
Published: (2025)
Annotating Scientific Uncertainty: A comprehensive model using linguistic patterns and comparison with existing approaches
by: Ningrum, Panggih Kusuma, et al.
Published: (2025)
by: Ningrum, Panggih Kusuma, et al.
Published: (2025)
Accelerating Scientific Discovery with Multi-Document Summarization of Impact-Ranked Papers
by: Koloveas, Paris, et al.
Published: (2025)
by: Koloveas, Paris, et al.
Published: (2025)
Big Tech-Funded AI Papers Have Higher Citation Impact, Greater Insularity, and Larger Recency Bias
by: Gnewuch, Max Martin, et al.
Published: (2025)
by: Gnewuch, Max Martin, et al.
Published: (2025)
Toward Purpose-oriented Topic Model Evaluation enabled by Large Language Models
by: Tan, Zhiyin, et al.
Published: (2025)
by: Tan, Zhiyin, et al.
Published: (2025)
Bridging the Evaluation Gap: Leveraging Large Language Models for Topic Model Evaluation
by: Tan, Zhiyin, et al.
Published: (2025)
by: Tan, Zhiyin, et al.
Published: (2025)
Context Selection for Hypothesis and Statistical Evidence Extraction from Full-Text Scientific Articles
by: Koneru, Sai, et al.
Published: (2026)
by: Koneru, Sai, et al.
Published: (2026)
Towards the AI Historian: Agentic Information Extraction from Primary Sources
by: Hufe, Lorenz, et al.
Published: (2026)
by: Hufe, Lorenz, et al.
Published: (2026)
SPACE-IDEAS: A Dataset for Salient Information Detection in Space Innovation
by: García-Silva, Andrés, et al.
Published: (2024)
by: García-Silva, Andrés, et al.
Published: (2024)
Scientific Statement Classification over arXiv.org
by: Ginev, Deyan, et al.
Published: (2019)
by: Ginev, Deyan, et al.
Published: (2019)
LLMs4Synthesis: Leveraging Large Language Models for Scientific Synthesis
by: Giglou, Hamed Babaei, et al.
Published: (2024)
by: Giglou, Hamed Babaei, et al.
Published: (2024)
NewsEdits 2.0: Learning the Intentions Behind Updating News
by: Spangher, Alexander, et al.
Published: (2024)
by: Spangher, Alexander, et al.
Published: (2024)
How Do LLMs Encode Scientific Quality? An Empirical Study Using Monosemantic Features from Sparse Autoencoders
by: McCoubrey, Michael, et al.
Published: (2026)
by: McCoubrey, Michael, et al.
Published: (2026)
HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
by: Sakai, Yusuke, et al.
Published: (2026)
by: Sakai, Yusuke, et al.
Published: (2026)
Textual Entailment for Effective Triple Validation in Object Prediction
by: García-Silva, Andrés, et al.
Published: (2024)
by: García-Silva, Andrés, et al.
Published: (2024)
NAIST Academic Travelogue Dataset
by: Ouchi, Hiroki, et al.
Published: (2023)
by: Ouchi, Hiroki, et al.
Published: (2023)
Exploring Scholarly Data by Semantic Query on Knowledge Graph Embedding Space
by: Tran, Hung Nghiep, et al.
Published: (2019)
by: Tran, Hung Nghiep, et al.
Published: (2019)
On the performativity of SDG classifications in large bibliometric databases
by: Ottaviani, Matteo, et al.
Published: (2024)
by: Ottaviani, Matteo, et al.
Published: (2024)
Giving Voice to the Constitution: Low-Resource Text-to-Speech for Quechua and Spanish Using a Bilingual Legal Corpus
by: Ortega, John E., et al.
Published: (2026)
by: Ortega, John E., et al.
Published: (2026)
HalluCiteChecker: A Lightweight Toolkit for Hallucinated Citation Detection and Verification in the Era of AI Scientists
by: Sakai, Yusuke, et al.
Published: (2026)
by: Sakai, Yusuke, et al.
Published: (2026)
Is there really a Citation Age Bias in NLP?
by: Nguyen, Hoa, et al.
Published: (2024)
by: Nguyen, Hoa, et al.
Published: (2024)
A Survey Forest Diagram : Gain a Divergent Insight View on a Specific Research Topic
by: Li, Jinghong, et al.
Published: (2024)
by: Li, Jinghong, et al.
Published: (2024)
Decoding AI and Human Authorship: Nuances Revealed Through NLP and Statistical Analysis
by: Akinwande, Mayowa, et al.
Published: (2024)
by: Akinwande, Mayowa, et al.
Published: (2024)
Hybrid X-Linker: Automated Data Generation and Extreme Multi-label Ranking for Biomedical Entity Linking
by: Ruas, Pedro, et al.
Published: (2024)
by: Ruas, Pedro, et al.
Published: (2024)
LitSearch: A Retrieval Benchmark for Scientific Literature Search
by: Ajith, Anirudh, et al.
Published: (2024)
by: Ajith, Anirudh, et al.
Published: (2024)
SemEval-2025 Task 5: LLMs4Subjects -- LLM-based Automated Subject Tagging for a National Technical Library's Open-Access Catalog
by: D'Souza, Jennifer, et al.
Published: (2025)
by: D'Souza, Jennifer, et al.
Published: (2025)
Similar Items
-
Beyond Public Access in LLM Pre-Training Data
by: Rosenblat, Sruly, et al.
Published: (2025) -
Real-World Gaps in AI Governance Research
by: Strauss, Ilan, et al.
Published: (2025) -
AI Blob! LLM-Driven Recontextualization of Italian Television Archives
by: Balestri, Roberto
Published: (2025) -
Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs
by: Seo, Yongsik, et al.
Published: (2026) -
Enhancing Research Idea Generation through Combinatorial Innovation and Multi-Agent Iterative Search Strategies
by: Chen, Shuai, et al.
Published: (2026)