Scaling In, Not Up? Testing Thick Citation Context Analysis with GPT-5 and Fragile Prompts
Fuente:
arXiv
Salvato in:
| Autore principale: | Simons, Arno |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Large Language Models for History, Philosophy, and Sociology of Science: Interpretive Uses, Methodological Challenges, and Critical Perspectives
di: Simons, Arno, et al.
Pubblicazione: (2025)
di: Simons, Arno, et al.
Pubblicazione: (2025)
LLMs Simulate Big Five Personality Traits: Further Evidence
di: Sorokovikova, Aleksandra, et al.
Pubblicazione: (2024)
di: Sorokovikova, Aleksandra, et al.
Pubblicazione: (2024)
BLT: Can Large Language Models Handle Basic Legal Text?
di: Blair-Stanek, Andrew, et al.
Pubblicazione: (2023)
di: Blair-Stanek, Andrew, et al.
Pubblicazione: (2023)
Scaling Laws for State Dynamics in Large Language Models
di: Li, Jacob X, et al.
Pubblicazione: (2025)
di: Li, Jacob X, et al.
Pubblicazione: (2025)
Can LLMs Identify Tax Abuse?
di: Blair-Stanek, Andrew, et al.
Pubblicazione: (2025)
di: Blair-Stanek, Andrew, et al.
Pubblicazione: (2025)
LegalCheck: Retrieval- and Context-Augmented Generation for Drafting Municipal Legal Advice Letters
di: van der Meer, Virgill, et al.
Pubblicazione: (2026)
di: van der Meer, Virgill, et al.
Pubblicazione: (2026)
OEMA: Ontology-Enhanced Multi-Agent Collaboration Framework for Zero-Shot Clinical Named Entity Recognition
di: Tao, Xinli, et al.
Pubblicazione: (2025)
di: Tao, Xinli, et al.
Pubblicazione: (2025)
Lossless Prompt Compression via Dictionary-Encoding and In-Context Learning: Enabling Cost-Effective LLM Analysis of Repetitive Data
di: de Campos, Andresa Rodrigues, et al.
Pubblicazione: (2026)
di: de Campos, Andresa Rodrigues, et al.
Pubblicazione: (2026)
Astro-HEP-BERT: A bidirectional language model for studying the meanings of concepts in astrophysics and high energy physics
di: Simons, Arno
Pubblicazione: (2024)
di: Simons, Arno
Pubblicazione: (2024)
Meaning at the Planck scale? Contextualized word embeddings for doing history, philosophy, and sociology of science
di: Simons, Arno
Pubblicazione: (2024)
di: Simons, Arno
Pubblicazione: (2024)
Comparative Analysis of AI Agent Architectures for Entity Relationship Classification
di: Berijanian, Maryam, et al.
Pubblicazione: (2025)
di: Berijanian, Maryam, et al.
Pubblicazione: (2025)
Beyond Imperfect Alternatives with Rulemapping: A Neuro-Symbolic Case Study on Online Hate Speech
di: von Cossel, Oskar
Pubblicazione: (2026)
di: von Cossel, Oskar
Pubblicazione: (2026)
AskSport: Web Application for Sports Question-Answering
di: Onofre, Enzo B, et al.
Pubblicazione: (2025)
di: Onofre, Enzo B, et al.
Pubblicazione: (2025)
How to Evaluate Medical AI
di: Kopanichuk, Ilia, et al.
Pubblicazione: (2025)
di: Kopanichuk, Ilia, et al.
Pubblicazione: (2025)
Comparing the Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature
di: Dayarathne, Ranul, et al.
Pubblicazione: (2025)
di: Dayarathne, Ranul, et al.
Pubblicazione: (2025)
Towards Safer Chatbots: Automated Policy Compliance Evaluation of Custom GPTs
di: Rodriguez, David, et al.
Pubblicazione: (2025)
di: Rodriguez, David, et al.
Pubblicazione: (2025)
ReTreVal: Reasoning Tree with Validation -- A Hybrid Framework for Enhanced LLM Multi-Step Reasoning
di: HS, Abhishek, et al.
Pubblicazione: (2026)
di: HS, Abhishek, et al.
Pubblicazione: (2026)
Challenges and Opportunities of NLP for HR Applications: A Discussion Paper
di: Leidner, Jochen L., et al.
Pubblicazione: (2024)
di: Leidner, Jochen L., et al.
Pubblicazione: (2024)
Evaluating Large Language Models on Historical Health Crisis Knowledge in Resource-Limited Settings: A Hybrid Multi-Metric Study
di: Hasan, Mohammed Rakibul
Pubblicazione: (2026)
di: Hasan, Mohammed Rakibul
Pubblicazione: (2026)
Discovering Differences in Strategic Behavior Between Humans and LLMs
di: Wang, Caroline, et al.
Pubblicazione: (2026)
di: Wang, Caroline, et al.
Pubblicazione: (2026)
Temporal Relation Extraction in Clinical Texts: A Span-based Graph Transformer Approach
di: Chaturvedi, Rochana, et al.
Pubblicazione: (2025)
di: Chaturvedi, Rochana, et al.
Pubblicazione: (2025)
Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes
di: Kenaston, Matthew W., et al.
Pubblicazione: (2025)
di: Kenaston, Matthew W., et al.
Pubblicazione: (2025)
From RAG to QA-RAG: Integrating Generative AI for Pharmaceutical Regulatory Compliance Process
di: Kim, Jaewoong, et al.
Pubblicazione: (2024)
di: Kim, Jaewoong, et al.
Pubblicazione: (2024)
LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization
di: Nguyen, Huyen, et al.
Pubblicazione: (2026)
di: Nguyen, Huyen, et al.
Pubblicazione: (2026)
A Graph-based RAG for Energy Efficiency Question Answering
di: Campi, Riccardo, et al.
Pubblicazione: (2025)
di: Campi, Riccardo, et al.
Pubblicazione: (2025)
EvoIdeator: Evolving Scientific Ideas through Checklist-Grounded Reinforcement Learning
di: Sauter, Andreas, et al.
Pubblicazione: (2026)
di: Sauter, Andreas, et al.
Pubblicazione: (2026)
The Good, the Bad, and the Hulk-like GPT: Analyzing Emotional Decisions of Large Language Models in Cooperation and Bargaining Games
di: Mozikov, Mikhail, et al.
Pubblicazione: (2024)
di: Mozikov, Mikhail, et al.
Pubblicazione: (2024)
Towards Democratized Flood Risk Management: An Advanced AI Assistant Enabled by GPT-4 for Enhanced Interpretability and Public Engagement
di: Martelo, Rafaela, et al.
Pubblicazione: (2024)
di: Martelo, Rafaela, et al.
Pubblicazione: (2024)
Unified Game Moderation: Soft-Prompting and LLM-Assisted Label Transfer for Resource-Efficient Toxicity Detection
di: Yang, Zachary, et al.
Pubblicazione: (2025)
di: Yang, Zachary, et al.
Pubblicazione: (2025)
An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs
di: Zhu, Qian, et al.
Pubblicazione: (2026)
di: Zhu, Qian, et al.
Pubblicazione: (2026)
GraphWalk: Enabling Reasoning in Large Language Models through Tool-Based Graph Navigation
di: Ghandi, Taraneh, et al.
Pubblicazione: (2026)
di: Ghandi, Taraneh, et al.
Pubblicazione: (2026)
AI-Powered Detection of Inappropriate Language in Medical School Curricula
di: Salavati, Chiman, et al.
Pubblicazione: (2025)
di: Salavati, Chiman, et al.
Pubblicazione: (2025)
Attention-based sequential recommendation system using multimodal data
di: Oh, Hyungtaik, et al.
Pubblicazione: (2024)
di: Oh, Hyungtaik, et al.
Pubblicazione: (2024)
Using a cognitive architecture to consider antiBlackness in design and development of AI systems
di: Dancy, Christopher L.
Pubblicazione: (2022)
di: Dancy, Christopher L.
Pubblicazione: (2022)
BabyReasoningBench: Generating Developmentally-Inspired Reasoning Tasks for Evaluating Baby Language Models
di: Dhole, Kaustubh D.
Pubblicazione: (2026)
di: Dhole, Kaustubh D.
Pubblicazione: (2026)
Categorical Perception in Large Language Model Hidden States: Structural Warping at Digit-Count Boundaries
di: Cacioli, Jon-Paul
Pubblicazione: (2026)
di: Cacioli, Jon-Paul
Pubblicazione: (2026)
Mitigating Structural Noise in Low-Resource S2TT: An Optimized Cascaded Nepali-English Pipeline with Punctuation Restoration
di: Chongbang, Tangsang, et al.
Pubblicazione: (2026)
di: Chongbang, Tangsang, et al.
Pubblicazione: (2026)
Exploring the Structure of AI-Induced Language Change in Scientific English
di: Galpin, Riley, et al.
Pubblicazione: (2025)
di: Galpin, Riley, et al.
Pubblicazione: (2025)
A Llama walks into the 'Bar': Efficient Supervised Fine-Tuning for Legal Reasoning in the Multi-state Bar Exam
di: Fernandes, Rean, et al.
Pubblicazione: (2025)
di: Fernandes, Rean, et al.
Pubblicazione: (2025)
VERITAS-NLI : Validation and Extraction of Reliable Information Through Automated Scraping and Natural Language Inference
di: Shah, Arjun, et al.
Pubblicazione: (2024)
di: Shah, Arjun, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Large Language Models for History, Philosophy, and Sociology of Science: Interpretive Uses, Methodological Challenges, and Critical Perspectives
di: Simons, Arno, et al.
Pubblicazione: (2025) -
LLMs Simulate Big Five Personality Traits: Further Evidence
di: Sorokovikova, Aleksandra, et al.
Pubblicazione: (2024) -
BLT: Can Large Language Models Handle Basic Legal Text?
di: Blair-Stanek, Andrew, et al.
Pubblicazione: (2023) -
Scaling Laws for State Dynamics in Large Language Models
di: Li, Jacob X, et al.
Pubblicazione: (2025) -
Can LLMs Identify Tax Abuse?
di: Blair-Stanek, Andrew, et al.
Pubblicazione: (2025)