Salvato in:
| Autori principali: | Jacas, Joshua, Winchester, Hana, Boyd, Alicia, Johnson, Brittany |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2503.09341 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HalluShift: Measuring Distribution Shifts towards Hallucination Detection in LLMs
di: Dasgupta, Sharanya, et al.
Pubblicazione: (2025)
di: Dasgupta, Sharanya, et al.
Pubblicazione: (2025)
Towards Evaluating Large Language Models for Graph Query Generation
di: Munir, Siraj, et al.
Pubblicazione: (2024)
di: Munir, Siraj, et al.
Pubblicazione: (2024)
Multi-Faceted Evaluation of Modeling Languages for Augmented Reality Applications -- The Case of ARWFML
di: Muff, Fabian, et al.
Pubblicazione: (2024)
di: Muff, Fabian, et al.
Pubblicazione: (2024)
Comparative Evaluation of Prompting and Fine-Tuning for Applying Large Language Models to Grid-Structured Geospatial Data
di: Dhruv, Akash, et al.
Pubblicazione: (2025)
di: Dhruv, Akash, et al.
Pubblicazione: (2025)
Benchmarking LLMs for Political Science: A United Nations Perspective
di: Liang, Yueqing, et al.
Pubblicazione: (2025)
di: Liang, Yueqing, et al.
Pubblicazione: (2025)
Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs
di: Huang, Yiming, et al.
Pubblicazione: (2026)
di: Huang, Yiming, et al.
Pubblicazione: (2026)
Zero-Shot Confidence Estimation for Small LLMs: When Supervised Baselines Aren't Worth Training
di: Nguyen, Luong N.
Pubblicazione: (2026)
di: Nguyen, Luong N.
Pubblicazione: (2026)
MolMetaLM: a Physicochemical Knowledge-Guided Molecular Meta Language Model
di: Wu, Yifan, et al.
Pubblicazione: (2024)
di: Wu, Yifan, et al.
Pubblicazione: (2024)
MiTa: A Hierarchical Multi-Agent Collaboration Framework with Memory-integrated and Task Allocation
di: Zhang, XiaoJie, et al.
Pubblicazione: (2026)
di: Zhang, XiaoJie, et al.
Pubblicazione: (2026)
Seq2Seq Model-Based Chatbot with LSTM and Attention Mechanism for Enhanced User Interaction
di: Benaddi, Lamya, et al.
Pubblicazione: (2024)
di: Benaddi, Lamya, et al.
Pubblicazione: (2024)
Transformer Models in Education: Summarizing Science Textbooks with AraBART, MT5, AraT5, and mBART
di: Masri, Sari, et al.
Pubblicazione: (2024)
di: Masri, Sari, et al.
Pubblicazione: (2024)
Introducing a new hyper-parameter for RAG: Context Window Utilization
di: Juvekar, Kush, et al.
Pubblicazione: (2024)
di: Juvekar, Kush, et al.
Pubblicazione: (2024)
AgriLLM: Harnessing Transformers for Farmer Queries
di: Didwania, Krish, et al.
Pubblicazione: (2024)
di: Didwania, Krish, et al.
Pubblicazione: (2024)
CultureVo: The Serious Game of Utilizing Gen AI for Enhancing Cultural Intelligence
di: Agarwala, Ajita, et al.
Pubblicazione: (2024)
di: Agarwala, Ajita, et al.
Pubblicazione: (2024)
LaScA: Language-Conditioned Scalable Modelling of Affective Dynamics
di: Pinitas, Kosmas, et al.
Pubblicazione: (2026)
di: Pinitas, Kosmas, et al.
Pubblicazione: (2026)
$C$-$ΔΘ$: Circuit-Restricted Weight Arithmetic for Selective Refusal
di: Kasliwal, Aditya, et al.
Pubblicazione: (2026)
di: Kasliwal, Aditya, et al.
Pubblicazione: (2026)
Systematic Review on Healthcare Systems Engineering utilizing ChatGPT
di: Kim, Jungwoo, et al.
Pubblicazione: (2024)
di: Kim, Jungwoo, et al.
Pubblicazione: (2024)
A Novel Nuanced Conversation Evaluation Framework for Large Language Models in Mental Health
di: Marrapese, Alexander, et al.
Pubblicazione: (2024)
di: Marrapese, Alexander, et al.
Pubblicazione: (2024)
The Future of MLLM Prompting is Adaptive: A Comprehensive Experimental Evaluation of Prompt Engineering Methods for Robust Multimodal Performance
di: Mohanty, Anwesha, et al.
Pubblicazione: (2025)
di: Mohanty, Anwesha, et al.
Pubblicazione: (2025)
Retrieve, Annotate, Evaluate, Repeat: Leveraging Multimodal LLMs for Large-Scale Product Retrieval Evaluation
di: Hosseini, Kasra, et al.
Pubblicazione: (2024)
di: Hosseini, Kasra, et al.
Pubblicazione: (2024)
Use of AI Tools: Guidelines to Maintain Academic Integrity in Computing Colleges
di: El-boghdadi, Hatem M., et al.
Pubblicazione: (2026)
di: El-boghdadi, Hatem M., et al.
Pubblicazione: (2026)
Graph Repairs with Large Language Models: An Empirical Study
di: Terdalkar, Hrishikesh, et al.
Pubblicazione: (2025)
di: Terdalkar, Hrishikesh, et al.
Pubblicazione: (2025)
On the Limitations of Compute Thresholds as a Governance Strategy
di: Hooker, Sara
Pubblicazione: (2024)
di: Hooker, Sara
Pubblicazione: (2024)
CoMoNM: A Cost Modeling Framework for Compute-Near-Memory Systems
di: Farzaneh, Hamid, et al.
Pubblicazione: (2025)
di: Farzaneh, Hamid, et al.
Pubblicazione: (2025)
LLMs and Agentic AI in Insurance Decision-Making: Opportunities and Challenges For Africa
di: Hill, Graham, et al.
Pubblicazione: (2025)
di: Hill, Graham, et al.
Pubblicazione: (2025)
ChatGPT is not A Man but Das Man: Representativeness and Structural Consistency of Silicon Samples Generated by Large Language Models
di: Li, Dai, et al.
Pubblicazione: (2025)
di: Li, Dai, et al.
Pubblicazione: (2025)
LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers
di: Li, Lingyao, et al.
Pubblicazione: (2026)
di: Li, Lingyao, et al.
Pubblicazione: (2026)
Evaluating Large Language Models for Anxiety and Depression Classification using Counseling and Psychotherapy Transcripts
di: Sun, Junwei, et al.
Pubblicazione: (2024)
di: Sun, Junwei, et al.
Pubblicazione: (2024)
Practical Insights into Fair Comparison and Evaluation Frame for Neutral-Atom Compilers
di: Khusainov, Emil, et al.
Pubblicazione: (2026)
di: Khusainov, Emil, et al.
Pubblicazione: (2026)
Automated Thematic Analyses Using LLMs: Xylazine Wound Management Social Media Chatter Use Case
di: Hairston, JaMor, et al.
Pubblicazione: (2025)
di: Hairston, JaMor, et al.
Pubblicazione: (2025)
"Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills
di: Liu, Yi, et al.
Pubblicazione: (2026)
di: Liu, Yi, et al.
Pubblicazione: (2026)
Multilingual Machine Translation with Quantum Encoder Decoder Attention-based Convolutional Variational Circuits
di: Dikshit, Subrit, et al.
Pubblicazione: (2025)
di: Dikshit, Subrit, et al.
Pubblicazione: (2025)
Iterative Prompting with Persuasion Skills in Jailbreaking Large Language Models
di: Ke, Shih-Wen, et al.
Pubblicazione: (2025)
di: Ke, Shih-Wen, et al.
Pubblicazione: (2025)
Large Language models for Time Series Analysis: Techniques, Applications, and Challenges
di: Shi, Feifei, et al.
Pubblicazione: (2025)
di: Shi, Feifei, et al.
Pubblicazione: (2025)
Bayesian Orchestration of Multi-LLM Agents for Cost-Aware Sequential Decision-Making
di: Amin, Danial
Pubblicazione: (2026)
di: Amin, Danial
Pubblicazione: (2026)
Are Large Language Models Reliable Argument Quality Annotators?
di: Mirzakhmedova, Nailia, et al.
Pubblicazione: (2024)
di: Mirzakhmedova, Nailia, et al.
Pubblicazione: (2024)
Machine Translation with Large Language Models: Decoder Only vs. Encoder-Decoder
di: M., Abhinav P., et al.
Pubblicazione: (2024)
di: M., Abhinav P., et al.
Pubblicazione: (2024)
Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning
di: Huang, Yiming, et al.
Pubblicazione: (2026)
di: Huang, Yiming, et al.
Pubblicazione: (2026)
How Many Bytes Can You Take Out Of Brain-To-Text Decoding?
di: Antonello, Richard, et al.
Pubblicazione: (2024)
di: Antonello, Richard, et al.
Pubblicazione: (2024)
Artificial Agency and Large Language Models
di: van Lier, Maud, et al.
Pubblicazione: (2024)
di: van Lier, Maud, et al.
Pubblicazione: (2024)
Documenti analoghi
-
HalluShift: Measuring Distribution Shifts towards Hallucination Detection in LLMs
di: Dasgupta, Sharanya, et al.
Pubblicazione: (2025) -
Towards Evaluating Large Language Models for Graph Query Generation
di: Munir, Siraj, et al.
Pubblicazione: (2024) -
Multi-Faceted Evaluation of Modeling Languages for Augmented Reality Applications -- The Case of ARWFML
di: Muff, Fabian, et al.
Pubblicazione: (2024) -
Comparative Evaluation of Prompting and Fine-Tuning for Applying Large Language Models to Grid-Structured Geospatial Data
di: Dhruv, Akash, et al.
Pubblicazione: (2025) -
Benchmarking LLMs for Political Science: A United Nations Perspective
di: Liang, Yueqing, et al.
Pubblicazione: (2025)