Salvato in:
| Autori principali: | Gupta, Rushil, Hartford, Jason, Liu, Bang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2509.21403 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MAC: Multi-Agent Constitution Learning
di: Thareja, Rushil, et al.
Pubblicazione: (2026)
di: Thareja, Rushil, et al.
Pubblicazione: (2026)
Are We There Yet? Revealing the Risks of Utilizing Large Language Models in Scholarly Peer Review
di: Ye, Rui, et al.
Pubblicazione: (2024)
di: Ye, Rui, et al.
Pubblicazione: (2024)
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
di: Atil, Berk, et al.
Pubblicazione: (2025)
di: Atil, Berk, et al.
Pubblicazione: (2025)
DP-Fusion: Token-Level Differentially Private Inference for Large Language Models
di: Thareja, Rushil, et al.
Pubblicazione: (2025)
di: Thareja, Rushil, et al.
Pubblicazione: (2025)
How Far Are We From AGI: Are LLMs All We Need?
di: Feng, Tao, et al.
Pubblicazione: (2024)
di: Feng, Tao, et al.
Pubblicazione: (2024)
$f$-Trajectory Balance: A Loss Family for Tuning GFlowNets, Generative Models, and LLMs with Off- and On-Policy Data
di: Fawkes, Jake, et al.
Pubblicazione: (2026)
di: Fawkes, Jake, et al.
Pubblicazione: (2026)
Pairing Analogy-Augmented Generation with Procedural Memory for Procedural Q&A
di: Roth, K, et al.
Pubblicazione: (2024)
di: Roth, K, et al.
Pubblicazione: (2024)
STEP: Scientific Time-Series Encoder Pretraining via Cross-Domain Distillation
di: Zhang, Chen, et al.
Pubblicazione: (2026)
di: Zhang, Chen, et al.
Pubblicazione: (2026)
Evaluating Embedding Frameworks for Scientific Domain
di: Ahmed, Nouman, et al.
Pubblicazione: (2025)
di: Ahmed, Nouman, et al.
Pubblicazione: (2025)
Enough Coin Flips Can Make LLMs Act Bayesian
di: Gupta, Ritwik, et al.
Pubblicazione: (2025)
di: Gupta, Ritwik, et al.
Pubblicazione: (2025)
Holistic Capability Preservation: Towards Compact Yet Comprehensive Reasoning Models
di: Ling Team, et al.
Pubblicazione: (2025)
di: Ling Team, et al.
Pubblicazione: (2025)
RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval
di: Wen, Kaiyue, et al.
Pubblicazione: (2024)
di: Wen, Kaiyue, et al.
Pubblicazione: (2024)
Sociodemographic Prompting is Not Yet an Effective Approach for Simulating Subjective Judgments with LLMs
di: Sun, Huaman, et al.
Pubblicazione: (2023)
di: Sun, Huaman, et al.
Pubblicazione: (2023)
Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data?
di: Wang, Shaobo, et al.
Pubblicazione: (2025)
di: Wang, Shaobo, et al.
Pubblicazione: (2025)
We're Different, We're the Same: Creative Homogeneity Across LLMs
di: Wenger, Emily, et al.
Pubblicazione: (2025)
di: Wenger, Emily, et al.
Pubblicazione: (2025)
SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding
di: Li, Sihang, et al.
Pubblicazione: (2024)
di: Li, Sihang, et al.
Pubblicazione: (2024)
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
di: Zhou, Yujun, et al.
Pubblicazione: (2024)
di: Zhou, Yujun, et al.
Pubblicazione: (2024)
S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
di: Ma, Ruotian, et al.
Pubblicazione: (2025)
di: Ma, Ruotian, et al.
Pubblicazione: (2025)
Dynamic Prompt Fusion for Multi-Task and Cross-Domain Adaptation in LLMs
di: Hu, Xin, et al.
Pubblicazione: (2025)
di: Hu, Xin, et al.
Pubblicazione: (2025)
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
di: Xu, Wanghan, et al.
Pubblicazione: (2025)
di: Xu, Wanghan, et al.
Pubblicazione: (2025)
Sci-LoRA: Mixture of Scientific LoRAs for Cross-Domain Lay Paraphrasing
di: Cheng, Ming, et al.
Pubblicazione: (2025)
di: Cheng, Ming, et al.
Pubblicazione: (2025)
Can LLMs Separate Instructions From Data? And What Do We Even Mean By That?
di: Zverev, Egor, et al.
Pubblicazione: (2024)
di: Zverev, Egor, et al.
Pubblicazione: (2024)
SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA
di: Xu, Haozhou, et al.
Pubblicazione: (2025)
di: Xu, Haozhou, et al.
Pubblicazione: (2025)
AutoDiscovery: Open-ended Scientific Discovery via Bayesian Surprise
di: Agarwal, Dhruv, et al.
Pubblicazione: (2025)
di: Agarwal, Dhruv, et al.
Pubblicazione: (2025)
Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility
di: Mondal, Ishani, et al.
Pubblicazione: (2026)
di: Mondal, Ishani, et al.
Pubblicazione: (2026)
Neural-Bayesian Program Learning for Few-shot Dialogue Intent Parsing
di: Hong, Mengze, et al.
Pubblicazione: (2024)
di: Hong, Mengze, et al.
Pubblicazione: (2024)
Distilled Self-Critique of LLMs with Synthetic Data: a Bayesian Perspective
di: Gallego, Victor
Pubblicazione: (2023)
di: Gallego, Victor
Pubblicazione: (2023)
Can We Infer Confidential Properties of Training Data from LLMs?
di: Huang, Pengrun, et al.
Pubblicazione: (2025)
di: Huang, Pengrun, et al.
Pubblicazione: (2025)
Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
di: Zhang, Xuan, et al.
Pubblicazione: (2024)
di: Zhang, Xuan, et al.
Pubblicazione: (2024)
REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
di: Hu, Jian, et al.
Pubblicazione: (2025)
di: Hu, Jian, et al.
Pubblicazione: (2025)
Advancing Semantic Caching for LLMs with Domain-Specific Embeddings and Synthetic Data
di: Gill, Waris, et al.
Pubblicazione: (2025)
di: Gill, Waris, et al.
Pubblicazione: (2025)
Robust and Efficient Fine-tuning of LLMs with Bayesian Reparameterization of Low-Rank Adaptation
di: Sengupta, Ayan, et al.
Pubblicazione: (2024)
di: Sengupta, Ayan, et al.
Pubblicazione: (2024)
Failure Modes of LLMs for Causal Reasoning on Narratives
di: Yamin, Khurram, et al.
Pubblicazione: (2024)
di: Yamin, Khurram, et al.
Pubblicazione: (2024)
SciMON: Scientific Inspiration Machines Optimized for Novelty
di: Wang, Qingyun, et al.
Pubblicazione: (2023)
di: Wang, Qingyun, et al.
Pubblicazione: (2023)
References Indeed Matter? Reference-Free Preference Optimization for Conversational Query Reformulation
di: Kim, Doyoung, et al.
Pubblicazione: (2025)
di: Kim, Doyoung, et al.
Pubblicazione: (2025)
Where Do We Go from Here? Multi-scale Allocentric Relational Inference from Natural Spatial Descriptions
di: Paz-Argaman, Tzuf, et al.
Pubblicazione: (2024)
di: Paz-Argaman, Tzuf, et al.
Pubblicazione: (2024)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
di: Li, Nathaniel, et al.
Pubblicazione: (2024)
di: Li, Nathaniel, et al.
Pubblicazione: (2024)
Can We Count on LLMs? The Fixed-Effect Fallacy and Claims of GPT-4 Capabilities
di: Ball, Thomas, et al.
Pubblicazione: (2024)
di: Ball, Thomas, et al.
Pubblicazione: (2024)
OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
di: Aggarwal, Pranjal, et al.
Pubblicazione: (2025)
di: Aggarwal, Pranjal, et al.
Pubblicazione: (2025)
Project Alexandria: Towards Freeing Scientific Knowledge from Copyright Burdens via LLMs
di: Schuhmann, Christoph, et al.
Pubblicazione: (2025)
di: Schuhmann, Christoph, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MAC: Multi-Agent Constitution Learning
di: Thareja, Rushil, et al.
Pubblicazione: (2026) -
Are We There Yet? Revealing the Risks of Utilizing Large Language Models in Scholarly Peer Review
di: Ye, Rui, et al.
Pubblicazione: (2024) -
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
di: Atil, Berk, et al.
Pubblicazione: (2025) -
DP-Fusion: Token-Level Differentially Private Inference for Large Language Models
di: Thareja, Rushil, et al.
Pubblicazione: (2025) -
How Far Are We From AGI: Are LLMs All We Need?
di: Feng, Tao, et al.
Pubblicazione: (2024)