Salvato in:
| Autori principali: | Zhang, Yuji, Wang, Qingyun, Qian, Cheng, Liu, Jiateng, Sun, Chenkai, Zhang, Denghui, Abdelzaher, Tarek, Zhai, Chengxiang, Nakov, Preslav, Ji, Heng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2506.06972 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation
di: Sun, Chenkai, et al.
Pubblicazione: (2025)
di: Sun, Chenkai, et al.
Pubblicazione: (2025)
Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models
di: Sahnan, Dhruv, et al.
Pubblicazione: (2026)
di: Sahnan, Dhruv, et al.
Pubblicazione: (2026)
From Chaos to Clarity: Claim Normalization to Empower Fact-Checking
di: Sundriyal, Megha, et al.
Pubblicazione: (2023)
di: Sundriyal, Megha, et al.
Pubblicazione: (2023)
The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination
di: Zhang, Yuji, et al.
Pubblicazione: (2025)
di: Zhang, Yuji, et al.
Pubblicazione: (2025)
Word Embeddings Are Steers for Language Models
di: Han, Chi, et al.
Pubblicazione: (2023)
di: Han, Chi, et al.
Pubblicazione: (2023)
PropaInsight: Toward Deeper Understanding of Propaganda in Terms of Techniques, Appeals, and Intent
di: Liu, Jiateng, et al.
Pubblicazione: (2024)
di: Liu, Jiateng, et al.
Pubblicazione: (2024)
EVEDIT: Event-based Knowledge Editing with Deductive Editing Boundaries
di: Liu, Jiateng, et al.
Pubblicazione: (2024)
di: Liu, Jiateng, et al.
Pubblicazione: (2024)
Context Engineering for Trustworthiness: Rescorla Wagner Steering Under Mixed and Inappropriate Contexts
di: Wang, Rushi, et al.
Pubblicazione: (2025)
di: Wang, Rushi, et al.
Pubblicazione: (2025)
Grounding Fallacies Misrepresenting Scientific Publications in Evidence
di: Glockner, Max, et al.
Pubblicazione: (2024)
di: Glockner, Max, et al.
Pubblicazione: (2024)
ISACL: Internal State Analyzer for Copyrighted Training Data Leakage
di: Zhang, Guangwei, et al.
Pubblicazione: (2025)
di: Zhang, Guangwei, et al.
Pubblicazione: (2025)
Generating Zero-shot Abstractive Explanations for Rumour Verification
di: Bilal, Iman Munire, et al.
Pubblicazione: (2024)
di: Bilal, Iman Munire, et al.
Pubblicazione: (2024)
How Does Prefix Matter in Reasoning Model Tuning?
di: Tomar, Raj Vardhan, et al.
Pubblicazione: (2026)
di: Tomar, Raj Vardhan, et al.
Pubblicazione: (2026)
UnsafeChain: Enhancing Reasoning Model Safety via Hard Cases
di: Tomar, Raj Vardhan, et al.
Pubblicazione: (2025)
di: Tomar, Raj Vardhan, et al.
Pubblicazione: (2025)
SciClaimEval: Cross-modal Claim Verification in Scientific Papers
di: Ho, Xanh, et al.
Pubblicazione: (2026)
di: Ho, Xanh, et al.
Pubblicazione: (2026)
MuSciClaims: Multimodal Scientific Claim Verification
di: Lal, Yash Kumar, et al.
Pubblicazione: (2025)
di: Lal, Yash Kumar, et al.
Pubblicazione: (2025)
Detecting Check-Worthy Claims in Political Debates, Speeches, and Interviews Using Audio Data
di: Ivanov, Petar, et al.
Pubblicazione: (2023)
di: Ivanov, Petar, et al.
Pubblicazione: (2023)
MuDRiC: Multi-Dialect Reasoning for Arabic Commonsense Validation
di: Elozeiri, Kareem, et al.
Pubblicazione: (2025)
di: Elozeiri, Kareem, et al.
Pubblicazione: (2025)
ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challenges
di: Qian, Cheng, et al.
Pubblicazione: (2025)
di: Qian, Cheng, et al.
Pubblicazione: (2025)
Veri-R1: Toward Precise and Faithful Claim Verification via Online Reinforcement Learning
di: He, Qi, et al.
Pubblicazione: (2025)
di: He, Qi, et al.
Pubblicazione: (2025)
TART: An Open-Source Tool-Augmented Framework for Explainable Table-based Reasoning
di: Lu, Xinyuan, et al.
Pubblicazione: (2024)
di: Lu, Xinyuan, et al.
Pubblicazione: (2024)
Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
di: Su, Jinyan, et al.
Pubblicazione: (2025)
di: Su, Jinyan, et al.
Pubblicazione: (2025)
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
di: Wang, Yuxia, et al.
Pubblicazione: (2024)
di: Wang, Yuxia, et al.
Pubblicazione: (2024)
Loki: An Open-Source Tool for Fact Verification
di: Li, Haonan, et al.
Pubblicazione: (2024)
di: Li, Haonan, et al.
Pubblicazione: (2024)
DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Models
di: Ren, Kaixuan, et al.
Pubblicazione: (2025)
di: Ren, Kaixuan, et al.
Pubblicazione: (2025)
Large Language Models are Few-Shot Training Example Generators: A Case Study in Fallacy Recognition
di: Alhindi, Tariq, et al.
Pubblicazione: (2023)
di: Alhindi, Tariq, et al.
Pubblicazione: (2023)
Rethinking STS and NLI in Large Language Models
di: Wang, Yuxia, et al.
Pubblicazione: (2023)
di: Wang, Yuxia, et al.
Pubblicazione: (2023)
CoDet-M4: Detecting Machine-Generated Code in Multi-Lingual, Multi-Generator and Multi-Domain Settings
di: Orel, Daniil, et al.
Pubblicazione: (2025)
di: Orel, Daniil, et al.
Pubblicazione: (2025)
VISPA: Pluralistic Alignment via Automatic Value Selection and Activation
di: Zheng, Shenyan, et al.
Pubblicazione: (2026)
di: Zheng, Shenyan, et al.
Pubblicazione: (2026)
Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification
di: Kumar, Sunisth, et al.
Pubblicazione: (2026)
di: Kumar, Sunisth, et al.
Pubblicazione: (2026)
Table-Text Alignment: Explaining Claim Verification Against Tables in Scientific Papers
di: Ho, Xanh, et al.
Pubblicazione: (2025)
di: Ho, Xanh, et al.
Pubblicazione: (2025)
Peerispect: Claim Verification in Scientific Peer Reviews
di: Ghorbanpour, Ali, et al.
Pubblicazione: (2026)
di: Ghorbanpour, Ali, et al.
Pubblicazione: (2026)
Adapting Fake News Detection to the Era of Large Language Models
di: Su, Jinyan, et al.
Pubblicazione: (2023)
di: Su, Jinyan, et al.
Pubblicazione: (2023)
SciMON: Scientific Inspiration Machines Optimized for Novelty
di: Wang, Qingyun, et al.
Pubblicazione: (2023)
di: Wang, Qingyun, et al.
Pubblicazione: (2023)
Knowledge Overshadowing Causes Amalgamated Hallucination in Large Language Models
di: Zhang, Yuji, et al.
Pubblicazione: (2024)
di: Zhang, Yuji, et al.
Pubblicazione: (2024)
SciClaimHunt: A Large Dataset for Evidence-based Scientific Claim Verification
di: Kumar, Sujit, et al.
Pubblicazione: (2025)
di: Kumar, Sujit, et al.
Pubblicazione: (2025)
Geometric-disentangelment Unlearning
di: Zhou, Duo, et al.
Pubblicazione: (2025)
di: Zhou, Duo, et al.
Pubblicazione: (2025)
DocCHA: Towards LLM-Augmented Interactive Online diagnosis System
di: Liu, Xinyi, et al.
Pubblicazione: (2025)
di: Liu, Xinyi, et al.
Pubblicazione: (2025)
Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models
di: Gao, Lang, et al.
Pubblicazione: (2024)
di: Gao, Lang, et al.
Pubblicazione: (2024)
If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents
di: Yang, Ke, et al.
Pubblicazione: (2024)
di: Yang, Ke, et al.
Pubblicazione: (2024)
ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety
di: Bates, Luke, et al.
Pubblicazione: (2025)
di: Bates, Luke, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation
di: Sun, Chenkai, et al.
Pubblicazione: (2025) -
Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models
di: Sahnan, Dhruv, et al.
Pubblicazione: (2026) -
From Chaos to Clarity: Claim Normalization to Empower Fact-Checking
di: Sundriyal, Megha, et al.
Pubblicazione: (2023) -
The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination
di: Zhang, Yuji, et al.
Pubblicazione: (2025) -
Word Embeddings Are Steers for Language Models
di: Han, Chi, et al.
Pubblicazione: (2023)