Propaganda AI: An Analysis of Semantic Divergence in Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Min, Nay Myat, Pham, Long H., Li, Yige, Sun, Jun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CROW: Eliminating Backdoors from Large Language Models via Internal Consistency Regularization
di: Min, Nay Myat, et al.
Pubblicazione: (2024)
di: Min, Nay Myat, et al.
Pubblicazione: (2024)
Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models
di: Min, Nay Myat, et al.
Pubblicazione: (2026)
di: Min, Nay Myat, et al.
Pubblicazione: (2026)
CORVUS: Red-Teaming Hallucination Detectors via Internal Signal Camouflage in Large Language Models
di: Min, Nay Myat, et al.
Pubblicazione: (2026)
di: Min, Nay Myat, et al.
Pubblicazione: (2026)
Unified Neural Backdoor Removal with Only Few Clean Samples through Unlearning and Relearning
di: Min, Nay Myat, et al.
Pubblicazione: (2024)
di: Min, Nay Myat, et al.
Pubblicazione: (2024)
Do Influence Functions Work on Large Language Models?
di: Li, Zhe, et al.
Pubblicazione: (2024)
di: Li, Zhe, et al.
Pubblicazione: (2024)
AUTOLAW: Enhancing Legal Compliance in Large Language Models via Case Law Generation and Jury-Inspired Deliberation
di: Nguyen, Tai D., et al.
Pubblicazione: (2025)
di: Nguyen, Tai D., et al.
Pubblicazione: (2025)
Unleashing the Unseen: Harnessing Benign Datasets for Jailbreaking Large Language Models
di: Zhao, Wei, et al.
Pubblicazione: (2024)
di: Zhao, Wei, et al.
Pubblicazione: (2024)
Graph Fusion Across Languages using Large Language Models
di: Kyaw, Kaung Myat, et al.
Pubblicazione: (2026)
di: Kyaw, Kaung Myat, et al.
Pubblicazione: (2026)
Adaptive Content Restriction for Large Language Models via Suffix Optimization
di: Li, Yige, et al.
Pubblicazione: (2025)
di: Li, Yige, et al.
Pubblicazione: (2025)
Are Large Language Models Good at Detecting Propaganda?
di: Jose, Julia, et al.
Pubblicazione: (2025)
di: Jose, Julia, et al.
Pubblicazione: (2025)
Large Language Models for Propaganda Span Annotation
di: Hasanain, Maram, et al.
Pubblicazione: (2023)
di: Hasanain, Maram, et al.
Pubblicazione: (2023)
LLMs Can Also Do Well! Breaking Barriers in Semantic Role Labeling via Large Language Models
di: Li, Xinxin, et al.
Pubblicazione: (2025)
di: Li, Xinxin, et al.
Pubblicazione: (2025)
Divergent Creativity in Humans and Large Language Models
di: Bellemare-Pepin, Antoine, et al.
Pubblicazione: (2024)
di: Bellemare-Pepin, Antoine, et al.
Pubblicazione: (2024)
"Understanding AI": Semantic Grounding in Large Language Models
di: Lyre, Holger
Pubblicazione: (2024)
di: Lyre, Holger
Pubblicazione: (2024)
Prompt-Response Semantic Divergence Metrics for Faithfulness Hallucination and Misalignment Detection in Large Language Models
di: Halperin, Igor
Pubblicazione: (2025)
di: Halperin, Igor
Pubblicazione: (2025)
D2LLM: Decomposed and Distilled Large Language Models for Semantic Search
di: Liao, Zihan, et al.
Pubblicazione: (2024)
di: Liao, Zihan, et al.
Pubblicazione: (2024)
Structured Semantic Cloaking for Jailbreak Attacks on Large Language Models
di: Sun, Xiaobing, et al.
Pubblicazione: (2026)
di: Sun, Xiaobing, et al.
Pubblicazione: (2026)
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
di: Li, Yige, et al.
Pubblicazione: (2025)
di: Li, Yige, et al.
Pubblicazione: (2025)
RKLD: Reverse KL-Divergence-based Knowledge Distillation for Unlearning Personal Information in Large Language Models
di: Wang, Bichen, et al.
Pubblicazione: (2024)
di: Wang, Bichen, et al.
Pubblicazione: (2024)
Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA
di: Huang, Sterling, et al.
Pubblicazione: (2026)
di: Huang, Sterling, et al.
Pubblicazione: (2026)
Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing
di: Li, Zhe, et al.
Pubblicazione: (2025)
di: Li, Zhe, et al.
Pubblicazione: (2025)
Enhancing Semantic Consistency of Large Language Models through Model Editing: An Interpretability-Oriented Approach
di: Yang, Jingyuan, et al.
Pubblicazione: (2025)
di: Yang, Jingyuan, et al.
Pubblicazione: (2025)
A Benchmark Dataset and Evaluation Framework for Vietnamese Large Language Models in Customer Support
di: Nguyen, Long S. T., et al.
Pubblicazione: (2025)
di: Nguyen, Long S. T., et al.
Pubblicazione: (2025)
From Voices to Validity: Leveraging Large Language Models (LLMs) for Textual Analysis of Policy Stakeholder Interviews
di: Liu, Alex, et al.
Pubblicazione: (2023)
di: Liu, Alex, et al.
Pubblicazione: (2023)
Semantic Consistency Regularization with Large Language Models for Semi-supervised Sentiment Analysis
di: Li, Kunrong, et al.
Pubblicazione: (2025)
di: Li, Kunrong, et al.
Pubblicazione: (2025)
Zero-Shot Defense Against Toxic Images via Inherent Multimodal Alignment in LVLMs
di: Zhao, Wei, et al.
Pubblicazione: (2025)
di: Zhao, Wei, et al.
Pubblicazione: (2025)
Structural Reformation of Large Language Model Neuron Encapsulation for Divergent Information Aggregation
di: Bakushev, Denis, et al.
Pubblicazione: (2025)
di: Bakushev, Denis, et al.
Pubblicazione: (2025)
Detecting Hallucinations in Large Language Models via Internal Attention Divergence Signals
di: van Dijk, Gijs
Pubblicazione: (2026)
di: van Dijk, Gijs
Pubblicazione: (2026)
Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
di: Liang, Tian, et al.
Pubblicazione: (2023)
di: Liang, Tian, et al.
Pubblicazione: (2023)
Semantic Pivots Enable Cross-Lingual Transfer in Large Language Models
di: He, Kaiyu, et al.
Pubblicazione: (2025)
di: He, Kaiyu, et al.
Pubblicazione: (2025)
Semantic and Structural Analysis of Implicit Biases in Large Language Models: An Interpretable Approach
di: Zhang, Renhan, et al.
Pubblicazione: (2025)
di: Zhang, Renhan, et al.
Pubblicazione: (2025)
How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective
di: Xiao, Teng, et al.
Pubblicazione: (2024)
di: Xiao, Teng, et al.
Pubblicazione: (2024)
Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
di: Li, Hang, et al.
Pubblicazione: (2025)
di: Li, Hang, et al.
Pubblicazione: (2025)
When Can Large Reasoning Models Save Thinking? Mechanistic Analysis of Behavioral Divergence in Reasoning
di: Zhu, Rongzhi, et al.
Pubblicazione: (2025)
di: Zhu, Rongzhi, et al.
Pubblicazione: (2025)
Generator-Assistant Stepwise Rollback Framework for Large Language Model Agent
di: Li, Xingzuo, et al.
Pubblicazione: (2025)
di: Li, Xingzuo, et al.
Pubblicazione: (2025)
XIFBench: Evaluating Large Language Models on Multilingual Instruction Following
di: Li, Zhenyu, et al.
Pubblicazione: (2025)
di: Li, Zhenyu, et al.
Pubblicazione: (2025)
Can GPT-4 Identify Propaganda? Annotation and Detection of Propaganda Spans in News Articles
di: Hasanain, Maram, et al.
Pubblicazione: (2024)
di: Hasanain, Maram, et al.
Pubblicazione: (2024)
Multi-Granularity Semantic Revision for Large Language Model Distillation
di: Liu, Xiaoyu, et al.
Pubblicazione: (2024)
di: Liu, Xiaoyu, et al.
Pubblicazione: (2024)
On the Semantics of Large Language Models
di: Schuele, Martin
Pubblicazione: (2025)
di: Schuele, Martin
Pubblicazione: (2025)
Beyond Divergent Creativity: A Human-Based Evaluation of Creativity in Large Language Models
di: Nakajima, Kumiko, et al.
Pubblicazione: (2026)
di: Nakajima, Kumiko, et al.
Pubblicazione: (2026)
Documenti analoghi
-
CROW: Eliminating Backdoors from Large Language Models via Internal Consistency Regularization
di: Min, Nay Myat, et al.
Pubblicazione: (2024) -
Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models
di: Min, Nay Myat, et al.
Pubblicazione: (2026) -
CORVUS: Red-Teaming Hallucination Detectors via Internal Signal Camouflage in Large Language Models
di: Min, Nay Myat, et al.
Pubblicazione: (2026) -
Unified Neural Backdoor Removal with Only Few Clean Samples through Unlearning and Relearning
di: Min, Nay Myat, et al.
Pubblicazione: (2024) -
Do Influence Functions Work on Large Language Models?
di: Li, Zhe, et al.
Pubblicazione: (2024)