Text2VLM: Adapting Text-Only Datasets to Evaluate Alignment Training in Visual Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Downer, Gabriel, Craven, Sean, Ruck, Damian, Thomas, Jake |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Text Embedding Inversion Security for Multilingual Language Models
by: Chen, Yiyi, et al.
Published: (2024)
by: Chen, Yiyi, et al.
Published: (2024)
Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
by: Liu, Ken Ziyu, et al.
Published: (2025)
by: Liu, Ken Ziyu, et al.
Published: (2025)
Poison Once, Control Anywhere: Clean-Text Visual Backdoors in VLM-based Mobile Agents
by: Wang, Xuan, et al.
Published: (2025)
by: Wang, Xuan, et al.
Published: (2025)
From Thinking to Output: Chain-of-Thought and Text Generation Characteristics in Reasoning Language Models
by: Liu, Junhao, et al.
Published: (2025)
by: Liu, Junhao, et al.
Published: (2025)
Finding a Wolf in Sheep's Clothing: Combating Adversarial Text-To-Image Prompts with Text Summarization
by: Cooper, Portia, et al.
Published: (2024)
by: Cooper, Portia, et al.
Published: (2024)
Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations
by: Teja, Lekkala Sai, et al.
Published: (2025)
by: Teja, Lekkala Sai, et al.
Published: (2025)
BitAbuse: A Dataset of Visually Perturbed Texts for Defending Phishing Attacks
by: Lee, Hanyong, et al.
Published: (2025)
by: Lee, Hanyong, et al.
Published: (2025)
Towards Privacy-Preserving Large Language Model: Text-free Inference Through Alignment and Adaptation
by: Yoon, Jeongho, et al.
Published: (2026)
by: Yoon, Jeongho, et al.
Published: (2026)
Adversarial Text Purification: A Large Language Model Approach for Defense
by: Moraffah, Raha, et al.
Published: (2024)
by: Moraffah, Raha, et al.
Published: (2024)
Safe Text-to-Image Generation: Simply Sanitize the Prompt Embedding
by: Qiu, Huming, et al.
Published: (2024)
by: Qiu, Huming, et al.
Published: (2024)
Waterfall: Framework for Robust and Scalable Text Watermarking and Provenance for LLMs
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
XMark: Reliable Multi-Bit Watermarking for LLM-Generated Texts
by: Xu, Jiahao, et al.
Published: (2026)
by: Xu, Jiahao, et al.
Published: (2026)
Less is More: Sparse Watermarking in LLMs with Enhanced Text Quality
by: Hoang, Duy C., et al.
Published: (2024)
by: Hoang, Duy C., et al.
Published: (2024)
ForensicsData: A Digital Forensics Dataset for Large Language Models
by: Chakir, Youssef, et al.
Published: (2025)
by: Chakir, Youssef, et al.
Published: (2025)
Unleashing the Unseen: Harnessing Benign Datasets for Jailbreaking Large Language Models
by: Zhao, Wei, et al.
Published: (2024)
by: Zhao, Wei, et al.
Published: (2024)
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
by: Wei, Zeming, et al.
Published: (2023)
by: Wei, Zeming, et al.
Published: (2023)
SastBench: A Benchmark for Testing Agentic SAST Triage
by: Feiglin, Jake, et al.
Published: (2026)
by: Feiglin, Jake, et al.
Published: (2026)
Where to Start Alignment? Diffusion Large Language Model May Demand a Distinct Position
by: Xie, Zhixin, et al.
Published: (2025)
by: Xie, Zhixin, et al.
Published: (2025)
AgentAlign: Navigating Safety Alignment in the Shift from Informative to Agentic Large Language Models
by: Zhang, Jinchuan, et al.
Published: (2025)
by: Zhang, Jinchuan, et al.
Published: (2025)
Threat Modelling using Domain-Adapted Language Models: Empirical Evaluation and Insights
by: Pourhanifeh, Saba, et al.
Published: (2026)
by: Pourhanifeh, Saba, et al.
Published: (2026)
SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness
by: Huo, Jiahao, et al.
Published: (2026)
by: Huo, Jiahao, et al.
Published: (2026)
SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks
by: Li, Tianhao, et al.
Published: (2024)
by: Li, Tianhao, et al.
Published: (2024)
Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
by: Wang, Haoran, et al.
Published: (2023)
by: Wang, Haoran, et al.
Published: (2023)
Acquiring Clean Language Models from Backdoor Poisoned Datasets by Downscaling Frequency Space
by: Wu, Zongru, et al.
Published: (2024)
by: Wu, Zongru, et al.
Published: (2024)
Primus: A Pioneering Collection of Open-Source Datasets for Cybersecurity LLM Training
by: Yu, Yao-Ching, et al.
Published: (2025)
by: Yu, Yao-Ching, et al.
Published: (2025)
C2RUST-BENCH: A Minimized, Representative Dataset for C-to-Rust Transpilation Evaluation
by: Sirlanci, Melih, et al.
Published: (2025)
by: Sirlanci, Melih, et al.
Published: (2025)
On Evaluating The Performance of Watermarked Machine-Generated Texts Under Adversarial Attacks
by: Liu, Zesen, et al.
Published: (2024)
by: Liu, Zesen, et al.
Published: (2024)
Imitate Before Detect: Aligning Machine Stylistic Preference for Machine-Revised Text Detection
by: Chen, Jiaqi, et al.
Published: (2024)
by: Chen, Jiaqi, et al.
Published: (2024)
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
by: Kumar, Anurakt, et al.
Published: (2024)
by: Kumar, Anurakt, et al.
Published: (2024)
Mark My Words: Analyzing and Evaluating Language Model Watermarks
by: Piet, Julien, et al.
Published: (2023)
by: Piet, Julien, et al.
Published: (2023)
Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization
by: Liu, Tiantian, et al.
Published: (2024)
by: Liu, Tiantian, et al.
Published: (2024)
FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
by: Gong, Yichen, et al.
Published: (2023)
by: Gong, Yichen, et al.
Published: (2023)
Lifelong Safety Alignment for Language Models
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action
by: Shao, Yijia, et al.
Published: (2024)
by: Shao, Yijia, et al.
Published: (2024)
CREBench: Evaluating Large Language Models in Cryptographic Binary Reverse Engineering
by: Chen, Baicheng, et al.
Published: (2026)
by: Chen, Baicheng, et al.
Published: (2026)
PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality
by: Li, Nanxi, et al.
Published: (2025)
by: Li, Nanxi, et al.
Published: (2025)
Text Steganography with Dynamic Codebook and Multimodal Large Language Model
by: Gao, Jianxin, et al.
Published: (2026)
by: Gao, Jianxin, et al.
Published: (2026)
From Texts to Shields: Convergence of Large Language Models and Cybersecurity
by: Li, Tao, et al.
Published: (2025)
by: Li, Tao, et al.
Published: (2025)
SeedPrints: Fingerprints Can Even Tell Which Seed Your Large Language Model Was Trained From
by: Tong, Yao, et al.
Published: (2025)
by: Tong, Yao, et al.
Published: (2025)
Leaking LoRa: An Evaluation of Password Leaks and Knowledge Storage in Large Language Models
by: Marinelli, Ryan, et al.
Published: (2025)
by: Marinelli, Ryan, et al.
Published: (2025)
Similar Items
-
Text Embedding Inversion Security for Multilingual Language Models
by: Chen, Yiyi, et al.
Published: (2024) -
Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
by: Liu, Ken Ziyu, et al.
Published: (2025) -
Poison Once, Control Anywhere: Clean-Text Visual Backdoors in VLM-based Mobile Agents
by: Wang, Xuan, et al.
Published: (2025) -
From Thinking to Output: Chain-of-Thought and Text Generation Characteristics in Reasoning Language Models
by: Liu, Junhao, et al.
Published: (2025) -
Finding a Wolf in Sheep's Clothing: Combating Adversarial Text-To-Image Prompts with Text Summarization
by: Cooper, Portia, et al.
Published: (2024)