LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents
Fuente:
arXiv
Guardado en:
| Autores principales: | Patel, Hitesh Laxmichand, Agarwal, Amit, Kumar, Bhargava, Gupta, Karan, Pattnayak, Priyaranjan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Tokenization Matters: Improving Zero-Shot NER for Indic Languages
por: Pattnayak, Priyaranjan, et al.
Publicado: (2025)
por: Pattnayak, Priyaranjan, et al.
Publicado: (2025)
Clinical QA 2.0: Multi-Task Learning for Answer Extraction and Categorization
por: Pattnayak, Priyaranjan, et al.
Publicado: (2025)
por: Pattnayak, Priyaranjan, et al.
Publicado: (2025)
Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts
por: Agarwal, Amit, et al.
Publicado: (2024)
por: Agarwal, Amit, et al.
Publicado: (2024)
LLM-Guided Lifecycle-Aware Clustering of Multi-Turn Customer Support Conversations
por: Pattnayak, Priyaranjan, et al.
Publicado: (2026)
por: Pattnayak, Priyaranjan, et al.
Publicado: (2026)
Hard Negative Mining for Domain-Specific Retrieval in Enterprise Systems
por: Meghwani, Hansa, et al.
Publicado: (2025)
por: Meghwani, Hansa, et al.
Publicado: (2025)
Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy
por: Pattnayak, Priyaranjan, et al.
Publicado: (2024)
por: Pattnayak, Priyaranjan, et al.
Publicado: (2024)
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
por: Kumar, Anurakt, et al.
Publicado: (2024)
por: Kumar, Anurakt, et al.
Publicado: (2024)
Hybrid AI for Responsive Multi-Turn Online Conversations with Novel Dynamic Routing and Feedback Adaptation
por: Pattnayak, Priyaranjan, et al.
Publicado: (2025)
por: Pattnayak, Priyaranjan, et al.
Publicado: (2025)
SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech Models
por: Dua, Karan, et al.
Publicado: (2025)
por: Dua, Karan, et al.
Publicado: (2025)
AccessEval: Benchmarking Disability Bias in Large Language Models
por: Panda, Srikant, et al.
Publicado: (2025)
por: Panda, Srikant, et al.
Publicado: (2025)
IndicSafe: A Benchmark for Evaluating Multilingual LLM Safety in South Asia
por: Pattnayak, Priyaranjan, et al.
Publicado: (2026)
por: Pattnayak, Priyaranjan, et al.
Publicado: (2026)
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use
por: Patel, Hitesh Laxmichand, et al.
Publicado: (2025)
por: Patel, Hitesh Laxmichand, et al.
Publicado: (2025)
EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
por: Banayeeanzade, Amin, et al.
Publicado: (2026)
por: Banayeeanzade, Amin, et al.
Publicado: (2026)
Prompt Leakage effect and defense strategies for multi-turn LLM interactions
por: Agarwal, Divyansh, et al.
Publicado: (2024)
por: Agarwal, Divyansh, et al.
Publicado: (2024)
SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning
por: Zhou, Kaiwen, et al.
Publicado: (2025)
por: Zhou, Kaiwen, et al.
Publicado: (2025)
Virus Infection Attack on LLMs: Your Poisoning Can Spread "VIA" Synthetic Data
por: Liang, Zi, et al.
Publicado: (2025)
por: Liang, Zi, et al.
Publicado: (2025)
LLM in the Shell: Generative Honeypots
por: Sladić, Muris, et al.
Publicado: (2023)
por: Sladić, Muris, et al.
Publicado: (2023)
Certifying LLM Safety against Adversarial Prompting
por: Kumar, Aounon, et al.
Publicado: (2023)
por: Kumar, Aounon, et al.
Publicado: (2023)
Synthetic Data for Veterinary EHR De-identification: Benefits, Limits, and Safety Trade-offs Under Fixed Compute
por: Brundage, David
Publicado: (2026)
por: Brundage, David
Publicado: (2026)
DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning
por: Zhang, Junbo, et al.
Publicado: (2026)
por: Zhang, Junbo, et al.
Publicado: (2026)
OneShield -- the Next Generation of LLM Guardrails
por: DeLuca, Chad, et al.
Publicado: (2025)
por: DeLuca, Chad, et al.
Publicado: (2025)
Token-level Data Selection for Safe LLM Fine-tuning
por: Li, Yanping, et al.
Publicado: (2026)
por: Li, Yanping, et al.
Publicado: (2026)
XMark: Reliable Multi-Bit Watermarking for LLM-Generated Texts
por: Xu, Jiahao, et al.
Publicado: (2026)
por: Xu, Jiahao, et al.
Publicado: (2026)
IndicJR: A Judge-Free Benchmark of Jailbreak Robustness in South Asian Languages
por: Pattnayak, Priyaranjan, et al.
Publicado: (2026)
por: Pattnayak, Priyaranjan, et al.
Publicado: (2026)
LLM for SoC Security: A Paradigm Shift
por: Saha, Dipayan, et al.
Publicado: (2023)
por: Saha, Dipayan, et al.
Publicado: (2023)
LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive Generation
por: Li, Xingyu, et al.
Publicado: (2025)
por: Li, Xingyu, et al.
Publicado: (2025)
Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation
por: Schwartz, Daniel, et al.
Publicado: (2025)
por: Schwartz, Daniel, et al.
Publicado: (2025)
What Was Your Prompt? A Remote Keylogging Attack on AI Assistants
por: Weiss, Roy, et al.
Publicado: (2024)
por: Weiss, Roy, et al.
Publicado: (2024)
When Reject Turns into Accept: Quantifying the Vulnerability of LLM-Based Scientific Reviewers to Indirect Prompt Injection
por: Sahoo, Devanshu, et al.
Publicado: (2025)
por: Sahoo, Devanshu, et al.
Publicado: (2025)
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
por: Benjamin, Victoria, et al.
Publicado: (2024)
por: Benjamin, Victoria, et al.
Publicado: (2024)
Contextualized Privacy Defense for LLM Agents
por: Wen, Yule, et al.
Publicado: (2026)
por: Wen, Yule, et al.
Publicado: (2026)
LLM Jailbreak Detection for (Almost) Free!
por: Chen, Guorui, et al.
Publicado: (2025)
por: Chen, Guorui, et al.
Publicado: (2025)
RedSage: A Cybersecurity Generalist LLM
por: Suryanto, Naufal, et al.
Publicado: (2026)
por: Suryanto, Naufal, et al.
Publicado: (2026)
FLAME: Flexible LLM-Assisted Moderation Engine
por: Bakulin, Ivan, et al.
Publicado: (2025)
por: Bakulin, Ivan, et al.
Publicado: (2025)
Geneshift: Impact of different scenario shift on Jailbreaking LLM
por: Wu, Tianyi, et al.
Publicado: (2025)
por: Wu, Tianyi, et al.
Publicado: (2025)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
por: Tong, Terry, et al.
Publicado: (2025)
por: Tong, Terry, et al.
Publicado: (2025)
SSG: Logit-Balanced Vocabulary Partitioning for LLM Watermarking
por: Gu, Chenxi, et al.
Publicado: (2026)
por: Gu, Chenxi, et al.
Publicado: (2026)
Beyond Jailbreaking: Auditing Contextual Privacy in LLM Agents
por: Das, Saswat, et al.
Publicado: (2025)
por: Das, Saswat, et al.
Publicado: (2025)
NeuroFilter: Privacy Guardrails for Conversational LLM Agents
por: Das, Saswat, et al.
Publicado: (2026)
por: Das, Saswat, et al.
Publicado: (2026)
Searching for Privacy Risks in LLM Agents via Simulation
por: Zhang, Yanzhe, et al.
Publicado: (2025)
por: Zhang, Yanzhe, et al.
Publicado: (2025)
Ejemplares similares
-
Tokenization Matters: Improving Zero-Shot NER for Indic Languages
por: Pattnayak, Priyaranjan, et al.
Publicado: (2025) -
Clinical QA 2.0: Multi-Task Learning for Answer Extraction and Categorization
por: Pattnayak, Priyaranjan, et al.
Publicado: (2025) -
Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts
por: Agarwal, Amit, et al.
Publicado: (2024) -
LLM-Guided Lifecycle-Aware Clustering of Multi-Turn Customer Support Conversations
por: Pattnayak, Priyaranjan, et al.
Publicado: (2026) -
Hard Negative Mining for Domain-Specific Retrieval in Enterprise Systems
por: Meghwani, Hansa, et al.
Publicado: (2025)