Beyond BeautifulSoup: Benchmarking LLM-Powered Web Scraping for Everyday Users
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bhardwaj, Arth, Diwan, Nirav, Wang, Gang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
von: Chu, Junjie, et al.
Veröffentlicht: (2026)
von: Chu, Junjie, et al.
Veröffentlicht: (2026)
PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design
von: Yang, Ruozhao, et al.
Veröffentlicht: (2025)
von: Yang, Ruozhao, et al.
Veröffentlicht: (2025)
SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers
von: Qin, Kaihua, et al.
Veröffentlicht: (2026)
von: Qin, Kaihua, et al.
Veröffentlicht: (2026)
AdaptiveGuard: Towards Adaptive Runtime Safety for LLM-Powered Software
von: Yang, Rui, et al.
Veröffentlicht: (2025)
von: Yang, Rui, et al.
Veröffentlicht: (2025)
AutoEG: Exploiting Known Third-Party Vulnerabilities in Black-Box Web Applications
von: Yang, Ruozhao, et al.
Veröffentlicht: (2026)
von: Yang, Ruozhao, et al.
Veröffentlicht: (2026)
Toward Patch Robustness Certification and Detection for Deep Learning Systems Beyond Consistent Samples
von: Zhou, Qilin, et al.
Veröffentlicht: (2025)
von: Zhou, Qilin, et al.
Veröffentlicht: (2025)
Implicit Patterns in LLM-Based Binary Analysis
von: Li, Qiang, et al.
Veröffentlicht: (2026)
von: Li, Qiang, et al.
Veröffentlicht: (2026)
Unveiling the Landscape of LLM Deployment in the Wild: An Empirical Study
von: Hou, Xinyi, et al.
Veröffentlicht: (2025)
von: Hou, Xinyi, et al.
Veröffentlicht: (2025)
Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs
von: Yuan, He Yang, et al.
Veröffentlicht: (2026)
von: Yuan, He Yang, et al.
Veröffentlicht: (2026)
VULPO: Context-Aware Vulnerability Detection via On-Policy LLM Optimization
von: Li, Youpeng, et al.
Veröffentlicht: (2025)
von: Li, Youpeng, et al.
Veröffentlicht: (2025)
LLM-enabled Applications Require System-Level Threat Monitoring
von: Zhang, Yedi, et al.
Veröffentlicht: (2026)
von: Zhang, Yedi, et al.
Veröffentlicht: (2026)
DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle
von: Tang, Yuheng, et al.
Veröffentlicht: (2026)
von: Tang, Yuheng, et al.
Veröffentlicht: (2026)
From SFT to RL: Demystifying the Post-Training Pipeline for LLM-based Vulnerability Detection
von: Li, Youpeng, et al.
Veröffentlicht: (2026)
von: Li, Youpeng, et al.
Veröffentlicht: (2026)
Beyond Embeddings: Interpretable Feature Extraction for Binary Code Similarity
von: Gagnon, Charles E., et al.
Veröffentlicht: (2025)
von: Gagnon, Charles E., et al.
Veröffentlicht: (2025)
Casting a SPELL: Sentence Pairing Exploration for LLM Limitation-breaking
von: Huang, Yifan, et al.
Veröffentlicht: (2025)
von: Huang, Yifan, et al.
Veröffentlicht: (2025)
Harnessing the Power of LLMs in Source Code Vulnerability Detection
von: Mahyari, Andrew A
Veröffentlicht: (2024)
von: Mahyari, Andrew A
Veröffentlicht: (2024)
Beyond Trusting Trust: Multi-Model Validation for Robust Code Generation
von: McDanel, Bradley
Veröffentlicht: (2025)
von: McDanel, Bradley
Veröffentlicht: (2025)
Zero-Permission Manipulation: Can We Trust Large Multimodal Model Powered GUI Agents?
von: Qian, Yi, et al.
Veröffentlicht: (2026)
von: Qian, Yi, et al.
Veröffentlicht: (2026)
Beyond Classification: Evaluating LLMs for Fine-Grained Automatic Malware Behavior Auditing
von: Zheng, Xinran, et al.
Veröffentlicht: (2025)
von: Zheng, Xinran, et al.
Veröffentlicht: (2025)
Benchmarking Prompt Engineering Techniques for Secure Code Generation with GPT Models
von: Bruni, Marc, et al.
Veröffentlicht: (2025)
von: Bruni, Marc, et al.
Veröffentlicht: (2025)
DUALGUAGE: Automated Joint Security-Functionality Benchmarking for Secure Code Generation
von: Pathak, Abhijeet, et al.
Veröffentlicht: (2025)
von: Pathak, Abhijeet, et al.
Veröffentlicht: (2025)
CASTLE: Benchmarking Dataset for Static Code Analyzers and LLMs towards CWE Detection
von: Dubniczky, Richard A., et al.
Veröffentlicht: (2025)
von: Dubniczky, Richard A., et al.
Veröffentlicht: (2025)
Security of LLM-generated Code: A Comparative Analysis
von: Morkonda, Srivathsan G, et al.
Veröffentlicht: (2026)
von: Morkonda, Srivathsan G, et al.
Veröffentlicht: (2026)
The potential of LLM-generated reports in DevSecOps
von: Lykousas, Nikolaos, et al.
Veröffentlicht: (2024)
von: Lykousas, Nikolaos, et al.
Veröffentlicht: (2024)
Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing
von: Peng, Jiaren, et al.
Veröffentlicht: (2026)
von: Peng, Jiaren, et al.
Veröffentlicht: (2026)
An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong Detection
von: Yan, Shenao, et al.
Veröffentlicht: (2024)
von: Yan, Shenao, et al.
Veröffentlicht: (2024)
Harnessing Large Language Models for Software Vulnerability Detection: A Comprehensive Benchmarking Study
von: Tamberg, Karl, et al.
Veröffentlicht: (2024)
von: Tamberg, Karl, et al.
Veröffentlicht: (2024)
SKILLS: Structured Knowledge Injection for LLM-Driven Telecommunications Operations
von: Brett, Ivo
Veröffentlicht: (2026)
von: Brett, Ivo
Veröffentlicht: (2026)
Measuring and Exploiting Contextual Bias in LLM-Assisted Security Code Review
von: Mitropoulos, Dimitris, et al.
Veröffentlicht: (2026)
von: Mitropoulos, Dimitris, et al.
Veröffentlicht: (2026)
Needles at Scale: LLM-Assisted Target Selection for Windows Vulnerability Research
von: Bommarito II, Michael J.
Veröffentlicht: (2026)
von: Bommarito II, Michael J.
Veröffentlicht: (2026)
Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
von: He, Ping, et al.
Veröffentlicht: (2025)
von: He, Ping, et al.
Veröffentlicht: (2025)
QLPro: Automated Code Vulnerability Discovery via LLM and Static Code Analysis Integration
von: Hu, Junze, et al.
Veröffentlicht: (2025)
von: Hu, Junze, et al.
Veröffentlicht: (2025)
Fortifying LLM-Based Code Generation with Graph-Based Reasoning on Secure Coding Practices
von: Patir, Rupam, et al.
Veröffentlicht: (2025)
von: Patir, Rupam, et al.
Veröffentlicht: (2025)
Semantic-Aware Fuzzing: An Empirical Framework for LLM-Guided, Reasoning-Driven Input Mutation
von: Lu, Mengdi, et al.
Veröffentlicht: (2025)
von: Lu, Mengdi, et al.
Veröffentlicht: (2025)
Taint-Style Vulnerability Detection and Confirmation for Node.js Packages Using LLM Agent Reasoning
von: Ni, Ronghao, et al.
Veröffentlicht: (2026)
von: Ni, Ronghao, et al.
Veröffentlicht: (2026)
Poisoned Identifiers Survive LLM Deobfuscation: A Case Study on Claude Opus 4.6
von: Lorenzo, Luis Guzmán
Veröffentlicht: (2026)
von: Lorenzo, Luis Guzmán
Veröffentlicht: (2026)
SOK: Exploring Hallucinations and Security Risks in AI-Assisted Software Development with Insights for LLM Deployment
von: Haque, Ariful, et al.
Veröffentlicht: (2025)
von: Haque, Ariful, et al.
Veröffentlicht: (2025)
SecureFixAgent: A Hybrid LLM Agent for Automated Python Static Vulnerability Repair
von: Gajjar, Jugal, et al.
Veröffentlicht: (2025)
von: Gajjar, Jugal, et al.
Veröffentlicht: (2025)
LLM4Vuln: A Unified Evaluation Framework for Decoupling and Enhancing LLMs' Vulnerability Reasoning
von: Sun, Yuqiang, et al.
Veröffentlicht: (2024)
von: Sun, Yuqiang, et al.
Veröffentlicht: (2024)
Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges
von: Al-Kaswan, Ali, et al.
Veröffentlicht: (2026)
von: Al-Kaswan, Ali, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
von: Chu, Junjie, et al.
Veröffentlicht: (2026) -
PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design
von: Yang, Ruozhao, et al.
Veröffentlicht: (2025) -
SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers
von: Qin, Kaihua, et al.
Veröffentlicht: (2026) -
AdaptiveGuard: Towards Adaptive Runtime Safety for LLM-Powered Software
von: Yang, Rui, et al.
Veröffentlicht: (2025) -
AutoEG: Exploiting Known Third-Party Vulnerabilities in Black-Box Web Applications
von: Yang, Ruozhao, et al.
Veröffentlicht: (2026)