Machine Generated Product Advertisements: Benchmarking LLMs Against Human Performance
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Ghosh, Sanjukta |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation
von: Limkonchotiwat, Peerat, et al.
Veröffentlicht: (2025)
von: Limkonchotiwat, Peerat, et al.
Veröffentlicht: (2025)
Human vs. Machine Deception: Distinguishing AI-Generated and Human-Written Fake News Using Ensemble Learning
von: Jaeger, Samuel, et al.
Veröffentlicht: (2026)
von: Jaeger, Samuel, et al.
Veröffentlicht: (2026)
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
von: Quan, Shanghaoran, et al.
Veröffentlicht: (2025)
von: Quan, Shanghaoran, et al.
Veröffentlicht: (2025)
Cards Against LLMs: Benchmarking Humor Alignment in Large Language Models
von: Fettach, Yousra, et al.
Veröffentlicht: (2026)
von: Fettach, Yousra, et al.
Veröffentlicht: (2026)
Visual Language Model as a Judge for Object Detection in Industrial Diagrams
von: Ghosh, Sanjukta
Veröffentlicht: (2025)
von: Ghosh, Sanjukta
Veröffentlicht: (2025)
Are You Human? An Adversarial Benchmark to Expose LLMs
von: Gressel, Gilad, et al.
Veröffentlicht: (2024)
von: Gressel, Gilad, et al.
Veröffentlicht: (2024)
Performance Gains of LLMs With Humans in a World of LLMs Versus Humans
von: McCullum, Lucas, et al.
Veröffentlicht: (2025)
von: McCullum, Lucas, et al.
Veröffentlicht: (2025)
AdTEC: A Unified Benchmark for Evaluating Text Quality in Search Engine Advertising
von: Zhang, Peinan, et al.
Veröffentlicht: (2024)
von: Zhang, Peinan, et al.
Veröffentlicht: (2024)
"Be My Cheese?": Cultural Nuance Benchmarking for Machine Translation in Multilingual LLMs
von: Van Doren, Madison, et al.
Veröffentlicht: (2026)
von: Van Doren, Madison, et al.
Veröffentlicht: (2026)
Striking Gold in Advertising: Standardization and Exploration of Ad Text Generation
von: Mita, Masato, et al.
Veröffentlicht: (2023)
von: Mita, Masato, et al.
Veröffentlicht: (2023)
Human-Level and Beyond: Benchmarking Large Language Models Against Clinical Pharmacists in Prescription Review
von: Yang, Yan, et al.
Veröffentlicht: (2025)
von: Yang, Yan, et al.
Veröffentlicht: (2025)
Generative AI Advertising as a Problem of Trustworthy Commercial Intervention
von: Qiu, Jingyi, et al.
Veröffentlicht: (2026)
von: Qiu, Jingyi, et al.
Veröffentlicht: (2026)
LLMs as Data Annotators: How Close Are We to Human Performance
von: Haq, Muhammad Uzair Ul, et al.
Veröffentlicht: (2025)
von: Haq, Muhammad Uzair Ul, et al.
Veröffentlicht: (2025)
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
von: Mavi, John, et al.
Veröffentlicht: (2024)
von: Mavi, John, et al.
Veröffentlicht: (2024)
HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
von: Li, Chance Jiajie, et al.
Veröffentlicht: (2025)
von: Li, Chance Jiajie, et al.
Veröffentlicht: (2025)
RiddleBench: A New Generative Reasoning Benchmark for LLMs
von: Halder, Deepon, et al.
Veröffentlicht: (2025)
von: Halder, Deepon, et al.
Veröffentlicht: (2025)
MATCHED: Multimodal Authorship-Attribution To Combat Human Trafficking in Escort-Advertisement Data
von: Saxena, Vageesh, et al.
Veröffentlicht: (2024)
von: Saxena, Vageesh, et al.
Veröffentlicht: (2024)
Improving the Performance of Radiology Report De-identification with Large-Scale Training and Benchmarking Against Cloud Vendor Methods
von: Prakash, Eva, et al.
Veröffentlicht: (2025)
von: Prakash, Eva, et al.
Veröffentlicht: (2025)
Teaching Values to Machines: Simulating Human-Like Behavior in LLMs
von: Yehudai, Asaf, et al.
Veröffentlicht: (2026)
von: Yehudai, Asaf, et al.
Veröffentlicht: (2026)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
Benchmarking Logistic Regression, SVM, and LightGBM Against BiLSTM with Attention for Sentiment Analysis on Indonesian Product Reviews
von: Hamdi, Razin Hafid, et al.
Veröffentlicht: (2026)
von: Hamdi, Razin Hafid, et al.
Veröffentlicht: (2026)
Defending Against Social Engineering Attacks in the Age of LLMs
von: Ai, Lin, et al.
Veröffentlicht: (2024)
von: Ai, Lin, et al.
Veröffentlicht: (2024)
Generative Data Augmentation using LLMs improves Distributional Robustness in Question Answering
von: Chowdhury, Arijit Ghosh, et al.
Veröffentlicht: (2023)
von: Chowdhury, Arijit Ghosh, et al.
Veröffentlicht: (2023)
Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis
von: Bhowmik, Shimanto, et al.
Veröffentlicht: (2025)
von: Bhowmik, Shimanto, et al.
Veröffentlicht: (2025)
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry
von: Wang, Shanshan, et al.
Veröffentlicht: (2025)
von: Wang, Shanshan, et al.
Veröffentlicht: (2025)
Beyond Traditional Benchmarks: Analyzing Behaviors of Open LLMs on Data-to-Text Generation
von: Kasner, Zdeněk, et al.
Veröffentlicht: (2024)
von: Kasner, Zdeněk, et al.
Veröffentlicht: (2024)
LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
von: Wu, Yuhao, et al.
Veröffentlicht: (2024)
von: Wu, Yuhao, et al.
Veröffentlicht: (2024)
RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems
von: Lin, Jingru, et al.
Veröffentlicht: (2025)
von: Lin, Jingru, et al.
Veröffentlicht: (2025)
HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue
von: Iyer, Laya, et al.
Veröffentlicht: (2026)
von: Iyer, Laya, et al.
Veröffentlicht: (2026)
CodeCipher: Learning to Obfuscate Source Code Against LLMs
von: Lin, Yalan, et al.
Veröffentlicht: (2024)
von: Lin, Yalan, et al.
Veröffentlicht: (2024)
How Far Are LLMs from Believable AI? A Benchmark for Evaluating the Believability of Human Behavior Simulation
von: Xiao, Yang, et al.
Veröffentlicht: (2023)
von: Xiao, Yang, et al.
Veröffentlicht: (2023)
Benchmarking Gaslighting Attacks Against Speech Large Language Models
von: Wu, Jinyang, et al.
Veröffentlicht: (2025)
von: Wu, Jinyang, et al.
Veröffentlicht: (2025)
Rethinking Machine Ethics -- Can LLMs Perform Moral Reasoning through the Lens of Moral Theories?
von: Zhou, Jingyan, et al.
Veröffentlicht: (2023)
von: Zhou, Jingyan, et al.
Veröffentlicht: (2023)
Grammaticality Judgments in Humans and Language Models: Revisiting Generative Grammar with LLMs
von: Johnsen, Lars G. B.
Veröffentlicht: (2025)
von: Johnsen, Lars G. B.
Veröffentlicht: (2025)
Human Bias in the Face of AI: Examining Human Judgment Against Text Labeled as AI Generated
von: Zhu, Tiffany, et al.
Veröffentlicht: (2024)
von: Zhu, Tiffany, et al.
Veröffentlicht: (2024)
Benchmarking Chinese Medical LLMs: A Medbench-based Analysis of Performance Gaps and Hierarchical Optimization Strategies
von: Jiang, Luyi, et al.
Veröffentlicht: (2025)
von: Jiang, Luyi, et al.
Veröffentlicht: (2025)
ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
von: Hua, Tianyu, et al.
Veröffentlicht: (2025)
von: Hua, Tianyu, et al.
Veröffentlicht: (2025)
LLM-Generated Negative News Headlines Dataset: Creation and Benchmarking Against Real Journalism
von: Babalola, Olusola, et al.
Veröffentlicht: (2025)
von: Babalola, Olusola, et al.
Veröffentlicht: (2025)
Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility
von: Mondal, Ishani, et al.
Veröffentlicht: (2026)
von: Mondal, Ishani, et al.
Veröffentlicht: (2026)
Estonian WinoGrande Dataset: Comparative Analysis of LLM Performance on Human and Machine Translation
von: Ojastu, Marii, et al.
Veröffentlicht: (2025)
von: Ojastu, Marii, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation
von: Limkonchotiwat, Peerat, et al.
Veröffentlicht: (2025) -
Human vs. Machine Deception: Distinguishing AI-Generated and Human-Written Fake News Using Ensemble Learning
von: Jaeger, Samuel, et al.
Veröffentlicht: (2026) -
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
von: Quan, Shanghaoran, et al.
Veröffentlicht: (2025) -
Cards Against LLMs: Benchmarking Humor Alignment in Large Language Models
von: Fettach, Yousra, et al.
Veröffentlicht: (2026) -
Visual Language Model as a Judge for Object Detection in Industrial Diagrams
von: Ghosh, Sanjukta
Veröffentlicht: (2025)