Machine Generated Product Advertisements: Benchmarking LLMs Against Human Performance
Fuente:
arXiv
Saved in:
| Main Author: | Ghosh, Sanjukta |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation
by: Limkonchotiwat, Peerat, et al.
Published: (2025)
by: Limkonchotiwat, Peerat, et al.
Published: (2025)
Human vs. Machine Deception: Distinguishing AI-Generated and Human-Written Fake News Using Ensemble Learning
by: Jaeger, Samuel, et al.
Published: (2026)
by: Jaeger, Samuel, et al.
Published: (2026)
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
by: Quan, Shanghaoran, et al.
Published: (2025)
by: Quan, Shanghaoran, et al.
Published: (2025)
Cards Against LLMs: Benchmarking Humor Alignment in Large Language Models
by: Fettach, Yousra, et al.
Published: (2026)
by: Fettach, Yousra, et al.
Published: (2026)
Visual Language Model as a Judge for Object Detection in Industrial Diagrams
by: Ghosh, Sanjukta
Published: (2025)
by: Ghosh, Sanjukta
Published: (2025)
Are You Human? An Adversarial Benchmark to Expose LLMs
by: Gressel, Gilad, et al.
Published: (2024)
by: Gressel, Gilad, et al.
Published: (2024)
Performance Gains of LLMs With Humans in a World of LLMs Versus Humans
by: McCullum, Lucas, et al.
Published: (2025)
by: McCullum, Lucas, et al.
Published: (2025)
AdTEC: A Unified Benchmark for Evaluating Text Quality in Search Engine Advertising
by: Zhang, Peinan, et al.
Published: (2024)
by: Zhang, Peinan, et al.
Published: (2024)
"Be My Cheese?": Cultural Nuance Benchmarking for Machine Translation in Multilingual LLMs
by: Van Doren, Madison, et al.
Published: (2026)
by: Van Doren, Madison, et al.
Published: (2026)
Striking Gold in Advertising: Standardization and Exploration of Ad Text Generation
by: Mita, Masato, et al.
Published: (2023)
by: Mita, Masato, et al.
Published: (2023)
Human-Level and Beyond: Benchmarking Large Language Models Against Clinical Pharmacists in Prescription Review
by: Yang, Yan, et al.
Published: (2025)
by: Yang, Yan, et al.
Published: (2025)
Generative AI Advertising as a Problem of Trustworthy Commercial Intervention
by: Qiu, Jingyi, et al.
Published: (2026)
by: Qiu, Jingyi, et al.
Published: (2026)
LLMs as Data Annotators: How Close Are We to Human Performance
by: Haq, Muhammad Uzair Ul, et al.
Published: (2025)
by: Haq, Muhammad Uzair Ul, et al.
Published: (2025)
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
by: Mavi, John, et al.
Published: (2024)
by: Mavi, John, et al.
Published: (2024)
HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
by: Li, Chance Jiajie, et al.
Published: (2025)
by: Li, Chance Jiajie, et al.
Published: (2025)
RiddleBench: A New Generative Reasoning Benchmark for LLMs
by: Halder, Deepon, et al.
Published: (2025)
by: Halder, Deepon, et al.
Published: (2025)
MATCHED: Multimodal Authorship-Attribution To Combat Human Trafficking in Escort-Advertisement Data
by: Saxena, Vageesh, et al.
Published: (2024)
by: Saxena, Vageesh, et al.
Published: (2024)
Improving the Performance of Radiology Report De-identification with Large-Scale Training and Benchmarking Against Cloud Vendor Methods
by: Prakash, Eva, et al.
Published: (2025)
by: Prakash, Eva, et al.
Published: (2025)
Teaching Values to Machines: Simulating Human-Like Behavior in LLMs
by: Yehudai, Asaf, et al.
Published: (2026)
by: Yehudai, Asaf, et al.
Published: (2026)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
by: Jing, Huihao, et al.
Published: (2026)
by: Jing, Huihao, et al.
Published: (2026)
Benchmarking Logistic Regression, SVM, and LightGBM Against BiLSTM with Attention for Sentiment Analysis on Indonesian Product Reviews
by: Hamdi, Razin Hafid, et al.
Published: (2026)
by: Hamdi, Razin Hafid, et al.
Published: (2026)
Defending Against Social Engineering Attacks in the Age of LLMs
by: Ai, Lin, et al.
Published: (2024)
by: Ai, Lin, et al.
Published: (2024)
Generative Data Augmentation using LLMs improves Distributional Robustness in Question Answering
by: Chowdhury, Arijit Ghosh, et al.
Published: (2023)
by: Chowdhury, Arijit Ghosh, et al.
Published: (2023)
Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis
by: Bhowmik, Shimanto, et al.
Published: (2025)
by: Bhowmik, Shimanto, et al.
Published: (2025)
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry
by: Wang, Shanshan, et al.
Published: (2025)
by: Wang, Shanshan, et al.
Published: (2025)
Beyond Traditional Benchmarks: Analyzing Behaviors of Open LLMs on Data-to-Text Generation
by: Kasner, Zdeněk, et al.
Published: (2024)
by: Kasner, Zdeněk, et al.
Published: (2024)
LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
by: Wu, Yuhao, et al.
Published: (2024)
by: Wu, Yuhao, et al.
Published: (2024)
RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems
by: Lin, Jingru, et al.
Published: (2025)
by: Lin, Jingru, et al.
Published: (2025)
HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue
by: Iyer, Laya, et al.
Published: (2026)
by: Iyer, Laya, et al.
Published: (2026)
CodeCipher: Learning to Obfuscate Source Code Against LLMs
by: Lin, Yalan, et al.
Published: (2024)
by: Lin, Yalan, et al.
Published: (2024)
How Far Are LLMs from Believable AI? A Benchmark for Evaluating the Believability of Human Behavior Simulation
by: Xiao, Yang, et al.
Published: (2023)
by: Xiao, Yang, et al.
Published: (2023)
Benchmarking Gaslighting Attacks Against Speech Large Language Models
by: Wu, Jinyang, et al.
Published: (2025)
by: Wu, Jinyang, et al.
Published: (2025)
Rethinking Machine Ethics -- Can LLMs Perform Moral Reasoning through the Lens of Moral Theories?
by: Zhou, Jingyan, et al.
Published: (2023)
by: Zhou, Jingyan, et al.
Published: (2023)
Grammaticality Judgments in Humans and Language Models: Revisiting Generative Grammar with LLMs
by: Johnsen, Lars G. B.
Published: (2025)
by: Johnsen, Lars G. B.
Published: (2025)
Human Bias in the Face of AI: Examining Human Judgment Against Text Labeled as AI Generated
by: Zhu, Tiffany, et al.
Published: (2024)
by: Zhu, Tiffany, et al.
Published: (2024)
Benchmarking Chinese Medical LLMs: A Medbench-based Analysis of Performance Gaps and Hierarchical Optimization Strategies
by: Jiang, Luyi, et al.
Published: (2025)
by: Jiang, Luyi, et al.
Published: (2025)
ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
by: Hua, Tianyu, et al.
Published: (2025)
by: Hua, Tianyu, et al.
Published: (2025)
LLM-Generated Negative News Headlines Dataset: Creation and Benchmarking Against Real Journalism
by: Babalola, Olusola, et al.
Published: (2025)
by: Babalola, Olusola, et al.
Published: (2025)
Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility
by: Mondal, Ishani, et al.
Published: (2026)
by: Mondal, Ishani, et al.
Published: (2026)
Estonian WinoGrande Dataset: Comparative Analysis of LLM Performance on Human and Machine Translation
by: Ojastu, Marii, et al.
Published: (2025)
by: Ojastu, Marii, et al.
Published: (2025)
Similar Items
-
Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation
by: Limkonchotiwat, Peerat, et al.
Published: (2025) -
Human vs. Machine Deception: Distinguishing AI-Generated and Human-Written Fake News Using Ensemble Learning
by: Jaeger, Samuel, et al.
Published: (2026) -
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
by: Quan, Shanghaoran, et al.
Published: (2025) -
Cards Against LLMs: Benchmarking Humor Alignment in Large Language Models
by: Fettach, Yousra, et al.
Published: (2026) -
Visual Language Model as a Judge for Object Detection in Industrial Diagrams
by: Ghosh, Sanjukta
Published: (2025)