Gemini Pro Defeated by GPT-4V: Evidence from Education
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Gyeong-Geon, Latif, Ehsan, Shi, Lehong, Zhai, Xiaoming |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Using GPT-4 to Augment Unbalanced Data for Automatic Scoring
by: Fang, Luyang, et al.
Published: (2023)
by: Fang, Luyang, et al.
Published: (2023)
Applying Large Language Models and Chain-of-Thought for Automatic Scoring
by: Lee, Gyeong-Geon, et al.
Published: (2023)
by: Lee, Gyeong-Geon, et al.
Published: (2023)
G-SciEdBERT: A Contextualized LLM for Science Assessment Tasks in German
by: Latif, Ehsan, et al.
Published: (2024)
by: Latif, Ehsan, et al.
Published: (2024)
Realizing Visual Question Answering for Education: GPT-4V as a Multimodal AI
by: Lee, Gyeong-Geon, et al.
Published: (2024)
by: Lee, Gyeong-Geon, et al.
Published: (2024)
Fine-tuning ChatGPT for Automatic Scoring of Written Scientific Explanations in Chinese
by: Yang, Jie, et al.
Published: (2025)
by: Yang, Jie, et al.
Published: (2025)
Knowledge Distillation of LLM for Automatic Scoring of Science Education Assessments
by: Latif, Ehsan, et al.
Published: (2023)
by: Latif, Ehsan, et al.
Published: (2023)
Privacy-Preserved Automated Scoring using Federated Learning for Educational Research
by: Latif, Ehsan, et al.
Published: (2025)
by: Latif, Ehsan, et al.
Published: (2025)
Collaborative Learning with Artificial Intelligence Speakers (CLAIS): Pre-Service Elementary Science Teachers' Responses to the Prototype
by: Lee, Gyeong-Geon, et al.
Published: (2023)
by: Lee, Gyeong-Geon, et al.
Published: (2023)
An evaluation of LLMs for generating movie reviews: GPT-4o, Gemini-2.0 and DeepSeek-V3
by: Sands, Brendan, et al.
Published: (2025)
by: Sands, Brendan, et al.
Published: (2025)
SketchMind: A Multi-Agent Cognitive Framework for Assessing Student-Drawn Scientific Sketches
by: Latif, Ehsan, et al.
Published: (2025)
by: Latif, Ehsan, et al.
Published: (2025)
Can OpenAI o1 outperform humans in higher-order cognitive thinking?
by: Latif, Ehsan, et al.
Published: (2024)
by: Latif, Ehsan, et al.
Published: (2024)
Efficient Multi-Task Inferencing with a Shared Backbone and Lightweight Task-Specific Adapters for Automatic Scoring
by: Latif, Ehsan, et al.
Published: (2024)
by: Latif, Ehsan, et al.
Published: (2024)
Using ChatGPT for Science Learning: A Study on Pre-service Teachers' Lesson Planning
by: Lee, Gyeong-Geon, et al.
Published: (2024)
by: Lee, Gyeong-Geon, et al.
Published: (2024)
Gemini Embedding: Generalizable Embeddings from Gemini
by: Lee, Jinhyuk, et al.
Published: (2025)
by: Lee, Jinhyuk, et al.
Published: (2025)
A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5
by: Ma, Xingjun, et al.
Published: (2026)
by: Ma, Xingjun, et al.
Published: (2026)
ChatGPT vs Gemini vs LLaMA on Multilingual Sentiment Analysis
by: Buscemi, Alessio, et al.
Published: (2024)
by: Buscemi, Alessio, et al.
Published: (2024)
Artificial Intelligence Bias on English Language Learners in Automatic Scoring
by: Guo, Shuchen, et al.
Published: (2025)
by: Guo, Shuchen, et al.
Published: (2025)
REFIND at SemEval-2025 Task 3: Retrieval-Augmented Factuality Hallucination Detection in Large Language Models
by: Lee, DongGeon, et al.
Published: (2025)
by: Lee, DongGeon, et al.
Published: (2025)
Does Tone Change the Answer? Evaluating Prompt Politeness Effects on Modern LLMs: GPT, Gemini, and LLaMA
by: Cai, Hanyu, et al.
Published: (2025)
by: Cai, Hanyu, et al.
Published: (2025)
Defeating the Training-Inference Mismatch via FP16
by: Qi, Penghui, et al.
Published: (2025)
by: Qi, Penghui, et al.
Published: (2025)
Can generative AI figure out figurative language? The influence of idioms on essay scoring by ChatGPT, Gemini, and Deepseek
by: Oğuz, Enis
Published: (2025)
by: Oğuz, Enis
Published: (2025)
DeepSeek-R1 Outperforms Gemini 2.0 Pro, OpenAI o1, and o3-mini in Bilingual Complex Ophthalmology Reasoning
by: Xu, Pusheng, et al.
Published: (2025)
by: Xu, Pusheng, et al.
Published: (2025)
Zero-Shot End-to-End Relation Extraction in Chinese: A Comparative Study of Gemini, LLaMA and ChatGPT
by: Du, Shaoshuai, et al.
Published: (2025)
by: Du, Shaoshuai, et al.
Published: (2025)
An Evaluation of GPT-4V and Gemini in Online VQA
by: Liu, Mengchen, et al.
Published: (2023)
by: Liu, Mengchen, et al.
Published: (2023)
Language-Dependent Political Bias in AI: A Study of ChatGPT and Gemini
by: Yuksel, Dogus, et al.
Published: (2025)
by: Yuksel, Dogus, et al.
Published: (2025)
Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece
by: Kumar, Anshul
Published: (2026)
by: Kumar, Anshul
Published: (2026)
Gemma: Open Models Based on Gemini Research and Technology
by: Gemma Team, et al.
Published: (2024)
by: Gemma Team, et al.
Published: (2024)
GPT-4V Cannot Generate Radiology Reports Yet
by: Jiang, Yuyang, et al.
Published: (2024)
by: Jiang, Yuyang, et al.
Published: (2024)
Accelerating Scientific Research with Gemini: Case Studies and Common Techniques
by: Woodruff, David P., et al.
Published: (2026)
by: Woodruff, David P., et al.
Published: (2026)
Who Would Chatbots Vote For? Political Preferences of ChatGPT and Gemini in the 2024 European Union Elections
by: Haman, Michael, et al.
Published: (2024)
by: Haman, Michael, et al.
Published: (2024)
Assessing the Effectiveness of GPT-4o in Climate Change Evidence Synthesis and Systematic Assessments: Preliminary Insights
by: Joe, Elphin Tom, et al.
Published: (2024)
by: Joe, Elphin Tom, et al.
Published: (2024)
How to Detect and Defeat Molecular Mirage: A Metric-Driven Benchmark for Hallucination in LLM-based Molecular Comprehension
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
GPTEval: A Survey on Assessments of ChatGPT and GPT-4
by: Mao, Rui, et al.
Published: (2023)
by: Mao, Rui, et al.
Published: (2023)
Exploring the Capability of ChatGPT to Reproduce Human Labels for Social Computing Tasks (Extended Version)
by: Zhu, Yiming, et al.
Published: (2024)
by: Zhu, Yiming, et al.
Published: (2024)
Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings
by: Hackl, Veronika, et al.
Published: (2023)
by: Hackl, Veronika, et al.
Published: (2023)
ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
by: Chen, Guiming Hardy, et al.
Published: (2024)
by: Chen, Guiming Hardy, et al.
Published: (2024)
GPT-4 Technical Report
by: OpenAI, et al.
Published: (2023)
by: OpenAI, et al.
Published: (2023)
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
by: Gemini Team, et al.
Published: (2024)
by: Gemini Team, et al.
Published: (2024)
Building Production-Ready Probes For Gemini
by: Kramár, János, et al.
Published: (2026)
by: Kramár, János, et al.
Published: (2026)
Fairness of Automatic Speech Recognition: Looking Through a Philosophical Lens
by: Choi, Anna Seo Gyeong, et al.
Published: (2025)
by: Choi, Anna Seo Gyeong, et al.
Published: (2025)
Similar Items
-
Using GPT-4 to Augment Unbalanced Data for Automatic Scoring
by: Fang, Luyang, et al.
Published: (2023) -
Applying Large Language Models and Chain-of-Thought for Automatic Scoring
by: Lee, Gyeong-Geon, et al.
Published: (2023) -
G-SciEdBERT: A Contextualized LLM for Science Assessment Tasks in German
by: Latif, Ehsan, et al.
Published: (2024) -
Realizing Visual Question Answering for Education: GPT-4V as a Multimodal AI
by: Lee, Gyeong-Geon, et al.
Published: (2024) -
Fine-tuning ChatGPT for Automatic Scoring of Written Scientific Explanations in Chinese
by: Yang, Jie, et al.
Published: (2025)