MiniGPT-Reverse-Designing: Predicting Image Adjustments Utilizing MiniGPT-4
Fuente:
arXiv
Saved in:
| Main Authors: | Azizi, Vahid, Koochaki, Fatemeh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens
by: Zheng, Kaizhi, et al.
Published: (2023)
by: Zheng, Kaizhi, et al.
Published: (2023)
MiniGPT-Pancreas: Multimodal Large Language Model for Pancreas Cancer Classification and Detection
by: Moglia, Andrea, et al.
Published: (2024)
by: Moglia, Andrea, et al.
Published: (2024)
MiniGPT-Med: Large Language Model as a General Interface for Radiology Diagnosis
by: Alkhaldi, Asma, et al.
Published: (2024)
by: Alkhaldi, Asma, et al.
Published: (2024)
MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors
by: Tang, Yuan, et al.
Published: (2024)
by: Tang, Yuan, et al.
Published: (2024)
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
by: Ataallah, Kirolos, et al.
Published: (2024)
by: Ataallah, Kirolos, et al.
Published: (2024)
M-MiniGPT4: Multilingual VLLM Alignment via Translated Data
by: Han, Seung Hun, et al.
Published: (2026)
by: Han, Seung Hun, et al.
Published: (2026)
MiniGPT: Rebuilding GPT from First Principles
by: Joseph, Jibin
Published: (2026)
by: Joseph, Jibin
Published: (2026)
LlamaRec-LKG-RAG: A Single-Pass, Learnable Knowledge Graph-RAG Framework for LLM-Based Ranking
by: Azizi, Vahid, et al.
Published: (2025)
by: Azizi, Vahid, et al.
Published: (2025)
Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities
by: Xie, Zhifei, et al.
Published: (2024)
by: Xie, Zhifei, et al.
Published: (2024)
GPT as Psychologist? Preliminary Evaluations for GPT-4V on Visual Affective Computing
by: Lu, Hao, et al.
Published: (2024)
by: Lu, Hao, et al.
Published: (2024)
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
by: Chen, Junying, et al.
Published: (2025)
by: Chen, Junying, et al.
Published: (2025)
TractoGPT: A GPT architecture for White Matter Segmentation
by: Goel, Anoushkrit, et al.
Published: (2025)
by: Goel, Anoushkrit, et al.
Published: (2025)
OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
by: Chen, Zhihong, et al.
Published: (2025)
by: Chen, Zhihong, et al.
Published: (2025)
Will GPT-4 Run DOOM?
by: de Wynter, Adrian
Published: (2024)
by: de Wynter, Adrian
Published: (2024)
Zero-shot Building Age Classification from Facade Image Using GPT-4
by: Zeng, Zichao, et al.
Published: (2024)
by: Zeng, Zichao, et al.
Published: (2024)
A Systematic Evaluation of GPT-4V's Multimodal Capability for Medical Image Analysis
by: Li, Yingshu, et al.
Published: (2023)
by: Li, Yingshu, et al.
Published: (2023)
VTG-GPT: Tuning-Free Zero-Shot Video Temporal Grounding with GPT
by: Xu, Yifang, et al.
Published: (2024)
by: Xu, Yifang, et al.
Published: (2024)
An Evaluation of GPT-4V and Gemini in Online VQA
by: Liu, Mengchen, et al.
Published: (2023)
by: Liu, Mengchen, et al.
Published: (2023)
GPT-4V Takes the Wheel: Promises and Challenges for Pedestrian Behavior Prediction
by: Huang, Jia, et al.
Published: (2023)
by: Huang, Jia, et al.
Published: (2023)
MiniCPM-V: A GPT-4V Level MLLM on Your Phone
by: Yao, Yuan, et al.
Published: (2024)
by: Yao, Yuan, et al.
Published: (2024)
MiniMaxAD: A Lightweight Autoencoder for Feature-Rich Anomaly Detection
by: Wang, Fengjie, et al.
Published: (2024)
by: Wang, Fengjie, et al.
Published: (2024)
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
by: Zhang, Shaolei, et al.
Published: (2025)
by: Zhang, Shaolei, et al.
Published: (2025)
Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
by: Ye, Junyan, et al.
Published: (2025)
by: Ye, Junyan, et al.
Published: (2025)
IQAGPT: Image Quality Assessment with Vision-language and ChatGPT Models
by: Chen, Zhihao, et al.
Published: (2023)
by: Chen, Zhihao, et al.
Published: (2023)
Evaluation of GPT-4o and GPT-4o-mini's Vision Capabilities for Compositional Analysis from Dried Solution Drops
by: Dangi, Deven B., et al.
Published: (2024)
by: Dangi, Deven B., et al.
Published: (2024)
RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding
by: Xu, Linrui, et al.
Published: (2024)
by: Xu, Linrui, et al.
Published: (2024)
Demystifying the Potential of ChatGPT-4 Vision for Construction Progress Monitoring
by: Ersoz, Ahmet Bahaddin
Published: (2024)
by: Ersoz, Ahmet Bahaddin
Published: (2024)
ForgeryGPT: A Multimodal LLM for Interpretable Image Forgery Detection and Localization
by: Zhang, Fanrui, et al.
Published: (2024)
by: Zhang, Fanrui, et al.
Published: (2024)
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
by: Li, Yanwei, et al.
Published: (2024)
by: Li, Yanwei, et al.
Published: (2024)
MedGemma vs GPT-4: Open-Source and Proprietary Zero-shot Medical Disease Classification from Images
by: Prottasha, Md. Sazzadul Islam, et al.
Published: (2025)
by: Prottasha, Md. Sazzadul Islam, et al.
Published: (2025)
Evaluating ChatGPT's Performance in Classifying Pneumonia from Chest X-Ray Images
by: Prahallad, Pragna, et al.
Published: (2025)
by: Prahallad, Pragna, et al.
Published: (2025)
Is ChatGPT-5 Ready for Mammogram VQA?
by: Li, Qiang, et al.
Published: (2025)
by: Li, Qiang, et al.
Published: (2025)
Video-GPT via Next Clip Diffusion
by: Zhuang, Shaobin, et al.
Published: (2025)
by: Zhuang, Shaobin, et al.
Published: (2025)
G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o
by: Tong, Tony Cheng, et al.
Published: (2024)
by: Tong, Tony Cheng, et al.
Published: (2024)
CollagePrompt: A Benchmark for Budget-Friendly Visual Recognition with GPT-4V
by: Xu, Siyu, et al.
Published: (2024)
by: Xu, Siyu, et al.
Published: (2024)
3DAxisPrompt: Promoting the 3D Grounding and Reasoning in GPT-4o
by: Liu, Dingning, et al.
Published: (2025)
by: Liu, Dingning, et al.
Published: (2025)
Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
by: Lai, Xin, et al.
Published: (2025)
by: Lai, Xin, et al.
Published: (2025)
Preliminary Explorations with GPT-4o(mni) Native Image Generation
by: Cao, Pu, et al.
Published: (2025)
by: Cao, Pu, et al.
Published: (2025)
GPTDrawer: Enhancing Visual Synthesis through ChatGPT
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
FlattenGPT: Depth Compression for Transformer with Layer Flattening
by: Xu, Ruihan, et al.
Published: (2026)
by: Xu, Ruihan, et al.
Published: (2026)
Similar Items
-
MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens
by: Zheng, Kaizhi, et al.
Published: (2023) -
MiniGPT-Pancreas: Multimodal Large Language Model for Pancreas Cancer Classification and Detection
by: Moglia, Andrea, et al.
Published: (2024) -
MiniGPT-Med: Large Language Model as a General Interface for Radiology Diagnosis
by: Alkhaldi, Asma, et al.
Published: (2024) -
MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors
by: Tang, Yuan, et al.
Published: (2024) -
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
by: Ataallah, Kirolos, et al.
Published: (2024)