MiniGPT-Reverse-Designing: Predicting Image Adjustments Utilizing MiniGPT-4
Fuente:
arXiv
Guardado en:
| Autores principales: | Azizi, Vahid, Koochaki, Fatemeh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens
por: Zheng, Kaizhi, et al.
Publicado: (2023)
por: Zheng, Kaizhi, et al.
Publicado: (2023)
MiniGPT-Pancreas: Multimodal Large Language Model for Pancreas Cancer Classification and Detection
por: Moglia, Andrea, et al.
Publicado: (2024)
por: Moglia, Andrea, et al.
Publicado: (2024)
MiniGPT-Med: Large Language Model as a General Interface for Radiology Diagnosis
por: Alkhaldi, Asma, et al.
Publicado: (2024)
por: Alkhaldi, Asma, et al.
Publicado: (2024)
MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors
por: Tang, Yuan, et al.
Publicado: (2024)
por: Tang, Yuan, et al.
Publicado: (2024)
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
por: Ataallah, Kirolos, et al.
Publicado: (2024)
por: Ataallah, Kirolos, et al.
Publicado: (2024)
M-MiniGPT4: Multilingual VLLM Alignment via Translated Data
por: Han, Seung Hun, et al.
Publicado: (2026)
por: Han, Seung Hun, et al.
Publicado: (2026)
MiniGPT: Rebuilding GPT from First Principles
por: Joseph, Jibin
Publicado: (2026)
por: Joseph, Jibin
Publicado: (2026)
LlamaRec-LKG-RAG: A Single-Pass, Learnable Knowledge Graph-RAG Framework for LLM-Based Ranking
por: Azizi, Vahid, et al.
Publicado: (2025)
por: Azizi, Vahid, et al.
Publicado: (2025)
Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities
por: Xie, Zhifei, et al.
Publicado: (2024)
por: Xie, Zhifei, et al.
Publicado: (2024)
GPT as Psychologist? Preliminary Evaluations for GPT-4V on Visual Affective Computing
por: Lu, Hao, et al.
Publicado: (2024)
por: Lu, Hao, et al.
Publicado: (2024)
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
por: Chen, Junying, et al.
Publicado: (2025)
por: Chen, Junying, et al.
Publicado: (2025)
TractoGPT: A GPT architecture for White Matter Segmentation
por: Goel, Anoushkrit, et al.
Publicado: (2025)
por: Goel, Anoushkrit, et al.
Publicado: (2025)
OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
por: Chen, Zhihong, et al.
Publicado: (2025)
por: Chen, Zhihong, et al.
Publicado: (2025)
Will GPT-4 Run DOOM?
por: de Wynter, Adrian
Publicado: (2024)
por: de Wynter, Adrian
Publicado: (2024)
Zero-shot Building Age Classification from Facade Image Using GPT-4
por: Zeng, Zichao, et al.
Publicado: (2024)
por: Zeng, Zichao, et al.
Publicado: (2024)
A Systematic Evaluation of GPT-4V's Multimodal Capability for Medical Image Analysis
por: Li, Yingshu, et al.
Publicado: (2023)
por: Li, Yingshu, et al.
Publicado: (2023)
VTG-GPT: Tuning-Free Zero-Shot Video Temporal Grounding with GPT
por: Xu, Yifang, et al.
Publicado: (2024)
por: Xu, Yifang, et al.
Publicado: (2024)
An Evaluation of GPT-4V and Gemini in Online VQA
por: Liu, Mengchen, et al.
Publicado: (2023)
por: Liu, Mengchen, et al.
Publicado: (2023)
GPT-4V Takes the Wheel: Promises and Challenges for Pedestrian Behavior Prediction
por: Huang, Jia, et al.
Publicado: (2023)
por: Huang, Jia, et al.
Publicado: (2023)
MiniCPM-V: A GPT-4V Level MLLM on Your Phone
por: Yao, Yuan, et al.
Publicado: (2024)
por: Yao, Yuan, et al.
Publicado: (2024)
MiniMaxAD: A Lightweight Autoencoder for Feature-Rich Anomaly Detection
por: Wang, Fengjie, et al.
Publicado: (2024)
por: Wang, Fengjie, et al.
Publicado: (2024)
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
por: Zhang, Shaolei, et al.
Publicado: (2025)
por: Zhang, Shaolei, et al.
Publicado: (2025)
Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
por: Ye, Junyan, et al.
Publicado: (2025)
por: Ye, Junyan, et al.
Publicado: (2025)
IQAGPT: Image Quality Assessment with Vision-language and ChatGPT Models
por: Chen, Zhihao, et al.
Publicado: (2023)
por: Chen, Zhihao, et al.
Publicado: (2023)
Evaluation of GPT-4o and GPT-4o-mini's Vision Capabilities for Compositional Analysis from Dried Solution Drops
por: Dangi, Deven B., et al.
Publicado: (2024)
por: Dangi, Deven B., et al.
Publicado: (2024)
RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding
por: Xu, Linrui, et al.
Publicado: (2024)
por: Xu, Linrui, et al.
Publicado: (2024)
Demystifying the Potential of ChatGPT-4 Vision for Construction Progress Monitoring
por: Ersoz, Ahmet Bahaddin
Publicado: (2024)
por: Ersoz, Ahmet Bahaddin
Publicado: (2024)
ForgeryGPT: A Multimodal LLM for Interpretable Image Forgery Detection and Localization
por: Zhang, Fanrui, et al.
Publicado: (2024)
por: Zhang, Fanrui, et al.
Publicado: (2024)
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
por: Li, Yanwei, et al.
Publicado: (2024)
por: Li, Yanwei, et al.
Publicado: (2024)
MedGemma vs GPT-4: Open-Source and Proprietary Zero-shot Medical Disease Classification from Images
por: Prottasha, Md. Sazzadul Islam, et al.
Publicado: (2025)
por: Prottasha, Md. Sazzadul Islam, et al.
Publicado: (2025)
Evaluating ChatGPT's Performance in Classifying Pneumonia from Chest X-Ray Images
por: Prahallad, Pragna, et al.
Publicado: (2025)
por: Prahallad, Pragna, et al.
Publicado: (2025)
Is ChatGPT-5 Ready for Mammogram VQA?
por: Li, Qiang, et al.
Publicado: (2025)
por: Li, Qiang, et al.
Publicado: (2025)
Video-GPT via Next Clip Diffusion
por: Zhuang, Shaobin, et al.
Publicado: (2025)
por: Zhuang, Shaobin, et al.
Publicado: (2025)
G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o
por: Tong, Tony Cheng, et al.
Publicado: (2024)
por: Tong, Tony Cheng, et al.
Publicado: (2024)
CollagePrompt: A Benchmark for Budget-Friendly Visual Recognition with GPT-4V
por: Xu, Siyu, et al.
Publicado: (2024)
por: Xu, Siyu, et al.
Publicado: (2024)
3DAxisPrompt: Promoting the 3D Grounding and Reasoning in GPT-4o
por: Liu, Dingning, et al.
Publicado: (2025)
por: Liu, Dingning, et al.
Publicado: (2025)
Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
por: Lai, Xin, et al.
Publicado: (2025)
por: Lai, Xin, et al.
Publicado: (2025)
Preliminary Explorations with GPT-4o(mni) Native Image Generation
por: Cao, Pu, et al.
Publicado: (2025)
por: Cao, Pu, et al.
Publicado: (2025)
GPTDrawer: Enhancing Visual Synthesis through ChatGPT
por: Li, Kun, et al.
Publicado: (2024)
por: Li, Kun, et al.
Publicado: (2024)
FlattenGPT: Depth Compression for Transformer with Layer Flattening
por: Xu, Ruihan, et al.
Publicado: (2026)
por: Xu, Ruihan, et al.
Publicado: (2026)
Ejemplares similares
-
MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens
por: Zheng, Kaizhi, et al.
Publicado: (2023) -
MiniGPT-Pancreas: Multimodal Large Language Model for Pancreas Cancer Classification and Detection
por: Moglia, Andrea, et al.
Publicado: (2024) -
MiniGPT-Med: Large Language Model as a General Interface for Radiology Diagnosis
por: Alkhaldi, Asma, et al.
Publicado: (2024) -
MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors
por: Tang, Yuan, et al.
Publicado: (2024) -
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
por: Ataallah, Kirolos, et al.
Publicado: (2024)