Generative AI for Vision: A Comprehensive Study of Frameworks and Applications
Fuente:
arXiv
Saved in:
| Main Author: | Bousetouane, Fouad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generative AI for Industrial Contour Detection: A Language-Guided Vision System
by: Gong, Liang, et al.
Published: (2025)
by: Gong, Liang, et al.
Published: (2025)
Evaluating the Efficacy of Prompt-Engineered Large Multimodal Models Versus Fine-Tuned Vision Transformers in Image-Based Security Applications
by: Trad, Fouad, et al.
Published: (2024)
by: Trad, Fouad, et al.
Published: (2024)
HVT: A Comprehensive Vision Framework for Learning in Non-Euclidean Space
by: Fein-Ashley, Jacob, et al.
Published: (2024)
by: Fein-Ashley, Jacob, et al.
Published: (2024)
Vision Mamba in Remote Sensing: A Comprehensive Survey of Techniques, Applications and Outlook
by: Bao, Muyi, et al.
Published: (2025)
by: Bao, Muyi, et al.
Published: (2025)
EasyRobust: A Comprehensive and Easy-to-use Toolkit for Robust and Generalized Vision
by: Mao, Xiaofeng, et al.
Published: (2025)
by: Mao, Xiaofeng, et al.
Published: (2025)
GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI
by: Li, Tianbin, et al.
Published: (2024)
by: Li, Tianbin, et al.
Published: (2024)
PDC-ViT : Source Camera Identification using Pixel Difference Convolution and Vision Transformer
by: Elharrouss, Omar, et al.
Published: (2025)
by: Elharrouss, Omar, et al.
Published: (2025)
Domain Adaptable Fine-Tune Distillation Framework For Advancing Farm Surveillance
by: Imam, Raza, et al.
Published: (2024)
by: Imam, Raza, et al.
Published: (2024)
How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study
by: Liu, Che, et al.
Published: (2025)
by: Liu, Che, et al.
Published: (2025)
CTForensics: A Comprehensive Dataset and Method for AI-Generated CT Image Detection
by: Li, Yiheng, et al.
Published: (2026)
by: Li, Yiheng, et al.
Published: (2026)
Explainable AI-Driven Detection of Human Monkeypox Using Deep Learning and Vision Transformers: A Comprehensive Analysis
by: Hossain, Md. Zahid, et al.
Published: (2025)
by: Hossain, Md. Zahid, et al.
Published: (2025)
OpenTAD: A Unified Framework and Comprehensive Study of Temporal Action Detection
by: Liu, Shuming, et al.
Published: (2025)
by: Liu, Shuming, et al.
Published: (2025)
MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
by: Hou, Ruibing, et al.
Published: (2025)
by: Hou, Ruibing, et al.
Published: (2025)
VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework
by: Kelly, Chris, et al.
Published: (2024)
by: Kelly, Chris, et al.
Published: (2024)
A Comprehensive Review of Knowledge Distillation in Computer Vision
by: Habib, Gousia, et al.
Published: (2024)
by: Habib, Gousia, et al.
Published: (2024)
GANji: A Framework for Introductory AI Image Generation
by: Hamel, Chandon, et al.
Published: (2025)
by: Hamel, Chandon, et al.
Published: (2025)
A Non-Invasive 3D Gait Analysis Framework for Quantifying Psychomotor Retardation in Major Depressive Disorder
by: Boutaleb, Fouad, et al.
Published: (2026)
by: Boutaleb, Fouad, et al.
Published: (2026)
A Comprehensive Evaluation Framework for the Study of the Effects of Facial Filters on Face Recognition Accuracy
by: Ozturk, Kagan, et al.
Published: (2025)
by: Ozturk, Kagan, et al.
Published: (2025)
Generative AI in Vision: A Survey on Models, Metrics and Applications
by: Raut, Gaurav, et al.
Published: (2024)
by: Raut, Gaurav, et al.
Published: (2024)
A Comprehensive Survey on Diffusion Models and Their Applications
by: Ahsan, Md Manjurul, et al.
Published: (2024)
by: Ahsan, Md Manjurul, et al.
Published: (2024)
RPNR: Robust-Perception Neural Reshading
by: Afiouni, Fouad, et al.
Published: (2024)
by: Afiouni, Fouad, et al.
Published: (2024)
Edge-Enhanced Vision Transformer Framework for Accurate AI-Generated Image Detection
by: Das, Dabbrata, et al.
Published: (2025)
by: Das, Dabbrata, et al.
Published: (2025)
UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding
by: Zhang, Da, et al.
Published: (2025)
by: Zhang, Da, et al.
Published: (2025)
ODExAI: A Comprehensive Object Detection Explainable AI Evaluation
by: Nguyen, Loc Phuc Truong, et al.
Published: (2025)
by: Nguyen, Loc Phuc Truong, et al.
Published: (2025)
Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation
by: Ling, Lu, et al.
Published: (2025)
by: Ling, Lu, et al.
Published: (2025)
OneVision: An End-to-End Generative Framework for Multi-view E-commerce Vision Search
by: Zheng, Zexin, et al.
Published: (2025)
by: Zheng, Zexin, et al.
Published: (2025)
MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities
by: Wu, Bizhu, et al.
Published: (2025)
by: Wu, Bizhu, et al.
Published: (2025)
M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation
by: Luo, Mingshuang, et al.
Published: (2024)
by: Luo, Mingshuang, et al.
Published: (2024)
Generative Physical AI in Vision: A Survey
by: Liu, Daochang, et al.
Published: (2025)
by: Liu, Daochang, et al.
Published: (2025)
PAI-Bench: A Comprehensive Benchmark For Physical AI
by: Zhou, Fengzhe, et al.
Published: (2025)
by: Zhou, Fengzhe, et al.
Published: (2025)
A Comprehensive Survey on World Models for Embodied AI
by: Li, Xinqing, et al.
Published: (2025)
by: Li, Xinqing, et al.
Published: (2025)
VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward Models
by: Ruan, Jiacheng, et al.
Published: (2025)
by: Ruan, Jiacheng, et al.
Published: (2025)
Deep Learning for Event-based Vision: A Comprehensive Survey and Benchmarks
by: Zheng, Xu, et al.
Published: (2023)
by: Zheng, Xu, et al.
Published: (2023)
Towards Application-Specific Evaluation of Vision Models: Case Studies in Ecology and Biology
by: Chan, Alex Hoi Hang, et al.
Published: (2025)
by: Chan, Alex Hoi Hang, et al.
Published: (2025)
Evaluating Attribute Comprehension in Large Vision-Language Models
by: Zhang, Haiwen, et al.
Published: (2024)
by: Zhang, Haiwen, et al.
Published: (2024)
Mamba in Vision: A Comprehensive Survey of Techniques and Applications
by: Rahman, Md Maklachur, et al.
Published: (2024)
by: Rahman, Md Maklachur, et al.
Published: (2024)
Adapter-X: A Novel General Parameter-Efficient Fine-Tuning Framework for Vision
by: Li, Minglei, et al.
Published: (2024)
by: Li, Minglei, et al.
Published: (2024)
Depth-Aware Rover: A Study of Edge AI and Monocular Vision for Real-World Implementation
by: Relia, Lomash, et al.
Published: (2026)
by: Relia, Lomash, et al.
Published: (2026)
INSIGHT: An Interpretable Neural Vision-Language Framework for Reasoning of Generative Artifacts
by: Bagaria, Anshul
Published: (2025)
by: Bagaria, Anshul
Published: (2025)
InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models
by: Deng, Nianchen, et al.
Published: (2025)
by: Deng, Nianchen, et al.
Published: (2025)
Similar Items
-
Generative AI for Industrial Contour Detection: A Language-Guided Vision System
by: Gong, Liang, et al.
Published: (2025) -
Evaluating the Efficacy of Prompt-Engineered Large Multimodal Models Versus Fine-Tuned Vision Transformers in Image-Based Security Applications
by: Trad, Fouad, et al.
Published: (2024) -
HVT: A Comprehensive Vision Framework for Learning in Non-Euclidean Space
by: Fein-Ashley, Jacob, et al.
Published: (2024) -
Vision Mamba in Remote Sensing: A Comprehensive Survey of Techniques, Applications and Outlook
by: Bao, Muyi, et al.
Published: (2025) -
EasyRobust: A Comprehensive and Easy-to-use Toolkit for Robust and Generalized Vision
by: Mao, Xiaofeng, et al.
Published: (2025)