Generative Artificial Intelligence: A Systematic Review and Applications
Fuente:
arXiv
Salvato in:
| Autori principali: | Sengar, Sandeep Singh, Hasan, Affan Bin, Kumar, Sanjay, Carroll, Fiona |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RepVGG-GELAN: Enhanced GELAN with VGG-STYLE ConvNets for Brain Tumour Detection
di: Balakrishnan, Thennarasi, et al.
Pubblicazione: (2024)
di: Balakrishnan, Thennarasi, et al.
Pubblicazione: (2024)
A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future
di: Sun, Shilin, et al.
Pubblicazione: (2024)
di: Sun, Shilin, et al.
Pubblicazione: (2024)
Medical Reasoning in the Era of LLMs: A Systematic Review of Enhancement Techniques and Applications
di: Wang, Wenxuan, et al.
Pubblicazione: (2025)
di: Wang, Wenxuan, et al.
Pubblicazione: (2025)
A Multimodal, Multitask System for Generating E Commerce Text Listings from Images
di: Singh, Nayan Kumar
Pubblicazione: (2025)
di: Singh, Nayan Kumar
Pubblicazione: (2025)
Foundations of Multisensory Artificial Intelligence
di: Liang, Paul Pu
Pubblicazione: (2024)
di: Liang, Paul Pu
Pubblicazione: (2024)
Toward More Reliable Artificial Intelligence: Reducing Hallucinations in Vision-Language Models
di: Sanogo, Kassoum, et al.
Pubblicazione: (2025)
di: Sanogo, Kassoum, et al.
Pubblicazione: (2025)
Fast & Efficient Normalizing Flows and Applications of Image Generative Models
di: Nagar, Sandeep
Pubblicazione: (2025)
di: Nagar, Sandeep
Pubblicazione: (2025)
VigilEye -- Artificial Intelligence-based Real-time Driver Drowsiness Detection
di: Sengar, Sandeep Singh, et al.
Pubblicazione: (2024)
di: Sengar, Sandeep Singh, et al.
Pubblicazione: (2024)
Zero-Shot Refinement of Buildings' Segmentation Models using SAM
di: Mayladan, Ali, et al.
Pubblicazione: (2023)
di: Mayladan, Ali, et al.
Pubblicazione: (2023)
Recent Trends in Artificial Intelligence Technology: A Scoping Review
di: Niskanen, Teemu, et al.
Pubblicazione: (2023)
di: Niskanen, Teemu, et al.
Pubblicazione: (2023)
Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review
di: Lisondra, Matthew, et al.
Pubblicazione: (2025)
di: Lisondra, Matthew, et al.
Pubblicazione: (2025)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
di: Shang, Chuyi, et al.
Pubblicazione: (2024)
di: Shang, Chuyi, et al.
Pubblicazione: (2024)
A Vision for Multisensory Intelligence: Sensing, Science, and Synergy
di: Liang, Paul Pu
Pubblicazione: (2026)
di: Liang, Paul Pu
Pubblicazione: (2026)
BetterNet: An Efficient CNN Architecture with Residual Learning and Attention for Precision Polyp Segmentation
di: Singh, Owen, et al.
Pubblicazione: (2024)
di: Singh, Owen, et al.
Pubblicazione: (2024)
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
di: Song, Lin, et al.
Pubblicazione: (2026)
di: Song, Lin, et al.
Pubblicazione: (2026)
Advancing Autonomous Vehicle Intelligence: Deep Learning and Multimodal LLM for Traffic Sign Recognition and Robust Lane Detection
di: Sah, Chandan Kumar, et al.
Pubblicazione: (2025)
di: Sah, Chandan Kumar, et al.
Pubblicazione: (2025)
Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
di: Li, Yulin, et al.
Pubblicazione: (2025)
di: Li, Yulin, et al.
Pubblicazione: (2025)
ViPRA: Video Prediction for Robot Actions
di: Routray, Sandeep, et al.
Pubblicazione: (2025)
di: Routray, Sandeep, et al.
Pubblicazione: (2025)
Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations
di: Srivastava, Archita, et al.
Pubblicazione: (2025)
di: Srivastava, Archita, et al.
Pubblicazione: (2025)
Mamba in Vision: A Comprehensive Survey of Techniques and Applications
di: Rahman, Md Maklachur, et al.
Pubblicazione: (2024)
di: Rahman, Md Maklachur, et al.
Pubblicazione: (2024)
TaxaBind: A Unified Embedding Space for Ecological Applications
di: Sastry, Srikumar, et al.
Pubblicazione: (2024)
di: Sastry, Srikumar, et al.
Pubblicazione: (2024)
Unaligning Everything: Or Aligning Any Text to Any Image in Multimodal Models
di: Salman, Shaeke, et al.
Pubblicazione: (2024)
di: Salman, Shaeke, et al.
Pubblicazione: (2024)
Generative AI for Character Animation: A Comprehensive Survey of Techniques, Applications, and Future Directions
di: Abootorabi, Mohammad Mahdi, et al.
Pubblicazione: (2025)
di: Abootorabi, Mohammad Mahdi, et al.
Pubblicazione: (2025)
Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities
di: Yang, Enneng, et al.
Pubblicazione: (2024)
di: Yang, Enneng, et al.
Pubblicazione: (2024)
Improve Academic Query Resolution through BERT-based Question Extraction from Images
di: Kamal, Nidhi, et al.
Pubblicazione: (2024)
di: Kamal, Nidhi, et al.
Pubblicazione: (2024)
POSESTITCH-SLT: Linguistically Inspired Pose-Stitching for End-to-End Sign Language Translation
di: Joshi, Abhinav, et al.
Pubblicazione: (2025)
di: Joshi, Abhinav, et al.
Pubblicazione: (2025)
Generative AI in Vision: A Survey on Models, Metrics and Applications
di: Raut, Gaurav, et al.
Pubblicazione: (2024)
di: Raut, Gaurav, et al.
Pubblicazione: (2024)
MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning
di: Kumar, Somnath, et al.
Pubblicazione: (2024)
di: Kumar, Somnath, et al.
Pubblicazione: (2024)
Generative Emotion Cause Explanation in Multimodal Conversations
di: Wang, Lin, et al.
Pubblicazione: (2024)
di: Wang, Lin, et al.
Pubblicazione: (2024)
Translation-Enhanced Multilingual Text-to-Image Generation
di: Li, Yaoyiran, et al.
Pubblicazione: (2023)
di: Li, Yaoyiran, et al.
Pubblicazione: (2023)
Interleaving Reasoning for Better Text-to-Image Generation
di: Huang, Wenxuan, et al.
Pubblicazione: (2025)
di: Huang, Wenxuan, et al.
Pubblicazione: (2025)
Impact of Layer Norm on Memorization and Generalization in Transformers
di: Singhal, Rishi, et al.
Pubblicazione: (2025)
di: Singhal, Rishi, et al.
Pubblicazione: (2025)
SIDE: Sparse Information Disentanglement for Explainable Artificial Intelligence
di: Dubovik, Viktar, et al.
Pubblicazione: (2025)
di: Dubovik, Viktar, et al.
Pubblicazione: (2025)
LMFusion: Adapting Pretrained Language Models for Multimodal Generation
di: Shi, Weijia, et al.
Pubblicazione: (2024)
di: Shi, Weijia, et al.
Pubblicazione: (2024)
Efficient Pre-training for Localized Instruction Generation of Videos
di: Batra, Anil, et al.
Pubblicazione: (2023)
di: Batra, Anil, et al.
Pubblicazione: (2023)
Rethinking Training Dynamics in Scale-wise Autoregressive Generation
di: Zhou, Gengze, et al.
Pubblicazione: (2025)
di: Zhou, Gengze, et al.
Pubblicazione: (2025)
Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging
di: Cai, Zhenyang, et al.
Pubblicazione: (2024)
di: Cai, Zhenyang, et al.
Pubblicazione: (2024)
Euclid's Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks
di: Lian, Shijie, et al.
Pubblicazione: (2025)
di: Lian, Shijie, et al.
Pubblicazione: (2025)
Seeing Justice Clearly: Handwritten Legal Document Translation with OCR and Vision-Language Models
di: Nigam, Shubham Kumar, et al.
Pubblicazione: (2025)
di: Nigam, Shubham Kumar, et al.
Pubblicazione: (2025)
LLMs can see and hear without any training
di: Ashutosh, Kumar, et al.
Pubblicazione: (2025)
di: Ashutosh, Kumar, et al.
Pubblicazione: (2025)
Documenti analoghi
-
RepVGG-GELAN: Enhanced GELAN with VGG-STYLE ConvNets for Brain Tumour Detection
di: Balakrishnan, Thennarasi, et al.
Pubblicazione: (2024) -
A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future
di: Sun, Shilin, et al.
Pubblicazione: (2024) -
Medical Reasoning in the Era of LLMs: A Systematic Review of Enhancement Techniques and Applications
di: Wang, Wenxuan, et al.
Pubblicazione: (2025) -
A Multimodal, Multitask System for Generating E Commerce Text Listings from Images
di: Singh, Nayan Kumar
Pubblicazione: (2025) -
Foundations of Multisensory Artificial Intelligence
di: Liang, Paul Pu
Pubblicazione: (2024)