MMIS: Multimodal Dataset for Interior Scene Visual Generation and Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Kassab, Hozaifa, Mahmoud, Ahmed, Bahaa, Mohamed, Mohamed, Ammar, Hamdi, Ali |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GCF: Graph Convolutional Networks for Facial Expression Recognition
by: Kassab, Hozaifa, et al.
Published: (2024)
by: Kassab, Hozaifa, et al.
Published: (2024)
RIRO: Reshaping Inputs, Refining Outputs Unlocking the Potential of Large Language Models in Data-Scarce Contexts
by: Hamdi, Ali, et al.
Published: (2024)
by: Hamdi, Ali, et al.
Published: (2024)
Advancing Automated Deception Detection: A Multimodal Approach to Feature Extraction and Analysis
by: Bahaa, Mohamed, et al.
Published: (2024)
by: Bahaa, Mohamed, et al.
Published: (2024)
Uncertainty-Guided Attention and Entropy-Weighted Loss for Precise Plant Seedling Segmentation
by: Ehab, Mohamed, et al.
Published: (2026)
by: Ehab, Mohamed, et al.
Published: (2026)
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
by: Abdelrahman, Eslam, et al.
Published: (2023)
by: Abdelrahman, Eslam, et al.
Published: (2023)
MMIS-Net for Retinal Fluid Segmentation and Detection
by: Ndipenocha, Nchongmaje, et al.
Published: (2025)
by: Ndipenocha, Nchongmaje, et al.
Published: (2025)
Data-Augmented Multimodal Feature Fusion for Multiclass Visual Recognition of Oral Cancer Lesions
by: Naoum, Joy, et al.
Published: (2025)
by: Naoum, Joy, et al.
Published: (2025)
DenTab: A Dataset for Table Recognition and Visual QA on Real-World Dental Estimates
by: Hamdi, Laziz, et al.
Published: (2026)
by: Hamdi, Laziz, et al.
Published: (2026)
Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
by: Ahmed, Mahmoud, et al.
Published: (2024)
by: Ahmed, Mahmoud, et al.
Published: (2024)
A Comprehensive Survey of Masked Faces: Recognition, Detection, and Unmasking
by: Mahmoud, Mohamed, et al.
Published: (2024)
by: Mahmoud, Mohamed, et al.
Published: (2024)
3DCoMPaT$^{++}$: An improved Large-scale 3D Vision Dataset for Compositional Recognition
by: Slim, Habib, et al.
Published: (2023)
by: Slim, Habib, et al.
Published: (2023)
ReceiptSense: Beyond Traditional OCR -- A Dataset for Receipt Understanding
by: Abdallah, Abdelrahman, et al.
Published: (2024)
by: Abdallah, Abdelrahman, et al.
Published: (2024)
Are Visual-Language Models Effective in Action Recognition? A Comparative Study
by: Ali, Mahmoud, et al.
Published: (2024)
by: Ali, Mahmoud, et al.
Published: (2024)
GNN-MoE: Context-Aware Patch Routing using GNNs for Parameter-Efficient Domain Generalization
by: Soliman, Mahmoud, et al.
Published: (2025)
by: Soliman, Mahmoud, et al.
Published: (2025)
QARI-OCR: High-Fidelity Arabic Text Recognition through Multimodal Large Language Model Adaptation
by: Wasfy, Ahmed, et al.
Published: (2025)
by: Wasfy, Ahmed, et al.
Published: (2025)
Dataset and Benchmark for Urdu Natural Scenes Text Detection, Recognition and Visual Question Answering
by: Maryam, Hiba, et al.
Published: (2024)
by: Maryam, Hiba, et al.
Published: (2024)
MedDChest: A Content-Aware Multimodal Foundational Vision Model for Thoracic Imaging
by: Soliman, Mahmoud, et al.
Published: (2025)
by: Soliman, Mahmoud, et al.
Published: (2025)
BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment
by: Mounis, Mohamed Darwish, et al.
Published: (2026)
by: Mounis, Mohamed Darwish, et al.
Published: (2026)
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
by: Kim, Younggun, et al.
Published: (2025)
by: Kim, Younggun, et al.
Published: (2025)
Neural Catalog: Scaling Species Recognition with Catalog of Life-Augmented Generation
by: Khan, Faizan Farooq, et al.
Published: (2025)
by: Khan, Faizan Farooq, et al.
Published: (2025)
Foundation Models as Class-Incremental Learners for Dermatological Image Classification
by: Elkhayat, Mohamed, et al.
Published: (2025)
by: Elkhayat, Mohamed, et al.
Published: (2025)
HieroGlyphTranslator: Automatic Recognition and Translation of Egyptian Hieroglyphs to English
by: Nasser, Ahmed, et al.
Published: (2025)
by: Nasser, Ahmed, et al.
Published: (2025)
A Data Efficiency Study of Synthetic Fog for Object Detection Using the Clear2Fog Pipeline
by: Mohamed, Mohamed Ahmed, et al.
Published: (2026)
by: Mohamed, Mohamed Ahmed, et al.
Published: (2026)
OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving
by: Zheng, Lianqing, et al.
Published: (2024)
by: Zheng, Lianqing, et al.
Published: (2024)
T2Vs Meet VLMs: A Scalable Multimodal Dataset for Visual Harmfulness Recognition
by: Yeh, Chen, et al.
Published: (2024)
by: Yeh, Chen, et al.
Published: (2024)
3DCoMPaT200: Language-Grounded Compositional Understanding of Parts and Materials of 3D Shapes
by: Ahmed, Mahmoud, et al.
Published: (2025)
by: Ahmed, Mahmoud, et al.
Published: (2025)
Real-Time On-the-Go Annotation Framework Using YOLO for Automated Dataset Generation
by: Salem, Mohamed Abdallah, et al.
Published: (2025)
by: Salem, Mohamed Abdallah, et al.
Published: (2025)
Exploring the Limits of Semantic Image Compression at Micro-bits per Pixel
by: Dotzel, Jordan, et al.
Published: (2024)
by: Dotzel, Jordan, et al.
Published: (2024)
Language-EXtended Indoor SLAM (LEXIS): A Versatile System for Real-time Visual Scene Understanding
by: Kassab, Christina, et al.
Published: (2023)
by: Kassab, Christina, et al.
Published: (2023)
The First Swahili Language Scene Text Detection and Recognition Dataset
by: Douamba, Fadila Wendigoundi, et al.
Published: (2024)
by: Douamba, Fadila Wendigoundi, et al.
Published: (2024)
Conditional Generative Models for High-Resolution Range Profiles: Capturing Geometry-Driven Trends in a Large-Scale Maritime Dataset
by: Brient, Edwyn, et al.
Published: (2026)
by: Brient, Edwyn, et al.
Published: (2026)
ORCA: Orchestrated Reasoning with Collaborative Agents for Document Visual Question Answering
by: Lassoued, Aymen, et al.
Published: (2026)
by: Lassoued, Aymen, et al.
Published: (2026)
CasaGPT: Cuboid Arrangement and Scene Assembly for Interior Design
by: Feng, Weitao, et al.
Published: (2025)
by: Feng, Weitao, et al.
Published: (2025)
IndicSTR12: A Dataset for Indic Scene Text Recognition
by: Lunia, Harsh, et al.
Published: (2024)
by: Lunia, Harsh, et al.
Published: (2024)
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
by: Imam, Mohamed Fazli, et al.
Published: (2025)
by: Imam, Mohamed Fazli, et al.
Published: (2025)
Generating Multimodal Driving Scenes via Next-Scene Prediction
by: Wu, Yanhao, et al.
Published: (2025)
by: Wu, Yanhao, et al.
Published: (2025)
ScenarioCLIP: Pretrained Transferable Visual Language Models and Action-Genome Dataset for Natural Scene Analysis
by: Sinha, Advik, et al.
Published: (2025)
by: Sinha, Advik, et al.
Published: (2025)
Two-Stage Quranic QA via Ensemble Retrieval and Instruction-Tuned Answer Extraction
by: Basem, Mohamed, et al.
Published: (2025)
by: Basem, Mohamed, et al.
Published: (2025)
SARD: A Large-Scale Synthetic Arabic OCR Dataset for Book-Style Text Recognition
by: Nacar, Omer, et al.
Published: (2025)
by: Nacar, Omer, et al.
Published: (2025)
FedPartWhole: Federated domain generalization via consistent part-whole hierarchies
by: Radwan, Ahmed, et al.
Published: (2024)
by: Radwan, Ahmed, et al.
Published: (2024)
Similar Items
-
GCF: Graph Convolutional Networks for Facial Expression Recognition
by: Kassab, Hozaifa, et al.
Published: (2024) -
RIRO: Reshaping Inputs, Refining Outputs Unlocking the Potential of Large Language Models in Data-Scarce Contexts
by: Hamdi, Ali, et al.
Published: (2024) -
Advancing Automated Deception Detection: A Multimodal Approach to Feature Extraction and Analysis
by: Bahaa, Mohamed, et al.
Published: (2024) -
Uncertainty-Guided Attention and Entropy-Weighted Loss for Precise Plant Seedling Segmentation
by: Ehab, Mohamed, et al.
Published: (2026) -
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
by: Abdelrahman, Eslam, et al.
Published: (2023)