Neural Catalog: Scaling Species Recognition with Catalog of Life-Augmented Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Khan, Faizan Farooq, Chen, Jun, Mohamed, Youssef, Feng, Chun-Mei, Elhoseiny, Mohamed |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AI Art Neural Constellation: Revealing the Collective and Contrastive State of AI-Generated and Human Art
por: Khan, Faizan Farooq, et al.
Publicado: (2024)
por: Khan, Faizan Farooq, et al.
Publicado: (2024)
Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
Step-by-step Layered Design Generation
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
por: Shen, Xiaoqian, et al.
Publicado: (2025)
por: Shen, Xiaoqian, et al.
Publicado: (2025)
How Well Can Vision Language Models See Image Details?
por: Gou, Chenhui, et al.
Publicado: (2024)
por: Gou, Chenhui, et al.
Publicado: (2024)
Document Haystacks: Vision-Language Reasoning Over Piles of 1000+ Documents
por: Chen, Jun, et al.
Publicado: (2024)
por: Chen, Jun, et al.
Publicado: (2024)
WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article Generation
por: Yang, Zhongyu, et al.
Publicado: (2025)
por: Yang, Zhongyu, et al.
Publicado: (2025)
StoryGPT-V: Large Language Models as Consistent Story Visualizers
por: Shen, Xiaoqian, et al.
Publicado: (2023)
por: Shen, Xiaoqian, et al.
Publicado: (2023)
Efficient Self-supervised Vision Pretraining with Local Masked Reconstruction
por: Chen, Jun, et al.
Publicado: (2022)
por: Chen, Jun, et al.
Publicado: (2022)
ReefNet: A Large-Scale Dataset and Benchmark for Fine-Grained Coral Reef Recognition
por: Felemban, Abdulwahab, et al.
Publicado: (2025)
por: Felemban, Abdulwahab, et al.
Publicado: (2025)
FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
Overcoming Generic Knowledge Loss with Selective Parameter Update
por: Zhang, Wenxuan, et al.
Publicado: (2023)
por: Zhang, Wenxuan, et al.
Publicado: (2023)
VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
por: Li, Xiang, et al.
Publicado: (2024)
por: Li, Xiang, et al.
Publicado: (2024)
Continual Learning on a Diet: Learning from Sparsely Labeled Streams Under Constrained Computation
por: Zhang, Wenxuan, et al.
Publicado: (2024)
por: Zhang, Wenxuan, et al.
Publicado: (2024)
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
por: Abdelrahman, Eslam, et al.
Publicado: (2023)
por: Abdelrahman, Eslam, et al.
Publicado: (2023)
iMotion-LLM: Instruction-Conditioned Trajectory Generation
por: Felemban, Abdulwahab, et al.
Publicado: (2024)
por: Felemban, Abdulwahab, et al.
Publicado: (2024)
CatalogStitch: Dimension-Aware and Occlusion-Preserving Object Compositing for Catalog Image Generation
por: Jain, Sanyam, et al.
Publicado: (2026)
por: Jain, Sanyam, et al.
Publicado: (2026)
Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
por: Ahmed, Mahmoud, et al.
Publicado: (2024)
por: Ahmed, Mahmoud, et al.
Publicado: (2024)
3DCoMPaT200: Language-Grounded Compositional Understanding of Parts and Materials of 3D Shapes
por: Ahmed, Mahmoud, et al.
Publicado: (2025)
por: Ahmed, Mahmoud, et al.
Publicado: (2025)
ToddlerDiffusion: Interactive Structured Image Generation with Cascaded Schrödinger Bridge
por: Abdelrahman, Eslam, et al.
Publicado: (2023)
por: Abdelrahman, Eslam, et al.
Publicado: (2023)
From Statics to Dynamics: Physics-Aware Image Editing with Latent Transition Priors
por: Zhao, Liangbing, et al.
Publicado: (2026)
por: Zhao, Liangbing, et al.
Publicado: (2026)
MMIS: Multimodal Dataset for Interior Scene Visual Generation and Recognition
por: Kassab, Hozaifa, et al.
Publicado: (2024)
por: Kassab, Hozaifa, et al.
Publicado: (2024)
Time Blindness: Why Video-Language Models Can't See What Humans Can?
por: Upadhyay, Ujjwal, et al.
Publicado: (2025)
por: Upadhyay, Ujjwal, et al.
Publicado: (2025)
Domain-Aware Continual Zero-Shot Learning
por: Yi, Kai, et al.
Publicado: (2021)
por: Yi, Kai, et al.
Publicado: (2021)
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
por: Shen, Xiaoqian, et al.
Publicado: (2025)
por: Shen, Xiaoqian, et al.
Publicado: (2025)
MiniGPT-Med: Large Language Model as a General Interface for Radiology Diagnosis
por: Alkhaldi, Asma, et al.
Publicado: (2024)
por: Alkhaldi, Asma, et al.
Publicado: (2024)
From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
por: Zhuo, Le, et al.
Publicado: (2025)
por: Zhuo, Le, et al.
Publicado: (2025)
3DCoMPaT$^{++}$: An improved Large-scale 3D Vision Dataset for Compositional Recognition
por: Slim, Habib, et al.
Publicado: (2023)
por: Slim, Habib, et al.
Publicado: (2023)
Mining Field Data for Tree Species Recognition at Scale
por: Gominski, Dimitri, et al.
Publicado: (2024)
por: Gominski, Dimitri, et al.
Publicado: (2024)
Fusion of Single and Integral Multispectral Aerial Images
por: Youssef, Mohamed, et al.
Publicado: (2023)
por: Youssef, Mohamed, et al.
Publicado: (2023)
Text to Image for Multi-Label Image Recognition with Joint Prompt-Adapter Learning
por: Feng, Chun-Mei, et al.
Publicado: (2025)
por: Feng, Chun-Mei, et al.
Publicado: (2025)
InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows
por: Ataallah, Kirolos, et al.
Publicado: (2024)
por: Ataallah, Kirolos, et al.
Publicado: (2024)
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
por: Ataallah, Kirolos, et al.
Publicado: (2024)
por: Ataallah, Kirolos, et al.
Publicado: (2024)
Stroke Locus Net: Occluded Vessel Localization from MRI Modalities
por: Hamad, Mohamed, et al.
Publicado: (2025)
por: Hamad, Mohamed, et al.
Publicado: (2025)
SPDGAN: A Generative Adversarial Network based on SPD Manifold Learning for Automatic Image Colorization
por: Mourchid, Youssef, et al.
Publicado: (2023)
por: Mourchid, Youssef, et al.
Publicado: (2023)
HAND: Hierarchical Attention Network for Multi-Scale Handwritten Document Recognition and Layout Analysis
por: Hamdan, Mohammed, et al.
Publicado: (2024)
por: Hamdan, Mohammed, et al.
Publicado: (2024)
Visual-Augmented Dynamic Semantic Prototype for Generative Zero-Shot Learning
por: Hou, Wenjin, et al.
Publicado: (2024)
por: Hou, Wenjin, et al.
Publicado: (2024)
Scaling Human Activity Recognition: A Comparative Evaluation of Synthetic Data Generation and Augmentation Techniques
por: Leng, Zikang, et al.
Publicado: (2025)
por: Leng, Zikang, et al.
Publicado: (2025)
Diffusion-Enhanced Test-time Adaptation with Text and Image Augmentation
por: Feng, Chun-Mei, et al.
Publicado: (2024)
por: Feng, Chun-Mei, et al.
Publicado: (2024)
Ejemplares similares
-
AI Art Neural Constellation: Revealing the Collective and Contrastive State of AI-Generated and Human Art
por: Khan, Faizan Farooq, et al.
Publicado: (2024) -
Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders
por: Khan, Faizan Farooq, et al.
Publicado: (2025) -
Step-by-step Layered Design Generation
por: Khan, Faizan Farooq, et al.
Publicado: (2025) -
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
por: Shen, Xiaoqian, et al.
Publicado: (2025) -
How Well Can Vision Language Models See Image Details?
por: Gou, Chenhui, et al.
Publicado: (2024)