DeepTaxon: An Interpretable Retrieval-Augmented Multimodal Framework for Unified Species Identification and Discovery
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jiawei, Lei, Ming, Yang, Yaning, Lin, Xinyan, Le, Yuquan, Ma, Qiwei, Xu, Zhiwei, Lv, Zheqi, Ang, Yuchen, Quan, Zhe, Chua, Tat-Seng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Seeing The Words: Evaluating AI-generated Biblical Art
by: Makimei, Hidde, et al.
Published: (2025)
by: Makimei, Hidde, et al.
Published: (2025)
Lens Distortion Encoding System Version 1.0
by: Fober, Jakub Maksymilian
Published: (2024)
by: Fober, Jakub Maksymilian
Published: (2024)
Evaluating gesture generation in a large-scale open challenge: The GENEA Challenge 2022
by: Kucherenko, Taras, et al.
Published: (2023)
by: Kucherenko, Taras, et al.
Published: (2023)
TAG-WM: Tamper-Aware Generative Image Watermarking via Diffusion Inversion Sensitivity
by: Chen, Yuzhuo, et al.
Published: (2025)
by: Chen, Yuzhuo, et al.
Published: (2025)
A Low-Latency 3D Live Remote Visualization System for Tourist Sites Integrating Dynamic and Pre-captured Static Point Clouds
by: Matsumoto, Takahiro, et al.
Published: (2025)
by: Matsumoto, Takahiro, et al.
Published: (2025)
Towards Robust and Controllable Text-to-Motion via Masked Autoregressive Diffusion
by: Zhang, Zongye, et al.
Published: (2025)
by: Zhang, Zongye, et al.
Published: (2025)
4Doodle: Two-handed Gestures for Immersive Sketching of Architectural Models
by: Fonseca, Fernando, et al.
Published: (2024)
by: Fonseca, Fernando, et al.
Published: (2024)
Annotating Satellite Images of Forests with Keywords from a Specialized Corpus in the Context of Change Detection
by: Neptune, Nathalie, et al.
Published: (2025)
by: Neptune, Nathalie, et al.
Published: (2025)
Racism in the Machine: Visualization Ethics in Digital Humanities Projects
by: Hepworth, K. J., et al.
Published: (2024)
by: Hepworth, K. J., et al.
Published: (2024)
ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use
by: Li, Kaixin, et al.
Published: (2025)
by: Li, Kaixin, et al.
Published: (2025)
MetaDigiHuman: Haptic Interfaces for Digital Humans in Metaverse
by: Jagatheesaperumal, Senthil Kumar, et al.
Published: (2024)
by: Jagatheesaperumal, Senthil Kumar, et al.
Published: (2024)
Graph-PiT: Enhancing Structural Coherence in Part-Based Image Synthesis via Graph Priors
by: Zhang, Junbin, et al.
Published: (2026)
by: Zhang, Junbin, et al.
Published: (2026)
SIMMER: Cross-Modal Food Image--Recipe Retrieval via MLLM-Based Embedding
by: Gomi, Keisuke, et al.
Published: (2026)
by: Gomi, Keisuke, et al.
Published: (2026)
HOME-KGQA: A Benchmark Dataset for Multimodal Knowledge Graph Question Answering on Household Daily Activities
by: Egami, Shusaku, et al.
Published: (2026)
by: Egami, Shusaku, et al.
Published: (2026)
Visual Style Prompt Learning Using Diffusion Models for Blind Face Restoration
by: Lu, Wanglong, et al.
Published: (2024)
by: Lu, Wanglong, et al.
Published: (2024)
FACEMUG: A Multimodal Generative and Fusion Framework for Local Facial Editing
by: Lu, Wanglong, et al.
Published: (2024)
by: Lu, Wanglong, et al.
Published: (2024)
Goal-Based Vision-Language Driving
by: Patapati, Santosh, et al.
Published: (2025)
by: Patapati, Santosh, et al.
Published: (2025)
Pinching Visuo-haptic Display: Investigating Cross-Modal Effects of Visual Textures on Electrostatic Cloth Tactile Sensations
by: Kitagishi, Takekazu, et al.
Published: (2025)
by: Kitagishi, Takekazu, et al.
Published: (2025)
EditIQ: Automated Cinematic Editing of Static Wide-Angle Videos via Dialogue Interpretation and Saliency Cues
by: Girmaji, Rohit, et al.
Published: (2025)
by: Girmaji, Rohit, et al.
Published: (2025)
3DGesPolicy: Phoneme-Aware Holistic Co-Speech Gesture Generation Based on Action Control
by: Sha, Xuanmeng, et al.
Published: (2026)
by: Sha, Xuanmeng, et al.
Published: (2026)
StyleMM: Stylized 3D Morphable Face Model via Text-Driven Aligned Image Translation
by: Lee, Seungmi, et al.
Published: (2025)
by: Lee, Seungmi, et al.
Published: (2025)
StyleID: A Perception-Aware Dataset and Metric for Stylization-Agnostic Facial Identity Recognition
by: Yun, Kwan, et al.
Published: (2026)
by: Yun, Kwan, et al.
Published: (2026)
Don't Splat your Gaussians: Volumetric Ray-Traced Primitives for Modeling and Rendering Scattering and Emissive Media
by: Condor, Jorge, et al.
Published: (2024)
by: Condor, Jorge, et al.
Published: (2024)
Improving Angular Speed Uniformity by Piecewise Radical Reparameterization
by: Hong, Hoon, et al.
Published: (2024)
by: Hong, Hoon, et al.
Published: (2024)
PhysHand: A Hand Simulation Model with Physiological Geometry, Physical Deformation, and Accurate Contact Handling
by: Sun, Mingyang, et al.
Published: (2024)
by: Sun, Mingyang, et al.
Published: (2024)
Under-Canopy Terrain Reconstruction in Dense Forests Using RGB Imaging and Neural 3D Reconstruction
by: Sheffer, Refael, et al.
Published: (2026)
by: Sheffer, Refael, et al.
Published: (2026)
Seeing in the Dark: A Teacher-Student Framework for Dark Video Action Recognition via Knowledge Distillation and Contrastive Learning
by: Dass, Sharana Dharshikgan Suresh, et al.
Published: (2025)
by: Dass, Sharana Dharshikgan Suresh, et al.
Published: (2025)
ActNetFormer: Transformer-ResNet Hybrid Method for Semi-Supervised Action Recognition in Videos
by: Dass, Sharana Dharshikgan Suresh, et al.
Published: (2024)
by: Dass, Sharana Dharshikgan Suresh, et al.
Published: (2024)
Interactive Visualization on Large High-Resolution Displays: A Survey
by: Belkacem, Ilyasse, et al.
Published: (2022)
by: Belkacem, Ilyasse, et al.
Published: (2022)
Finite Boolean Algebras for Solid Geometry using Julia's Sparse Arrays
by: Paoluzzi, Alberto, et al.
Published: (2019)
by: Paoluzzi, Alberto, et al.
Published: (2019)
GTA-HDR: A Large-Scale Synthetic Dataset for HDR Image Reconstruction
by: Barua, Hrishav Bakul, et al.
Published: (2024)
by: Barua, Hrishav Bakul, et al.
Published: (2024)
A Scalable Pipeline Combining Procedural 3D Graphics and Guided Diffusion for Photorealistic Synthetic Training Data Generation in White Button Mushroom Segmentation
by: Károly, Artúr I., et al.
Published: (2025)
by: Károly, Artúr I., et al.
Published: (2025)
Inpaint4Drag: Repurposing Inpainting Models for Drag-Based Image Editing via Bidirectional Warping
by: Lu, Jingyi, et al.
Published: (2025)
by: Lu, Jingyi, et al.
Published: (2025)
AnyMoLe: Any Character Motion In-betweening Leveraging Video Diffusion Models
by: Yun, Kwan, et al.
Published: (2025)
by: Yun, Kwan, et al.
Published: (2025)
OCMG-Net: Neural Oriented Normal Refinement for Unstructured Point Clouds
by: Wu, Yingrui, et al.
Published: (2024)
by: Wu, Yingrui, et al.
Published: (2024)
Sketch2Manga: Shaded Manga Screening from Sketch with Diffusion Models
by: Lin, Jian, et al.
Published: (2024)
by: Lin, Jian, et al.
Published: (2024)
CLM: Removing the GPU Memory Barrier for 3D Gaussian Splatting
by: Zhao, Hexu, et al.
Published: (2025)
by: Zhao, Hexu, et al.
Published: (2025)
Experimental Evaluation of Static Image Sub-Region-Based Search Models Using CLIP
by: Jäckl, Bastian, et al.
Published: (2025)
by: Jäckl, Bastian, et al.
Published: (2025)
Vectorized Region Based Brush Strokes for Artistic Rendering
by: Prudviraj, Jeripothula, et al.
Published: (2025)
by: Prudviraj, Jeripothula, et al.
Published: (2025)
Sketch & Paint: Stroke-by-Stroke Evolution of Visual Artworks
by: Prudviraj, Jeripothula, et al.
Published: (2025)
by: Prudviraj, Jeripothula, et al.
Published: (2025)
Similar Items
-
Seeing The Words: Evaluating AI-generated Biblical Art
by: Makimei, Hidde, et al.
Published: (2025) -
Lens Distortion Encoding System Version 1.0
by: Fober, Jakub Maksymilian
Published: (2024) -
Evaluating gesture generation in a large-scale open challenge: The GENEA Challenge 2022
by: Kucherenko, Taras, et al.
Published: (2023) -
TAG-WM: Tamper-Aware Generative Image Watermarking via Diffusion Inversion Sensitivity
by: Chen, Yuzhuo, et al.
Published: (2025) -
A Low-Latency 3D Live Remote Visualization System for Tourist Sites Integrating Dynamic and Pre-captured Static Point Clouds
by: Matsumoto, Takahiro, et al.
Published: (2025)