BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Shinoda, Risa, Shiohara, Kaede, Inoue, Nakamasa, Saito, Kuniaki, Santo, Hiroaki, Okura, Fumio |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AnimalCLAP: Taxonomy-Aware Language-Audio Pretraining for Species Recognition and Trait Inference
by: Shinoda, Risa, et al.
Published: (2026)
by: Shinoda, Risa, et al.
Published: (2026)
PetFace: A Large-Scale Dataset and Benchmark for Animal Identification
by: Shinoda, Risa, et al.
Published: (2024)
by: Shinoda, Risa, et al.
Published: (2024)
OpenAnimalTracks: A Dataset for Animal Track Recognition
by: Shinoda, Risa, et al.
Published: (2024)
by: Shinoda, Risa, et al.
Published: (2024)
GaussianPlant: Structure-aligned Gaussian Splatting for 3D Reconstruction of Plants
by: Yang, Yang, et al.
Published: (2025)
by: Yang, Yang, et al.
Published: (2025)
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2026)
by: Saito, Kuniaki, et al.
Published: (2026)
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2025)
by: Saito, Kuniaki, et al.
Published: (2025)
AgroBench: Vision-Language Model Benchmark in Agriculture
by: Shinoda, Risa, et al.
Published: (2025)
by: Shinoda, Risa, et al.
Published: (2025)
Interaction-via-Actions: Cattle Interaction Detection with Joint Learning of Action-Interaction Latent Space
by: Nakagawa, Ren, et al.
Published: (2025)
by: Nakagawa, Ren, et al.
Published: (2025)
Zero-shot Hierarchical Plant Segmentation via Foundation Segmentation Models and Text-to-image Attention
by: Xing, Junhao, et al.
Published: (2025)
by: Xing, Junhao, et al.
Published: (2025)
Pyramid Coder: Hierarchical Code Generator for Compositional Visual Question Answering
by: Shen, Ruoyue, et al.
Published: (2024)
by: Shen, Ruoyue, et al.
Published: (2024)
Unified Vector Floorplan Generation via Markup Representation
by: Shiohara, Kaede, et al.
Published: (2026)
by: Shiohara, Kaede, et al.
Published: (2026)
Face2Diffusion for Fast and Editable Face Personalization
by: Shiohara, Kaede, et al.
Published: (2024)
by: Shiohara, Kaede, et al.
Published: (2024)
AnimalClue: Recognizing Animals by their Traces
by: Shinoda, Risa, et al.
Published: (2025)
by: Shinoda, Risa, et al.
Published: (2025)
TreeFormer: Single-view Plant Skeleton Estimation via Tree-constrained Graph Generation
by: Liu, Xinpeng, et al.
Published: (2024)
by: Liu, Xinpeng, et al.
Published: (2024)
PlantPose: Universal Plant Skeleton Estimation via Tree-constrained Graph Generation
by: Liu, Xinpeng, et al.
Published: (2026)
by: Liu, Xinpeng, et al.
Published: (2026)
SBS Figures: Pre-training Figure QA from Stage-by-Stage Synthesized Images
by: Shinoda, Risa, et al.
Published: (2024)
by: Shinoda, Risa, et al.
Published: (2024)
ExposeAnyone: Personalized Audio-to-Expression Diffusion Models Are Robust Zero-Shot Face Forgery Detectors
by: Shiohara, Kaede, et al.
Published: (2026)
by: Shiohara, Kaede, et al.
Published: (2026)
ControlVP: Interactive Geometric Refinement of AI-Generated Images with Consistent Vanishing Points
by: Okumura, Ryota, et al.
Published: (2025)
by: Okumura, Ryota, et al.
Published: (2025)
NeuraLeaf: Neural Parametric Leaf Models with Shape and Deformation Disentanglement
by: Yang, Yang, et al.
Published: (2025)
by: Yang, Yang, et al.
Published: (2025)
DP-SfM: Dual-Pixel Structure-from-Motion without Scale Ambiguity
by: Makabe, Lilika, et al.
Published: (2026)
by: Makabe, Lilika, et al.
Published: (2026)
Spectral Sensitivity Estimation with an Uncalibrated Diffraction Grating
by: Makabe, Lilika, et al.
Published: (2025)
by: Makabe, Lilika, et al.
Published: (2025)
Near-light Photometric Stereo with Symmetric Lights
by: Makabe, Lilika, et al.
Published: (2026)
by: Makabe, Lilika, et al.
Published: (2026)
STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models
by: Ukai, Mahiro, et al.
Published: (2025)
by: Ukai, Mahiro, et al.
Published: (2025)
CAMOT: Camera Angle-aware Multi-Object Tracking
by: Limanta, Felix, et al.
Published: (2024)
by: Limanta, Felix, et al.
Published: (2024)
NeRSP: Neural 3D Reconstruction for Reflective Objects with Sparse Polarized Images
by: Han, Yufei, et al.
Published: (2024)
by: Han, Yufei, et al.
Published: (2024)
Gaussian Mesh Renderer for Lightweight Differentiable Rendering
by: Liu, Xinpeng, et al.
Published: (2026)
by: Liu, Xinpeng, et al.
Published: (2026)
PowerCLIP: Powerset Alignment for Contrastive Pre-Training
by: Kawamura, Masaki, et al.
Published: (2025)
by: Kawamura, Masaki, et al.
Published: (2025)
Robust Deepfake Detection for Electronic Know Your Customer Systems Using Registered Images
by: Amada, Takuma, et al.
Published: (2025)
by: Amada, Takuma, et al.
Published: (2025)
Training-Free Label Space Alignment for Universal Domain Adaptation
by: Lee, Dujin, et al.
Published: (2025)
by: Lee, Dujin, et al.
Published: (2025)
AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question Answering
by: Ukai, Mahiro, et al.
Published: (2024)
by: Ukai, Mahiro, et al.
Published: (2024)
Learning Contrastive Multimodal Fusion with Improved Modality Dropout for Disease Detection and Prediction
by: Gu, Yi, et al.
Published: (2025)
by: Gu, Yi, et al.
Published: (2025)
Towards Safer Mobile Agents: Scalable Generation and Evaluation of Diverse Scenarios for VLMs
by: Taniguchi, Takara, et al.
Published: (2026)
by: Taniguchi, Takara, et al.
Published: (2026)
Formula-Supervised Visual-Geometric Pre-training
by: Yamada, Ryosuke, et al.
Published: (2024)
by: Yamada, Ryosuke, et al.
Published: (2024)
EC-Bench: Enumeration and Counting Benchmark for Ultra-Long Videos
by: Tsuchiya, Fumihiko, et al.
Published: (2026)
by: Tsuchiya, Fumihiko, et al.
Published: (2026)
Unsupervised 3D Human Pose Estimation via Conditional Multi-view Ancestral Sampling
by: Goto, Ryohei, et al.
Published: (2026)
by: Goto, Ryohei, et al.
Published: (2026)
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
by: Ohi, Masanari, et al.
Published: (2024)
by: Ohi, Masanari, et al.
Published: (2024)
Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and Benchmark
by: Chen, Ziyang, et al.
Published: (2024)
by: Chen, Ziyang, et al.
Published: (2024)
Weak-to-Strong Compositional Learning from Generative Models for Language-based Object Detection
by: Park, Kwanyong, et al.
Published: (2024)
by: Park, Kwanyong, et al.
Published: (2024)
VISTA: A Visual and Textual Attention Dataset for Interpreting Multimodal Models
by: Harshit, et al.
Published: (2024)
by: Harshit, et al.
Published: (2024)
HoGS: Unified Near and Far Object Reconstruction via Homogeneous Gaussian Splatting
by: Liu, Xinpeng, et al.
Published: (2025)
by: Liu, Xinpeng, et al.
Published: (2025)
Similar Items
-
AnimalCLAP: Taxonomy-Aware Language-Audio Pretraining for Species Recognition and Trait Inference
by: Shinoda, Risa, et al.
Published: (2026) -
PetFace: A Large-Scale Dataset and Benchmark for Animal Identification
by: Shinoda, Risa, et al.
Published: (2024) -
OpenAnimalTracks: A Dataset for Animal Track Recognition
by: Shinoda, Risa, et al.
Published: (2024) -
GaussianPlant: Structure-aligned Gaussian Splatting for 3D Reconstruction of Plants
by: Yang, Yang, et al.
Published: (2025) -
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2026)