Saved in:
| Main Authors: | Takizawa, Ryo, Kodera, Satoshi, Kabayama, Tempei, Matsuoka, Ryo, Ando, Yuta, Nakamura, Yuto, Settai, Haruki, Takeda, Norihiko |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2504.18800 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CAG-VLM: Fine-Tuning of a Large-Scale Model to Recognize Angiographic Images for Next-Generation Diagnostic Systems
by: Nakamura, Yuto, et al.
Published: (2025)
by: Nakamura, Yuto, et al.
Published: (2025)
MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer
by: Yamane, Taiga, et al.
Published: (2025)
by: Yamane, Taiga, et al.
Published: (2025)
MVTrajecter: Multi-View Pedestrian Tracking with Trajectory Motion Cost and Trajectory Appearance Cost
by: Yamane, Taiga, et al.
Published: (2025)
by: Yamane, Taiga, et al.
Published: (2025)
MoireDB: Formula-generated Interference-fringe Image Dataset
by: Matsuo, Yuto, et al.
Published: (2025)
by: Matsuo, Yuto, et al.
Published: (2025)
EchoPrime: A Multi-Video View-Informed Vision-Language Model for Comprehensive Echocardiography Interpretation
by: Vukadinovic, Milos, et al.
Published: (2024)
by: Vukadinovic, Milos, et al.
Published: (2024)
SANER: Annotation-free Societal Attribute Neutralizer for Debiasing CLIP
by: Hirota, Yusuke, et al.
Published: (2024)
by: Hirota, Yusuke, et al.
Published: (2024)
Exploration-assisted Bottleneck Transition Toward Robust and Data-efficient Deformable Object Manipulation
by: Onishi, Yujiro, et al.
Published: (2026)
by: Onishi, Yujiro, et al.
Published: (2026)
Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation
by: Baba, Kaito, et al.
Published: (2026)
by: Baba, Kaito, et al.
Published: (2026)
VIOLA: Towards Video In-Context Learning with Minimal Annotations
by: Fujii, Ryo, et al.
Published: (2026)
by: Fujii, Ryo, et al.
Published: (2026)
Weakly Semi-supervised Tool Detection in Minimally Invasive Surgery Videos
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
Enhancing Reusability of Learned Skills for Robot Manipulation via Gaze Information and Motion Bottlenecks
by: Takizawa, Ryo, et al.
Published: (2025)
by: Takizawa, Ryo, et al.
Published: (2025)
M-PhyGs: Multi-Material Object Dynamics from Video
by: Wada, Norika, et al.
Published: (2025)
by: Wada, Norika, et al.
Published: (2025)
RealTraj: Towards Real-World Pedestrian Trajectory Forecasting
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
Automated Interpretable 2D Video Extraction from 3D Echocardiography
by: Vukadinovic, Milos, et al.
Published: (2025)
by: Vukadinovic, Milos, et al.
Published: (2025)
Prompt-driven Universal Model for View-Agnostic Echocardiography Analysis
by: Kim, Sekeun, et al.
Published: (2024)
by: Kim, Sekeun, et al.
Published: (2024)
From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment
by: Hirota, Yusuke, et al.
Published: (2024)
by: Hirota, Yusuke, et al.
Published: (2024)
Attention-Guided Integration of CLIP and SAM for Precise Object Masking in Robotic Manipulation
by: Muttaqien, Muhammad A., et al.
Published: (2025)
by: Muttaqien, Muhammad A., et al.
Published: (2025)
MV-CLIP: Multi-View CLIP for Zero-shot 3D Shape Recognition
by: Song, Dan, et al.
Published: (2023)
by: Song, Dan, et al.
Published: (2023)
EMAG: Ego-motion Aware and Generalizable 2D Hand Forecasting from Egocentric Videos
by: Hatano, Masashi, et al.
Published: (2024)
by: Hatano, Masashi, et al.
Published: (2024)
Whom to Respond To? A Transformer-Based Model for Multi-Party Social Robot Interaction
by: Zhu, He, et al.
Published: (2025)
by: Zhu, He, et al.
Published: (2025)
Multimodal Cross-Domain Few-Shot Learning for Egocentric Action Recognition
by: Hatano, Masashi, et al.
Published: (2024)
by: Hatano, Masashi, et al.
Published: (2024)
Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation
by: Ishikawa, Reina, et al.
Published: (2025)
by: Ishikawa, Reina, et al.
Published: (2025)
Piggyback Camera: Easy-to-Deploy Visual Surveillance by Mobile Sensing on Commercial Robot Vacuums
by: Yonetani, Ryo
Published: (2025)
by: Yonetani, Ryo
Published: (2025)
MSMVD: Exploiting Multi-scale Image Features via Multi-scale BEV Features for Multi-view Pedestrian Detection
by: Yamane, Taiga, et al.
Published: (2025)
by: Yamane, Taiga, et al.
Published: (2025)
Primitive Geometry Segment Pre-training for 3D Medical Image Segmentation
by: Tadokoro, Ryu, et al.
Published: (2024)
by: Tadokoro, Ryu, et al.
Published: (2024)
Computer-Aided Multi-Stroke Character Simplification by Stroke Removal
by: Ishiyama, Ryo, et al.
Published: (2025)
by: Ishiyama, Ryo, et al.
Published: (2025)
CLIP-Guided Multi-Task Regression for Multi-View Plant Phenotyping
by: Warmers, Simon, et al.
Published: (2026)
by: Warmers, Simon, et al.
Published: (2026)
LiDAR Data Synthesis with Denoising Diffusion Probabilistic Models
by: Nakashima, Kazuto, et al.
Published: (2023)
by: Nakashima, Kazuto, et al.
Published: (2023)
NeuralLabeling: A versatile toolset for labeling vision datasets using Neural Radiance Fields
by: Erich, Floris, et al.
Published: (2023)
by: Erich, Floris, et al.
Published: (2023)
CrowdMAC: Masked Crowd Density Completion for Robust Crowd Density Forecasting
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
Interpretable Debiasing of Vision-Language Models for Social Fairness
by: An, Na Min, et al.
Published: (2026)
by: An, Na Min, et al.
Published: (2026)
Beyond Independent Frames: Latent Attention Masked Autoencoders for Multi-View Echocardiography
by: Böhi, Simon, et al.
Published: (2026)
by: Böhi, Simon, et al.
Published: (2026)
Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models
by: Suzuki, Satoshi, et al.
Published: (2025)
by: Suzuki, Satoshi, et al.
Published: (2025)
Adversarial Robustness for Deep Learning-based Wildfire Prediction Models
by: Ide, Ryo, et al.
Published: (2024)
by: Ide, Ryo, et al.
Published: (2024)
AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models
by: Zhou, Yutong, et al.
Published: (2024)
by: Zhou, Yutong, et al.
Published: (2024)
Physics-Free Spectrally Multiplexed Photometric Stereo under Unknown Spectral Composition
by: Ikehata, Satoshi, et al.
Published: (2024)
by: Ikehata, Satoshi, et al.
Published: (2024)
Duoduo CLIP: Efficient 3D Understanding with Multi-View Images
by: Lee, Han-Hung, et al.
Published: (2024)
by: Lee, Han-Hung, et al.
Published: (2024)
EchoAgent: Towards Reliable Echocardiography Interpretation with "Eyes","Hands" and "Minds"
by: Wang, Qin, et al.
Published: (2026)
by: Wang, Qin, et al.
Published: (2026)
From Global to Local: Social Bias Transfer in CLIP
by: Ramos, Ryan, et al.
Published: (2025)
by: Ramos, Ryan, et al.
Published: (2025)
Action Motifs: Self-Supervised Hierarchical Representation of Human Body Movements
by: Kinoshita, Genki, et al.
Published: (2026)
by: Kinoshita, Genki, et al.
Published: (2026)
Similar Items
-
CAG-VLM: Fine-Tuning of a Large-Scale Model to Recognize Angiographic Images for Next-Generation Diagnostic Systems
by: Nakamura, Yuto, et al.
Published: (2025) -
MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer
by: Yamane, Taiga, et al.
Published: (2025) -
MVTrajecter: Multi-View Pedestrian Tracking with Trajectory Motion Cost and Trajectory Appearance Cost
by: Yamane, Taiga, et al.
Published: (2025) -
MoireDB: Formula-generated Interference-fringe Image Dataset
by: Matsuo, Yuto, et al.
Published: (2025) -
EchoPrime: A Multi-Video View-Informed Vision-Language Model for Comprehensive Echocardiography Interpretation
by: Vukadinovic, Milos, et al.
Published: (2024)