Saved in:
| Main Authors: | Ishikawa, Reina, Fujii, Ryo, Saito, Hideo, Hachiuma, Ryo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2509.03385 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Predicting Any Human Trajectory In Context
by: Fujii, Ryo, et al.
Published: (2025)
by: Fujii, Ryo, et al.
Published: (2025)
RealTraj: Towards Real-World Pedestrian Trajectory Forecasting
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
Weakly Semi-supervised Tool Detection in Minimally Invasive Surgery Videos
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
VIOLA: Towards Video In-Context Learning with Minimal Annotations
by: Fujii, Ryo, et al.
Published: (2026)
by: Fujii, Ryo, et al.
Published: (2026)
CrowdMAC: Masked Crowd Density Completion for Robust Crowd Density Forecasting
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
Multimodal Cross-Domain Few-Shot Learning for Egocentric Action Recognition
by: Hatano, Masashi, et al.
Published: (2024)
by: Hatano, Masashi, et al.
Published: (2024)
EMAG: Ego-motion Aware and Generalizable 2D Hand Forecasting from Egocentric Videos
by: Hatano, Masashi, et al.
Published: (2024)
by: Hatano, Masashi, et al.
Published: (2024)
Learning from Synthetic Data via Provenance-Based Input Gradient Guidance
by: Nagano, Koshiro, et al.
Published: (2026)
by: Nagano, Koshiro, et al.
Published: (2026)
EgoSurgery-Tool: A Dataset of Surgical Tool and Hand Detection from Egocentric Open Surgery Videos
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
EgoSurgery-HTS: A Dataset for Egocentric Hand-Tool Segmentation in Open Surgery Videos
by: Darjana, Nathan, et al.
Published: (2025)
by: Darjana, Nathan, et al.
Published: (2025)
EgoSurgery-Phase: A Dataset of Surgical Phase Recognition from Egocentric Open Surgery Videos
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment
by: Hirota, Yusuke, et al.
Published: (2024)
by: Hirota, Yusuke, et al.
Published: (2024)
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
by: Shen, Xiaoqian, et al.
Published: (2025)
by: Shen, Xiaoqian, et al.
Published: (2025)
Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
by: Lee, Byung-Kwan, et al.
Published: (2025)
by: Lee, Byung-Kwan, et al.
Published: (2025)
F-Bench: Rethinking Human Preference Evaluation Metrics for Benchmarking Face Generation, Customization, and Restoration
by: Liu, Lu, et al.
Published: (2024)
by: Liu, Lu, et al.
Published: (2024)
Interpretable Debiasing of Vision-Language Models for Social Fairness
by: An, Na Min, et al.
Published: (2026)
by: An, Na Min, et al.
Published: (2026)
Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation
by: Hirota, Yusuke, et al.
Published: (2025)
by: Hirota, Yusuke, et al.
Published: (2025)
Unified Reinforcement and Imitation Learning for Vision-Language Models
by: Lee, Byung-Kwan, et al.
Published: (2025)
by: Lee, Byung-Kwan, et al.
Published: (2025)
VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language Models
by: Lee, Byung-Kwan, et al.
Published: (2024)
by: Lee, Byung-Kwan, et al.
Published: (2024)
4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation
by: Yang, Chiao-An, et al.
Published: (2025)
by: Yang, Chiao-An, et al.
Published: (2025)
SANER: Annotation-free Societal Attribute Neutralizer for Debiasing CLIP
by: Hirota, Yusuke, et al.
Published: (2024)
by: Hirota, Yusuke, et al.
Published: (2024)
Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks
by: Heo, Miran, et al.
Published: (2025)
by: Heo, Miran, et al.
Published: (2025)
OP-Align: Object-level and Part-level Alignment for Self-supervised Category-level Articulated Object Pose Estimation
by: Che, Yuchen, et al.
Published: (2024)
by: Che, Yuchen, et al.
Published: (2024)
PC-Diffusion: Aligning Diffusion Models with Human Preferences via Preference Classifier
by: Wang, Shaomeng, et al.
Published: (2025)
by: Wang, Shaomeng, et al.
Published: (2025)
Comprehensive Evaluation of Rule-Based, Machine Learning, and Deep Learning in Human Estimation Using Radio Wave Sensing: Accuracy, Spatial Generalization, and Output Granularity Trade-offs
by: Tanaka, Tomoya, et al.
Published: (2025)
by: Tanaka, Tomoya, et al.
Published: (2025)
Autoregressive Universal Video Segmentation Model
by: Heo, Miran, et al.
Published: (2025)
by: Heo, Miran, et al.
Published: (2025)
AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models
by: Zhou, Yutong, et al.
Published: (2024)
by: Zhou, Yutong, et al.
Published: (2024)
MVReward: Better Aligning and Evaluating Multi-View Diffusion Models with Human Preferences
by: Wang, Weitao, et al.
Published: (2024)
by: Wang, Weitao, et al.
Published: (2024)
V2V-LLM: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models
by: Chiu, Hsu-kuang, et al.
Published: (2025)
by: Chiu, Hsu-kuang, et al.
Published: (2025)
Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences
by: Lu, Yunhong, et al.
Published: (2025)
by: Lu, Yunhong, et al.
Published: (2025)
Piggyback Camera: Easy-to-Deploy Visual Surveillance by Mobile Sensing on Commercial Robot Vacuums
by: Yonetani, Ryo
Published: (2025)
by: Yonetani, Ryo
Published: (2025)
Profile-Specific 3DMM Regression from a Single Lateral Face Image
by: Kanaya, Taiki, et al.
Published: (2026)
by: Kanaya, Taiki, et al.
Published: (2026)
Aligning Multimodal LLM with Human Preference: A Survey
by: Yu, Tao, et al.
Published: (2025)
by: Yu, Tao, et al.
Published: (2025)
Aligning Object Detector Bounding Boxes with Human Preference
by: Strafforello, Ombretta, et al.
Published: (2024)
by: Strafforello, Ombretta, et al.
Published: (2024)
ADHMR: Aligning Diffusion-based Human Mesh Recovery via Direct Preference Optimization
by: Shen, Wenhao, et al.
Published: (2025)
by: Shen, Wenhao, et al.
Published: (2025)
Benchmarking XAI Explanations with Human-Aligned Evaluations
by: Kazmierczak, Rémi, et al.
Published: (2024)
by: Kazmierczak, Rémi, et al.
Published: (2024)
AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation
by: Liang, Chao, et al.
Published: (2025)
by: Liang, Chao, et al.
Published: (2025)
LOTUS: A Leaderboard for Detailed Image Captioning from Quality to Societal Bias and User Preferences
by: Hirota, Yusuke, et al.
Published: (2025)
by: Hirota, Yusuke, et al.
Published: (2025)
OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
by: Zhao, Xiangyu, et al.
Published: (2025)
by: Zhao, Xiangyu, et al.
Published: (2025)
PrefPaint: Aligning Image Inpainting Diffusion Model with Human Preference
by: Liu, Kendong, et al.
Published: (2024)
by: Liu, Kendong, et al.
Published: (2024)
Similar Items
-
Towards Predicting Any Human Trajectory In Context
by: Fujii, Ryo, et al.
Published: (2025) -
RealTraj: Towards Real-World Pedestrian Trajectory Forecasting
by: Fujii, Ryo, et al.
Published: (2024) -
Weakly Semi-supervised Tool Detection in Minimally Invasive Surgery Videos
by: Fujii, Ryo, et al.
Published: (2024) -
VIOLA: Towards Video In-Context Learning with Minimal Annotations
by: Fujii, Ryo, et al.
Published: (2026) -
CrowdMAC: Masked Crowd Density Completion for Robust Crowd Density Forecasting
by: Fujii, Ryo, et al.
Published: (2024)