Saved in:
| Main Authors: | Sha, Xuanmeng, Zhang, Liyun, Mashita, Tomohiro, Chiba, Naoya, Uranishi, Yuki |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.18451 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
3DFacePolicy: Audio-Driven 3D Facial Animation Based on Action Control
by: Sha, Xuanmeng, et al.
Published: (2024)
by: Sha, Xuanmeng, et al.
Published: (2024)
AVControl: Efficient Framework for Training Audio-Visual Controls
by: Ben-Yosef, Matan, et al.
Published: (2026)
by: Ben-Yosef, Matan, et al.
Published: (2026)
Geo2Sound: A Scalable Geo-Aligned Framework for Soundscape Generation from Satellite Imagery
by: Wu, Kunlin, et al.
Published: (2026)
by: Wu, Kunlin, et al.
Published: (2026)
Light Future: Multimodal Action Frame Prediction via InstructPix2Pix
by: Zhong, Zesen, et al.
Published: (2025)
by: Zhong, Zesen, et al.
Published: (2025)
Leum-VL Technical Report
by: He, Yuxuan, et al.
Published: (2026)
by: He, Yuxuan, et al.
Published: (2026)
Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion Reasoning
by: Han, Zhiyuan, et al.
Published: (2025)
by: Han, Zhiyuan, et al.
Published: (2025)
Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis
by: Korolkov, Vasilii
Published: (2025)
by: Korolkov, Vasilii
Published: (2025)
Evaluating Voice Command Pipelines for Drone Control: From STT and LLM to Direct Classification and Siamese Networks
by: Simões, Lucca Emmanuel Pineli, et al.
Published: (2024)
by: Simões, Lucca Emmanuel Pineli, et al.
Published: (2024)
EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction
by: Su, Qile, et al.
Published: (2025)
by: Su, Qile, et al.
Published: (2025)
A Low-Latency 3D Live Remote Visualization System for Tourist Sites Integrating Dynamic and Pre-captured Static Point Clouds
by: Matsumoto, Takahiro, et al.
Published: (2025)
by: Matsumoto, Takahiro, et al.
Published: (2025)
Cross-Modal Transfer from Memes to Videos: Addressing Data Scarcity in Hateful Video Detection
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
SIA-OVD: Shape-Invariant Adapter for Bridging the Image-Region Gap in Open-Vocabulary Detection
by: Wang, Zishuo, et al.
Published: (2024)
by: Wang, Zishuo, et al.
Published: (2024)
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
by: Tu, Songjun, et al.
Published: (2025)
by: Tu, Songjun, et al.
Published: (2025)
Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos
by: Zhang, Junbin, et al.
Published: (2022)
by: Zhang, Junbin, et al.
Published: (2022)
Hierarchical Image-Guided 3D Point Cloud Segmentation in Industrial Scenes via Multi-View Bayesian Fusion
by: Zhu, Yu, et al.
Published: (2025)
by: Zhu, Yu, et al.
Published: (2025)
MemeCraft: Contextual and Stance-Driven Multimodal Meme Generation
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
A Hybrid Deterministic Framework for Named Entity Extraction in Broadcast News Video
by: Lucas, Andrea Filiberto, et al.
Published: (2026)
by: Lucas, Andrea Filiberto, et al.
Published: (2026)
ActAlign: Zero-Shot Fine-Grained Video Classification via Language-Guided Sequence Alignment
by: Aghdam, Amir, et al.
Published: (2025)
by: Aghdam, Amir, et al.
Published: (2025)
MetaErr: Towards Predicting Error Patterns in Deep Neural Networks
by: Totakura, Varun, et al.
Published: (2026)
by: Totakura, Varun, et al.
Published: (2026)
Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-aware Spatio-Temporal Sampling
by: Liu, Xinhang, et al.
Published: (2024)
by: Liu, Xinhang, et al.
Published: (2024)
4Doodle: Two-handed Gestures for Immersive Sketching of Architectural Models
by: Fonseca, Fernando, et al.
Published: (2024)
by: Fonseca, Fernando, et al.
Published: (2024)
AIM 2024 Challenge on Video Saliency Prediction: Methods and Results
by: Moskalenko, Andrey, et al.
Published: (2024)
by: Moskalenko, Andrey, et al.
Published: (2024)
NTIRE 2026 Challenge on Video Saliency Prediction: Methods and Results
by: Moskalenko, Andrey, et al.
Published: (2026)
by: Moskalenko, Andrey, et al.
Published: (2026)
A Roadmap for Multilingual, Multimodal Domain Independent Deception Detection
by: Boumber, Dainis, et al.
Published: (2024)
by: Boumber, Dainis, et al.
Published: (2024)
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
by: Ge, Shiping, et al.
Published: (2024)
by: Ge, Shiping, et al.
Published: (2024)
Multi-level SSL Feature Gating for Audio Deepfake Detection
by: Tran, Hoan My, et al.
Published: (2025)
by: Tran, Hoan My, et al.
Published: (2025)
Cora: Correspondence-aware image editing using few step diffusion
by: Alimohammadi, Amirhossein, et al.
Published: (2025)
by: Alimohammadi, Amirhossein, et al.
Published: (2025)
Bridging Knowledge Gap Between Image Inpainting and Large-Area Visible Watermark Removal
by: Leng, Yicheng, et al.
Published: (2025)
by: Leng, Yicheng, et al.
Published: (2025)
3DGEER: 3D Gaussian Rendering Made Exact and Efficient for Generic Cameras
by: Huang, Zixun, et al.
Published: (2025)
by: Huang, Zixun, et al.
Published: (2025)
Appearance-Invariant Detection of Suggestive Motion via Laban Movement Descriptors on SMPL Skeletons
by: Ahn, Jaehoon, et al.
Published: (2026)
by: Ahn, Jaehoon, et al.
Published: (2026)
SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
by: Melechovsky, Jan, et al.
Published: (2025)
by: Melechovsky, Jan, et al.
Published: (2025)
A Real-Time, Vision-Based System for Badminton Smash Speed Estimation on Mobile Devices
by: Huang, Diwen
Published: (2025)
by: Huang, Diwen
Published: (2025)
A Scalable Pipeline Combining Procedural 3D Graphics and Guided Diffusion for Photorealistic Synthetic Training Data Generation in White Button Mushroom Segmentation
by: Károly, Artúr I., et al.
Published: (2025)
by: Károly, Artúr I., et al.
Published: (2025)
SSD-GS: Scattering and Shadow Decomposition for Relightable 3D Gaussian Splatting
by: Zheng, Iris, et al.
Published: (2026)
by: Zheng, Iris, et al.
Published: (2026)
See-through: Single-image Layer Decomposition for Anime Characters
by: Lin, Jian, et al.
Published: (2026)
by: Lin, Jian, et al.
Published: (2026)
MRD: Using Physically Based Differentiable Rendering to Probe Vision Models for 3D Scene Understanding
by: Beilharz, Benjamin, et al.
Published: (2025)
by: Beilharz, Benjamin, et al.
Published: (2025)
MSGS: Multispectral 3D Gaussian Splatting
by: Zheng, Iris, et al.
Published: (2026)
by: Zheng, Iris, et al.
Published: (2026)
Automatic Detection of Intro and Credits in Video using CLIP and Multihead Attention
by: Korolkov, Vasilii, et al.
Published: (2025)
by: Korolkov, Vasilii, et al.
Published: (2025)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
by: Raoufi, Behnam, et al.
Published: (2025)
by: Raoufi, Behnam, et al.
Published: (2025)
Graph-PiT: Enhancing Structural Coherence in Part-Based Image Synthesis via Graph Priors
by: Zhang, Junbin, et al.
Published: (2026)
by: Zhang, Junbin, et al.
Published: (2026)
Similar Items
-
3DFacePolicy: Audio-Driven 3D Facial Animation Based on Action Control
by: Sha, Xuanmeng, et al.
Published: (2024) -
AVControl: Efficient Framework for Training Audio-Visual Controls
by: Ben-Yosef, Matan, et al.
Published: (2026) -
Geo2Sound: A Scalable Geo-Aligned Framework for Soundscape Generation from Satellite Imagery
by: Wu, Kunlin, et al.
Published: (2026) -
Light Future: Multimodal Action Frame Prediction via InstructPix2Pix
by: Zhong, Zesen, et al.
Published: (2025) -
Leum-VL Technical Report
by: He, Yuxuan, et al.
Published: (2026)