Saved in:
| Main Authors: | Dagli, Rishit, Berger, Guillaume, Materzynska, Joanna, Bax, Ingo, Memisevic, Roland |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2410.02921 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Vision-Language Models Answer Face to Face Questions in the Real-World?
by: Pourreza, Reza, et al.
Published: (2025)
by: Pourreza, Reza, et al.
Published: (2025)
DiffuseRAW: End-to-End Generative RAW Image Processing for Low-Light Images
by: Dagli, Rishit
Published: (2023)
by: Dagli, Rishit
Published: (2023)
What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction
by: Panchal, Sunny, et al.
Published: (2024)
by: Panchal, Sunny, et al.
Published: (2024)
NeRF-US: Removing Ultrasound Imaging Artifacts from Neural Radiance Fields in the Wild
by: Dagli, Rishit, et al.
Published: (2024)
by: Dagli, Rishit, et al.
Published: (2024)
SEE-2-SOUND: Zero-Shot Spatial Environment-to-Spatial Sound
by: Dagli, Rishit, et al.
Published: (2024)
by: Dagli, Rishit, et al.
Published: (2024)
FairyGen: Storied Cartoon Video from a Single Child-Drawn Character
by: Zheng, Jiayi, et al.
Published: (2025)
by: Zheng, Jiayi, et al.
Published: (2025)
Squeeze3D: Your 3D Generation Model is Secretly an Extreme Neural Compressor
by: Dagli, Rishit, et al.
Published: (2025)
by: Dagli, Rishit, et al.
Published: (2025)
Opt-In Art: Learning Art Styles Only from Few Examples
by: Ren, Hui, et al.
Published: (2024)
by: Ren, Hui, et al.
Published: (2024)
Unified Concept Editing in Diffusion Models
by: Gandikota, Rohit, et al.
Published: (2023)
by: Gandikota, Rohit, et al.
Published: (2023)
NewMove: Customizing text-to-video models with novel motions
by: Materzynska, Joanna, et al.
Published: (2023)
by: Materzynska, Joanna, et al.
Published: (2023)
OpenCOOD-Air: Prompting Heterogeneous Ground-Air Collaborative Perception with Spatial Conversion and Offset Prediction
by: Wu, Xianke, et al.
Published: (2026)
by: Wu, Xianke, et al.
Published: (2026)
FreeForm: Reduced-Order Deformable Simulation from Particle-Based Skinning Eigenmodes
by: Xiang, Donglai, et al.
Published: (2026)
by: Xiang, Donglai, et al.
Published: (2026)
Model-Agnostic Open-Set Air-to-Air Visual Object Detection for Reliable UAV Perception
by: Loukovitis, Spyridon, et al.
Published: (2025)
by: Loukovitis, Spyridon, et al.
Published: (2025)
SurgOnAir: Hierarchy-Aware Real-Time Surgical Video Commentary
by: He, Jingyi, et al.
Published: (2026)
by: He, Jingyi, et al.
Published: (2026)
Look, Remember and Reason: Grounded reasoning in videos with language models
by: Bhattacharyya, Apratim, et al.
Published: (2023)
by: Bhattacharyya, Apratim, et al.
Published: (2023)
Common Corruptions for Enhancing and Evaluating Robustness in Air-to-Air Visual Object Detection
by: Arsenos, Anastasios, et al.
Published: (2024)
by: Arsenos, Anastasios, et al.
Published: (2024)
Character Mixing for Video Generation
by: Liao, Tingting, et al.
Published: (2025)
by: Liao, Tingting, et al.
Published: (2025)
PM25Vision: A Large-Scale Benchmark Dataset for Visual Estimation of Air Quality
by: Han, Yang
Published: (2025)
by: Han, Yang
Published: (2025)
SCT-MOT: Enhancing Air-to-Air Multiple UAVs Tracking with Swarm-Coupled Motion and Trajectory Guidance
by: Chu, Zhaochen, et al.
Published: (2026)
by: Chu, Zhaochen, et al.
Published: (2026)
MovieCharacter: A Tuning-Free Framework for Controllable Character Video Synthesis
by: Qiu, Di, et al.
Published: (2024)
by: Qiu, Di, et al.
Published: (2024)
AirZoo: A Unified Large-Scale Dataset for Grounding Aerial Geometric 3D Vision
by: Cheng, Xiaoya, et al.
Published: (2026)
by: Cheng, Xiaoya, et al.
Published: (2026)
Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
by: Bhattacharyya, Apratim, et al.
Published: (2025)
by: Bhattacharyya, Apratim, et al.
Published: (2025)
Detecting Pen-In-Air States from Video: A Proof-of-Concept Toward Complementary Handwriting Analysis
by: Sismeiro, Lauren, et al.
Published: (2026)
by: Sismeiro, Lauren, et al.
Published: (2026)
AirRoom: Objects Matter in Room Reidentification
by: Yao, Runmao, et al.
Published: (2025)
by: Yao, Runmao, et al.
Published: (2025)
AirCast: Improving Air Pollution Forecasting Through Multi-Variable Data Alignment
by: Nedungadi, Vishal, et al.
Published: (2025)
by: Nedungadi, Vishal, et al.
Published: (2025)
Do Generalised Classifiers really work on Human Drawn Sketches?
by: Bandyopadhyay, Hmrishav, et al.
Published: (2024)
by: Bandyopadhyay, Hmrishav, et al.
Published: (2024)
Rotate Your Character: Revisiting Video Diffusion Models for High-Quality 3D Character Generation
by: Wang, Jin, et al.
Published: (2026)
by: Wang, Jin, et al.
Published: (2026)
Understanding Depth and Height Perception in Large Visual-Language Models
by: Azad, Shehreen, et al.
Published: (2024)
by: Azad, Shehreen, et al.
Published: (2024)
Exploring the Efficacy of Modified Transfer Learning in Identifying Parkinson's Disease Through Drawn Image Patterns
by: Daiyan, Nabil, et al.
Published: (2025)
by: Daiyan, Nabil, et al.
Published: (2025)
AirV2X: Unified Air-Ground Vehicle-to-Everything Collaboration
by: Gao, Xiangbo, et al.
Published: (2025)
by: Gao, Xiangbo, et al.
Published: (2025)
Enhancing 3D-Air Signature by Pen Tip Tail Trajectory Awareness: Dataset and Featuring by Novel Spatio-temporal CNN
by: Atreya, Saurabh, et al.
Published: (2024)
by: Atreya, Saurabh, et al.
Published: (2024)
OpenMarcie: Dataset for Multimodal Action Recognition in Industrial Environments
by: Bello, Hymalai, et al.
Published: (2026)
by: Bello, Hymalai, et al.
Published: (2026)
VoMP: Predicting Volumetric Mechanical Property Fields
by: Dagli, Rishit, et al.
Published: (2025)
by: Dagli, Rishit, et al.
Published: (2025)
AirShot: Efficient Few-Shot Detection for Autonomous Exploration
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling
by: Men, Yifang, et al.
Published: (2024)
by: Men, Yifang, et al.
Published: (2024)
AnimationBench: Are Video Models Good at Character-Centric Animation?
by: Wu, Leyi, et al.
Published: (2026)
by: Wu, Leyi, et al.
Published: (2026)
Gloria: Consistent Character Video Generation via Content Anchors
by: Yang, Yuhang, et al.
Published: (2026)
by: Yang, Yuhang, et al.
Published: (2026)
Transport Network, Graph, and Air Pollution
by: Xu, Nan
Published: (2025)
by: Xu, Nan
Published: (2025)
Oracle-MNIST: a Dataset of Oracle Characters for Benchmarking Machine Learning Algorithms
by: Wang, Mei, et al.
Published: (2022)
by: Wang, Mei, et al.
Published: (2022)
Mobile Robot Navigation Using Hand-Drawn Maps: A Vision Language Model Approach
by: Tan, Aaron Hao, et al.
Published: (2025)
by: Tan, Aaron Hao, et al.
Published: (2025)
Similar Items
-
Can Vision-Language Models Answer Face to Face Questions in the Real-World?
by: Pourreza, Reza, et al.
Published: (2025) -
DiffuseRAW: End-to-End Generative RAW Image Processing for Low-Light Images
by: Dagli, Rishit
Published: (2023) -
What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction
by: Panchal, Sunny, et al.
Published: (2024) -
NeRF-US: Removing Ultrasound Imaging Artifacts from Neural Radiance Fields in the Wild
by: Dagli, Rishit, et al.
Published: (2024) -
SEE-2-SOUND: Zero-Shot Spatial Environment-to-Spatial Sound
by: Dagli, Rishit, et al.
Published: (2024)