MonSTeR: a Unified Model for Motion, Scene, Text Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Collorone, Luca, Gioia, Matteo, Pappa, Massimiliano, Leoni, Paolo, Ficarra, Giovanni, Litany, Or, Spinelli, Indro, Galasso, Fabio |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoDiPO: text-to-motion alignment via AI-feedback-driven Direct Preference Optimization
by: Pappa, Massimiliano, et al.
Published: (2024)
by: Pappa, Massimiliano, et al.
Published: (2024)
PhysTalk: Language-driven Real-time Physics in 3D Gaussian Scenes
by: Collorone, Luca, et al.
Published: (2025)
by: Collorone, Luca, et al.
Published: (2025)
ANTHROPOS-V: benchmarking the novel task of Crowd Volume Estimation
by: Collorone, Luca, et al.
Published: (2025)
by: Collorone, Luca, et al.
Published: (2025)
Length-Aware Motion Synthesis via Latent Diffusion
by: Sampieri, Alessio, et al.
Published: (2024)
by: Sampieri, Alessio, et al.
Published: (2024)
Human Motion Unlearning
by: De Matteis, Edoardo, et al.
Published: (2025)
by: De Matteis, Edoardo, et al.
Published: (2025)
Social EgoMesh Estimation
by: Scofano, Luca, et al.
Published: (2024)
by: Scofano, Luca, et al.
Published: (2024)
Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models
by: Pappa, Massimiliano, et al.
Published: (2026)
by: Pappa, Massimiliano, et al.
Published: (2026)
Following the Human Thread in Social Navigation
by: Scofano, Luca, et al.
Published: (2024)
by: Scofano, Luca, et al.
Published: (2024)
OVOSE: Open-Vocabulary Semantic Segmentation in Event-Based Cameras
by: Rahman, Muhammad Rameez Ur, et al.
Published: (2024)
by: Rahman, Muhammad Rameez Ur, et al.
Published: (2024)
Video Unlearning via Low-Rank Refusal Vector
by: Facchiano, Simone, et al.
Published: (2025)
by: Facchiano, Simone, et al.
Published: (2025)
Quantifying Self-Preservation Bias in Large Language Models
by: Migliarini, Matteo, et al.
Published: (2026)
by: Migliarini, Matteo, et al.
Published: (2026)
UnScene3D: Unsupervised 3D Instance Segmentation for Indoor Scenes
by: Rozenberszki, David, et al.
Published: (2023)
by: Rozenberszki, David, et al.
Published: (2023)
Bimanual Robot Manipulation via Multi-Agent In-Context Learning
by: Palma, Alessio, et al.
Published: (2026)
by: Palma, Alessio, et al.
Published: (2026)
Adaptive Point Transformer
by: Baiocchi, Alessandro, et al.
Published: (2024)
by: Baiocchi, Alessandro, et al.
Published: (2024)
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
by: Zhang, Junyi, et al.
Published: (2024)
by: Zhang, Junyi, et al.
Published: (2024)
SeRpEnt: Selective Resampling for Expressive State Space Models
by: Rando, Stefano, et al.
Published: (2025)
by: Rando, Stefano, et al.
Published: (2025)
About latent roles in forecasting players in team sports
by: Scofano, Luca, et al.
Published: (2023)
by: Scofano, Luca, et al.
Published: (2023)
Gaussian See, Gaussian Do: Semantic 3D Motion Transfer from Multiview Video
by: Bekor, Yarin, et al.
Published: (2025)
by: Bekor, Yarin, et al.
Published: (2025)
Partial Scene Text Retrieval
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
InstructMix2Mix: Consistent Sparse-View Editing Through Multi-View Model Personalization
by: Gilo, Daniel, et al.
Published: (2025)
by: Gilo, Daniel, et al.
Published: (2025)
Multi-Track Timeline Control for Text-Driven 3D Human Motion Generation
by: Petrovich, Mathis, et al.
Published: (2024)
by: Petrovich, Mathis, et al.
Published: (2024)
SceneTeract: Agentic Functional Affordances and VLM Grounding in 3D Scenes
by: Maillard, Léopold, et al.
Published: (2026)
by: Maillard, Léopold, et al.
Published: (2026)
ZDySS -- Zero-Shot Dynamic Scene Stylization using Gaussian Splatting
by: Saroha, Abhishek, et al.
Published: (2025)
by: Saroha, Abhishek, et al.
Published: (2025)
Hyperbolic Active Learning for Semantic Segmentation under Domain Shift
by: Franco, Luca, et al.
Published: (2023)
by: Franco, Luca, et al.
Published: (2023)
Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy
by: Jingyu, Gong, et al.
Published: (2025)
by: Jingyu, Gong, et al.
Published: (2025)
PaSTe: Improving the Efficiency of Visual Anomaly Detection at the Edge
by: Barusco, Manuel, et al.
Published: (2024)
by: Barusco, Manuel, et al.
Published: (2024)
OccSTeP: Benchmarking 4D Occupancy Spatio-Temporal Persistence
by: Zheng, Yu, et al.
Published: (2025)
by: Zheng, Yu, et al.
Published: (2025)
Generating Human Interaction Motions in Scenes with Text Control
by: Yi, Hongwei, et al.
Published: (2024)
by: Yi, Hongwei, et al.
Published: (2024)
MotionRFT: Unified Reinforcement Fine-Tuning for Text-to-Motion Generation
by: Tan, Xiaofeng, et al.
Published: (2026)
by: Tan, Xiaofeng, et al.
Published: (2026)
DIMA: DIffusing Motion Artifacts for unsupervised correction in brain MRI images
by: Angella, Paolo, et al.
Published: (2025)
by: Angella, Paolo, et al.
Published: (2025)
Zero-to-Hero: Enhancing Zero-Shot Novel View Synthesis via Attention Map Filtering
by: Sobol, Ido, et al.
Published: (2024)
by: Sobol, Ido, et al.
Published: (2024)
Appreciate the View: A Task-Aware Evaluation Framework for Novel View Synthesis
by: Stern, Saar, et al.
Published: (2025)
by: Stern, Saar, et al.
Published: (2025)
STeInFormer: Spatial-Temporal Interaction Transformer Architecture for Remote Sensing Change Detection
by: Ma, Xiaowen, et al.
Published: (2024)
by: Ma, Xiaowen, et al.
Published: (2024)
Time-to-Move: Training-Free Motion Controlled Video Generation via Dual-Clock Denoising
by: Singer, Assaf, et al.
Published: (2025)
by: Singer, Assaf, et al.
Published: (2025)
FreeMotion: A Unified Framework for Number-free Text-to-Motion Synthesis
by: Fan, Ke, et al.
Published: (2024)
by: Fan, Ke, et al.
Published: (2024)
Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting
by: Colombo, Antonio, et al.
Published: (2026)
by: Colombo, Antonio, et al.
Published: (2026)
FlowSeek: Optical Flow Made Easier with Depth Foundation Models and Motion Bases
by: Poggi, Matteo, et al.
Published: (2025)
by: Poggi, Matteo, et al.
Published: (2025)
MonSter++: Unified Stereo Matching, Multi-view Stereo, and Real-time Stereo with Monodepth Priors
by: Cheng, Junda, et al.
Published: (2025)
by: Cheng, Junda, et al.
Published: (2025)
Generating Human Motion in 3D Scenes from Text Descriptions
by: Cen, Zhi, et al.
Published: (2024)
by: Cen, Zhi, et al.
Published: (2024)
STeP: A Framework for Solving Scientific Video Inverse Problems with Spatiotemporal Diffusion Priors
by: Zhang, Bingliang, et al.
Published: (2025)
by: Zhang, Bingliang, et al.
Published: (2025)
Similar Items
-
MoDiPO: text-to-motion alignment via AI-feedback-driven Direct Preference Optimization
by: Pappa, Massimiliano, et al.
Published: (2024) -
PhysTalk: Language-driven Real-time Physics in 3D Gaussian Scenes
by: Collorone, Luca, et al.
Published: (2025) -
ANTHROPOS-V: benchmarking the novel task of Crowd Volume Estimation
by: Collorone, Luca, et al.
Published: (2025) -
Length-Aware Motion Synthesis via Latent Diffusion
by: Sampieri, Alessio, et al.
Published: (2024) -
Human Motion Unlearning
by: De Matteis, Edoardo, et al.
Published: (2025)