Multi-scale Temporal Prediction via Incremental Generation and Multi-agent Collaboration
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Zhitao, Yuan, Guojian, Mao, Junyuan, Wang, Yuxuan, Jia, Xiaoshuang, Jin, Yueming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Surg$Σ$: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence
by: Zeng, Zhitao, et al.
Published: (2026)
by: Zeng, Zhitao, et al.
Published: (2026)
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
by: Zeng, Zhitao, et al.
Published: (2025)
by: Zeng, Zhitao, et al.
Published: (2025)
Learning Association via Track-Detection Matching for Multi-Object Tracking
by: Adžemović, Momir
Published: (2025)
by: Adžemović, Momir
Published: (2025)
Zero-Shot Multi-Criteria Visual Quality Inspection for Semi-Controlled Industrial Environments via Real-Time 3D Digital Twin Simulation
by: Araya-Martinez, Jose Moises, et al.
Published: (2025)
by: Araya-Martinez, Jose Moises, et al.
Published: (2025)
Video-CoE: Reinforcing Video Event Prediction via Chain of Events
by: Su, Qile, et al.
Published: (2026)
by: Su, Qile, et al.
Published: (2026)
Force-Aware 3D Contact Modeling for Stable Grasp Generation
by: Chen, Zhuo, et al.
Published: (2025)
by: Chen, Zhuo, et al.
Published: (2025)
Meaning over Motion: A Semantic-First Approach to 360° Viewport Prediction
by: Khah, Arman Nik, et al.
Published: (2026)
by: Khah, Arman Nik, et al.
Published: (2026)
RefineFormer3D: Efficient 3D Medical Image Segmentation via Adaptive Multi-Scale Transformer with Cross Attention Fusion
by: Tyagi, Kavyansh, et al.
Published: (2026)
by: Tyagi, Kavyansh, et al.
Published: (2026)
Predicting and Analyzing Pedestrian Crossing Behavior at Unsignalized Crossings
by: Zhang, Chi, et al.
Published: (2024)
by: Zhang, Chi, et al.
Published: (2024)
Predicting Pedestrian Crossing Behavior in Germany and Japan: Insights into Model Transferability
by: Zhang, Chi, et al.
Published: (2024)
by: Zhang, Chi, et al.
Published: (2024)
PlaneSAM: Multimodal Plane Instance Segmentation Using the Segment Anything Model
by: Deng, Zhongchen, et al.
Published: (2024)
by: Deng, Zhongchen, et al.
Published: (2024)
A Survey on Dynamic Neural Networks: from Computer Vision to Multi-modal Sensor Fusion
by: Montello, Fabio, et al.
Published: (2025)
by: Montello, Fabio, et al.
Published: (2025)
DiffYOLO: Object Detection for Anti-Noise via YOLO and Diffusion Models
by: Liu, Yichen, et al.
Published: (2024)
by: Liu, Yichen, et al.
Published: (2024)
Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think
by: Tian, Jie, et al.
Published: (2025)
by: Tian, Jie, et al.
Published: (2025)
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
by: Sáez, Arnau Igualde, et al.
Published: (2025)
by: Sáez, Arnau Igualde, et al.
Published: (2025)
Fast 3D point clouds retrieval for Large-scale 3D Place Recognition
by: Zede, Chahine-Nicolas, et al.
Published: (2025)
by: Zede, Chahine-Nicolas, et al.
Published: (2025)
DOD-SA: Infrared-Visible Decoupled Object Detection with Single-Modality Annotations
by: Jin, Hang, et al.
Published: (2025)
by: Jin, Hang, et al.
Published: (2025)
When Less is Enough: Adaptive Token Reduction for Efficient Image Representation
by: Allakhverdov, Eduard, et al.
Published: (2025)
by: Allakhverdov, Eduard, et al.
Published: (2025)
Image Reconstruction as a Tool for Feature Analysis
by: Allakhverdov, Eduard, et al.
Published: (2025)
by: Allakhverdov, Eduard, et al.
Published: (2025)
Sequence Matters: Harnessing Video Models in 3D Super-Resolution
by: Ko, Hyun-kyu, et al.
Published: (2024)
by: Ko, Hyun-kyu, et al.
Published: (2024)
Learning Discriminative Spatio-temporal Representations for Semi-supervised Action Recognition
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
by: Zinnen, Mathias, et al.
Published: (2025)
by: Zinnen, Mathias, et al.
Published: (2025)
A large-scale, physically-based synthetic dataset for satellite pose estimation
by: Velkei, Szabolcs, et al.
Published: (2025)
by: Velkei, Szabolcs, et al.
Published: (2025)
Do Generative Metrics Predict YOLO Performance? An Evaluation Across Models, Augmentation Ratios, and Dataset Complexity
by: Marian, Vasile, et al.
Published: (2026)
by: Marian, Vasile, et al.
Published: (2026)
ClustViT: Clustering-based Token Merging for Semantic Segmentation
by: Montello, Fabio, et al.
Published: (2025)
by: Montello, Fabio, et al.
Published: (2025)
Deep Learning Approaches for Human Action Recognition in Video Data
by: Xie, Yufei
Published: (2024)
by: Xie, Yufei
Published: (2024)
LatentForensics: Towards frugal deepfake detection in the StyleGAN latent space
by: Delmas, Matthieu, et al.
Published: (2023)
by: Delmas, Matthieu, et al.
Published: (2023)
Synthetic Industrial Object Detection: GenAI vs. Feature-Based Methods
by: Araya-Martinez, Jose Moises, et al.
Published: (2025)
by: Araya-Martinez, Jose Moises, et al.
Published: (2025)
One-to-Normal: Anomaly Personalization for Few-shot Anomaly Detection
by: Li, Yiyue, et al.
Published: (2025)
by: Li, Yiyue, et al.
Published: (2025)
UrbanAlign: Post-hoc Semantic Calibration for VLM-Human Preference Alignment
by: Zhang, Yecheng, et al.
Published: (2026)
by: Zhang, Yecheng, et al.
Published: (2026)
SynthRender and IRIS: Open-Source Framework and Dataset for Bidirectional Sim-Real Transfer in Industrial Object Perception
by: Araya-Martinez, Jose Moises, et al.
Published: (2026)
by: Araya-Martinez, Jose Moises, et al.
Published: (2026)
Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos
by: Zhang, Junbin, et al.
Published: (2022)
by: Zhang, Junbin, et al.
Published: (2022)
LVP-CLIP:Revisiting CLIP for Continual Learning with Label Vector Pool
by: Ma, Yue, et al.
Published: (2024)
by: Ma, Yue, et al.
Published: (2024)
VA-$π$: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
by: Liao, Xinyao, et al.
Published: (2025)
by: Liao, Xinyao, et al.
Published: (2025)
Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models
by: Li, Jinhao, et al.
Published: (2024)
by: Li, Jinhao, et al.
Published: (2024)
DSER: Spectral Epipolar Representation for Efficient Light Field Depth Estimation
by: Mohammad, Noor Islam S., et al.
Published: (2025)
by: Mohammad, Noor Islam S., et al.
Published: (2025)
Hierarchical Spatial Algorithms for High-Resolution Image Quantization and Feature Extraction
by: Mohammad, Noor Islam S.
Published: (2025)
by: Mohammad, Noor Islam S.
Published: (2025)
OpenFusion++: An Open-vocabulary Real-time Scene Understanding System
by: Jin, Xiaofeng, et al.
Published: (2025)
by: Jin, Xiaofeng, et al.
Published: (2025)
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
Graph-PiT: Enhancing Structural Coherence in Part-Based Image Synthesis via Graph Priors
by: Zhang, Junbin, et al.
Published: (2026)
by: Zhang, Junbin, et al.
Published: (2026)
Similar Items
-
Surg$Σ$: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence
by: Zeng, Zhitao, et al.
Published: (2026) -
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
by: Zeng, Zhitao, et al.
Published: (2025) -
Learning Association via Track-Detection Matching for Multi-Object Tracking
by: Adžemović, Momir
Published: (2025) -
Zero-Shot Multi-Criteria Visual Quality Inspection for Semi-Controlled Industrial Environments via Real-Time 3D Digital Twin Simulation
by: Araya-Martinez, Jose Moises, et al.
Published: (2025) -
Video-CoE: Reinforcing Video Event Prediction via Chain of Events
by: Su, Qile, et al.
Published: (2026)