CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Grover, Shresth, Pathak, Priyank, Kumar, Akash, Vineet, Vibhav, Rawat, Yogesh S |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Navigating Hallucinations for Reasoning of Unintentional Activities
von: Grover, Shresth, et al.
Veröffentlicht: (2024)
von: Grover, Shresth, et al.
Veröffentlicht: (2024)
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
OmViD: Omni-supervised active learning for video action detection
von: Rana, Aayush, et al.
Veröffentlicht: (2025)
von: Rana, Aayush, et al.
Veröffentlicht: (2025)
Colors See Colors Ignore: Clothes Changing ReID with Color Disentanglement
von: Pathak, Priyank, et al.
Veröffentlicht: (2025)
von: Pathak, Priyank, et al.
Veröffentlicht: (2025)
Coarse Attribute Prediction with Task Agnostic Distillation for Real World Clothes Changing ReID
von: Pathak, Priyank, et al.
Veröffentlicht: (2025)
von: Pathak, Priyank, et al.
Veröffentlicht: (2025)
HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
von: Azad, Shehreen, et al.
Veröffentlicht: (2025)
von: Azad, Shehreen, et al.
Veröffentlicht: (2025)
StreamReady: Learning What to Answer and When in Long Streaming Videos
von: Azad, Shehreen, et al.
Veröffentlicht: (2026)
von: Azad, Shehreen, et al.
Veröffentlicht: (2026)
On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes
von: Modi, Rajat, et al.
Veröffentlicht: (2024)
von: Modi, Rajat, et al.
Veröffentlicht: (2024)
LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models
von: Pathak, Priyank, et al.
Veröffentlicht: (2025)
von: Pathak, Priyank, et al.
Veröffentlicht: (2025)
Understanding Depth and Height Perception in Large Visual-Language Models
von: Azad, Shehreen, et al.
Veröffentlicht: (2024)
von: Azad, Shehreen, et al.
Veröffentlicht: (2024)
STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding
von: Garg, Aaryan, et al.
Veröffentlicht: (2025)
von: Garg, Aaryan, et al.
Veröffentlicht: (2025)
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
Stable Mean Teacher for Semi-supervised Video Action Detection
von: Kumar, Akash, et al.
Veröffentlicht: (2024)
von: Kumar, Akash, et al.
Veröffentlicht: (2024)
Robustness Analysis on Foundational Segmentation Models
von: Schiappa, Madeline Chantry, et al.
Veröffentlicht: (2023)
von: Schiappa, Madeline Chantry, et al.
Veröffentlicht: (2023)
Physics Knowledge in Frontier Models: A Diagnostic Study of Failure Modes
von: Bagdonaviciute, Ieva, et al.
Veröffentlicht: (2025)
von: Bagdonaviciute, Ieva, et al.
Veröffentlicht: (2025)
RobustGait: Robustness Analysis for Appearance Based Gait Recognition
von: Sayera, Reeshoon, et al.
Veröffentlicht: (2025)
von: Sayera, Reeshoon, et al.
Veröffentlicht: (2025)
DIFFER: Disentangling Identity Features via Semantic Cues for Clothes-Changing Person Re-ID
von: Liang, Xin, et al.
Veröffentlicht: (2025)
von: Liang, Xin, et al.
Veröffentlicht: (2025)
Semi-supervised Active Learning for Video Action Detection
von: Singh, Ayush, et al.
Veröffentlicht: (2023)
von: Singh, Ayush, et al.
Veröffentlicht: (2023)
How Do Inpainting Artifacts Propagate to Language?
von: Yashwante, Pratham, et al.
Veröffentlicht: (2026)
von: Yashwante, Pratham, et al.
Veröffentlicht: (2026)
ProDiG: Progressive Diffusion-Guided Gaussian Splatting for Aerial to Ground Reconstruction
von: Mitra, Sirshapan, et al.
Veröffentlicht: (2026)
von: Mitra, Sirshapan, et al.
Veröffentlicht: (2026)
DisenQ: Disentangling Q-Former for Activity-Biometrics
von: Azad, Shehreen, et al.
Veröffentlicht: (2025)
von: Azad, Shehreen, et al.
Veröffentlicht: (2025)
Towards Scene Graph Anticipation
von: Peddi, Rohith, et al.
Veröffentlicht: (2024)
von: Peddi, Rohith, et al.
Veröffentlicht: (2024)
PEEKABOO: Interactive Video Generation via Masked-Diffusion
von: Jain, Yash, et al.
Veröffentlicht: (2023)
von: Jain, Yash, et al.
Veröffentlicht: (2023)
Towards Unbiased and Robust Spatio-Temporal Scene Graph Generation and Anticipation
von: Peddi, Rohith, et al.
Veröffentlicht: (2024)
von: Peddi, Rohith, et al.
Veröffentlicht: (2024)
Adaptive Visual Scene Understanding: Incremental Scene Graph Generation
von: Khandelwal, Naitik, et al.
Veröffentlicht: (2023)
von: Khandelwal, Naitik, et al.
Veröffentlicht: (2023)
Activity-Biometrics: Person Identification from Daily Activities
von: Azad, Shehreen, et al.
Veröffentlicht: (2024)
von: Azad, Shehreen, et al.
Veröffentlicht: (2024)
Towards Spatio-Temporal World Scene Graph Generation from Monocular Videos
von: Peddi, Rohith, et al.
Veröffentlicht: (2026)
von: Peddi, Rohith, et al.
Veröffentlicht: (2026)
GaitCrafter: Diffusion Model for Biometric Preserving Gait Synthesis
von: Mitra, Sirshapan, et al.
Veröffentlicht: (2025)
von: Mitra, Sirshapan, et al.
Veröffentlicht: (2025)
Scaling Open-Vocabulary Action Detection
von: Sia, Zhen Hao, et al.
Veröffentlicht: (2025)
von: Sia, Zhen Hao, et al.
Veröffentlicht: (2025)
Streamlining Video Analysis for Efficient Violence Detection
von: Pathak, Gourang, et al.
Veröffentlicht: (2024)
von: Pathak, Gourang, et al.
Veröffentlicht: (2024)
VISTA: Video Interaction Spatio-Temporal Analysis Benchmark
von: Aparcedo, Alejandro, et al.
Veröffentlicht: (2026)
von: Aparcedo, Alejandro, et al.
Veröffentlicht: (2026)
MolVision: Molecular Property Prediction with Vision Language Models
von: Adak, Deepan, et al.
Veröffentlicht: (2025)
von: Adak, Deepan, et al.
Veröffentlicht: (2025)
iSafetyBench: A video-language benchmark for safety in industrial environment
von: Abdullah, Raiyaan, et al.
Veröffentlicht: (2025)
von: Abdullah, Raiyaan, et al.
Veröffentlicht: (2025)
Asynchronous Perception Machine For Efficient Test-Time-Training
von: Modi, Rajat, et al.
Veröffentlicht: (2024)
von: Modi, Rajat, et al.
Veröffentlicht: (2024)
Grounding Task Assistance with Multimodal Cues from a Single Demonstration
von: Sarch, Gabriel, et al.
Veröffentlicht: (2025)
von: Sarch, Gabriel, et al.
Veröffentlicht: (2025)
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames
von: Ravi, Sahithya, et al.
Veröffentlicht: (2025)
von: Ravi, Sahithya, et al.
Veröffentlicht: (2025)
EZ-CLIP: Efficient Zeroshot Video Action Recognition
von: Ahmad, Shahzad, et al.
Veröffentlicht: (2023)
von: Ahmad, Shahzad, et al.
Veröffentlicht: (2023)
Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
von: Kumar, Yogesh
Veröffentlicht: (2025)
von: Kumar, Yogesh
Veröffentlicht: (2025)
Belief Scene Graphs: Expanding Partial Scenes with Objects through Computation of Expectation
von: Saucedo, Mario A. V., et al.
Veröffentlicht: (2024)
von: Saucedo, Mario A. V., et al.
Veröffentlicht: (2024)
OS-Marathon: Benchmarking Computer-Use Agents on Long-Horizon Repetitive Tasks
von: Wu, Jing, et al.
Veröffentlicht: (2026)
von: Wu, Jing, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Navigating Hallucinations for Reasoning of Unintentional Activities
von: Grover, Shresth, et al.
Veröffentlicht: (2024) -
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
von: Kumar, Akash, et al.
Veröffentlicht: (2025) -
OmViD: Omni-supervised active learning for video action detection
von: Rana, Aayush, et al.
Veröffentlicht: (2025) -
Colors See Colors Ignore: Clothes Changing ReID with Color Disentanglement
von: Pathak, Priyank, et al.
Veröffentlicht: (2025) -
Coarse Attribute Prediction with Task Agnostic Distillation for Real World Clothes Changing ReID
von: Pathak, Priyank, et al.
Veröffentlicht: (2025)