DynaStride: Dynamic Stride Windowing with MMCoT for Instructional Multi-Scene Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Pham, Eddison, Priyadarshini, Prisha, Maliackel, Adrian, Bandi, Kanishk, Meo, Cristian, Zhu, Kevin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StrideNET: Swin Transformer for Terrain Recognition with Dynamic Roughness Extraction
by: Shelare, Maitreya, et al.
Published: (2024)
by: Shelare, Maitreya, et al.
Published: (2024)
Hybrid of DiffStride and Spectral Pooling in Convolutional Neural Networks
by: Rafif, Sulthan, et al.
Published: (2024)
by: Rafif, Sulthan, et al.
Published: (2024)
MapTracker: Tracking with Strided Memory Fusion for Consistent Vector HD Mapping
by: Chen, Jiacheng, et al.
Published: (2024)
by: Chen, Jiacheng, et al.
Published: (2024)
Markerless Stride Length estimation in Athletic using Pose Estimation with monocular vision
by: Skorupski, Patryk, et al.
Published: (2025)
by: Skorupski, Patryk, et al.
Published: (2025)
Stride-Net: Fairness-Aware Disentangled Representation Learning for Chest X-Ray Diagnosis
by: Rashid, Darakshan, et al.
Published: (2026)
by: Rashid, Darakshan, et al.
Published: (2026)
ZACH-ViT: A Zero-Token Vision Transformer with ShuffleStrides Data Augmentation for Robust Lung Ultrasound Classification
by: Angelakis, Athanasios, et al.
Published: (2025)
by: Angelakis, Athanasios, et al.
Published: (2025)
SasMamba: A Lightweight Structure-Aware Stride State Space Model for 3D Human Pose Estimation
by: Cui, Hu, et al.
Published: (2025)
by: Cui, Hu, et al.
Published: (2025)
Britain. Striding clear
Published: (1995)
Published: (1995)
Strides in Development of Medical Education
Published: (2021)
Published: (2021)
Strided Difference Bound Matrices
by: Pitchanathan, Arjun, et al.
Published: (2024)
by: Pitchanathan, Arjun, et al.
Published: (2024)
Aiding Medical Diagnosis through Image Synthesis and Classification
by: Choudhary, Kanishk
Published: (2025)
by: Choudhary, Kanishk
Published: (2025)
The Inductive Bottleneck: Data-Driven Emergence of Representational Sparsity in Vision Transformers
by: Awadhiya, Kanishk
Published: (2025)
by: Awadhiya, Kanishk
Published: (2025)
DynaVol: Unsupervised Learning for Dynamic Scenes through Object-Centric Voxelization
by: Zhao, Yanpeng, et al.
Published: (2023)
by: Zhao, Yanpeng, et al.
Published: (2023)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
by: Yao, Linli, et al.
Published: (2026)
by: Yao, Linli, et al.
Published: (2026)
Multi-Strided Access Patterns to Boost Hardware Prefetching
by: Blom, Miguel O., et al.
Published: (2024)
by: Blom, Miguel O., et al.
Published: (2024)
Shifted Window Fourier Transform And Retention For Image Captioning
by: Hu, Jia Cheng, et al.
Published: (2024)
by: Hu, Jia Cheng, et al.
Published: (2024)
DynaSplat: Dynamic-Static Gaussian Splatting with Hierarchical Motion Decomposition for Scene Reconstruction
by: Deng, Junli, et al.
Published: (2025)
by: Deng, Junli, et al.
Published: (2025)
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
by: Zheng, Guangcong, et al.
Published: (2025)
by: Zheng, Guangcong, et al.
Published: (2025)
Grounding Foundational Vision Models with 3D Human Poses for Robust Action Recognition
by: Babey, Nicholas, et al.
Published: (2025)
by: Babey, Nicholas, et al.
Published: (2025)
Instruction-guided Multi-Granularity Segmentation and Captioning with Large Multimodal Model
by: Zhou, Li, et al.
Published: (2024)
by: Zhou, Li, et al.
Published: (2024)
Retrieval, Refinement, and Ranking for Text-to-Video Generation via Prompt Optimization and Test-Time Scaling
by: Rahman, Zillur, et al.
Published: (2026)
by: Rahman, Zillur, et al.
Published: (2026)
Quantum Convolutional Neural Network with Flexible Stride
by: Yu, Kai, et al.
Published: (2024)
by: Yu, Kai, et al.
Published: (2024)
ROSER: Few-Shot Robotic Sequence Retrieval for Scalable Robot Learning
by: Rahman, Zillur, et al.
Published: (2026)
by: Rahman, Zillur, et al.
Published: (2026)
DynaGSLAM: Real-Time Gaussian-Splatting SLAM for Online Rendering, Tracking, Motion Predictions of Moving Objects in Dynamic Scenes
by: Li, Runfa Blark, et al.
Published: (2025)
by: Li, Runfa Blark, et al.
Published: (2025)
X-Dyna: Expressive Dynamic Human Image Animation
by: Chang, Di, et al.
Published: (2025)
by: Chang, Di, et al.
Published: (2025)
Scene Graph-guided SegCaptioning Transformer with Fine-grained Alignment for Controllable Video Segmentation and Captioning
by: Zhang, Xu, et al.
Published: (2026)
by: Zhang, Xu, et al.
Published: (2026)
Evaluation of Vision-LLMs in Surveillance Video
by: Benschop, Pascal, et al.
Published: (2025)
by: Benschop, Pascal, et al.
Published: (2025)
TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
by: Jin, Bu, et al.
Published: (2024)
by: Jin, Bu, et al.
Published: (2024)
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
An Alternative to Stride-Based RNG for Monte Carlo Transport
by: Cuneo, Braxton S., et al.
Published: (2024)
by: Cuneo, Braxton S., et al.
Published: (2024)
Dyna3DGR: 4D Cardiac Motion Tracking with Dynamic 3D Gaussian Representation
by: Fu, Xueming, et al.
Published: (2025)
by: Fu, Xueming, et al.
Published: (2025)
IF-VidCap: Can Video Caption Models Follow Instructions?
by: Li, Shihao, et al.
Published: (2025)
by: Li, Shihao, et al.
Published: (2025)
Fine-Grained Captioning of Long Videos through Scene Graph Consolidation
by: Chu, Sanghyeok, et al.
Published: (2025)
by: Chu, Sanghyeok, et al.
Published: (2025)
DynaDrag: Dynamic Drag-Style Image Editing by Motion Prediction
by: Sui, Jiacheng, et al.
Published: (2026)
by: Sui, Jiacheng, et al.
Published: (2026)
Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning
by: Yu, Hanxun, et al.
Published: (2025)
by: Yu, Hanxun, et al.
Published: (2025)
Slide-SAM: Medical SAM Meets Sliding Window
by: Quan, Quan, et al.
Published: (2023)
by: Quan, Quan, et al.
Published: (2023)
BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving
by: Brandstaetter, Felix, et al.
Published: (2025)
by: Brandstaetter, Felix, et al.
Published: (2025)
SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioning
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
Autonomous Character-Scene Interaction Synthesis from Text Instruction
by: Jiang, Nan, et al.
Published: (2024)
by: Jiang, Nan, et al.
Published: (2024)
DynaGuide: A Generalizable Dynamic Guidance Framework for Unsupervised Semantic Segmentation
by: Guermazi, Boujemaa, et al.
Published: (2026)
by: Guermazi, Boujemaa, et al.
Published: (2026)
Similar Items
-
StrideNET: Swin Transformer for Terrain Recognition with Dynamic Roughness Extraction
by: Shelare, Maitreya, et al.
Published: (2024) -
Hybrid of DiffStride and Spectral Pooling in Convolutional Neural Networks
by: Rafif, Sulthan, et al.
Published: (2024) -
MapTracker: Tracking with Strided Memory Fusion for Consistent Vector HD Mapping
by: Chen, Jiacheng, et al.
Published: (2024) -
Markerless Stride Length estimation in Athletic using Pose Estimation with monocular vision
by: Skorupski, Patryk, et al.
Published: (2025) -
Stride-Net: Fairness-Aware Disentangled Representation Learning for Chest X-Ray Diagnosis
by: Rashid, Darakshan, et al.
Published: (2026)