iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Manyi, Zhuang, Bingbing, Garg, Sparsh, Roy-Chowdhury, Amit, Shelton, Christian, Chandraker, Manmohan, Aich, Abhishek |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Image-Specific Adaptation of Transformer Encoders for Compute-Efficient Segmentation
by: Yao, Manyi, et al.
Published: (2024)
by: Yao, Manyi, et al.
Published: (2024)
What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs
by: Aich, Abhishek, et al.
Published: (2026)
by: Aich, Abhishek, et al.
Published: (2026)
Mapillary Vistas Validation for Fine-Grained Traffic Signs: A Benchmark Revealing Vision-Language Model Limitations
by: Garg, Sparsh, et al.
Published: (2025)
by: Garg, Sparsh, et al.
Published: (2025)
Progressive Token Length Scaling in Transformer Encoders for Efficient Universal Segmentation
by: Aich, Abhishek, et al.
Published: (2024)
by: Aich, Abhishek, et al.
Published: (2024)
HorizonWeaver: Generalizable Multi-Level Semantic Editing for Driving Scenes
by: Soroco, Mauricio, et al.
Published: (2026)
by: Soroco, Mauricio, et al.
Published: (2026)
AutoScape: Geometry-Consistent Long-Horizon Scene Generation
by: Chen, Jiacheng, et al.
Published: (2025)
by: Chen, Jiacheng, et al.
Published: (2025)
Instantaneous Perception of Moving Objects in 3D
by: Liu, Di, et al.
Published: (2024)
by: Liu, Di, et al.
Published: (2024)
A Minimalist Prompt for Zero-Shot Policy Learning
by: Song, Meng, et al.
Published: (2024)
by: Song, Meng, et al.
Published: (2024)
AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving
by: Liang, Mingfu, et al.
Published: (2024)
by: Liang, Mingfu, et al.
Published: (2024)
Tuned Contrastive Learning
by: Animesh, Chaitanya, et al.
Published: (2023)
by: Animesh, Chaitanya, et al.
Published: (2023)
LidaRF: Delving into Lidar for Neural Radiance Field on Street Scenes
by: Sun, Shanlin, et al.
Published: (2024)
by: Sun, Shanlin, et al.
Published: (2024)
Drive-1-to-3: Enriching Diffusion Priors for Novel View Synthesis of Real Vehicles
by: Lin, Chuang, et al.
Published: (2024)
by: Lin, Chuang, et al.
Published: (2024)
LLM-Assist: Enhancing Closed-Loop Planning with Language-Based Reasoning
by: Sharan, S P, et al.
Published: (2023)
by: Sharan, S P, et al.
Published: (2023)
RAD-LAD: Rule and Language Grounded Autonomous Driving in Real-Time
by: Ghosh, Anurag, et al.
Published: (2026)
by: Ghosh, Anurag, et al.
Published: (2026)
Tell, Don't Show!: Language Guidance Eases Transfer Across Domains in Images and Videos
by: Kalluri, Tarun, et al.
Published: (2024)
by: Kalluri, Tarun, et al.
Published: (2024)
Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion
by: Li, Haodong, et al.
Published: (2026)
by: Li, Haodong, et al.
Published: (2026)
DC-Gaussian: Improving 3D Gaussian Splatting for Reflective Dash Cam Videos
by: Wang, Linhan, et al.
Published: (2024)
by: Wang, Linhan, et al.
Published: (2024)
UDA-Bench: Revisiting Common Assumptions in Unsupervised Domain Adaptation Using a Standardized Framework
by: Kalluri, Tarun, et al.
Published: (2024)
by: Kalluri, Tarun, et al.
Published: (2024)
Locally Orderless Images for Optimization in Differentiable Rendering
by: Mehta, Ishit, et al.
Published: (2025)
by: Mehta, Ishit, et al.
Published: (2025)
RDC‐GS: Enhanced 3D Gaussian Splatting for Robust Dash Cam Video Reconstruction
by: Yunong Mao, et al.
Published: (2026)
by: Yunong Mao, et al.
Published: (2026)
Depth Any Camera: Zero-Shot Metric Depth Estimation from Any Camera
by: Guo, Yuliang, et al.
Published: (2025)
by: Guo, Yuliang, et al.
Published: (2025)
DashCam Video: A complementary low-cost data stream for on-demand forest-infrastructure system monitoring
by: Joshi, Durga, et al.
Published: (2025)
by: Joshi, Durga, et al.
Published: (2025)
NERFIFY: A Multi-Agent Framework for Turning NeRF Papers into Code
by: Jain, Seemandhar, et al.
Published: (2026)
by: Jain, Seemandhar, et al.
Published: (2026)
PhyCo: Learning Controllable Physical Priors for Generative Motion
by: Narayanan, Sriram, et al.
Published: (2026)
by: Narayanan, Sriram, et al.
Published: (2026)
ST-VLM: Kinematic Instruction Tuning for Spatio-Temporal Reasoning in Vision-Language Models
by: Ko, Dohwan, et al.
Published: (2025)
by: Ko, Dohwan, et al.
Published: (2025)
SplatSim: Zero-Shot Sim2Real Transfer of RGB Manipulation Policies Using Gaussian Splatting
by: Qureshi, Mohammad Nomaan, et al.
Published: (2024)
by: Qureshi, Mohammad Nomaan, et al.
Published: (2024)
Mitigating Participation Imbalance Bias in Asynchronous Federated Learning
by: Chang, Xiangyu, et al.
Published: (2025)
by: Chang, Xiangyu, et al.
Published: (2025)
ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
by: Khalifi, Omar El, et al.
Published: (2026)
by: Khalifi, Omar El, et al.
Published: (2026)
TruthLens: Visual Grounding for Universal DeepFake Reasoning
by: Kundu, Rohit, et al.
Published: (2025)
by: Kundu, Rohit, et al.
Published: (2025)
AnyCamVLA: Zero-Shot Camera Adaptation for Viewpoint Robust Vision-Language-Action Models
by: Heo, Hyeongjun, et al.
Published: (2026)
by: Heo, Hyeongjun, et al.
Published: (2026)
ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search
by: Yu, Tao, et al.
Published: (2026)
by: Yu, Tao, et al.
Published: (2026)
Fast Dash
by: Kedar Dabhadkar
Published: (2025)
by: Kedar Dabhadkar
Published: (2025)
SAModified: A Foundation Model-Based Zero-Shot Approach for Refining Noisy Land-Use Land-Cover Maps
by: Pekhale, Sparsh, et al.
Published: (2024)
by: Pekhale, Sparsh, et al.
Published: (2024)
MSNav: Zero-Shot Vision-and-Language Navigation with Dynamic Memory and LLM Spatial Reasoning
by: Liu, Chenghao, et al.
Published: (2025)
by: Liu, Chenghao, et al.
Published: (2025)
LANGTRAJ: Diffusion Model and Dataset for Language-Conditioned Trajectory Simulation
by: Chang, Wei-Jer, et al.
Published: (2025)
by: Chang, Wei-Jer, et al.
Published: (2025)
Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement
by: Khan, Zaid, et al.
Published: (2024)
by: Khan, Zaid, et al.
Published: (2024)
SAFE-SIM: Safety-Critical Closed-Loop Traffic Simulation with Diffusion-Controllable Adversaries
by: Chang, Wei-Jer, et al.
Published: (2023)
by: Chang, Wei-Jer, et al.
Published: (2023)
DashCop: Automated E-ticket Generation for Two-Wheeler Traffic Violations Using Dashcam Videos
by: Rawat, Deepti, et al.
Published: (2025)
by: Rawat, Deepti, et al.
Published: (2025)
Smart Fast Finish: Preventing Overdelivery via Daily Budget Pacing at DoorDash
by: Garg, Rohan, et al.
Published: (2025)
by: Garg, Rohan, et al.
Published: (2025)
SwasthLLM: a Unified Cross-Lingual, Multi-Task, and Meta-Learning Zero-Shot Framework for Medical Diagnosis Using Contrastive Representations
by: Sar, Ayan, et al.
Published: (2025)
by: Sar, Ayan, et al.
Published: (2025)
Similar Items
-
Image-Specific Adaptation of Transformer Encoders for Compute-Efficient Segmentation
by: Yao, Manyi, et al.
Published: (2024) -
What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs
by: Aich, Abhishek, et al.
Published: (2026) -
Mapillary Vistas Validation for Fine-Grained Traffic Signs: A Benchmark Revealing Vision-Language Model Limitations
by: Garg, Sparsh, et al.
Published: (2025) -
Progressive Token Length Scaling in Transformer Encoders for Efficient Universal Segmentation
by: Aich, Abhishek, et al.
Published: (2024) -
HorizonWeaver: Generalizable Multi-Level Semantic Editing for Driving Scenes
by: Soroco, Mauricio, et al.
Published: (2026)