Saved in:
| Main Authors: | Agarwal, Anmol, Meshram, Pranay, Singh, Sumer, Suman, Saurav, Lapp, Andrew, Matiana, Shahbuland, Castricato, Louis, Frazier, Spencer |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.00825 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
QAL: A Loss for Recall Precision Balance in 3D Reconstruction
by: Meshram, Pranay, et al.
Published: (2025)
by: Meshram, Pranay, et al.
Published: (2025)
Empir3D : A Framework for Multi-Dimensional Point Cloud Assessment
by: Turkar, Yash, et al.
Published: (2023)
by: Turkar, Yash, et al.
Published: (2023)
Humanoid World Models: Open World Foundation Models for Humanoid Robotics
by: Ali, Muhammad Qasim, et al.
Published: (2025)
by: Ali, Muhammad Qasim, et al.
Published: (2025)
Low-Compute Watermark Removal via Dual-Domain Natural Projection
by: Meshram, Pragati Shuddhodhan, et al.
Published: (2025)
by: Meshram, Pragati Shuddhodhan, et al.
Published: (2025)
Exploiting Contextual Uncertainty of Visual Data for Efficient Training of Deep Models
by: Agarwal, Sharat
Published: (2024)
by: Agarwal, Sharat
Published: (2024)
Explainable Gait Abnormality Detection Using Dual-Dataset CNN-LSTM Models
by: Agarwal, Parth, et al.
Published: (2025)
by: Agarwal, Parth, et al.
Published: (2025)
Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
by: Nilaksh, et al.
Published: (2026)
by: Nilaksh, et al.
Published: (2026)
WorldBench: Disambiguating Physics for Diagnostic Evaluation of World Models
by: Upadhyay, Rishi, et al.
Published: (2026)
by: Upadhyay, Rishi, et al.
Published: (2026)
Probing the effectiveness of World Models for Spatial Reasoning through Test-time Scaling
by: Jha, Saurav, et al.
Published: (2025)
by: Jha, Saurav, et al.
Published: (2025)
Attention Isn't All You Need for Emotion Recognition:Domain Features Outperform Transformers on the EAV Dataset
by: Guragain, Anmol
Published: (2026)
by: Guragain, Anmol
Published: (2026)
Class Confidence Aware Reweighting for Long Tailed Learning
by: Jagati, Brainard Philemon, et al.
Published: (2026)
by: Jagati, Brainard Philemon, et al.
Published: (2026)
TIDE: Training Locally Interpretable Domain Generalization Models Enables Test-time Correction
by: Agarwal, Aishwarya, et al.
Published: (2024)
by: Agarwal, Aishwarya, et al.
Published: (2024)
WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks
by: Bai, Hao, et al.
Published: (2026)
by: Bai, Hao, et al.
Published: (2026)
WorldAgents: Can Foundation Image Models be Agents for 3D World Models?
by: Erkoç, Ziya, et al.
Published: (2026)
by: Erkoç, Ziya, et al.
Published: (2026)
PatrolVision: Automated License Plate Recognition in the wild
by: Singhal, Anmol Singhal Navya
Published: (2025)
by: Singhal, Anmol Singhal Navya
Published: (2025)
CLAP4CLIP: Continual Learning with Probabilistic Finetuning for Vision-Language Models
by: Jha, Saurav, et al.
Published: (2024)
by: Jha, Saurav, et al.
Published: (2024)
3D Universal Lesion Detection and Tagging in CT with Self-Training
by: Frazier, Jared, et al.
Published: (2025)
by: Frazier, Jared, et al.
Published: (2025)
What's in the Flow? Exploiting Temporal Motion Cues for Unsupervised Generic Event Boundary Detection
by: Gothe, Sourabh Vasant, et al.
Published: (2024)
by: Gothe, Sourabh Vasant, et al.
Published: (2024)
Transfer Learning-Based CNN Models for Plant Species Identification Using Leaf Venation Patterns
by: Bharadwaj, Bandita, et al.
Published: (2025)
by: Bharadwaj, Bandita, et al.
Published: (2025)
SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model
by: Du, Jiayuan, et al.
Published: (2025)
by: Du, Jiayuan, et al.
Published: (2025)
fine-CLIP: Enhancing Zero-Shot Fine-Grained Surgical Action Recognition with Vision-Language Models
by: Sharma, Saurav, et al.
Published: (2025)
by: Sharma, Saurav, et al.
Published: (2025)
Post-Training and Test-Time Scaling of Generative Agent Behavior Models for Interactive Autonomous Driving
by: Seong, Hyunki, et al.
Published: (2025)
by: Seong, Hyunki, et al.
Published: (2025)
World Guidance: World Modeling in Condition Space for Action Generation
by: Su, Yue, et al.
Published: (2026)
by: Su, Yue, et al.
Published: (2026)
Snowy Scenes,Clear Detections: A Robust Model for Traffic Light Detection in Adverse Weather Conditions
by: Garg, Shivank, et al.
Published: (2024)
by: Garg, Shivank, et al.
Published: (2024)
Surgical Vision World Model
by: Koju, Saurabh, et al.
Published: (2025)
by: Koju, Saurabh, et al.
Published: (2025)
Hybrid deep learning-based strategy for the hepatocellular carcinoma cancer grade classification of H&E stained liver histopathology images
by: Deshpande, Ajinkya, et al.
Published: (2024)
by: Deshpande, Ajinkya, et al.
Published: (2024)
Optimizing Multitask Industrial Processes with Predictive Action Guidance
by: Mehta, Naval Kishore, et al.
Published: (2025)
by: Mehta, Naval Kishore, et al.
Published: (2025)
ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering?
by: Meshram, Pragati Shuddhodhan, et al.
Published: (2024)
by: Meshram, Pragati Shuddhodhan, et al.
Published: (2024)
ECoDepth: Effective Conditioning of Diffusion Models for Monocular Depth Estimation
by: Patni, Suraj, et al.
Published: (2024)
by: Patni, Suraj, et al.
Published: (2024)
Automatic Report Generation for Histopathology images using pre-trained Vision Transformers and BERT
by: Sengupta, Saurav, et al.
Published: (2023)
by: Sengupta, Saurav, et al.
Published: (2023)
Towards Driver Behavior Understanding: Weakly-Supervised Risk Perception in Driving Scenes
by: Agarwal, Nakul, et al.
Published: (2026)
by: Agarwal, Nakul, et al.
Published: (2026)
A Multimodal Dataset for Enhancing Industrial Task Monitoring and Engagement Prediction
by: Mehta, Naval Kishore, et al.
Published: (2025)
by: Mehta, Naval Kishore, et al.
Published: (2025)
MultiWorld: Scalable Multi-Agent Multi-View Video World Models
by: Wu, Haoyu, et al.
Published: (2026)
by: Wu, Haoyu, et al.
Published: (2026)
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
by: Liu, Fangfu, et al.
Published: (2026)
by: Liu, Fangfu, et al.
Published: (2026)
World2Act: Latent Action Post-Training from World Model Dynamics
by: Vuong, An Dinh, et al.
Published: (2026)
by: Vuong, An Dinh, et al.
Published: (2026)
Training-free Color-Style Disentanglement for Constrained Text-to-Image Synthesis
by: Agarwal, Aishwarya, et al.
Published: (2024)
by: Agarwal, Aishwarya, et al.
Published: (2024)
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features
by: Sengupta, Saurav, et al.
Published: (2025)
by: Sengupta, Saurav, et al.
Published: (2025)
Multimodal Engagement Analysis from Facial Videos in the Classroom
by: Sümer, Ömer, et al.
Published: (2021)
by: Sümer, Ömer, et al.
Published: (2021)
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions
by: Sengupta, Saurav, et al.
Published: (2025)
by: Sengupta, Saurav, et al.
Published: (2025)
Language-Conditioned World Modeling for Visual Navigation
by: Dong, Yifei, et al.
Published: (2026)
by: Dong, Yifei, et al.
Published: (2026)
Similar Items
-
QAL: A Loss for Recall Precision Balance in 3D Reconstruction
by: Meshram, Pranay, et al.
Published: (2025) -
Empir3D : A Framework for Multi-Dimensional Point Cloud Assessment
by: Turkar, Yash, et al.
Published: (2023) -
Humanoid World Models: Open World Foundation Models for Humanoid Robotics
by: Ali, Muhammad Qasim, et al.
Published: (2025) -
Low-Compute Watermark Removal via Dual-Domain Natural Projection
by: Meshram, Pragati Shuddhodhan, et al.
Published: (2025) -
Exploiting Contextual Uncertainty of Visual Data for Efficient Training of Deep Models
by: Agarwal, Sharat
Published: (2024)