Towards Visual Syntactical Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Chowdhury, Sayeed Shafayet, Chandra, Soumyadeep, Roy, Kaushik |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ViTALS: Vision Transformer for Action Localization in Surgical Nephrectomy
by: Chandra, Soumyadeep, et al.
Published: (2024)
by: Chandra, Soumyadeep, et al.
Published: (2024)
REMAP: Regularized Matching and Partial Alignment of Video Embeddings
by: Chandra, Soumyadeep, et al.
Published: (2025)
by: Chandra, Soumyadeep, et al.
Published: (2025)
2D-ThermAl: Physics-Informed Framework for Thermal Analysis of Circuits using Generative AI
by: Chandra, Soumyadeep, et al.
Published: (2025)
by: Chandra, Soumyadeep, et al.
Published: (2025)
Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
by: Alwis, Praditha, et al.
Published: (2026)
by: Alwis, Praditha, et al.
Published: (2026)
Semantic-Syntactic Discrepancy in Images (SSDI): Learning Meaning and Order of Features from Natural Images
by: Tao, Chun, et al.
Published: (2024)
by: Tao, Chun, et al.
Published: (2024)
Towards Scalable Modeling of Compressed Videos for Efficient Action Recognition
by: Biswas, Shristi Das, et al.
Published: (2025)
by: Biswas, Shristi Das, et al.
Published: (2025)
Dataset and Benchmark for Urdu Natural Scenes Text Detection, Recognition and Visual Question Answering
by: Maryam, Hiba, et al.
Published: (2024)
by: Maryam, Hiba, et al.
Published: (2024)
Towards More Unified In-context Visual Understanding
by: Sheng, Dianmo, et al.
Published: (2023)
by: Sheng, Dianmo, et al.
Published: (2023)
CURE: Concept Unlearning via Orthogonal Representation Editing in Diffusion Models
by: Biswas, Shristi Das, et al.
Published: (2025)
by: Biswas, Shristi Das, et al.
Published: (2025)
HEART: Hyperspherical Embedding Alignment via Kent-Representation Traversal in Diffusion Models
by: Roy, Arani, et al.
Published: (2026)
by: Roy, Arani, et al.
Published: (2026)
Now You See It, Now You Don't - Instant Concept Erasure for Safe Text-to-Image and Video Generation
by: Biswas, Shristi Das, et al.
Published: (2025)
by: Biswas, Shristi Das, et al.
Published: (2025)
ICSVR: Investigating Compositional and Syntactic Understanding in Video Retrieval Models
by: Madasu, Avinash, et al.
Published: (2023)
by: Madasu, Avinash, et al.
Published: (2023)
From Bands to Depth: Understanding Bathymetry Decisions on Sentinel-2
by: Chowdhury, Satyaki Roy, et al.
Published: (2026)
by: Chowdhury, Satyaki Roy, et al.
Published: (2026)
MAGIC: Multimodal Alignment & Grounding-aware Instruction Coreset for Vision-Language Models
by: Biswas, Shristi Das, et al.
Published: (2026)
by: Biswas, Shristi Das, et al.
Published: (2026)
EgoAVU: Egocentric Audio-Visual Understanding
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
Variance-Penalized MC-Dropout as a Learned Smoothing Prior for Brain Tumour Segmentation
by: Chowdhury, Satyaki Roy, et al.
Published: (2026)
by: Chowdhury, Satyaki Roy, et al.
Published: (2026)
Towards Two-Stream Foveation-based Active Vision Learning
by: Ibrayev, Timur, et al.
Published: (2024)
by: Ibrayev, Timur, et al.
Published: (2024)
TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy
by: Zhao, Weichao, et al.
Published: (2024)
by: Zhao, Weichao, et al.
Published: (2024)
Towards Understanding Visual Grounding in Visual Language Models
by: Pantazopoulos, Georgios, et al.
Published: (2025)
by: Pantazopoulos, Georgios, et al.
Published: (2025)
Panoptic Diffusion Models: co-generation of images and segmentation maps
by: Long, Yinghan, et al.
Published: (2024)
by: Long, Yinghan, et al.
Published: (2024)
SNAP: Towards Segmenting Anything in Any Point Cloud
by: Gupta, Aniket, et al.
Published: (2025)
by: Gupta, Aniket, et al.
Published: (2025)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
by: Song, Wei, et al.
Published: (2025)
by: Song, Wei, et al.
Published: (2025)
DCT-CryptoNets: Scaling Private Inference in the Frequency Domain
by: Roy, Arjun, et al.
Published: (2024)
by: Roy, Arjun, et al.
Published: (2024)
Real-Time Detection and Analysis of Vehicles and Pedestrians using Deep Learning
by: Sadik, Md Nahid, et al.
Published: (2024)
by: Sadik, Md Nahid, et al.
Published: (2024)
IPNET:Influential Prototypical Networks for Few Shot Learning
by: Chowdhury, Ranjana Roy, et al.
Published: (2022)
by: Chowdhury, Ranjana Roy, et al.
Published: (2022)
Towards a Universal Synthetic Video Detector: From Face or Background Manipulations to Fully AI-Generated Content
by: Kundu, Rohit, et al.
Published: (2024)
by: Kundu, Rohit, et al.
Published: (2024)
Visual Alignment of Medical Vision-Language Models for Grounded Radiology Report Generation
by: Bose, Sarosij, et al.
Published: (2025)
by: Bose, Sarosij, et al.
Published: (2025)
SlimDiff: Training-Free, Activation-Guided Hands-free Slimming of Diffusion Models
by: Roy, Arani, et al.
Published: (2025)
by: Roy, Arani, et al.
Published: (2025)
TM-PATHVQA:90000+ Textless Multilingual Questions for Medical Visual Question Answering
by: Rajkhowa, Tonmoy, et al.
Published: (2024)
by: Rajkhowa, Tonmoy, et al.
Published: (2024)
BoMuDANet: Unsupervised Adaptation for Visual Scene Understanding in Unstructured Driving Environments
by: Kothandaraman, Divya, et al.
Published: (2020)
by: Kothandaraman, Divya, et al.
Published: (2020)
Understanding the Risks of Asphalt Art to the Reliability of Vision-Based Perception Systems
by: Ma, Jin, et al.
Published: (2025)
by: Ma, Jin, et al.
Published: (2025)
QORT-Former: Query-optimized Real-time Transformer for Understanding Two Hands Manipulating Objects
by: Ismayilzada, Elkhan, et al.
Published: (2025)
by: Ismayilzada, Elkhan, et al.
Published: (2025)
RA-SSU: Towards Fine-Grained Audio-Visual Learning with Region-Aware Sound Source Understanding
by: Sun, Muyi, et al.
Published: (2026)
by: Sun, Muyi, et al.
Published: (2026)
Advancing Compressed Video Action Recognition through Progressive Knowledge Distillation
by: Soufleri, Efstathia, et al.
Published: (2024)
by: Soufleri, Efstathia, et al.
Published: (2024)
On Inherent Adversarial Robustness of Active Vision Systems
by: Mukherjee, Amitangshu, et al.
Published: (2024)
by: Mukherjee, Amitangshu, et al.
Published: (2024)
TruthLens: Visual Grounding for Universal DeepFake Reasoning
by: Kundu, Rohit, et al.
Published: (2025)
by: Kundu, Rohit, et al.
Published: (2025)
Learning Sparse Label Couplings for Multilabel Chest X-Ray Diagnosis
by: Srivastava, Utkarsh Prakash, et al.
Published: (2025)
by: Srivastava, Utkarsh Prakash, et al.
Published: (2025)
Enhancing Deepfake Detection using SE Block Attention with CNN
by: Dasgupta, Subhram, et al.
Published: (2025)
by: Dasgupta, Subhram, et al.
Published: (2025)
VUGEN: Visual Understanding priors for GENeration
by: Chen, Xiangyi, et al.
Published: (2025)
by: Chen, Xiangyi, et al.
Published: (2025)
Sanvaad: A Multimodal Accessibility Framework for ISL Recognition and Voice-Based Interaction
by: Revankar, Kush, et al.
Published: (2025)
by: Revankar, Kush, et al.
Published: (2025)
Similar Items
-
ViTALS: Vision Transformer for Action Localization in Surgical Nephrectomy
by: Chandra, Soumyadeep, et al.
Published: (2024) -
REMAP: Regularized Matching and Partial Alignment of Video Embeddings
by: Chandra, Soumyadeep, et al.
Published: (2025) -
2D-ThermAl: Physics-Informed Framework for Thermal Analysis of Circuits using Generative AI
by: Chandra, Soumyadeep, et al.
Published: (2025) -
Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
by: Alwis, Praditha, et al.
Published: (2026) -
Semantic-Syntactic Discrepancy in Images (SSDI): Learning Meaning and Order of Features from Natural Images
by: Tao, Chun, et al.
Published: (2024)